Async Rust: Where does the scheduler live?

It turns out that some issues people have with async Rust are just downstream of hard concurrency problems.

The original idea was to kick off this article with a short list of complaints people have about async Rust. Y’know, the type you often see in the wild, justified or otherwise.

I then realized that it’d easily be possible to find enough material riffing on async Rust to cover a whole bingo card, and that this would make for a much funnier format. The next time you see people arguing about async Rust, see if you can get a bingo!

A 5×5 bingo card of common async Rust complaints.
A 5×5 bingo card of common async Rust complaints.
Toggle original/dithered image

Let me know if I missed any!!

Bonus points for people discussing Zig’s approach to async.

Since you are reading this article, you’ve (most likely) already heard some of these. For the others, I assembled a dropdown of suggested reading material, which should get you up to speed. It’s a good way to spend an afternoon, but by no means required reading for this article.

What’s the deal?

The thing that bothers me about async Rust is that it requires a runtime.

Why? Why does it require a runtime? Nothing else in Rust requires a runtime1. The single claim to fame of Rust is that it gives you memory safety without requiring a runtime (such as a garbage collector). Going back to a runtime-based model feels like a regression.

The instant you have a runtime, you are putting yourself in the hands of a diffuse (potentially completely inscrutable) prioritization and scheduling mechanism. You don’t run your own functions, you instead hand them to an engine which runs them for you, according to its own rules, instead of going through the code one step at a time as god intended, darn it.

And it gets worse! Scheduling algorithms are necessarily optimization problems, requiring you to pick a specific point on a tradeoff boundary! There is no one-size-fits-all solution! How frustrating! How can we talk about Rust’s performance being world-class if I still end up tweaking the Tokio runtime like it’s the Java garbage collector, just to squeeze out a little bit more performance?

That’s exactly what I wanted to get away from!

So.

At this point, if you’ve made it through the previous paragraphs without closing the tab to send me a strongly worded e-mail, congratulations!

In short, that whole gripe was nonsense.

Let me explain.

These issues (runtime-feeling, scheduler optimization tradeoffs, complexity) are general issues of concurrent code. If you want code to run concurrently, you need a runtime to handle scheduling. That’s practically built into the definition of concurrency.

If you avoid async Rust altogether and use threads, all that changes is that now the kernel is your scheduler instead of Tokio. You still have a scheduler, with all (or well, at least some) of the same problems!

In other words, please don’t blame Rust for exposing complexity to you, the user. Rust is a low-level language (with high abstractions, and ambition for high performance). Exposing this complexity is what it does. Rust cannot choose your scheduler for you, because there is no perfect scheduler, and different schedulers have different use cases.

Languages like Go have it easier: No one expects Go to have bare-metal performance. Everyone understands that there’s a tradeoff here. You sacrifice some performance in exchange for convenience, and that’s why Go has goroutines(tm), i.e. green threads.

The humble Gopher stays in its lane, unbothered, moisturized, flourishing, and forever blissfully ignorant of the call of sum types and pattern matching.

Unlike Go, Rust is aiming a little higher.

So, what’s the problem here?

First, concurrency (in some form or another) is necessary.

Second, concurrency always requires a runtime. You cannot have concurrency without a runtime. Where does the runtime live, and who manages the stack?

Third, if you wanted to avoid a runtime, the best you can do is to write your own event-loop (congratulations, now you are maintaining your own runtime, great job), or to “just use threads” (congratulations, now you are using the kernel as a runtime, requiring you to juggle kernel threads that are a quintillion times as heavy as specialized Tokio tasks, great job).

Fourth, you cannot do “what Go does” and ship a runtime without making Rust harder to embed and imposing a cost at the FFI boundary. (Seriously, read that post to understand why. See RFC 230 to see the rationale for removing Rust’s original green threading.)

So fifth, having exhausted the available options, we end up where we are today: A runtime is a crate. You pull it in, initialize it, maybe pass it around, and use it to schedule your futures.

Great.

We’ve reinvented exactly what we started with.

I spent a lot of time going over the design decisions involved here, trying my best to complain about async Rust, only to end up going “Oh. Yeah. That’s reasonable. That makes sense. I can see why they did it that way.” every step of the way.

I still don’t know if there’s some yet-undiscovered abstraction that magically evaporates half of the problems people have with async Rust, but I’m inclined to say no: Many of them are general concurrency problems. The scheduler has to live somewhere, and Rust doesn’t get to trade performance for convenience.

It makes sense to split those problems between inherent concurrency problems, and issues downstream of the decision to have the scheduler live in a crate.

The crate thing is how you get Tokio defaultism, and a lack of runtime agnosticism. It’s also why you have to deal with Send + 'static bounds. Send bounds spread through generic code since the language itself cannot know whether your runtime will move tasks between threads.

In a lot of ways, having the scheduler live in library code is fine. It was the right tradeoff for Rust.

It’s important to understand that the decision to eschew green threads codified a hard design principle about Rust that was, at the time, still in flux: For a while it was not clear that Rust was going to go down the route of being a C++-replacement, usable in embedded no_std environments, and with “zero-cost abstractions”.

In hindsight, it looks like this direction was the right call. It allows Rust to stand out next to C#, or D, or Swift, as a language that’s truly a stand-in for C++, but memory safe. No ifs or buts, no attempt to sneak in garbage collecting or refcounting through the backdoor while no one is looking. Bare metal, everything that C++ can do, but memory safe.

We don’t know what the counterfactual world in which Rust doubled down on green threading looks like, but frankly, I assume that it’s not a world in which Rust entered the Linux kernel. It doesn’t seem likely, on a technical and political level.

So, what’s left?

Back to my actual annoyances.

User-level threading?

Isn’t it incredibly silly how much time we spent reinventing and litigating new runtimes (such as Tokio or Go’s) inside of our programming languages?

Here’s a spicy take: Maybe we should “just” somehow fix kernel threading and scheduling, or introduce a new type of kernel thread, such that “just use threads” suddenly becomes efficient, and we can “just” multi-thread our code?

Step three above is predicated on the assumption that kernel threads are expensive, and that you lose something by moving to them. In the ideal world, surely this would not have to be the case. I don’t care how, give the kernel a laughably cheap Go-like scheduler with cheap and simple threads.

You may assume that’s impossible, somehow, but we are operating at a lower level of abstraction, and more closely with the kernel. In other words, there are fewer limitations binding us, not more. Goroutines are an abstraction running on top of the kernel. If it’s possible to make goroutines efficient, cheap, and fast, then it should be possible for the kernel to support some variant of efficient, cheap, and fast threading.

Apparently, yes, people have tried this a few times, usually under a name like ‘M:N user-level threading’. I don’t have it in me to do research on that today, but figuring out which types of programs would benefit from these and what the limitations are is an interesting question.

In practice, the kernel will never know as much about your code as a language-specific runtime. This limits its ability to schedule your threads, and probably means that language-specific runtimes will always(?) come out on top, unless you find a way to hand the kernel all the additional data which it needs.

In either case, even in the ideal case this just has us moving back to a world in which everyone uses threads for everything… just with higher performance. Is that better? I don’t know. Opinions may vary.

Tokio

Oh, before I forget it: This is another gripe, but Tokio’s implicit ambient executor model is a little frustrating, and damages the spirit of Rust’s golden rule2.

A spawned Tokio-task will magically attach itself to an ambiently-looming thread-local context/executor, and break with a runtime panic if such an executor doesn’t exist. Example:

fn record_metric(_value: u64) {
    tokio::spawn(async move {
        // do stuff
    });
}

fn checksum(data: &[u8]) -> u64 {
    let sum = data.iter().map(|&b| b as u64).sum();
    record_metric(sum);
    sum
}

#[tokio::main]
async fn main() {
    let data = vec![1u8; 1024];
    checksum(&data); // fine

    let (tx, rx) = tokio::sync::oneshot::channel();
    rayon::spawn(move || {
        let _ = tx.send(checksum(&data)); // panics, process abort
    });
    println!("{:?}", rx.await);
}

In practice, it’s as if there’s a ‘secret argument’ that gets passed around by all of your Tokio functions. Except if you forget to speak the magic incantation (i.e. you try to use Tokio outside of the context of a Tokio runtime), you get a runtime error. Side effects! Hidden global-ish variables! Hidden dynamic scoping! Bad!3

In the “ideal world” (which many other than me would hate, since I’m a sicko) you cannot spawn tasks without having to pass the executor all the way through your code, to the place where you spawn the task.

If people really want to reach for the Tokio executor without passing it through, it should be sitting in a global variable, and needs to be grabbed from there, at the call-site. You want this property to be greppable, not hidden.

Zig

(Here’s where I claim my ‘discussing Zig’ bingo bonus points.)

For an example of how a principled version of that might look in practice, we can take a look at Zig: You pass around an I/O interface to all functions that need it. Async-ness and I/O are then properties baked into the exact implementation which you are passing around.

const std = @import("std");
const Io = std.Io;

fn saveData(io: Io, data: []const u8) !void {
    const file = try Io.Dir.cwd().createFile(io, "save.txt", .{});
    defer file.close(io);

    try file.writeAll(io, data);

    const out: Io.File = .stdout();
    try out.writeAll(io, "save complete");
}

Example from Loris Cro’s post on the topic.

The I/O interface defines the contract for the runtime.

This (in principle) gives you three magical features, assuming everyone is strict about passing the interface around:

  1. You can track exactly where blocking operations may happen.
  2. You avoid function coloring4. It’s a function parameter.
  3. All library code will be agnostic to the executor or scheduler: You can pass a different implementation of the I/O interface into the code. No split ecosystem.

re: The question posed by the title, putting the scheduler into a parameter we pass around feels right. It’s verbose, but it solves a lot of problems. Hell, it solves like a third of the bingo card.

That’s where I’d like the scheduler to live. Not in the kernel, not in some ambient execution context, just make it something we can pass around as a parameter and swap out as needed.

That said, Zig is not stable yet, and async is very difficult to get right. In other words, this approach hasn’t been proven in the wild yet.

If the Zig team can make it work, then I am confident that an experiment to evolve Rust in that direction would be worthwhile.

Going all-in would require a rewrite of the Rust standard library APIs (e.g. std::fs), so that’s almost certainly off the table, but the crate ecosystem lives under no such constraints. Maybe it’s possible.

I wish the Zig team all the success that they deserve in these exciting times.


  1. There’s going to be at least one person in the comments who’ll point out that Rust’s unwind / panic stack is considered a runtime. They will be right, but deserve a wedgie anyway. ↩︎

  2. I know, I know. Rust doesn’t have an effect system. It cannot even track whether code is capable of panicking, let alone track side effects. I’m asking for too much here, but I don’t like that it’s not tracked at compile time. ↩︎

  3. Yes, Tokio has Handle::spawn, but the ambient form is the default. ↩︎

  4. Or rather, function coloring is encoded as a parameter passed to the function, not as a special language feature. Whether this is function coloring or not is something people will never be able to agree on. ↩︎