Gavin Gray and his co-authors open their new paper with a sentence that sounds almost boring: many programming languages now provide async and await for expressing concurrency. Then they spend the rest of the paper explaining that those keywords do not mean the same thing everywhere. Nine design dimensions, by their count, shape how a task actually behaves once you write it.
My reaction, as somebody who spends most of his week installing agent toolkits and watching them misbehave: finally. I have been filing bug reports against this problem for a year without knowing it had a name.
Same keyword, different machine
The paper is “A Design Space Exploration of Async/Await” by Gray, Shriram Krishnamurthi, and Will Crichton out of Brown’s Cognitive Engineering Lab, headed to OOPSLA 2026, preprint posted to arXiv on August 21. It landed on Hacker News and picked up a modest nine points, which tells you something about how programming language research travels compared to a model release.
The pitch of async/await has always been straight-line asynchrony. You write code that reads top to bottom, the runtime handles the waiting, and you skip the callback pyramid. The paper’s argument is that the reading-top-to-bottom part is the only thing that generalizes. Underneath it, languages disagree on task lifecycle and on cancellation, and those disagreements are where your code breaks.
Their running example is about as small as an example gets. An async function prints “A”, awaits a two-second sleep meant to stand in for a log write, then prints “B”. Another async function calls it. That is the whole setup, and it is enough to expose the fault lines. When does the body start running, at the call or at the await? What happens to the pending sleep if the caller goes away? Does “B” ever print? You cannot answer any of that from the source text alone.
Why toolkit reviewers should care
Almost every AI agent framework I test is async underneath. Tool calls, model requests, retries, parallel fan-out across sub-agents, streaming responses. Concurrency is not a feature of these toolkits, it is the substrate.
And the failures I keep hitting cluster around exactly the dimensions the paper names. A user cancels a run and the framework reports it stopped, but a tool call keeps going and writes to a database. A parent task returns while a child is still in flight and the child’s error surfaces nowhere. A wrapper library ported from one language’s async model to another looks like a faithful translation and quietly changes when work begins.
I have written those up as toolkit bugs. Some of them are. But a good share are the framework author inheriting semantics from their language and assuming everyone shares them. Cross-language SDKs are the worst offenders, because the Python client and the TypeScript client can present identical APIs while disagreeing about what cancellation means.
What I am adding to my review checklist
Reading the paper’s framing changed how I plan to poke at async-heavy toolkits. Concretely:
- Cancel a run mid-tool-call and check for side effects that landed after the cancel returned.
- Kill a parent task with children in flight and see whether errors and cleanup propagate or vanish.
- Compare the same operation across a toolkit’s language bindings before trusting that behavior transfers.
- Check whether “timeout” in the docs means the task stopped or just that the caller stopped waiting.
- Look for whether cancellation is documented at all. Most of the time it is not, and that absence is itself a finding.
None of that is new advice in principle. What is new is having a vocabulary for it, so a bug report can say which dimension the toolkit got wrong instead of “cancellation is weird.”
Honest limits
I should be clear about what I have and have not verified. I have read the paper’s framing and its opening example, not run its artifact against every language it covers. I am not going to list which language sits where on which axis, because getting that wrong would be worse than not saying it. If you want the per-language detail, the arXiv preprint and the artifact are the source, and they go considerably deeper than a review blog needs to.
I also do not think this fixes anything on its own. Design space papers describe; they do not force library authors to document their cancellation story. The realistic outcome is that a few framework maintainers read it, recognize their own ambiguity, and write it down. That would already be an improvement over the current state, where the answer to “what happens if I cancel this” is usually an experiment.
The keywords promised that concurrent code would read like ordinary code. It does. The paper’s contribution is showing how much that pleasant surface hides, and giving us nine specific places to look. For anyone reviewing tools built on top of it, that is a useful map.
đź•’ Published: