\n\n\n\n Nine Dimensions of Waiting and Why Your Agent Toolkit Keeps Tripping Over Them - AgntBox Nine Dimensions of Waiting and Why Your Agent Toolkit Keeps Tripping Over Them - AgntBox \n

Nine Dimensions of Waiting and Why Your Agent Toolkit Keeps Tripping Over Them

📖 5 min read•810 words•Updated Sep 13, 2026

Gavin Gray, Shriram Krishnamurthi, and Will Crichton open their OOPSLA 2026 paper with a claim that sounds almost too modest for how much trouble it explains: many languages now hand you async and await, and those keywords look interchangeable across languages, but they are not. The authors lay out a design space with nine dimensions, and different languages pick different points in it.

My first reaction, as someone who spends most of his week wiring AI toolkits together and writing up what breaks: finally, somebody wrote down the thing that has been costing me afternoons.

The promise is straight-line code

The pitch for async/await has always been readability. You write what looks like sequential code, sprinkle in await where you need to wait, and the runtime handles the rest. The paper’s framing of this as straight-line asynchrony is the right one. That is the whole selling point. Callbacks and explicit state machines are hard to read, and async/await makes concurrent code look like the code you already know how to read.

The paper’s example is about as small as an example can get. An async function prints “A”, awaits a two-second sleep meant to stand in for a log write, then prints “B”. You can read it top to bottom without thinking. That is the point, and it is also where the trouble starts, because reading it top to bottom tells you nothing about when the function actually starts running, who owns the task, or what happens if somebody decides they no longer want the result.

Where toolkit reviewers get burned

I test a lot of agent frameworks, SDKs, and orchestration layers. Almost all of them are async. Almost all of them are wrappers around someone else’s async runtime, in a language whose async semantics were chosen by someone else again. When the paper points at task lifecycle and cancellation as places where languages diverge, that lands hard for anyone who has debugged an agent loop.

Think about what an AI toolkit actually does with async. It streams tokens. It fans out parallel tool calls. It races a model response against a timeout. It abandons work when a user closes a tab or a supervisor agent changes its mind. Every one of those is a cancellation story, and cancellation is precisely the dimension where “it works the same everywhere” stops being true.

The failure mode I keep hitting is not a crash. Crashes are easy. It is the toolkit that keeps a request in flight after you thought you killed it, or the cleanup handler that never runs, or the parallel batch where one failure quietly leaves siblings running. The code reads like straight-line logic. The behavior is something else. The paper’s contribution, as I read it, is giving that mismatch a vocabulary instead of leaving it as folklore passed between engineers who got bitten.

What nine dimensions means for how you evaluate tools

I have not read every dimension in detail yet, and I am not going to pretend otherwise or invent a list the authors did not write. What I can say is what the existence of nine axes implies for tool evaluation, which is my job here.

  • Async/await support is not a checkbox. “Supports async” tells you almost nothing about how a library behaves under cancellation or partial failure.
  • Porting a pattern across languages is a design decision, not a translation. The idiom that is correct in one runtime can leak tasks in another.
  • Docs that never mention cancellation are a warning sign. If a toolkit’s README shows only happy-path streaming, assume the hard cases were not designed, just deferred.
  • Your tests probably do not cover this. Timeout and abort paths are the least tested part of most integrations I look at, mine included.

Why this deserved more than nine points

The Hacker News thread for this paper sat at nine points when I looked. That is a fair signal of how programming languages research usually travels: quietly, until the ideas show up years later in a language’s release notes and everyone treats them as obvious.

My honest take is that this work is more useful to practitioners than its ranking suggests. Not because it hands you a fix, but because it gives you the questions to ask. When I review the next agent framework and the async behavior feels off, I now have a better move than shrugging and adding a timeout. I can ask which point in the design space this library assumed, and whether that matches the runtime I am dropping it into.

The paradigm was supposed to make concurrent programming simpler. It did, at the level of reading code. Underneath, the choices got moved rather than removed. Knowing where they went is the difference between a toolkit you can trust in production and one that only looks trustworthy in a demo.

đź•’ Published:

đź§°
Written by Jake Chen

Software reviewer and AI tool expert. Independently tests and benchmarks AI products. No sponsored reviews — ever.

Learn more →
Browse Topics: AI & Automation | Comparisons | Dev Tools | Infrastructure | Security & Monitoring
Scroll to Top