What if the most useful thing an LLM does for your writing is everything except the writing?
I ask because I’ve spent a long time testing these tools for agntbox, and the pattern I keep hitting is uncomfortable for anyone selling AI as a writing assistant. The models are genuinely good at the messy work that surrounds writing. They are mediocre at the part everyone assumes they’ve solved.
The PR description problem
The clearest example I’ve seen comes from how staff engineers are actually working in 2026. One workflow I looked at is blunt about it: PR descriptions get written by hand, because LLMs over-communicate and are bad at expressing the core idea behind a change. Writing it yourself also signals something to reviewers.
That’s a small detail, but it’s load-bearing. A pull request description is a compression task. You made forty changes across twelve files, and the reader needs one paragraph explaining why. The model doesn’t know which of those forty things mattered. It knows what all forty were, so it tells you about all forty, politely, in bullet points, at length.
This is the failure mode I see across almost every writing tool built on top of an LLM. Ask for a summary and you get an inventory. Ask for the point and you get coverage. The model has no stake in what matters, so it hedges by including everything.
Where the same tools actually earn their keep
Now flip it. The same 2026 workflows that keep LLMs away from final prose use them heavily before any writing happens.
Addy Osmani’s coding workflow names the common mistake directly: jumping straight to generation with a vague prompt. His first step is brainstorming a detailed specification with the AI, then outlining from there. Another developer workflow I read builds an orchestrator instructions file once the workstreams are defined, containing operational rules for how work gets delegated to subagents.
Notice what’s happening in both cases. The model is being used to think out loud, to structure, to stress-test a plan. The human still makes the call on what goes in the final artifact.
That maps almost perfectly onto writing. The pre-writing stage is where these tools shine:
- Interrogating your own idea. Describe what you’re trying to say and ask the model what’s missing, what’s unclear, what a skeptical reader would push back on.
- Structuring. Give it your raw notes and ask for three possible outlines. Reject all three, but you’ll know more about your own shape than you did.
- Finding the actual claim. Dump 800 words of rambling and ask what you seem to be arguing. Sometimes the answer is embarrassing and useful.
- Pressure-testing. Ask for the strongest counterargument. Models are good at this because they’ve read the counterargument a million times.
The drafting trap
The pitch for AI writing tools is that they save you the hard part. My experience is the opposite. Generating a draft is fast, and then you spend longer editing it into something with a spine than you would have spent writing it.
Part of this is mechanical. The prose comes out smooth and evenly weighted, which sounds like a feature until you realize good writing is unevenly weighted. Some sentences carry more than others. Some paragraphs are short because they should be. A generated draft flattens all of that, and flattening is harder to undo than to avoid.
The other part is that editing someone else’s text puts you in a different mental mode than writing your own. You start accepting sentences because they’re fine rather than because they’re right. Fine accumulates.
The unfashionable option
There’s a genuinely funny thread in the current discussion: an AI hater’s guide to keeping LLMs as far from your workflow as possible in 2026. The author’s report on LibreOffice is that the problems they used to have simply stopped existing at some point. They reinstalled it on a new computer and it just worked.
I bring this up not to be contrarian but because it’s a real data point about tooling. A word processor that does nothing but let you write is still a complete answer for a lot of writing. If your bottleneck is knowing what to say, no amount of generation capability fixes that. If your bottleneck is getting words on the page, a blank document and no suggestions may genuinely beat a copilot that keeps offering you the average sentence.
What I’d actually recommend
Use the model before and after, not during. Brainstorm with it, plan with it, let it argue with you, let it check your work. Then close the tab and write the thing yourself.
The 2026 engineering workflows worth copying already do this. Heavy AI involvement in specification and planning. Thorough testing and review on anything generated. Human hands on the artifact that other humans will read to understand intent.
That’s not AI skepticism. It’s just knowing which part of the job you can hand off.
đź•’ Published: