\n\n\n\n Your LLM Writes Code Fine, Just Don't Let It Write About It - AgntBox Your LLM Writes Code Fine, Just Don't Let It Write About It - AgntBox \n

Your LLM Writes Code Fine, Just Don’t Let It Write About It

📖 4 min read•784 words•Updated Sep 19, 2026

Stop asking the model for comments.

That’s the single biggest shift I’ve made in my own workflow this year, and I’ve spent enough time testing tools for agntbox to say it with some confidence. In 2026, LLMs assist with coding by generating both code and documentation. The generation part works. The documentation part is where I keep pulling the plug and doing it myself.

I’m not alone. A widely shared piece titled “2x, not 10x: coding with LLMs in 2026” puts it about as bluntly as you can: “Never write READMEs, docstrings, or comments. I will write those myself later. And yes, I really mean this.” That’s a rule in a system prompt, not a hot take. And the response to it has been a lot of developers nodding along because they’d quietly arrived at the same place.

Why the documentation ban makes sense

The reasoning is less about model quality than about what comments are for. A comment explains intent. It answers why this weird branch exists, why the retry count is three instead of five, why you didn’t use the obvious library. A model generating a comment from the code it just wrote is describing what the code does, which is information the code already contains. You end up with a file twice as long and no better understood.

There’s a second problem that bites later. Generated docstrings and READMEs go stale in a specific, annoying way. They read as authoritative because they’re well-formatted, so nobody questions them, and the drift between the prose and the behavior becomes a trap for whoever touches the file in six months. Sparse code with no comments is honest about what it doesn’t tell you. Confidently wrong documentation is not.

So the working split looks something like this: the model produces implementation, you produce explanation. The model is fast at the part that has a verifiable right answer. You’re better at the part that requires knowing what you were trying to do.

Plan first, generate second

The other consistent finding across serious workflows is that prompt quality dominates everything else. Addy Osmani, describing his own coding workflow going into 2026, names the common failure directly: starting code generation from a vague prompt. His first step is brainstorming a detailed specification with the AI, then outlining from there. Best practices in general point the same way, toward detailed planning and example-driven prompts.

This inverts what a lot of people expect from these tools. The appeal was supposed to be typing a sentence and getting a feature. What actually works is spending real time on the spec, the constraints, and a couple of concrete examples of the input and output you want, and then letting generation be the cheap step. The thinking doesn’t get outsourced. It gets front-loaded.

Which brings up a use case I find underrated. Models are genuinely good at reviewing a document you wrote against a checklist. As one talk on LLM architecture in 2026 framed it, you hand over your article or your list as markdown and say, instead of me checking all these things, here it is, make sure it holds up. That’s a review pass, not a generation pass. Different task, much higher hit rate.

Where the current models land

Capability isn’t the constraint anymore, at least not on length or coherence. Early models like GPT-1 would fall apart into nonsense after a few sentences. Today’s leading models, GPT-6 and MiniMax M3 among them, can produce tens of thousands of words that hold together, or write working code. The failure mode has moved. It’s no longer incoherence, it’s confident output that’s subtly wrong about your specific situation, which is harder to catch precisely because it reads well.

That’s also why the evaluation tooling space has filled out. There are roundups of LLM evaluation tools worth knowing in 2026 for a reason: once output looks good by default, you need something other than vibes to tell whether it’s correct.

What I’d actually do

If you take one thing from this, make it the sequencing. Spec, then examples, then generation, then your own prose on top. Turn off the documentation generation, in your system prompt if your tool supports it, and accept that you’ll write the README yourself. It’s twenty minutes you’d spend anyway re-reading generated text and deciding whether to trust it.

The honest framing of all this is the one in that title. Two times, not ten times. These tools are a solid multiplier on work you already understand how to structure. They are not a replacement for understanding it. The people getting the most out of them in 2026 aren’t the ones prompting the least. They’re the ones who plan hardest before they let the model type.

🕒 Published:

🧰
Written by Jake Chen

Software reviewer and AI tool expert. Independently tests and benchmarks AI products. No sponsored reviews — ever.

Learn more →
Browse Topics: AI & Automation | Comparisons | Dev Tools | Infrastructure | Security & Monitoring
Scroll to Top