Rahul, writing on X about building agents in 2026, made a point that stuck with me: the future of AI isn’t prompt to output. It’s a goal, and then a long chain of steps underneath it. The people who understand that shift, he argues, will build things that previously felt impossible.
I agree with him, and that’s exactly why I’m skeptical that “everyone” is going to use the agents OpenAI is shipping. Not because the agents are bad. Because a goal with five invisible steps under it is a fundamentally different product to own than a chat box, and most teams are not set up to own it.
What OpenAI actually shipped
In July, OpenAI launched ChatGPT Work, an agentic platform aimed at automating workplace tasks, timed alongside the wider rollout of GPT-5. The company also disclosed that enterprise now accounts for more than 40% of its revenue and is tracking toward parity with consumer.
That second number is the one worth sitting with. OpenAI built its reputation on consumer scale, and the money is quietly rotating toward companies. Product roadmaps follow revenue. If enterprise is heading to half the business, expect agents built for procurement committees and audit logs, not for the person who wants their inbox sorted.
The infrastructure argument is real
The case for why 2026 is different, and I think it holds up, is that the plumbing finally caught up. Agents as an idea have been around for years. What changed is that models reason more reliably across multiple steps, tool integrations stopped being duct tape, and enterprise data access is less of an archaeological dig.
You can see it in deployment numbers. Surveys of the space this June found organizations running agents in production, with roughly another 30% actively building agents and concrete plans to ship them. That’s not a pilot-purgatory pattern. That’s a pipeline.
So the technical objection I would have raised in 2024, that these things fall apart on step four, is weaker now. The objections that remain are less exciting and harder to fix.
Three reasons adoption won’t be universal
- Agents create maintenance debt that nobody budgets for. A prompt fails visibly and you retry it. An agent fails halfway through a workflow, having already sent the email or updated the record. Somebody has to own the retry logic, the rollback path, and the on-call rotation. In most companies, that person does not exist yet.
- Verification costs eat the savings. If a human has to check every output, you’ve automated the typing and kept the thinking. The teams getting real gains are the ones who found tasks where a wrong answer is cheap. That’s a narrower set of work than the marketing implies.
- Platform gravity cuts both ways. Buying agents from the same vendor that supplies your model is convenient and it concentrates risk. Pricing changes, deprecations, and capability shifts all land at once. Some organizations will accept that. Plenty of regulated ones won’t.
Who this genuinely fits
My honest read, from the reviewing chair: ChatGPT Work-style platforms will do well in companies that already have clean internal systems and a tolerance for process change. Those two things together are rarer than they sound. A company with tidy data and messy politics won’t ship. Neither will the reverse.
The best early candidates are repetitive, well-documented, low-stakes-if-wrong workflows. Ticket triage. Data reconciliation with a human sign-off at the end. Drafting that a person edits anyway. Boring stuff. Boring stuff is where automation has always paid off first, and 2026 has not repealed that.
What I’d test before signing anything
If you’re evaluating an agent platform this year, the demo is not the test. Ask for the failure behavior. What happens when a tool call times out mid-workflow? Can you see the full step trace after the fact, or just the final output? Who gets paged? Can you cap what an agent is allowed to touch, and is that cap enforced by permissions or by a prompt asking nicely?
Also ask what it costs when the agent is wrong at scale. Vendors quote per-token or per-seat pricing. Nobody quotes the cost of a hundred incorrect updates that a human now has to unwind.
So, will everyone use them
No. And I don’t think that’s the interesting question. Agents are following the trajectory of every serious infrastructure tool: adopted unevenly, deepest where the underlying systems are already in decent shape, abandoned quietly where they were bought as a strategy rather than a fix for a specific problem.
OpenAI is clearly positioning for the enterprise half of that split, and the revenue mix says the bet is working. Whether it works for you depends less on GPT-5 and more on whether your team can describe a workflow precisely enough to hand it off. That’s the part no model ships with.
🕒 Published: