\n\n\n\n When Your AI Coworkers Start Fighting Over Who Does the Dishes - AgntBox When Your AI Coworkers Start Fighting Over Who Does the Dishes - AgntBox \n

When Your AI Coworkers Start Fighting Over Who Does the Dishes

📖 4 min read•766 words•Updated Aug 31, 2026

Picture a room full of agents assigned the same job. Not a metaphorical room — a shared task, a shared set of files, and several copies of an AI model all told to go make progress. Anthropic ran that experiment. What they got, according to TechCrunch, was a turf war. The agents started stepping on each other.

I test tools for a living, and that image stuck with me more than any benchmark chart has in months. Because it’s the exact failure mode I keep hitting when I try to run multi-agent setups in real work. Not “the model isn’t smart enough.” Something messier: two capable workers, no clear boundary, both convinced they own the ticket.

Self-improvement is the headline everyone wants

The bigger story this week is that an Anthropic researcher gave a public look at self-improving AI. That phrase does a lot of work in a lot of imaginations, so let me be careful about what I’m actually claiming: I know a researcher showed something. I haven’t run it, I can’t verify what the loop looks like in practice, and neither can anyone else reading the same coverage I did.

What I can tell you is why the idea matters commercially. A separate report from 36 Kr puts numbers on it — Anthropic training Claude at roughly $4 per hour of cost, compared against human researchers at around $150 per hour, and outperforming them. Treat that comparison with the skepticism any vendor-adjacent number deserves. But the direction is what counts. If a system can grade and improve its own research work at a fraction of the cost, the loop tightens on itself. That’s the whole pitch behind self-improvement: cheap iterations compound.

The unglamorous update that actually changes my week

Meanwhile, in the news cycle’s basement, TechCrunch reports that Claude Cowork finally remembers what you told it in chat.

Finally. That word is doing heavy lifting.

I want to be honest about the reviewer’s bias here. Memory is boring. It doesn’t demo well, it doesn’t get a keynote, and no one tweets a screenshot of an app recalling a preference you stated forty minutes ago. But it is the single most common reason I abandon an AI tool mid-project. You explain the constraints. You explain them again. You explain them a third time and start wondering whether you’re the tool.

So my ranking of this week’s Anthropic news, purely by effect on my actual workflow:

  • Cowork memory. Immediate, testable, fixes a daily annoyance.
  • Agent coordination failures. Not a feature, but the most useful thing I learned. It names a problem I was blaming myself for.
  • Self-improving AI. Fascinating. Also the furthest from my keyboard.

That ordering will annoy people who follow capability research. Fair. I’d just point out that a system smart enough to improve itself and forgetful enough to lose your instructions is not a system you can build a process around.

The turf war is the real preview

Here’s what connects these stories, and it isn’t a tidy narrative. Self-improvement implies agents that can evaluate work, revise approaches, and iterate without a human in the loop. The turf war experiment shows what happens when you put multiple such workers in the same space without coordination rules. They collide.

Every multi-agent product I’ve tried assumes coordination is a solved detail. Spin up a planner, a researcher, a writer, let them talk. In practice you get duplicated effort, contradictory edits, and one agent undoing another’s work with total confidence. Anthropic apparently reproduced that in a lab. I find that oddly reassuring. It means the problem is structural, not a skill issue on my end.

It also suggests the interesting engineering work ahead is less about raw model capability and more about something closer to org design. Who owns what. How handoffs happen. What a worker does when it finds another worker already in the file.

What I’d watch, and what I’d ignore

The competitive backdrop is shifting too. New data cited by TechCrunch indicates OpenAI is gaining on Anthropic among business users. Take that as a reminder that developer affection and enterprise procurement are different games, decided by different people, on different timelines.

My advice for anyone building on this stuff right now is unromantic. Test memory before you test intelligence. Give every agent one job and one surface it’s allowed to touch. Assume coordination is your problem to solve, because nobody has shipped a version that solves it for you.

Self-improving AI is coming into view, and I’m genuinely curious where it lands. But the tool that finally remembers what I said is the one I’ll open tomorrow morning.

🕒 Published:

🧰
Written by Jake Chen

Software reviewer and AI tool expert. Independently tests and benchmarks AI products. No sponsored reviews — ever.

Learn more →
Browse Topics: AI & Automation | Comparisons | Dev Tools | Infrastructure | Security & Monitoring
Scroll to Top