\n\n\n\n A DeepMind Spinout Says It Beat the Giants at Their Own Homework - AgntBox A DeepMind Spinout Says It Beat the Giants at Their Own Homework - AgntBox \n

A DeepMind Spinout Says It Beat the Giants at Their Own Homework

📖 4 min read•719 words•Updated Aug 23, 2026

When was the last time a startup claimed to outperform Anthropic and OpenAI and you actually believed it? If your answer is “never,” you’re not cynical — you’re experienced. I’ve reviewed enough AI tools on this site to know that benchmark press releases and real-world performance often live in different zip codes. So when Inherent, a British AI lab founded by DeepMind alumni, announced in 2026 that its AI “teammate” Faraday outperformed both Anthropic and OpenAI at replicating research, my first instinct wasn’t excitement. It was a raised eyebrow.

But let me be fair, because there are a few reasons this one deserves more than a reflexive shrug.

Why Research Replication Is an Interesting Test

Most AI benchmarks measure things that are easy to game. Multiple-choice science questions. Coding puzzles with clean pass/fail criteria. Labs train against these targets, scores creep up, and users like you and me discover that the shiny number doesn’t translate to the messy work we actually do.

Research replication is a different kind of challenge. To replicate a study, an agent has to read a paper, understand the methodology, reconstruct the experiment, run it, and evaluate whether the results hold. That’s a long chain of interconnected tasks, and a failure at any link breaks the whole thing. You can’t shortcut your way through it with pattern matching. As benchmarks go, it’s one of the harder ones to fake, which is exactly why Inherent choosing this battlefield caught my attention.

If Faraday genuinely does this better than the models from Anthropic and OpenAI — the two labs that have dominated agentic AI conversations — that’s a meaningful signal, not just marketing noise.

Notice the Word “Teammate”

Inherent isn’t calling Faraday an assistant, a copilot, or a chatbot. They’re calling it a teammate. That framing matters more than it might seem.

An assistant waits for instructions. A teammate carries a workload. The positioning suggests Inherent wants Faraday judged on whether it can own a task end to end — the way a junior researcher would — rather than whether it can autocomplete your sentences. Research replication fits that pitch neatly: it’s the kind of grunt work that consumes enormous amounts of human researcher time and that most people would happily hand off.

As someone who tests these tools for a living, I’ve noticed the “teammate” framing is where the industry is heading. The question is always whether the product lives up to the noun.

What We Don’t Know Yet

And this is where my reviewer instincts kick back in. The announcement tells us Faraday outperformed the big labs at research replication. It doesn’t tell us a lot of other things I’d want answered before recommending anyone build a workflow around it:

  • What was the margin? “Outperformed” can mean a decisive lead or a rounding error. Those are very different stories.
  • Which models did it beat? Anthropic and OpenAI ship a range of systems. Beating a flagship is one thing; beating a mid-tier model is another.
  • Who ran the evaluation? Self-reported benchmarks deserve scrutiny until independent parties can reproduce them — which, given the subject matter here, would be fitting.
  • Can regular teams use it? A tool that shines in a lab evaluation and a tool I can put in front of readers are not the same product.

None of these gaps mean Inherent is bluffing. They mean the story is incomplete, and incomplete stories are where hype grows.

DeepMind’s Diaspora Keeps Delivering Headlines

Context helps here. DeepMind alumni have founded dozens of European startups in recent months, and Inherent is part of that wave. There’s a pattern forming: researchers leave one of the most credentialed AI labs on the planet, start something focused and narrow, and go after the incumbents in a specific domain rather than trying to build another everything-model.

That strategy makes sense to me. The giant labs are spread across consumer products, enterprise deals, safety research, and infrastructure. A small team betting everything on one capability — like autonomous research replication — can plausibly out-execute them in that lane. Whether that lane leads to a durable business is a separate question.

My Take as a Reviewer

Faraday goes on my watchlist, not my recommendation list. The claim is credible enough to take seriously: a hard benchmark, a focused team, real pedigree. But I’ve watched too many announcement-day champions stumble once actual

đź•’ Published:

đź§°
Written by Jake Chen

Software reviewer and AI tool expert. Independently tests and benchmarks AI products. No sponsored reviews — ever.

Learn more →
Browse Topics: AI & Automation | Comparisons | Dev Tools | Infrastructure | Security & Monitoring
Scroll to Top