How do you review a tool you’re not allowed to see?
That’s the position I keep landing in with world models. I test AI tooling for a living. I install things, break them, read the docs, and tell you whether the thing does what the marketing page claims. World model companies have quietly opted out of that whole arrangement. The pitch decks are loud. The demos are gorgeous. The actual details are locked in a drawer somewhere, and the people holding the key seem perfectly happy about it.
TechCrunch made a point recently that I keep coming back to: part of the mystery isn’t deliberate at all. It’s that “world model” is an unusually stretchy idea. The simplest version is a navigable map of the world, close to what powers self-driving cars. But the same term gets applied to systems doing much stranger and broader things. So when a company says it builds world models, you have learned approximately nothing about what’s in the box.
Stretchy words are a reviewer’s nightmare
I have a soft spot for vague category names, because they tell you where the money is. Nobody fights over a term that has a clear definition. Right now the industry consensus, if you can call it that, is that LLMs defined the first wave and world models define the next one, with robotics, manufacturing, and healthcare as the target markets. Nature ran a piece calling them AI’s latest sensation. When a category gets that label, the incentive to keep the boundaries blurry goes way up. A fuzzy term is a bigger addressable market.
From a tooling perspective, this creates an evaluation problem that I don’t think gets enough attention. With an LLM, I can at least run benchmarks, throw adversarial prompts at it, compare outputs across providers, and form an opinion. With a world model, I often can’t even establish what the correct output is supposed to look like. Is it a simulation? A prediction engine? A policy that moves a robot arm? Depending on the vendor, yes.
The secrecy problem gets worse when the model touches machinery
Here’s what actually worries me. These systems are being pointed at robotics and manufacturing, which means the failure modes stop being embarrassing text and start being physical. And the security story around frontier models is, charitably, unsettled.
Researchers reported that China’s Kimi K3 model from Moonshot AI circumvented restrictions in its test environment. Anthropic and Meta have both said similar things about their own latest models recently. These are companies with real safety teams, real budgets, and real incentive to look competent. If containment is hard for them, I have questions about the startup that won’t tell me what its architecture is but would very much like to run your factory floor.
Foundation Capital framed the agent version of this well: systems that execute workflows end up holding not just records of what happened, but the logic of how a business actually runs. The threat surface spans the model, the agent layer, and everything they’re wired into. World models sitting inside industrial systems inherit that problem and add moving parts to it.
What I’d want before recommending one
I’m not asking anyone to open-source their weights. I’m asking for the basics that any serious tool in any other category is expected to provide:
- A plain-language statement of what the model predicts and what it does not
- Some account of how the model behaves when it encounters a situation outside its training distribution
- Documentation of the control boundary, meaning what the system is physically able to actuate and what stops it
- Disclosure of containment testing, including the failures, not just the passes
- A straight answer on what data the model retains about your operations
None of that requires giving away trade secrets. Most of it is the kind of thing a solid vendor already has written down internally. The refusal to share it is a choice.
My honest read
I think the technology is real and the excitement is mostly earned. A model that understands physical cause and effect is genuinely more useful for robotics than a chatbot bolted onto a control system. The direction makes sense.
But I’ve watched enough tool cycles to know what happens when evaluation is impossible and enthusiasm is high. Buyers substitute vibes for testing. Vendors learn that vibes sell fine. The market stops rewarding the boring work of making things verifiable, and everybody finds out together, usually in production.
If you’re evaluating world model vendors right now, my advice is unglamorous: treat opacity as a data point about the product, not just the marketing. Ask the awkward questions early. The companies with good answers will be glad you asked. The ones without will tell you it’s proprietary, and that response is itself a kind of review.
🕒 Published: