ChatGPT broke Diogo Almeida’s heart. That’s the framing TechCrunch used, and it’s a strange thing to read about a person who helped build the thing. Almeida was an OpenAI researcher who worked on the chatbot and helped invent reinforcement learning from human feedback, the training technique that arguably made modern assistants usable at all. Then he went quiet for two years, and came back arguing that today’s large language models are structurally inefficient.
Not wrong. Not overhyped. Structurally inefficient. That’s an engineer’s complaint, not a marketer’s, and it’s the reason I’m paying attention.
The result is Jev, announced September 15, 2026, from a company called TypeSafe AI. It’s a frontier model that cannot write a sentence. What it does instead is return typed probabilistic decisions for software. No prose. No chat window. No personality. You ask it something with a defined answer shape, and it hands back a decision with a probability attached, in a type your program already understands.
Why a model that can’t talk is interesting
Most of what I test in my day job is not conversation. It’s classification dressed up as conversation. Is this support ticket a refund request or a bug report? Does this transaction look fraudulent? Should this document get routed to legal? Every one of those is a decision with maybe five valid outcomes, and the standard approach right now is to ask a text-generating model to produce JSON and then pray.
Anyone who has shipped that pattern knows the failure modes:
- The model returns a field name you didn’t ask for
- The model returns valid JSON wrapped in a markdown fence, or a polite preamble
- The model invents a category that isn’t in your enum
- Your confidence score is a number the model wrote because you asked for a number, not because it means anything
That last one is the quiet problem nobody talks about enough. When you ask a text model how confident it is, you get a token sequence that looks like confidence. It’s vibes in a decimal costume. A model designed from the start to return probabilities instead of words is attacking a real gap, not a hypothetical one.
The other pitch is efficiency. If you only need a decision, generating a few hundred tokens of explanation around that decision is wasted compute. TypeSafe AI claims Jev is faster and more efficient than previous models. Directionally that makes sense to me: less output means less work. I have not seen benchmark numbers I’d stand behind, so I’m treating the specific magnitude as an open question until someone independent measures it.
Where my skepticism sits
The “can’t hallucinate” framing, which showed up in coverage the day after launch, is doing a lot of work. A model that can only return values from a defined set cannot produce a made-up category. That’s real and it matters. It is not the same as being right. Constrained output means the wrong answer will now be a well-typed wrong answer, which is easier to handle in code and harder to spot in review. If your classifier confidently returns the wrong enum value with a clean 0.91 attached, nothing in your pipeline is going to raise an eyebrow.
So the question I’d want answered before putting this in a production path is calibration. Does a 0.7 from Jev actually mean it’s right about 70% of the time? That’s the whole value proposition. Typed output is a developer-experience win. Calibrated probability is an engineering win, because it lets you set real thresholds and route the uncertain cases to a human. One of those you can verify in an afternoon. The other takes a labeled evaluation set and some patience.
I’d also want to know how it handles the cases that don’t fit. Real inputs are messy. Sometimes a ticket is both a refund request and a bug report. Sometimes it’s neither. A system that must return one of your defined types needs a clean way to say “none of these” or “I’m not sure,” and how gracefully it does that separates a useful tool from an expensive coin flip.
What I’d actually do with it
If I were evaluating this for a team, I’d pick one high-volume classification job currently running through a chat model, shadow it with Jev, and compare on three things: agreement rate, latency, and whether the confidence scores track reality. That’s a week of work and it tells you everything. Don’t rewrite your stack around a September launch.
What I like most here is the underlying argument. Almeida is saying we bolted general-purpose text generation onto problems that never needed text, and paid for it in speed, cost, and reliability. Whether or not Jev is the right answer, that critique is worth taking seriously. The developer excitement around it suggests a lot of people have been quietly annoyed by the same thing.
đź•’ Published: