\n\n\n\n Small Models, Big Claims, and a GitHub Repo Called Kev - AgntBox Small Models, Big Claims, and a GitHub Repo Called Kev - AgntBox \n

Small Models, Big Claims, and a GitHub Repo Called Kev

📖 4 min read•794 words•Updated Sep 22, 2026

Kev is the most interesting thing I can’t yet recommend, and the reason is that almost nobody writing about it has read the code.

Let me back that up, because it matters more than the model sizes do.

What Kev actually is

Kev is a small family of decision models built on top of Qwen3.5, released in 2026 by Jared Palmer. It comes in three sizes: 0.8B, 4B, and 9B parameters. The pitch is that you can train and run the whole thing yourself, on your own hardware. It’s described as Jev-like, meaning it follows the same general idea as Jev but arrives there through different training methods.

That’s the verified core. Three sizes, one base model, local training and inference, a different training approach than the thing it resembles. Everything past that point in most of the coverage I’ve read is inference dressed up as reporting.

The sourcing problem

Here’s what pushed me from curious to cautious. explainx.ai published a Kev explainer and, to their credit, admitted something unusual in it: their earlier coverage of six Jev clones that shipped inside 48 hours was itself assembled from digest headlines, with no primary source to check against. That’s a publication telling you its own prior work was built on secondhand summaries.

I respect the disclosure. I also think it should change how you read every other Kev writeup, including the ones with confident benchmark talk. When a story about a model can be traced back to a headline about a headline, the numbers in it are decoration.

Meanwhile the repo itself contains planning documents — PLAN.md, PLAN_Qwen35.md — with notes referencing things like a corrected 35B MMLU-Pro figure and an autoresearch run. Those are working notes from someone building in public, dated September 2026. They’re not a model card and they’re not a claim aimed at reviewers. Treating scratch files as marketing copy is how bad numbers get laundered into articles.

Why the attention is still earned

Kev hit 370 points and 164 comments on Hacker News. That’s real interest, not bot noise, and I don’t think it’s hard to explain.

Jev’s whole premise, as described by people who’ve actually reproduced it, is giving up free-form text generation in favor of decisions. That’s a narrower job than chat, and narrow jobs are exactly where small models stop being toys. If your model’s output space is a choice rather than a paragraph, you don’t need 70B parameters to be useful. A 0.8B model that reliably picks the right branch is worth more in production than a 9B model that writes charming prose about the branch it might pick.

So the size ladder here makes sense to me as an engineering decision:

  • 0.8B — small enough to run on a laptop CPU or an edge box, the size where local actually means local
  • 4B — the size that fits comfortably on consumer GPUs with room for a training loop
  • 9B — the ceiling for a single-card setup, where you’d go if the small ones fall short on your task

And building on Qwen3.5 rather than training from scratch is the unglamorous right call. You inherit a solid base, you spend your effort on the decision-specific training, and anyone who wants to audit your work has a known starting point to compare against.

What I’d want before putting this in a stack

Three things, none of which I can confirm from what’s public and verified right now.

First, evaluation I can reproduce. Not a number in a planning file, but a script I can run against a task I care about, with the seed and the data.

Second, clarity on what “different training methods” means in practice. That phrase is doing a lot of work in every description of Kev, and the difference between Jev and Kev lives entirely inside it. If the methods differ, the failure modes differ, and failure modes are what you actually deploy around.

Third, honest cost numbers for the training path. “Train it yourself” is the headline feature. Whether that means an afternoon on one card or a weekend on four changes who this is for.

My read

Kev looks like the kind of project I usually end up liking: modest scope, open weights and code, built by someone who ships and keeps their notes in the open. The shape of it is right. Small decision models trained locally are a genuinely practical direction, and there’s a real appetite for it.

But I review tools, not vibes, and right now Kev’s public reputation is running well ahead of its verified record. Clone the repo, read the code, run it on your own task, and form your own view. Just don’t let a secondhand explainer do that thinking for you — one of them already admitted it wasn’t equipped to.

đź•’ Published:

đź§°
Written by Jake Chen

Software reviewer and AI tool expert. Independently tests and benchmarks AI products. No sponsored reviews — ever.

Learn more →
Browse Topics: AI & Automation | Comparisons | Dev Tools | Infrastructure | Security & Monitoring
Scroll to Top