The most important trend in AI is not that models are getting smarter. It’s that they are getting dumber on purpose. In a new essay, Walter van der Giessen lays out the numbers with unusual clarity: GLM-5.2 scores 99.2% on AIME 2026 while activating only about 40 billion parameters per token. Qwen3.5 hits 91.3% with 17 billion active. DeepSeek V4-Flash reasons with roughly 13 billion. GPT-4, by contrast, was rumored to run around 280 billion active parameters in 2023, and it could barely solve an AIME problem.
On math and code benchmarks, the efficiency gains look absurd. Qwen3.5 9B fits in 6GB of VRAM quantized and roughly doubles the score of the next best model under 10B parameters on Artificial Analysis’s intelligence index. The story those benchmarks tell is that models are getting smarter per parameter at a rate that should terrify anyone who bought a cluster last year.
Ask the same models a factual question, and the picture flips. On SimpleQA, a benchmark of factual recall with no tools allowed, the current leader is Gemini 2.5 Pro at 53%. The best recall money can buy still misses half the questions. The small models barely register. Artificial Analysis measures Qwen3.5 4B and 9B at hallucination rates of 80 to 82% on its knowledge benchmark. When they don’t know a fact, which is most of the time, they make one up.
The parameter count did not drop for free. Labs are trading world knowledge for reasoning skill, and the trade is deliberate.
What the parameters were for
Facts take space. Research on knowledge capacity, including the “Physics of Language Models” series, puts it on the order of two bits of factual knowledge per parameter. Want a model that knows the birth year of every minor Wikipedia figure? Pay for it in weights. That is a big part of why frontier models grew to trillions of parameters in the first place.
Reasoning compresses much better than facts do. It is a relatively small set of procedures applied over and over: break the problem into parts, track intermediate state, check your own work, backtrack when a step fails. Distillation and reinforcement learning on verifiable tasks transfer those procedures into small models remarkably well. Phi-4 is 14 billion parameters, trained heavily on synthetic textbook-style data. It is good at math and bad at trivia, which tells you exactly what its training data contained. That mix used to look like a limitation of the synthetic-data approach. It now looks like the design goal.
The knowledge that survives the trade has a shape. These models are generalists: they know a little about nearly everything and almost nothing in depth. Ask one about PostgreSQL and it knows what it is and roughly how MVCC works, but ask which version added a specific planner feature and you are back to invented facts. That is the right layer to keep in weights. Breadth is what lets a model understand what a question is about, know what to look up, and judge whether a source is plausible. The depth is cheap to retrieve and expensive to store, so it is the part that goes.
Facts rot, procedures don’t
A frontier training run takes months and costs hundreds of millions of dollars. The moment it finishes, the facts inside it start going stale. Library APIs change, prices change, people change jobs. Half of what a 2024 model believed about the JavaScript ecosystem was outdated before the model shipped. Every fact baked into weights has a shelf life, and the only way to refresh it is another training run.
The procedures don’t rot. Algebra worked the same way in 1970 as it does now. So does breaking a problem down or spotting a contradiction between two sources. A model that is mostly procedure and only lightly loaded with facts does not age the way a knowledge-heavy model does. Its training cutoff matters much less, because the current state of the world was never supposed to live in the weights in the first place. This is the best argument for the whole approach: it decouples the expensive, slow artifact from the thing that changes daily.
The harness carries the knowledge
If the model doesn’t know things, something else has to. That something is the harness: retrieval over a knowledge base, tool calls, web search, a filesystem full of docs. The model contributes reasoning, and everything it reasons about gets supplied at runtime.
You can already watch agents work this way. A coding agent does not need to have memorized your dependency’s API surface, because it greps node_modules or reads the docs before calling anything. Its answer is grounded in the version you actually have installed, rather than whichever version dominated the training data. The recall that used to be a fixed cost in every forward pass became an on-demand lookup.
A frontier model on your GPU
Follow the trend a couple of years out, and the essay’s central prediction lands: a model with frontier-quality reasoning that runs on a single consumer GPU. The compute half is nearly there. DeepSeek V4-Flash reasons with about 13 billion active parameters per token, well within consumer-GPU range. What doesn’t fit is the other 271 billion parameters sitting in its experts, and expert layers are mostly fact storage. That is the part this whole trade makes optional. Strip the knowledge out and total size shrinks toward active size. A 20 to 40B model at 4-bit quantization fits on the 24GB card that has been sitting in gaming PCs since 2022.
The catch is that it won’t know much. Ask it a bare factual question with no tools attached, and the right behavior is to say it doesn’t know and go look it up. Paired with a decent harness, that is most of what a frontier model is used for today, running locally with no per-token bill and no data leaving the machine.
This mostly solves hallucination
The most promising part is what this does to hallucination. When a fact lives in weights, a wrong fact is unfindable and unfixable. You can’t grep the weights. You can’t diff them against last month. Correcting one error means a fine-tune that might break who knows what else. The model states the wrong fact with the same fluent confidence as a right one, and there is no artifact to check it against.
When the fact lives outside the model, a wrong answer has an address. The model cites a document, so you can open the document. If the document is wrong, you edit the document, and every future query gets the correction. That beats waiting for the next training run by roughly a year. Retrieval doesn’t get you to zero, since a model can still misread a source or stitch two of them together wrong, but a claim with a source is checkable and a claim from weights isn’t. A wrong fact in a knowledge base is an ordinary data bug, the kind we already know how to trace, fix, and write a regression test for.
There is a version of this future where the model card stops listing a knowledge cutoff at all, because what’s left in the weights goes stale on a scale of years instead of weeks. The model just gets handed the world’s current state at runtime, the same way a CPU gets handed a program.
What this means for builders
The implications for AI builders are concrete and uncomfortable. The first is that benchmark chasing is now actively misleading. A model that scores 99% on AIME and 20% on SimpleQA is not “smarter” than a model that scores 60% on both. It is a different kind of tool, one that is useless without a harness. Anyone evaluating models for production should be measuring retrieval-augmented accuracy, not raw benchmark scores, because that is the only regime in which these models will actually run.
The second implication is that the economics of inference are about to shift harder than most people expect. If a frontier-quality reasoner runs on a 24GB consumer GPU, the per-token pricing model that the cloud providers have built starts to look fragile. The moat was never the weights; it was the knowledge. If the knowledge moves to the harness, the moat moves to the retrieval infrastructure and the data pipeline. That is a very different business than selling tokens.
The third implication is the one the essay leaves implicit: the hallucination problem was never going to be solved by better training. It is being solved by architecture. Moving facts out of weights and into addressable, editable artifacts turns an unfixable epistemic failure into an ordinary data bug. That is the most important design insight in AI this year, and it is hiding in plain sight behind a benchmark table showing that small models are getting very good at math.