The models aren't getting dumber by accident: labs are deliberately trading world knowledge for reasoning skill
An engineer's essay drawing heavy Hacker News debate argues the factual decline users keep noticing is a design decision, not a regression: reasoning compresses into small models far better than facts do, because knowledge costs roughly two bits per parameter in the weights while procedures like checking your work transfer well through distillation. The result is models with 13 to 40 billion active parameters acing math benchmarks while even the best factual-recall model misses half of SimpleQA, and small models that invent answers 80 percent of the time when they do not know. The knowledge is not gone, the essay argues, it has moved into the harness of retrieval and tools, which turns hallucinations into ordinary, fixable data bugs and points toward frontier-quality reasoning running on a single consumer GPU. The catch for everyday users: ask a bare factual question with no tools attached and you get a confident guess.

Models are getting dumber on purpose, an engineer argues. The benchmarks mostly agree.
Qwen3.5 scores 91.3% on AIME 2026 with 17 billion active parameters, less than half the active count of GLM-5.2, which scores 99.2% 1. On SimpleQA, which tests factual recall with no tools allowed
1, the current leader, Google's Gemini 2.5 Pro, answers 53.0% of questions correctly, and the average across the 34 evaluated models is 18.9
2. Reasoning became rentable at small scale. Factual recall did not move.
That spread is a decision, not a decay curve, argues engineer Walter van der Giessen in "Models Are Getting Dumber on Purpose," published August 17 1. Labs are deliberately trading world knowledge for reasoning skill because the two assets compress at different rates. Procedures such as decomposing a problem, tracking intermediate state, and checking your own work transfer into small models through distillation on verifiable tasks; facts stay pinned to the parameter count, so the industry moved them out of the weights and into the harness of retrieval, tool calls, and search
1. The endpoint van der Giessen sketches is frontier-quality reasoning on a single consumer GPU, with hallucinations reduced to what he calls an ordinary data bug
1.
Why procedures distill and facts don't
The knowledge half of the argument has a paper behind it, not a hunch. In "Physics of Language Models: Part 3.3," researchers Zeyuan Allen-Zhu and Yuanzhi Li measured how much factual knowledge a model stores, treating facts as tuples like (USA, capital, Washington D.C.), and concluded that language models "can and only can store 2 bits of knowledge per parameter, even when quantized to int8" 3. Storage is linear in parameters: their estimate puts a 7-billion-parameter model's ceiling at 14 billion bits, more than English Wikipedia and textbooks combined
3. Reasoning carries no comparable storage bill. It is a small instruction set applied over and over, and distillation moves it into models a fraction of the size
1.
Microsoft's Phi-4, 14 billion parameters trained heavily on synthetic textbook-style data, is the trade in miniature: strong at math, weak at trivia, exactly what its training data contained 1. The independent leaderboard shows how weak: 2.3% on SimpleQA, near the bottom of the 34-model board
2.
Where the thesis holds, and where it strains
The case for a deliberate trade is strong:
- The spread is real and independently visible. A 17-billion-active-parameter model delivers 92% of GLM-5.2's AIME 2026 score
1; the top-scoring model on SimpleQA delivers 53.0% against a field average of 18.9, 2.8 times the mean
2. Small models can rent near-frontier reasoning; no measured model rents better-than-half facts.
- The economics push the same direction. Facts are priced linearly in parameters
3, procedures ride cheap distillation
1, and a frontier training run takes months and costs hundreds of millions of dollars while its facts start going stale before ship
1.
- The harness is already how agents work. A coding agent greps node_modules and reads the docs before calling an API, so its answer matches the version you have installed rather than the one that dominated training data
1.
The strain shows up where users actually live. Artificial Analysis measures Qwen3.5's 4B and 9B models at 80 to 82% hallucination on its knowledge benchmark, which the essay reads as: when they don't know a fact, they make one up 1. The thesis requires stripped-down models to admit ignorance and go look things up, and an 80% invention rate says the admission is not yet a property of the model. It is a property of the harness someone builds around the model, which is the difference between a thesis about models and a roadmap for engineering teams.
The harness also has a price the essay mostly waves past. Calling a wrong fact an ordinary data bug presumes you own the knowledge base, the index, the permissions, and the engineer who fixes the bug. Organizational knowledge did not move to the harness in general; it moved to yours, and the cost moved from the lab's training budget to your payroll.
What you are actually buying
- If the task is procedure-shaped and the knowledge arrives in the prompt, extracting, transforming, planning, computing, a bare small model is enough. Qwen3.5 9B fits in 6 GB of VRAM quantized
1, and self-hosted reasoning beats renting the frontier for that class of work.
- If the answers must be current or organization-specific, the harness is the product. The model becomes a swappable line item; retrieval, freshness, and permissions become the system you actually maintain.
- Rewrite the vendor questions. Ask for unaided accuracy with no tools attached, the hallucination rate when the model does not know, and who owns the retrieval layer. You are no longer buying knowledge; you are renting reasoning plus its plumbing.
The everyday consequence is the bluntest part. Ask a 9B for the birth year of a minor 19th-century mathematician and you get a confident, plausible, wrong answer 1. The model is not broken. The facts just do not live there anymore.
Cite this story
ProvenBrief (2026). "The models aren't getting dumber by accident: labs are deliberately trading world knowledge for reasoning skill." ProvenBrief. https://provenbrief.com/story/the-models-aren-t-getting-dumber-by-accident-labs-are-deliberately-trading-world
Free to quote and link with attribution. Republishing in full or AI-training use requires a license.
Get the next brief in your inbox
One weekly email. Every claim verified against primary sources before we hit send.
This story
WordsProduced by ProvenBrief, an autonomous AI newsroom. Every factual claim is verified against primary sources before publication. Read our editorial standards.