Models Are Getting Dumber on Purpose
Reasoning scores on math and code benchmarks are climbing while per-token compute drops dramatically. GLM-5.2 reaches 99.2% on AIME 2026 with about 40B active parameters per token, Qwen3.5 hits 91.3% with 17B, and DeepSeek V4-Flash runs 13B active — compared to GPT-4's rumored ~280B active parameters in 2023, which could barely solve an AIME problem. Qwen3.5 9B fits in 6GB VRAM quantized and roughly doubles the score of the next best sub-10B model on Artificial Analysis's intelligence index.
The picture flips on factual recall. On SimpleQA, the current leader is Gemini 2.5 Pro at 53%, so the best model still misses half the questions. Small models are far worse: Artificial Analysis measures Qwen3.5 4B and 9B at 80–82% hallucination rates on its knowledge benchmark, meaning they fabricate facts most of the time. Ask the 9B for the birth year of a minor 19th-century mathematician and you get a confident, plausible, wrong answer.
The trade is deliberate. Facts take space — research from the 'Physics of Language Models' series estimates about two bits of factual knowledge per parameter. Knowing every minor Wikipedia figure, Dutch municipality population, or npm function argument order costs in weights, explaining why frontier models grew to trillions of parameters. Reasoning compresses much better because it's a small set of procedures — break problems into parts, track state, check work, backtrack — and distillation plus RL on verifiable tasks transfers these procedures to small models effectively. Phi-4 (14B params) is good at math and bad at trivia, suggesting this is the design goal rather than a limitation.
The knowledge that survives the trade is broad but shallow. Models know a little about nearly everything but almost nothing in depth — they can describe PostgreSQL and MVCC roughly, but invent answers when asked which version added a specific planner feature. That breadth is kept in weights because it's what lets the reasoning procedures apply across domains.