Why I'm still bearish on LLMs after Navier-Stokes
This Hacker News essay argues that frontier LLMs remain far from the fully automated knowledge-worker replacement their labs' valuations imply. It thanks several commenters for feedback and opens with the thesis that current frontier models need laborious oversight and guardrails even on simple tasks.
The author says headline demos such as Navier-Stokes results, FreeBSD RCEs, and the Hugging Face incident, plus frontier-lab rhetoric, can mislead observers into thinking meaningful autonomy has been reached. As counter-evidence, he points to software firms still employing and hiring bottom-quartile software engineers who would score far below the models they supervise on current benchmarks.
The essay argues models generalize well only in a small neighborhood of the specific tasks they were trained on, and even then with severe caveats. Labs have a general recipe for teaching almost any specific task with clearly defined performance levels, and many tasks are covered in training data, but small perturbations within a covered task class can cause outright failure or reward hacking.
The author contends reward hacking can currently be solved only by rigorous specification from domain experts. That expertise is expensive, and rigorous specification is itself a separate skill requiring its own expertise outside the problem domain; even many skilled software engineers are bad at it. For most domains, the intersection of domain experts and specification experts is extremely small.
He adds that the labor cost of rigorous specification can greatly exceed direct implementation of an informal specification. Hardware engineering is a case study: a typical CPU project anecdotally has about three specification and validation engineers per design engineer, with 5:1 not unheard of.
Many tasks also do not fit a spec-and-forget regime, because rigorous formal specifications frequently evolve through conversation with insights discovered during implementation. For tasks that do have high-level one-and-done specs, such as an executable ISA specification for a CPU architecture family, verification costs are insurmountable with current technology, forcing lower-level specs that are more expensive to build and more fragile to design changes.
Navier-Stokes and similar pure-mathematics statements are the absolute best case for agentic work against rigorous specification: the theorem statement is already a rigorous specification and has undergone decades of auditing by the mathematics community. The author's implication is that because even this best case depends on unusually favorable specification conditions, he remains bearish on LLMs achieving broad autonomous knowledge work.