Don’t be fooled—LLMs don’t reason
The piece revisits a March 2016 match in Seoul, where the author watched a program he helped build play Move 37 in game two against Lee Sedol. The move looked so absurd that some commentators suspected a programming glitch, but it was not: AlphaGo won the game and the five-game match 4-1 against Lee, one of the greatest professional Go players. Lee said afterward that he had thought AlphaGo was based on probability calculation and was merely a machine, but the move changed his mind: 'Surely, AlphaGo is creative.'
The essay contrasts Go with chess. When Deep Blue beat Garry Kasparov in 1997, it looked six to eight moves ahead per player and evaluated 200 million chess positions per second using rules hard-coded by humans. Go is vastly more complex: a stone's worth depends on how distant groups and territory unfold over dozens of moves, and computing even a fraction of possible outcomes would take a supercomputer billions of years. To win, AlphaGo had to sense who was ahead at a glance and invent moves no human had thought to play. Many accounts therefore portray Move 37 as pure machine intuition, but the author says that is a misunderstanding: it was AlphaGo's powers of reasoning that produced the creative choice—powers today's AI lacks. The author argues that future AI systems need genuine reasoning of this kind to produce trustworthy results and really novel insights in fields like science and medicine.
AlphaGo consisted of two systems. Its policy network was trained to guess what move a strong human would play; this intuitive part saw Move 37 as nothing special, a play an expert human had roughly a one-in-10,000 chance of making. What made AlphaGo choose it was its search machinery, which looked beyond immediate plausibility and weighed the future consequences of proposed moves. It explicitly constructed and searched a game tree with thousands of branches, each representing a possible future.
The essay invokes Daniel Kahneman's distinction between two modes of human thought: System 1 is fast, gut-level, and effortless; System 2 is slow, step-by-step, and deliberative. AlphaGo offered a machine analogue: its networks supplied hunches—this move looks promising, this position looks won—while its search supplied deliberation, testing those hunches against ensuing moves and countermoves. As in human cognition, neither half works alone: intuition alone would never have chosen Move 37, and brute-force search would have struggled to sieve through all possible moves. This is strikingly different from today's AI models, the author says: a large language model picks the next token over and over—a process he contrasts with the deliberation AlphaGo used.