Other MIT Technology Review (AI)

AI models flub these intelligence tests. Can you fare any better?

puzzlesspatial reasoningLLM limitationsAI evaluation

Puzzles and games have been central to AI development since the field's earliest days, from Arthur Samuel's 1959 checkers-playing algorithm that popularized the term 'machine learning' to chess and Go as benchmark test beds. The article highlights how quickly AI's puzzle-solving skills are improving: a Columbia University study found that in late 2024, even the best models could solve only 18% of New York Times Connections puzzles, yet by early 2025 some models could solve them near perfectly.

The piece offers seven tests designed to expose persistent gaps in machine cognition. Spatial reasoning is one major weak spot: mental rotation problems, standard on IQ tests, still trip up language models even when they can analyze visual inputs, because they cannot manipulate 3D objects the way architects and mechanical engineers do. Memory and adaptability also pose a challenge — LLMs have extraordinary memory from training data, but when a puzzle closely resembles something they saw before, they may ignore key differences and default to memorized responses, as seen in a 2024 study from Google and the University of Illinois Urbana-Champaign.

The collection illustrates specific ways machine and human cognition differ, from problems that are trivially easy for humans to ones that stump both. Readers are invited to try the puzzles and see whether they can out-perform the AI at its own game, providing a window into both the technology's strengths and its remaining weaknesses.

Read original →

← Back to home