Other MIT Technology Review (AI)

Kids outlearn AI—and we still don’t know why

data efficiency gapLLMscognitive sciencelanguage acquisition

For at least 100,000 years, the only thing capable of learning a human language to perfect fluency was a human child. Four years after ChatGPT's release, LLMs like Claude, DeepSeek, and OpenAI's GPT models can now converse fluently, but they need an inhuman amount of data—an LLM can easily churn through a hundred thousand times more words than a person hears while mastering their mother tongue. As Stanford cognitive scientist Michael C. Frank puts it, 'We still have to burn down a forest and scrape the entire sum of all human knowledge to re-create this milestone that happens in our living rooms over the course of a year.'

This 'data efficiency gap'—the divide between children and machines—poses a challenge for AI architects and a tantalizing question for cognitive scientists. For the past decade, language models have improved mainly by scaling up: Meta's Llama 3.1, released two years ago, used 15 trillion tokens in pretraining, and frontier models may be pretraining on 10 times that amount, says Georgetown cognitive scientist and linguist Ethan Gotlieb Wilcox. But there is only so much internet to train on, and the well of easily available data could run dry as early as the 2030s.

Kids demonstrate that it's possible to learn far more with far less. A preteen in a linguistically rich home may have heard around 100 million words; adding literacy, that could reach 300 million words by age 20. The scale difference is stark: 'Claude has seen the amount of language that an entire city will experience in one generation,' says Wilcox. Understanding how children achieve this could help build more efficient models and reveal more about how the human mind develops.

Read original →

← Back to home