Research arXiv cs.LG

Scaling of Capability and Efficiency at Inference Time in Large Reasoning Models

large reasoning modelsinference scalingDeepSeek-R1scaling laws

Capability and efficiency are two key dimensions of reasoning in large language models (LLMs): capability is the ability to solve a problem correctly, while efficiency is the ability to do so with limited resources. When LLMs use Chain-of-Thought (CoT) reasoning on problems of controlled hardness, both the number of problems solved correctly and the number of tokens needed to reach a correct answer depend on problem hardness and model size, but how these factors jointly shape capability and efficiency has remained poorly understood.

To investigate this, the authors use hierarchical Bayesian models to evaluate LLMs from the DeepSeek-R1-Distill model family across four classes of arithmetic and algorithmic reasoning problems.

The results show that at a fixed model size, the probability of correctly solving an instance decays approximately exponentially with instance size, the study's proxy for problem hardness. The decay scale grows sublinearly with model size, indicating that larger models are more capable but that capability gains diminish as scale increases. Output length grows as a power law with instance size, but the parameters of this power law do not vary systematically with model size, suggesting that larger models do not become more efficient.

Together, these findings reveal potential limitations of naive scaling as a strategy for developing more capable AI systems: capability improves with diminishing returns, while efficiency shows little to no improvement.

Read original →

← Back to home