Research arXiv cs.AI

Math Reasoning in LLMs is Organized by Approach, Not Topic

interpretabilitymath reasoningLLM internalsbenchmarking

Mathematical reasoning benchmarks are conventionally organized by topic, but language models may instead organize their internal computation around reusable reasoning approaches. This paper investigates whether open math-capable LLMs organize internally by topical sub-skill or by reasoning approach, and presents evidence that approach is the key.

The authors introduce a generation-replay protocol: a model first generates a solution, then the exact prompt-plus-generation trajectory is replayed and activation-importance signatures are extracted over the reasoning tokens. These signatures are clustered without supervision across eight models and five mathematical reasoning sources, and the recovered structure is evaluated with structural, semantic, and intervention tests.

Across all 40 model-source cells, the recovered clusters outperform matched-size random baselines. Two independent frontier-LLM judges find approach-level coherence in 77–82% of real clusters versus only 6–11% in within-source controls, and topic-pure clusters usually receive labels finer than the topic itself. In approach-controlled prompting, changing the requested reasoning approach shifts cluster assignment in seven of eight model conditions, whereas paraphrases largely preserve it.

These results indicate that math-capable LLMs organize internal mathematical computation by reasoning approach rather than benchmark topic. The implication is that topic-stratified benchmarks and topic-balanced training corpora can still miss the axis that matters — even deliberately topic-balanced corpora may remain imbalanced over reasoning approaches.

Read original →

← Back to home