Research arXiv cs.CL

Solving Every Step Is Not Enough: Milestone Oracles Reveal a Composition Gap in LLM Math Reasoning

LLM math reasoningbenchmarkcomposition gapOracleLadder

Large language models can solve every intermediate step of a multi-step math problem on its own and still fail the full problem — even when handed a roadmap of the steps and all of their answers. To pin down where that breakdown occurs, the authors introduce OracleLadder, a diagnostic evaluation that supplies increasing levels of oracle help.

For each problem, a teacher model writes a fixed roadmap of intermediate sub-goals (milestones), and a deterministic symbolic verifier grades every answer. Models are then tested with no help, with the roadmap, with the roadmap plus the milestone answers, and on each milestone alone, which sorts each failure into one of five reasoning gaps.

On 354 NuminaMath problems and six models ranging from 8B to 671B parameters (Qwen3, gpt-oss, Llama 3.3, DeepSeek-V3.1), the largest gap for every model is the composition gap — a stricter form of the compositionality gap. It accounts for 33-48% of problems, or 24-37% after excluding problems that an LLM review flags as grading errors. Accuracy and milestone-help recovery rank the two strongest models differently, and two RLVR runs with similar accuracy gains shift problems in different ways.

The roadmap effect replicates on MATH500 and AIME 2024/25, per-problem recovery agrees for 83-87% of problems under an independent second teacher, and the help ladder carries over to code generation. The data, roadmaps, prompts, and code are released at https://github.com/slark-prime/OracleLadder.

Read original →

← Back to home