Recursive Self-Improvement via On-Policy Distillation for Reasoning
On-policy distillation (OPD) trains a student model by having it generate trajectories and then matching its next-token predictions against an external teacher's predictions, which supplies dense, token-level supervision. On-policy self-distillation (OPSD) removes the need for an external teacher: a second frozen copy of the student, given the ground truth in its context, acts as a privileged teacher, while the student sees only the problem and learns to mimic it. Prior work found that freezing this teacher helps training stability, but the authors argue it also stops the teacher from incorporating improvements the student makes during training.
The paper's main contribution is a recursive framework with two complementary components addressing that limitation. First, Dynamic Co-Evolution (DCE) lets the privileged teacher co-evolve with the student, so revision learned in one round can guide the next. Second, because stronger revision can make responses overly verbose and self-critical, they add Self-Refined Concise Learning (SRCL), which trains on shorter, verified rewrites of the model's own on-policy responses.
Evaluations across multiple model scales and four competition-level mathematics benchmarks show DCE+SRCL outperforming OPSD. On Qwen3-8B, DCE+SRCL reaches 65.97% Average@12, exceeding OPSD by 35.62 percentage points, while also reducing mean output length by 7.80% relative to DCE alone — suggesting co-evolved self-distillation can deliver large reasoning gains without the verbosity penalty.