On-Policy Delta Distillation for Multilingual Math Reasoning
The authors explore OPD, a post-training technique that uses teacher-generated on-policy samples, and OPD^2, which improves upon it by using the probability gap between a post-trained teacher and its base model as the learning signal. Their experiments cover mathematical reasoning tasks in English, Korean, and Japanese, an area previously underexplored for OPD. The results indicate that OPD and OPD^2 can effectively improve multilingual math reasoning, potentially offering a more efficient and stable alternative to reinforcement learning. This work highlights how distillation-based methods can be adapted for multilingual and reasoning-focused scenarios, with implications for cost-effective LLM alignment.