Research arXiv cs.LG

A Shared Learning Rate Is Not a Neutral Control in Selective On-Policy Distillation

on-policy distillationlearning-rate sensitivityGSM8KLoRA

Selective on-policy distillation trains a student only at the token positions a selector scores highest, and the literature typically compares selectors under one shared learning rate chosen as a neutral control. The paper shows this control is not neutral: with LoRA on GSM8K using a Qwen2.5-1.5B student and 7B teacher, an 8x learning-rate grid reveals that dense supervision is statistically flat, with a swing of only 1.8 pp (p=0.26), while every selective arm moves with the rate—5.4 pp for a random 5% subset, 6.7 pp for a total-variation selector, and up to 17.7 pp for a teachability selector.

This rate sensitivity changes the conclusions drawn from selector comparisons. The dense-versus-selective verdict is 10.1 pp at lr=1e-4 but 5.1 pp at 5e-5, a 2.0x difference decided by a parameter the protocol treats as scenery, and two of six pairwise significance calls between selectors flip between adjacent rates without any rank inversion. The authors call this selector-rate entanglement and trace it to selection itself rather than step size: AdamW update magnitudes track the rate to within 2.2% despite 15.5x gradient-norm differences across arms.

A preregistered frozen-scoring ablation—where selection is scored by the initial student while criterion, budget, and on-policy rollouts are unchanged, with 12 seeds per cell—shows that live scoring adds 3.79+/-1.69 pp of rate sensitivity (p=0.035), while the frozen arm remains significantly entangled (p=0.015). The feedback loop therefore aggravates the phenomenon rather than causing it. Under full fine-tuning at the rates this literature actually uses (1e-6 to 1e-5), the pattern grows: dense itself swings 19.8 pp, the selective arm swings 49.5 pp, and the verdict ranges from a non-significant +3.6 pp at the published operating point to +34 pp (p=0.005) one notch hotter.

The rate dependence does not reproduce on MATH-500 under LoRA, which scopes the result, though the roughly 10 pp cost of selective training does persist there. The paper prescribes reporting the arm x rate matrix rather than a shared-rate column as a precondition for selector comparisons.

Read original →

← Back to home