Research arXiv cs.LG

On-Policy Attention Linearization

linear attentionknowledge distillationlong-contextQwen3-4B

Hybrid transformer architectures replace most softmax attention layers with linear attention to offer transformer-level quality at a fraction of the memory cost. Rather than pretraining such models, a growing body of work distills them from already trained full-attention transformers. However, these distilled models often collapse on long-context retrieval and reasoning tasks, particularly in thinking mode, where the efficiency gains of hybrid architectures matter most.

The problem is structural: linear attention layers must compress context into a fixed-size state, so their errors compound over long sequences. Off-policy distillation never teaches the student model to recover from this drift, which makes tasks that necessitate longer sequence lengths especially challenging.

OPAL, or On-Policy Attention Linearization, addresses this by having the hybrid attention student sample its own long-context trajectories and receive dense supervision from the frozen full-attention teacher.

Applying OPAL to Qwen3-4B and MiMo-7B-RL-0530 recovers 87–94% of full-attention performance on commonsense reasoning, 100% on needle-in-a-haystack (NIAH) retrieval, and 83–93% on mathematical reasoning with only 3B training tokens. These results are achieved without supervised fine-tuning (SFT) or reinforcement learning with verifiable rewards (RLVR).

Compared with the strongest prior linearization method, which recovers 68% of its teacher's retrieval performance and 21.6% absolute average mathematical reasoning accuracy, OPAL fully recovers retrieval and achieves 67.6–72.2% on math reasoning. The results suggest on-policy distillation can substantially close the long-context gap in efficient hybrid attention models.

Read original →

← Back to home