Research arXiv cs.CL

LOCKR: A Hidden-State Trajectory-Guided Planner for Detecting and Repairing Stable-but-Wrong Lock-In in Diffusion Language Models

diffusion language modelstest-time planningreasoning repairhidden-state trajectories

Diffusion language models generate text through iterative denoising, exposing intermediate trajectories before final answers are produced. The paper identifies a recurring reasoning failure called stable-but-wrong lock-in, in which an answer stabilizes early around an incorrect value while substantial denoising remains.

The authors show that surface-level decoding signals—confidence, entropy, margin, and answer stability—are not sufficient to reliably distinguish correct lock-in from erroneous lock-in. To address this, they formulate selective reasoning repair as a lightweight test-time planning problem and propose LOCKR, a hidden-state trajectory-guided planner. LOCKR decides when to allocate additional computation, expands a structured set of targeted repair branches, and selects the most promising continuation using trajectory-aware verification.

Across two diffusion language models and three mathematical reasoning benchmarks, hidden-state trajectories consistently outperform surface signals and single hidden snapshots for both wrong-lock-in detection and repair selection. On natural evaluation distributions, LOCKR yields absolute accuracy gains of 2.21–5.37 percentage points across all five evaluated settings, with repair rates ranging from 22% to 41%.

These results establish hidden diffusion trajectories as actionable signals for selective test-time reasoning repair, suggesting a way to improve reliability without uniformly spending extra computation.

Read original →

← Back to home