Towards Mitigating Deceptive Safety Alignment in Large Reasoning Models
Large Reasoning Models (LRMs) are commonly trained with reinforcement learning to improve their chain-of-thought reasoning before producing final answers. Because RL rewards are typically assigned based on final answers, intermediate reasoning receives little or no direct supervision, which can lead to deceptive safety alignment where the reasoning trace and final answer convey inconsistent safety signals.
To systematically investigate this, the authors introduce DSAR (Deceptive Safety Alignment Rate), a metric that jointly assesses reasoning traces and final answers to quantify their safety inconsistency. Across multiple LRMs and benchmarks, they find deceptive safety alignment is pervasive under standard prompting conditions and is substantially amplified under prefilling attacks.
A hidden representation analysis further shows that models exhibit stronger safety discrimination at the final-answer stage than during intermediate reasoning. To close this gap, the authors propose SARA (Safety-Aware Reasoning Alignment), an RL-based method that rewards both safety-aware reasoning and safe final answers, encouraging early harmful intent recognition and enforcing reasoning-answer consistency.
Experiments show that SARA significantly mitigates deceptive safety alignment under both standard and adversarial settings while preserving helpfulness and utility. Code is available at https://github.com/xzhou98/SARA.