SafeTune: A Unified Faithful Library for Auditing and Repairing Safety Drift in Fine-Tuned LLMs
Methods for addressing safety drift in fine-tuned large language models are scattered across incompatible implementations, lifecycle stages, and evaluation protocols, which makes them difficult to adopt and compare.
SafeTune is introduced as a source-available library that unifies four intervention paradigms: post-hoc weight recovery, safety-constrained fine-tuning, gradient-based unlearning, and inference-time steering. It also provides shared interpretability, evaluation, and deployment utilities, and offers a consistent configuration-driven workflow while preserving the distinct inputs and intervention points each paradigm requires.
The library's modular registry is designed to support new methods, benchmarks, judges, models, and fine-tuning domains without requiring the surrounding pipeline to be redesigned.
SafeTune is demonstrated through controlled comparisons and finance and medical deployment case studies, showing how it characterizes safety drift, evaluates feasible interventions on common refusal-behavior and capability evaluations, and supports calibrated or layered mitigation.