Research arXiv cs.AI

PASTABench: Proactive Assessment of Sequential Trajectories for Agent Safety

agent safetybenchmarkLLM agentsproactive monitoring

As LLMs evolve into autonomous agents that alter real-world states, ensuring safety across multi-step workflows has become a critical challenge. Existing approaches fall short in two ways: step-level methods evaluate actions in isolation and miss how risk accumulates across a trajectory, while trajectory-level evaluations are post-hoc and leave no room for timely intervention.

To address this, the authors formalize Decoupled Proactive Safety Monitoring along three dimensions: whether to intervene, when to intervene, and what the risk is. They introduce PASTABench, a benchmark of 1,139 multi-turn trajectories spanning 5 risk categories and 13 subcategories, along with the Optimal Intervention Window (OIW), anchored by annotated Earliest-Signal and Trigger turns, to quantify how timely an intervention is.

Evaluating 16 LLMs shows proactive intervention remains largely unsolved: the best model achieves only 40.74% optimal-timing interventions. Fine-grained diagnosis uncovers pervasive lexical overfitting — competitive safety scores from smaller models mask keyword hypersensitivity rather than genuine risk comprehension, with their proactive capability largely collapsing once hazard vocabulary is neutralized.

Read original →

← Back to home