Research arXiv cs.AI

OrchestraBench: Evaluating Multi-Agent Orchestration Failure Modes, Recovery, and Decomposition Quality

multi-agent orchestrationbenchmarkfailure injectionarXiv

OrchestraBench addresses a gap in multi-agent orchestration evaluation by focusing on failure diagnostics rather than just end-task accuracy. The benchmark employs a controlled, seed-reproducible failure-injection harness integrated with templated enterprise workflows, enabling systematic testing of pipeline resilience. It introduces cascade radius, a metric that measures how far a failure propagates through the system, as well as per-failure-mode recovery analysis. This allows developers to identify the exact routing decision or decomposition step that causes a cascade. The work targets the reliability gap between research demos and production deployments, offering a practical tool for comparing orchestration frameworks and improving their fault tolerance. Future work may extend the benchmark to more diverse workflow types and real-world failure distributions.

Read original →

← Back to home