StreamDecisionBench: Evaluating Decisions in Force on Evolving Language Streams
Language models increasingly make real-time decisions in applications that apply the latest answer until a newer one arrives, so a late answer can prolong an outdated decision. The paper gives the example of a call recorder still running while a customer reads out card details, an error that offline accuracy does not capture.
The authors release StreamDecisionBench (SDB), a dataset of eight streaming scenarios spanning four application families, with executable reference decisions derived from public rules. They also propose an evaluation protocol and metric, in-force accuracy: the share of time the applied decision is correct across update intervals of 0.5-8 s. This metric jointly reflects accuracy and latency and attributes each error to judgment, latency, or both.
Evaluating thirteen single-model settings, the attribution separates speed-limited from judgment-limited models: slower, more accurate models lose 42-51% of the time to outdated answers, while a fast model loses 34% to wrong ones.
The authors therefore test hybrids in which a slow model corrects a fast one; with the right pairing and configuration, a hybrid outperforms every single model. However, even the best evaluated system keeps a correct decision in force only about two-thirds of the time, leaving a substantial gap for real-time use.