Research arXiv cs.CL

Vox-Infinity: Benchmarking the Limits of Long-Context Spoken Language Models

spoken language modelslong-contextbenchmarkrecency effect

Long-context understanding remains a fundamental challenge for large language models, because excessively long inputs often cause models to forget salient information. The problem is even more pronounced in the speech domain, where audio as a low-compression modality requires substantially more embeddings than text to preserve both semantic content and acoustic cues.

To address this, the authors introduce Vox-Infinity, described as the first benchmark specifically designed to evaluate long-context understanding in spoken language models. Vox-Infinity systematically extends audio history along two dimensions, turn count and turn duration. It covers a diverse range of representative scenarios with varying interaction structures and semantic complexity. Crucially, it provides explicit answer-provenance annotations and organizes samples according to the amount of historical context required to resolve each query, enabling precise and length-aware evaluation.

Extensive evaluations of seven representative spoken language models reveal a clear overall recency effect: models generally achieve higher accuracy when answer-supporting evidence is closer to the query, but struggle to retrieve and use evidence located farther back in the dialogue history. Cases and datasets are available at https://vox-infinity.github.io.

Read original →

← Back to home