Beyond Symmetric Agents: Cognitive Diversity and Multi-Agent Debate in Small Language Models
Multi-agent debate (MAD) has been reported to improve reasoning and factuality over single-model inference, but prior work generally treats agents as symmetric peers, leaving unclear what actually drives the gains. This study tests the hypothesis that cognitive diversity among agents is the driver, choosing a measurable setting: small open-weight models that still have benchmark headroom.
Across 23 models from eleven vendor families, five tasks, and more than 5,500 debate and control runs, the authors varied diversity along three axes—personas, sampling temperature, and model identity—and paired every debate configuration with a generation-budget-matched majority-vote control.
The diversity hypothesis is rejected on every axis. Debate does beat single-agent inference by 3–7 points where tasks have headroom, but under matched-budget conditions it ties or even loses to self-consistency sampling while using about 1.6× the wall-clock time and 3.4× the token cost. Persona prompting reduces accuracy, and a dose-response experiment across each model's full combinatorial persona space shows this is a persona tax rather than a diversity tax: redundant personas hurt most, while maximally diverse teams recover only part of the loss. Mixed-model teams also lose to majority votes over their own rosters, with accuracy tracking member capability rather than heterogeneity, and nearly all of debate's benefit comes from the first exchange of answers.
The paper additionally identifies a pervasive measurement hazard: debate transcripts silently overflow serving context windows. Correcting for this alone moves the debate-versus-sampling comparison from −1.8 points to parity. The authors conclude that reported MAD gains should be recast as an ensemble-sampling effect, and they provide a budget-matched, contamination-checked baseline that future debate mechanisms should be required to clear.