Research arXiv cs.AI

Reasoning Concentrates Errors, and Self-Consistency Never Notices

self-consistencyreasoningconfidence estimationLLM evaluation

Self-consistency assumes that when a model is unsure, independent samples will disagree, making agreement evidence of correctness. Holding model weights fixed and toggling only a reasoning mode, the authors test this over five benchmarks and 74,944 samples. They find that reasoning concentrates errors: the probability that two independently drawn wrong answers coincide rises in all ten dataset-scale comparisons (p = 0.00098), and in nine of nine comparisons after restricting both arms to the problems each gets wrong.

The mechanism depends on answer-space structure. Where the answer space is unbounded, reasoning reduces the number of distinct answers produced to 0.43-0.65 of the non-reasoning count. Where it is bounded, both arms have an identical option set and reasoning instead concentrates probability mass on that set; the authors say no positional prior can explain this at fixed weights.

The aggregate cost is smaller than the mechanism alone predicts, because reasoning also shrinks the set of problems where answer diversity can decide anything—in ten of ten cells, by 2.7x. Normalized for available headroom, both arms convert about a quarter of it in domain. Confidence weighting does not recover what is left.

Across 280 method-dataset-model combinations on eight models and five benchmarks, not one beats plain majority voting after correction. Weighted voting agrees with majority voting on 98.5% of problem-method pairs and is right on only 56.3% of the rest. A signal’s direction can invert within fixed weights: answer log-probability predicts correctness when reasoning is off but error when it is on, and a learned six-signal combination gains nothing out of domain. The authors conclude confidence signals should be evaluated on decisions, not on discrimination.

Read original →

← Back to home