Research arXiv cs.AI

When Scientific Contradictions Are Lost in Translation

scientific verificationLLM reasoningbenchmarkcontradiction detection

Two scientific findings can disagree without actually contradicting each other, and deciding whether they conflict requires knowing whether they describe comparable measurements. The paper studies how language models behave at exactly this decision point.

The authors use a controlled task: they generate an unsatisfiable XOR constraint system and translate its constraints into scientific reports attributed to different laboratories. One assignment satisfies more of the constraints, while another satisfies fewer but better matches expected biology — creating a dilemma over whether the model picks the assignment that best fits the constraints or the one that better matches biological expectations.

When constraints are stated directly, GPT-5.6 Sol and Claude Opus 5 recover the best-supported assignment in 90% and 96% of cases respectively. In scientific prose the models diverge: Claude Opus 5 often prefers the biologically expected assignment, and removing that biological preference raises recovery of the better-supported assignment from 27% to 79% (p<.001). Recovery reaches 92% when the same Biology-favored record is accompanied by a formalization request and an explicit paired-design cue (p<.001). GPT-5.6 Sol is less sensitive to these manipulations, with neither corresponding change reaching statistical significance.

The results suggest that reliable scientific verification depends not only on formal reasoning capability, but also on how models decide which findings should be compared and what relations they imply.

Read original →

← Back to home