Science or Slop?: Benchmarking and Mitigating Scientific Slop in AI-Generated Papers
AI-generated content, or AI slop, is increasingly common, especially in academia. In scientific papers, slop has more complex patterns that token-based AI detectors struggle to catch: individual parts can look plausible while the scientific reasoning connecting them breaks down, misleading readers about the work's quality.
The authors benchmark these failures as scientific slop using six measures spanning Structure, Argument, and Artifacts. They build SciSlopBench, containing 390 AI-generated papers, mostly in computer science but also spanning the life, social, and natural sciences, each paired with a human-written paper matched by research problem and contribution type.
The measures identify the AI paper in each pair with 85.9% accuracy, compared with 68.7% for Binoculars. Higher scientific slop also tracks lower ICLR ratings and distinguishes rejected from accepted papers above chance in every year from 2017 to 2025.
Simply optimizing the measures does not easily reduce slop. The authors propose SciSlopHarness, a harness-level framework that guides a fixed LLM to revise slop only where experiment records support the change. Standard revisions leave residual slop and direct slop-aware prompting triggers reward hacking, but SciSlopHarness reduces the remaining AI-human gap by 63% over the strongest revision baseline without requiring human reference targets. The paper concludes that AI-generated scientific papers leave fundamental traces in their global reasoning, so responsible mitigation requires strict evidentiary grounding rather than mere prose refinement.