CleanScore: Black-Box Benchmark Audits with Negative Controls and Sensitivity Bounds
Public benchmark scores may reflect a model's skill, prior exposure to the questions, or both, and for most models the training data are unknown. CleanScore addresses this by performing a black-box audit using only scored outputs.
Each benchmark question becomes a parent item with one public form and two independently written fresh forms that preserve its numbers, facts, and answer. The audit reports an interval for the public-form advantage rather than a verdict, and uses a private negative-control bank with an explicit transport radius to separate exposure from ordinary form mismatch. A registered controlled-exposure experiment detects planted exposure and stays quiet under fresh-form exposure.
A registered audit of five open models on 200 GSM8K and 200 ARC-Challenge items finds no exposure-consistent advantage, bounding surface-form inflation below five points. Registered positive controls then bound what such a null can mean. Leaking an item raises accuracy on paraphrases the model never saw almost as much as on the leaked wording, leaving 52% to 110% of the effect invisible to a paraphrase audit. On ARC, a planted 49-point advantage shows an observable gap of -0.020, and about 20 points survive rewriting stem and options across four training seeds.
A surface-form null therefore bounds far less than the phrase contamination audit implies.