Research arXiv cs.CL

Beyond Overlap: Estimating the Causal Effect of Benchmark Exposure

benchmark contaminationevaluationcausal inferenceLLM benchmarks

The paper argues that evidence benchmark material entered a model's training data does not by itself reveal how much that exposure affected the resulting evaluation score. Provenance can establish contact between training data and a benchmark, but only a counterfactual comparison can quantify the performance attributable to that contact — leaving contaminated scores hard to interpret.

To supply that missing quantity, the authors present LeakScale, an interventional framework. It creates fresh executable tasks that require private, family-specific information absent from the public task and non-derivable from it, then controls access to that information and estimates the resulting control-adjusted change in executable accuracy.

The framework was tested across 2,048 unique families, two model families, two executable domains, and 262,144 generations. Exposure improved accuracy in every model-by-domain combination, with gains ranging from +7.17 to +27.31 percentage points.

The authors frame the contribution as separating two empirical questions that are often conflated: whether benchmark contact occurred, and how strongly a reported score depends on it. LeakScale is intended to make the second question directly measurable.

Read original →

← Back to home