Research arXiv cs.LG

Measuring the Checker: Mutation Analysis for GPU-Kernel Benchmark Oracles

GPU kernelsmutation analysisKernelBenchbenchmark oracles

Benchmarks for LLM-generated GPU kernels decide correctness using only a few random inputs and a loose floating-point tolerance, even though their verdicts feed leaderboards and reinforcement-learning rewards. Prior work agrees these checkers are weak and patches them by hand—adding input distributions, fuzzing recipes, or tighter tolerances—but there has been no way to measure whether a patch actually suffices.

The paper introduces mutation analysis as an adequacy metric for kernel-benchmark oracles. Deterministic rules inject 10,303 compilable faults into verified CUDA implementations of 188 KernelBench problems, 7,384 of which have an independent kill witness. Any test protocol is then scored by the fraction of those faults it detects.

The official check misses one in six witnessed faults (16.9%) deterministically, and the misses are family-skewed: 8.7% of arithmetic faults escape, but 78.6% of precision faults do. The metric explains why: there is a tolerance blind band that grows with reduction size, and a measured ceiling on input aggressiveness set by legitimate floating-point variance. It also audits the strongest existing patch: KernelBench-Verified's gain splits into +4.0 points from hidden inputs and +4.5 from tighter tolerance, a split its authors could not compute. A published fuzzing recipe is exposed as rejecting correct kernels 107 times. Optimizing suites over the kill matrix reaches 98.0% detection with two inputs per problem (94.8% held-out), and the fault taxonomy teaches a test generator more than the raw faults themselves. Across 48 whole architectures, blindness grows with scale and concentrates in deep homogeneous pipelines, and two problems prove unrefereeable because their official references violate the benchmark's own tolerance against fp64.

The authors release everything as KernelBench-M at https://huggingface.co/datasets/Elfsong/KernelBench-M, providing both a measurement standard for checker adequacy and a practical basis for stronger kernel-benchmark test protocols.

Read original →

← Back to home