How UK AISI and EvalEval Are Making Benchmark Results Reproducible
The EvalEval Coalition announced that the UK AI Security Institute (AISI) is using EvalEval's infrastructure to openly share evaluation results, supporting more reproducible and verifiable evaluation science. The collaboration began at a joint workshop alongside NeurIPS 2025, and feedback from AISI helped shape the Every Eval Ever (EEE) schema. This next phase puts that shared infrastructure into practice.
Reproducible evaluation reporting matters because, as AI deployment accelerates, evaluations are becoming increasingly important sources of evidence about model and system performance. Yet results are reported across many formats, platforms, and outlets, often without enough information to reproduce them, and running evaluations again may itself be prohibitively expensive. EvalEval's mission is to improve this ecosystem through the shared Every Eval Ever reporting schema and the open Evaluation Cards platform, which brings evaluation results and the information needed to interpret them into a common structure. This builds on AISI's work to make evaluation more efficient through OptStop, more statistically rigorous through HiBayES, and more standardised in areas including transcript analysis and capability elicitation. Together, AISI and EvalEval are working to diagnose gaps in evaluation reporting and build shared infrastructure to close them.
In the new phase, AISI is making publicly reported evaluation methods and findings available through Evaluation Cards where appropriate, with transcript-level transparency valued for reproducibility, analysis, and diagnosis. The release includes verified results, context, and configuration information for the five benchmarks in the paper's main experiment: HealthBench, FrontierMath, Humanity's Last Exam, SWE-Bench Pro, and Terminal-Bench 2.0. These results cover six frontier models: Claude Opus 4, Claude Opus 4.5, Claude Opus 4.6, GPT-5, GPT-5.2, and GPT-5.4. The release also includes results from two related cyber evaluations—Cyber CTFs and The Last Ones—which use a different, partially overlapping set of models.
The data accompany AISI's paper, How Inference Compute Shapes Frontier LLM Evaluation, which studies how benchmark performance depends on inference-time compute and evaluation protocol. Performance on Humanity's Last Exam changes with evaluation protocol and inference compute: each curve shows the cumulative share of attempted tasks solved within a given token count, using the earliest observed success per task. When models received correctness feedback from an oracle after each attempt, they continued to solve additional tasks as token use increased. When results are openly released with setup information, researchers and practitioners can examine individual studies more closely and compare findings across the wider ecosystem.