Evaluating Coding Agents on Kernel Exploit Generation
Coding agents can already find real vulnerabilities in production software, but bug-discovery results do not measure whether agents can construct exploit primitives—the building blocks needed to turn a vulnerability or crash into a usable exploit. KEX-bench is introduced to evaluate coding agents specifically on exploit primitive generation against real operating-system kernels.
The benchmark contains 45 task instances across 40 real Linux and Windows CVEs, covering kernel address leak, instruction-pointer control, heap read, heap write, and arbitrary address write. Each task runs in an isolated virtual machine, exposes controlled tools, and uses a deterministic verifier to check primitive-specific success. The authors evaluate state-of-the-art coding agents paired with frontier and open-weight models under fixed tool-call budgets.
Without a reference proof of concept (PoC), the strongest configuration solves 1 of 20 Windows tasks (5.0%) and 14 of 25 Linux tasks (56.0%). With a reference PoC, the strongest configuration solves 31 of 45 tasks (68.9%). This highlights the gap where agents reach kernel crashes but fail to shape kernel state into exploit primitives. KEX-bench is released for reproducible research on AI-assisted exploitation at https://kex-bench.github.io.