Control the Harness, Control the Cost: Routing and Governing AI Coding Agents in the Enterprise
Enterprises are scaling AI coding-agent harnesses from pilots with a few hundred seats to tens of thousands, but most buy rather than build them, relying on vendor products such as Anthropic's Claude Code or OpenAI's Codex. A harness controls which model answers, what the model reads, how the prompt cache is used, and which subagents run, so it effectively selects the rate and volume on the price sheet. Enterprises that leave a proprietary or untuned harness at its defaults inherit those choices and the resulting bill.
To address this, the authors build a fast, customisable router in which Jev, a classifier with calibrated probabilities, labels every prompt against a bring-your-own taxonomy of agentic requests. Because a single user turn comprises many requests over a prompt cache that belongs to one model, the router only moves work where no running conversation would have to rebuild its cache—at session start, in side lanes, and at subagent launch. From the price sheet, the paper derives when a mid-task switch pays back and identifies a crossover.
Repricing about 10,000 real sessions from public datasets confirms the crossover: on long, tool-heavy sessions, the highest-priced model costs less than the next tier. In an emulated enterprise of 10,000 seats with user behaviour taken from those datasets, the router recovers 14–21% of model spend at Anthropic's list prices of 21 September 2026, amounting to $3.3M to $5.0M a year.
The paper also maps risks across twenty harnesses, prices the dependence on one vendor's models, and proposes a control plane that enterprises can run from within starting now. It further offers a ladder for deciding later whether to own the harness.