Agents Can Use Base Models to Evade AI Detection
Prior research has shown that base language models can evade commercial AI-text detectors, but conventional "humanization" techniques rely on using those models to paraphrase AI outputs over several iterations, which invariably causes semantic drift. This paper instead explores having coding agents directly orchestrate the writing process.
The method equips a Claude Opus 5 agent operating in a Claude Code harness to orchestrate a local 32B-parameter OLMo-2 base language model, stitching together text samples from that base model. Across benchmarks spanning creative writing, factual grounding, health QA, and instruction following, the agent sacrifices little task accuracy while using up to 90% base LM tokens.
Responses constructed this way reduce the effectiveness of post-hoc detectors and watermarking. Pangram v4 detection rate drops from 77% to 24%, and soft watermarking applied a priori to the agent's generations falls to a simulated 10% detection at low false-positive rates. However, the evasion requires significantly more input and output tokens from the agent, increasing the dollar cost per query up to 30x at API pricing.
Overall, the work demonstrates a new class of adversarial attacks against AI text detection and urges post-hoc detection providers to include outputs of base models in their training data.