Research arXiv cs.AI

CodeMidas: Scaling Agentic Coding RL Environments from Code Itself

agentic codingreinforcement learningRL environmentsGRPO

Training capable coding agents with reinforcement learning requires diverse tasks and reliable verifiers. Open-source codebases are a rich source for such tasks, but existing methods typically rely on development artifacts such as issues and commits, which limits the range of tasks that can be extracted. CodeMidas is designed to better scale RL environments by treating source code as the only task-specific input and turning already-implemented functionality into executable environments.

CodeMidas is an agentic pipeline that allocates agentic compute to every stage of environment construction. Agents explore implemented functionality to formulate behavioral specifications, construct tests grounded in execution of the original code, and validate and filter candidate tasks through execution checks and repeated solution rollouts.

The resulting dataset contains 5,545 training tasks from 3,185 open-source codebases, spanning 23 programming languages and 15 technical domains. Training MiMo-V2.5 on these tasks with GRPO improves performance on all five diverse benchmarks, covering issue repair (DeepSWE +11.7%), whole-program construction (ProgramBench +17%), and terminal work (Terminal-Bench v2.1 +8.5%). Ablations show that increasing the number of high-quality training tasks improves performance, and trajectory analysis shows the RL-trained agent exhibits better behaviors such as increased codebase exploration and more diverse self-verification.

These results establish source code as a scalable foundation for constructing RL environments that improve coding agents across diverse software tasks.

Read original →

← Back to home