ZGCM-1: A Fully Open and Extremely Efficient Foundation Model for Math and Agentic Search
ZGCM-1 is a fully open 7B dense foundation model trained from scratch with emphasis on extreme data, system, and algorithmic efficiency. Its core premise is that compact models cannot passively memorize the open web, but can overcome parametric capacity limits by coupling deliberate internal thinking with active external tool use — a paradigm the team supports across a 256K context window.
To enable this, the authors develop an end-to-end, high-efficiency open training recipe. It includes architecture and system co-design: interleaved gated sliding-window and full attention, plus a stable FP8 Muon optimizer. Training also uses a progressive curriculum and MDP mid-training: context scaling across 16K, 64K, and 256K, and reformulating interaction traces into Markov Decision Processes.
The team further establishes an AI-native R&D workflow in which agent swarms autonomously manage cluster operations, data curation, and rapid diagnostic evaluation.
Extensive evaluations show ZGCM-1-7B is competitive across the 7B model family on general benchmarks. On several challenging mathematical reasoning and agentic search suites, it remains competitive with frontier models orders of magnitude larger, such as Qwen3-235B-A22B and GLM-5.1. The pre-training design also offers roughly a 4.2x efficiency improvement in 16K pre-training time-to-loss. Across the full development lifecycle, the authors distill eight actionable empirical findings spanning architectural scaling, SFT quality pruning, long-context generalization, and agentic co-training dynamics.
To support community research, the project open-sources model weights from the pre-training, mid-training, and post-training stages, intermediate checkpoints, training code, per-stage data and data recipes, and W&B logs.