Model Releases arXiv cs.CL

Kimi K3: Open Frontier Intelligence

Kimi K3Mixture-of-Expertsopen weightsscaling efficiency

Kimi K3 is a 2.8T-parameter Mixture-of-Experts model with 104 billion activated parameters, native vision capabilities, and a 1-million-token context window. It is built on Kimi Delta Attention and Attention Residuals, architectural advances that improve information flow across sequence length and model depth, plus Stable LatentMoE, which effectively activates 16 of 896 routed experts per token. Together with refined training and data recipes, these yield approximately a 2.5x improvement in overall scaling efficiency over Kimi K2.

Post-training applies reinforcement learning across general, agentic, and coding domains and multiple reasoning-effort levels, enabling compositional generalization and robust long-horizon execution. Infrastructure advances include algorithm-system co-design for Kimi Delta Attention, perfectly balanced expert-parallel training with efficient memory management, million-token agentic RL with persistent rollout and sandbox states, and deployment innovations.

Extensive evaluations show frontier-level performance on long-horizon coding, agentic, knowledge, reasoning, and vision tasks. While overall performance still trails Claude Fable 5 and GPT-5.6 Sol, Kimi K3 consistently outperforms other open and proprietary models in the evaluation suite. The full model weights are released to facilitate future research and broader deployment of frontier intelligence.

Read original →

← Back to home