Model Releases Hugging Face Blog

Granite 4.2 LLMs: How They're Built

Granite 4.2IBMreasoning LLMagentic RL

Granite 4.2 is IBM's reasoning-focused release of the Granite language-model family, adding explicit reasoning to earlier instruction-following assistants. Each model can produce a chain of thought before answering and supports thinking or non-thinking modes, plus a low-effort mode that spends a short reasoning budget on easy questions.

All three dense decoder-only models are pretrained from scratch on roughly 15T tokens using a five-phase strategy that extends the context window to 512K tokens, then supervised fine-tuned on chain-of-thought, reasoning, and agentic-trajectory data. Post-training uses a multi-stage reinforcement learning pipeline; the 8B and 30B models additionally undergo agentic RL, learning to call tools, edit and run code, drive a terminal, and search the web inside real sandboxed environments. Every model supports native tool calling, served through OpenAI-compatible endpoints (e.g., vLLM) or SGLang, emitting tool calls in OpenAI function-calling format without extra glue.

Architecturally, the models use Grouped Query Attention with 40 attention heads and 8 KV heads, Rotary Position Embedding with θ=10,000,000, SwiGLU-activated MLPs, RMSNorm, separate input/output embeddings, and bfloat16 precision. The 3B has embedding size 2560, 40 layers, and attention head size 64; the 8B has embedding size 4096, 40 layers, and head size 128; the 30B has embedding size 4096, 64 layers, and head size 128. The 3B and 8B use MLP hidden sizes of 8192 and 12800, respectively.

All Granite 4.2 models are released under the Apache 2.0 license, with the Hugging Face collection, GitHub repository, and docs linked in the post.

Read original →

← Back to home