IBM releases Granite 4.2 reasoning LLMs (3B/8B/30B, Apache 2.0, 512K context) trained on ~15T tokens with GB200 NVL72 GRPO RL

IBM

Models / LLM official 3 src. ~1 min

On Aug 25, 2026, IBM released Granite 4.2 — its first dense decoder-only reasoning LLM family in 3B/8B/30B sizes under Apache 2.0. Pre-trained from scratch on ~15T tokens via a five-phase strategy that extends context to 512K; SFT on ~7.2M samples (31.6% agentic, 68.4% non-agentic); multi-stage asynchronous GRPO RL — foundational RLVR for all sizes, agentic RL (SWE → Terminal → Search) for 8B/30B, RLHF for safety and reasoning-length. Thinking / non-thinking / low-effort modes; native OpenAI-compatible tool-calling API; 12 languages. Trained on an NVIDIA GB200 NVL72 cluster on CoreWeave using NeMo-RL and NeMo-Gym. Quantized FP8/NVFP4/MXFP4 and 14 GGUF variants. 30B reaches 89.17 on AIME25, 57.00 on SWE-Bench Verified, 62.00 on τ³-bench.

Why it matters

First Apache-2.0 dense reasoning LLM family from a Western frontier lab at three sizes, all with native tool-calling and 512K context — competitive with closed-source agentic APIs on agent benchmarks and ships with the full RL stack.

Importance: 4/5

first Apache-2.0 dense reasoning LLM family from a Western lab at 3 sizes, 512K context, agentic RL

Sources