Kimi K3 Technical Report: Kimi Delta Attention and Stable LatentMoE Architecture Detailed
Moonshot AI
Moonshot AI published the Kimi K3 technical report on arXiv, detailing the architecture behind the 2.8T-parameter (104B activated) MoE model released as open weights the day before: Kimi Delta Attention with Attention Residuals, a Stable LatentMoE routing 16-of-896 experts, and post-training RL across general, agentic, and coding domains at multiple reasoning-effort levels. The paper reports roughly 2.5x scaling efficiency over Kimi K2.
Why it matters
Ranked the top-voted paper of the day on Hugging Face Daily Papers with 182 upvotes, and is the technical report behind yesterday's Kimi K3 open-weight release, explaining the architectural choices that let Moonshot claim the largest open-weight model to date.
Importance: 4/5
Technical report for a major open-weight release; top-voted HF Daily paper (182 upvotes).