anagnorisis.cloudSign in

← Hourlies

Hourly ·

Moonshot AI Releases Kimi K3: First Open 2.8 Trillion-Parameter Model

Moonshot AI drops the world's first open-weight 3T-class model — 2.8T parameters, 896 experts, native multimodal, NoPE architecture, competitive with GPT-5.5 and Claude Fable 5 on major benchmarks.

Moonshot AI Releases Kimi K3: First Open 2.8 Trillion-Parameter Model

Moonshot AI has released Kimi K3, the world's first open-weight model at the 3-trillion-parameter class — 2.8 trillion total parameters with 104 billion activated per forward pass. The model uses a novel Mixture-of-Experts architecture built on Kimi Delta Attention (KDA) and Attention Residuals, along with a Stable LatentMoE framework that activates 16 of 896 experts per token, yielding roughly 2.5× better scaling efficiency than Kimi K2.

As Sebastian Raschka notes in his architecture breakdown, K3 is the first frontier-level model to drop positional embeddings entirely. Where most large models use RoPE in local attention layers, K3 uses NoPE across all 93 layers — a design inherited from Kimi Linear, scaled from 48B to 2.8T parameters. The model also features native multimodal support, handling text, images, and video within a single architecture, with a 1-million-token context window.

Benchmark results place K3 in elite territory. On GPQA Diamond it scores 93.5%, matching GPT-5.5. On BrowseComp it reaches 91.2%, ahead of Claude Fable 5 at 88.0%. On coding benchmarks, it posts 77.8% on ProgramBench and 67.5% on DeepSWE — competitive with the best closed models. Full weights are available on GitHub under the Kimi K3 License.

Sources: Sebastian Raschka, MoonshotAI GitHub

More Hourlies Stories

Content on Anagnorisis is summarized, paraphrased, and editorialized from publicly available sources for length and clarity. Original sources are linked where available. All trademarks belong to their respective owners.

More from Anagnorisis