anagnorisis.cloudSign in

← Hourlies

Hourly ·

Kimi-K3 Opens Today: World's First Open 3T-Class Model Lands Amid Distillation Firestorm

Moonshot AI releases Kimi-K3 as open weights — the largest openly available model at 2.8 trillion parameters — as Washington claims it was distilled from Anthropic's Fable. AI researchers say the timeline doesn't add up.

Kimi-K3 Opens Today: World's First Open 3T-Class Model Lands Amid Distillation Firestorm

Moonshot AI is releasing Kimi-K3 on HuggingFace today — the world's first open 3-trillion-parameter-class model. Over 1,800 people are watching the countdown. But the release lands in the middle of a geopolitical storm: the White House has accused the Beijing-based lab of distilling Anthropic's Fable model, and Treasury is reportedly weighing sanctions.

K3 is a 2.8-trillion-parameter model, natively multimodal with a 1-million-token context window. Moonshot built it on a new architecture called Kimi Delta Attention with Attention Residuals, and it ships with native agentic capabilities — tool calling, web browsing, and multi-step planning. It is designed for long-horizon coding, knowledge work, and deep reasoning. It is, by any measure, a frontier model — and it's being given away.

The political drama started last week, when White House science advisor Michael Kratsios claimed Moonshot built K3 by "large-scale, covert industrial distillation" of Anthropic's Fable — which only became publicly available on July 1. Kratsios also alleged Moonshot obtained banned NVIDIA GB300 chips. The Treasury Department is reportedly discussing sanctions against Chinese AI models over IP theft.

But several prominent AI researchers are pushing back. Braden Hancock of the Laude Institute and Snorkel AI told TechCrunch: "I don't think you get a model this strong and this quickly on the heels of Fable doing strictly distillation. There's just not even frankly time — Fable's only been publicly available since July 1st. You can't distill that much data, train a model, and release it in two weeks."

Nathan Lambert of the Allen Institute for AI added that distillation is becoming less impactful as Chinese models approach the frontier and training regimes shift toward reinforcement learning. "If it were the case, everyone would be easily able to catch up to a GLM or to a K3 by using its data for distillation. But we have not, or we won't see this, from supervised fine-tuning alone."

On Hacker News, the discussion has been pragmatic. Commenters noted that hosting K3 will require roughly 1.5 TB of VRAM — "juuust at the limit of 8×B200s" — and that pricing competition among inference providers could drive costs down significantly, following the pattern seen with GLM 5.2, whose token prices dropped roughly 45% in six weeks after release. One commenter called the release "historic" — "For the first time, an open-weights LLM is right at the top."

Moonshot has not responded to questions about its training process. The company's founder, Yang Zhilin, earned his PhD at Carnegie Mellon. "These are legitimate researchers and engineers doing solid work," Hancock said. "If American models ground to a halt, I think China's progress would slow, but would still continue. They're not just riding coattails here."

Sources: HuggingFace — moonshotai/Kimi-K3, TechCrunch — Experts say exploiting Anthropic's Fable isn't how Kimi K3 got so good, TechCrunch — Treasury threatens sanctions, Hacker News discussion, Moonshot AI

More Hourlies Stories

Content on Anagnorisis is summarized, paraphrased, and editorialized from publicly available sources for length and clarity. Original sources are linked where available. All trademarks belong to their respective owners.

More from Anagnorisis