AI
Moonshot's Kimi K3 Puts Open-Weight AI Within Months of the Frontier
China's latest flagship model concedes it trails America's best, but its price and openness have reset the race
5 min read
By Timmy
China's open-weight AI labs have spent the past two years chasing the American frontier. On July 16, Beijing-based Moonshot AI released the clearest evidence yet that the chase is working: Kimi K3, a 2.8-trillion-parameter system the company describes as the world's first open 3T-class model.
The model went live across Kimi.com, the Kimi Work desktop app, the Kimi Code programming tool and Moonshot's API. Full open weights are promised by July 27 alongside a technical report. The launch package includes a one-million-token context window, native image and video input, and reasoning that runs at maximum effort by default.
In an unusual gesture for a model launch, Moonshot's own materials concede that K3 trails Anthropic's Claude Fable 5 and OpenAI's GPT-5.6 Sol overall. The fine print still makes uneasy reading for US labs.
Despite being a highly competitive model overall, K3 nonetheless exhibits a noticeable gap in user experience compared with Claude Fable 5 and GPT 5.6 Sol.
-From Kimi K3 Tech Blog
Inside the machine
K3 is a mixture-of-experts design that activates 16 of 896 experts for each token, putting roughly 50 billion parameters to work on every pass. This cycle adds two architectural changes, called Kimi Delta Attention and Attention Residuals, on top of a routing framework named Stable LatentMoE.
Moonshot also trained the model to live with 4-bit MXFP4 weights and 8-bit activations, which shrinks weight storage to about 1.4 terabytes rather than the 5.6 terabytes full precision would demand. That is still far beyond consumer hardware. The company recommends server "supernodes" with at least 64 accelerators, and as an analysis from Binary Verse AI put it, "that's not a multi-GPU workstation problem, it's a small cluster problem."
Independent evaluators put K3 near the top. Artificial Analysis scored it 57 on its composite Intelligence Index, a point ahead of Anthropic's Claude Opus 4.8, and the model took first place on one of LMArena's coding boards within hours of release. The numbers come with a caveat, though: each model ran inside its own maker's tooling, so third-party verification, still under way, will be the cleaner test.

Moonshot flags quirks of its own. K3 was trained to keep its full reasoning history, and quality can wobble if a deployment drops that history mid-task. The company also warns that the model can be "excessively proactive" on ambiguous instructions, making decisions a user never asked for.
Wins, losses and an honest chart
On Moonshot's published benchmarks, K3 leads every compared model on SWE Marathon, a long-horizon coding evaluation, and edges Fable 5 on Terminal Bench 2.1 by 88.3 points to 84.6. It also posted the best scores on BrowseComp, which tests web-browsing agents, and OmniDocBench, which tests document understanding.
The losses are real. Fable 5 and GPT-5.6 Sol pulled clear on DeepSWE and FrontierSWE, and Fable 5 kept a wide lead on Humanity's Last Exam, a test designed to stump frontier models. "Kimi K3 did not end the frontier's lead," wrote BenchLM, an AI benchmarking publication. "It put the countdown on the open side of the ledger."
The price is the point
K3's API charges US$3 per million input tokens, US$0.30 when the input is cached, and US$15 per million output tokens, flat across the full million-token window. Fable 5 charges US$50 per million output tokens. "The pricing is the real signal," noted DataCamp in its launch coverage: if a Chinese model charges frontier rates and performs near the frontier, the premium attached to American models starts to look arbitrary.
Within China, the positioning is aggressive. Zhipu's GLM-5.2 undercuts the field at US$1.40 per million input tokens under an MIT licence, while Alibaba's Qwen 3.7 Max, from the same company that backs Moonshot, has gone closed. Moonshot is betting that scale plus openness beats both plays.
Moonshot says its Mooncake serving system, which splits the heavy stages of inference across separate machine pools, reaches a reported 90 per cent cache hit rate on coding workloads, which is what makes the cheap cached-input price possible. Artificial Analysis adds a practical warning: K3 reasons at length, burning 130 million output tokens across its evaluation suite against a median near 63 million, and it calls the model "more expensive and more verbose than comparable models in its price tier." A high solve rate can coexist with a high bill.
Months, not years
The launch lands on top of a measurable trend. Britain's AI Security Institute recently put the gap between open and closed frontier models at four to seven months on its cyber ranges, down from six to ten months through most of 2025. The labs closing that gap, Moonshot, Zhipu and DeepSeek among them, are Chinese, and they ship weights while the Western frontier ships APIs.
K3 also arrives with baggage. Researchers on Reddit and X have noted that the model occasionally echoes Anthropic's safety policies inside its reasoning traces, reviving speculation, trailing Moonshot since February, that a rival's outputs fed its training data. The questions remain unresolved.
The next checkpoint is dated. When the weights and technical report arrive on July 27, outside labs can try to reproduce Moonshot's numbers, fine-tune the model and probe the architecture for themselves. If the results hold, the assumption that frontier intelligence stays expensive, and American, will get harder to defend.







