On July 16, the developer Simon Willison did what he does to every new model: asked it to draw a pelican riding a bicycle as an SVG. Kimi K3 obliged, burning through 13,241 reasoning tokens to get there and costing him 25 cents for a single cartoon bird. The bird is not the story. The thing that drew it is: 2.8 trillion parameters, released by a Chinese lab called Moonshot AI, and the largest open-weight model anyone has ever shipped.
Sit with that number for a second. K3 is roughly a 3-trillion-parameter model built as a mixture-of-experts, which means it activates only 16 of its 896 expert sub-networks per token, so it runs far leaner than the headline size suggests. It reads a million tokens of context at once. It handles text, images and video frames natively rather than bolting vision on as an afterthought. Moonshot says the architecture is about 2.5 times more efficient at converting compute into intelligence than the previous Kimi, and the full open weights are promised by July 27.

The benchmarks aren’t a vanity lap either. Artificial Analysis clocked K3 at an overall Elo of 1547, a jump of 732 points over Kimi K2.6, and it currently tops Arena.ai’s Frontend Code leaderboard. On independent testing it lands around fourth among all frontier models, beating Claude Opus 4.8 and GPT-5.5, trailing only Claude Fable 5 and GPT-5.6 Sol. Read that again and notice what’s unusual: of the models in that top tier, K3 is the only one you can download and run yourself.
That single fact is what turns a product launch into a geopolitics story. For two years the comfortable assumption in Washington was that the frontier would stay closed and American, that export controls on advanced chips would keep China a generation behind, and that “open” would mean smaller, weaker, safe-to-release models. Moonshot built a top-five model under exactly those compute restrictions and then handed the weights to anyone who wants them. The open-versus-closed debate just stopped being theoretical.
It stopped being theoretical in Delhi faster than almost anywhere. Indian companies burned through the honeymoon with American APIs quickly, because the bill arrives in dollars and the revenue arrives in rupees. “The token bills are a serious issue, it is increasingly becoming unsustainable,” Nikhil Narendran, a partner at the law firm Trilegal, told KrASIA. GPT-5.5 runs roughly $5 to $12 per million input tokens; a self-hosted open-weight model, once you own the weights, is a hosting bill instead of a metered toll. For a Bengaluru startup running millions of calls a day, that gap is the difference between a viable unit economic and a slow bleed.
So they switched. KrASIA reports Chinese open-weight models racked up 25 trillion tokens of usage in late June, 78 percent more than American models over the same stretch, with names like Coinbase, DoorDash and Airbnb in the mix alongside Indian startups. Puneet Kumar of Mirae Asset’s India venture arm says consumer-tech founders have leaned on Chinese open weights heavily since mid-2025. The pattern founders describe is pragmatic, not ideological: prototype on Google or OpenAI to move fast, then quietly shift the expensive production workload onto open weights they can tune and host themselves.
There’s a second reason that lands harder in India than in the Valley. A model you host yourself keeps your data inside the country, which matters when the data is Indian users’ financial records or health information and the alternative is shipping it to a server in Virginia. That’s the part policy people at places like the Observer Research Foundation keep circling: real AI resilience means not depending on infrastructure a foreign government could throttle. An open model you’ve already downloaded can’t be revoked by an export-control memo.
None of this makes K3 a clean win. “Open weights” is not the same as open source: you get the model file, not necessarily the training data or a permissive license, and Trilegal’s lawyers have been warning clients to scan these downloads for the kind of hidden nasties a trojaned model could carry. K3 itself is rough in spots, with what Willison calls “only one reasoning effort right now,” which is why a cartoon pelican cost a quarter. Early-adopter tax, paid in reasoning tokens.
But the shape of the argument has changed for good. The most capable downloadable AI on the planet was trained in China, released for free, and is already doing production work in Indian startups that can’t afford the American alternative. For a founder in Gurugram or Chennai, the question is no longer whether open models are good enough. K3 answered that on July 16. The question now is who they’d rather owe, and the honest answer increasingly is: nobody they have to pay per token.
Sources
- LLM Stats — Model updates, July 2026
- Simon Willison — Kimi K3, and what we can still learn from the pelican benchmark (16 July 2026)
- Tom’s Hardware — Moonshot releases 2.8-trillion-parameter Kimi K3, largest open-weight model ever (July 2026)
- Automatio — Kimi K3: 2.8T MoE AI with 1M context and native vision
- KrASIA — Indian companies look to Chinese LLMs as AI costs bite
SEO
- SEO title: Kimi K3: The World’s Largest Open AI Model Is Chinese
- Meta description: Kimi K3, Moonshot AI’s 2.8-trillion-parameter release, is the largest open model ever — and Indian startups are already running on it to escape dollar API bills.
- Focus keyword: Kimi K3 largest open model
- Secondary keywords: Moonshot AI Kimi K3, open-weight AI India, Kimi K3 2.8 trillion parameters
- Slug: kimi-k3-largest-open-model
- Category: Culture & Diaspora
- Tags: Kimi K3, Moonshot AI, open-weight AI, Indian startups, AI geopolitics, South Asian diaspora