LLM
Kimi K3 Review
Moonshot AI · LLM
Kimi K3 is a 2.8-trillion-parameter mixture-of-experts model released by Beijing-based Moonshot AI on 16 July 2026, and the largest open-weight model anyon
Kimi K3 is a 2.8-trillion-parameter mixture-of-experts model released by Beijing-based Moonshot AI on 16 July 2026, and the largest open-weight model anyone has published. The headline count is less alarming than it sounds: roughly 104 billion parameters are active per token, so serving cost tracks the active path rather than the full network. The context window is a genuine 1,048,576 tokens. On coding benchmarks it took first place on SWE Marathon at 42.0 against GPT-5.6 Sol's 40.0, and on Program Bench at 77.8 against 77.6, while trailing narrowly on Terminal Bench 2.1 at 88.3 to 88.8. Two of those three margins are small enough to sit inside normal benchmark variance, so treat the ranking as a signal rather than a verdict. What holds up regardless of whose leaderboard you trust is the commercial position: $2.90 per million input tokens and $14.00 per million output, with open weights. For anyone with data-residency constraints or a large agentic coding workload, that combination changes the calculation in a way a marginal benchmark win does not.
Pros
- Open weights at genuine frontier coding quality
- Roughly a fifth the output price of comparable closed frontier models
- Real 1,048,576-token context window
- Only ~104B of 2.8T parameters active per token, so serving cost stays sane
- Self-hostable, which resolves data-residency problems outright
Cons
- Benchmark leads of 0.2 points are inside measurement noise, not real advantages
- Coding benchmarks are the most contamination-prone category in the field
- Self-hosting a 2.8T model needs serious infrastructure — most teams will use the API
- Independent verification of the efficiency claims is still thin
- No free tier for evaluation
Use cases
- High-volume agentic coding where output token cost dominates the bill
- Deployments with data-residency or sovereignty requirements
- Whole-repository analysis that needs the million-token window
- Teams wanting a credible open fallback to closed frontier providers
Pricing
paid — from $2.90/M input tokens. No free tier.