ANALYSIS2026-07-16·Moonshot AI (subject)

Kimi K3 — Moonshot AI's 2.8T-Parameter Open-Weight Model

Simon Willison (independent analysis; corroborated by Bloomberg/Fortune/Axios/CNBC/Forbes/Tom's Hardware coverage of Moonshot AI's announcement)
COMPILED NOTES

Moonshot AI released Kimi K3, a 2.8T-parameter MoE (896 experts, 16 active/token), 1M context, native vision — largest open-weight model to date; API live at $3/$15 per Mtok (Sonnet-tier pricing, most expensive Chinese-lab release yet); open weights due 2026-07-27; tops Arena.ai Frontend Code leaderboard ahead of Claude Fable 5.

Kimi K3 — Moonshot AI's 2.8T-Parameter Open-Weight Model

Fetch note: Moonshot AI's own blog (moonshot.ai/blog) did not resolve to announcement content (nginx default page) at ingest time. This entry is built from Simon Willison's independent hands-on technical review — a respected, non-vendor source that ran his own pelican-SVG benchmark against the model — cross-checked against wire coverage (Bloomberg, Fortune, Axios, CNBC, Forbes, Tom's Hardware) that quotes Moonshot's own press release and Arena.ai / Artificial Analysis third-party benchmark results. Treat pricing and benchmark numbers as vendor/third-party reported, not independently verified by MenFem.

Core Facts

Released 2026-07-16 by Moonshot AI (Beijing). Successor to the Kimi K2 family (K2, K2.5, K2.6, K2.7 Code).

  • Architecture: Mixture-of-Experts, ~2.8 trillion total parameters, 896 experts, only 16 active per token (~1.8% of the pool) — the forward-pass compute cost is far below what the headline parameter count implies. Moonshot calls the architecture "LatentMoE" and cites two new attention mechanisms, Kimi Delta Attention and Attention Residuals, aimed at compute efficiency and long, low-oversight coding/agent sessions.
  • Context window: 1,000,000 tokens. Native vision (multimodal) input.
  • Reasoning: launched with a single reasoning-effort setting, "max" — no lower-effort tier yet, which drives up token consumption on simple tasks (Willison's pelican-SVG test burned 13,241 reasoning tokens, ~$0.25, on a "hello world"-level prompt).
  • Pricing (API, live now): $3 / million input tokens (cache miss; $0.30 on cache hit), $15 / million output tokens — the same price tier as Anthropic's Claude Sonnet family, and a sharp increase from Kimi K2.6's $0.95/$4. Notably the most expensive model a Chinese lab has shipped to date.
  • Open-weight release: promised for 2026-07-27 (not yet available at ingest time — API-only for now).

Benchmark Results (third-party reported)

  • Moonshot's own claim: K3 trails Claude Fable 5 and OpenAI's GPT-5.6 Sol on overall performance, but "mostly beats" Claude Opus 4.8 and GPT-5.5, and substantially outperforms other tested models on coding/agentic tasks.
  • Arena.ai Frontend Code arena: K3 ranks #1, ahead of Claude Fable 5 — the first time an open-weight model has topped this leaderboard.
  • Vercel professional web-engineering evaluation: ranked #1 overall.
  • Artificial Analysis (long-horizon knowledge work Elo): 1547 (+732 over K2.6) — "behind only Claude Fable 5"; ~21% fewer output tokens than K2.6 on the same eval; ~$0.94 cost per task.

Limitations / Caveats

  • Full open weights not yet released at ingest time (due 2026-07-27) — independent third-party reproduction of benchmark claims is not yet possible.
  • Willison's own methodological note: the pelican-SVG benchmark "has mostly been severed" from actual model-quality correlation and should not be used comparatively — it's a "hello world" smoke test, not a capability measure.
  • Unusual tokenization observed (a 10-word prompt consumed 95 tokens, implying an ~85-token hidden system prompt) — a data point on the model's real per-request token overhead that headline pricing doesn't capture.
  • No primary Moonshot source was directly fetchable; all figures are vendor-press-release-via-journalism or third-party benchmark aggregator (Arena.ai, Artificial Analysis), not independently re-run by MenFem.

Why this matters for the KB

Direct hit on the AI LENS's "Open vs. closed" required chapter — K3 is the clearest 2026 instance (alongside DeepSeek V4, already in this KB) of an open-weight Chinese lab compressing the gap to the closed frontier while raising, not cutting, price versus its own prior generation. That's a reversal of the DeepSeek V4 story (which competed on cost) — K3 competes on capability at Sonnet-tier pricing, a genuinely different competitive posture worth tracking against the content-focus.md inference-economics spine.


Sources: Kimi K3, and what we can still learn from the pelican benchmark — Simon Willison, 2026-07-16. Corroborated by: Bloomberg, Fortune, Axios, CNBC, Tom's Hardware.

RELATED · IN THE BASE
Kimi K3 — Moonshot AI's 2.8T-Parameter Open-Weight Model | Knowledge Base | MenFem