The arena · AI model review
Kimi K3
by Moonshot AI
Moonshot's 2.8T-parameter MoE that sparked the 'new DeepSeek moment': frontier-adjacent scores at $3/$15
$15/1M out
Official Moonshot API list price for kimi-k3: $3/1M input (cache miss), $0.30/1M on cache hit, $15/1M output including reasoning trace, flat across the 1M context (no length tiers). Same $3/$15 on OpenRouter (no cache pricing exposed there). A 3x increase over Kimi K2.6's $0.95/$4. Verified against platform.kimi.ai/docs/pricing/chat-k3, 2026-07.
Moonshot AI
1M tokens
$3/1M in
$15/1M out
text, vision, video input
No
What is Kimi K3?
Released 2026-07-16 by Beijing-based Moonshot AI and unveiled at the World AI Conference in Shanghai. Kimi K3 is a 2.8-trillion-parameter Mixture-of-Experts model (896 experts, 16 active per token) with a 1M-token context window and native text, image and video input. API ID kimi-k3 at $3/1M input ($0.30 on cache hit) and $15/1M output, flat across the whole context window, also on OpenRouter. It reached #3 on the Artificial Analysis Intelligence Index at launch, comparable to Claude Opus 4.8 and GPT-5.5, triggering 'new DeepSeek moment' headlines and a semiconductor sell-off. Open weights (~594 GB, MXFP4) were announced for 2026-07-27; not yet published at the time of this fiche, with a Modified-MIT-style license expected but unconfirmed.
Kimi K3 pros & cons
Pros
- #3 on the Artificial Analysis Intelligence Index (57) at launch, comparable to Claude Opus 4.8 and GPT-5.5: the closest a Chinese lab has come to the closed US frontier
- 93.4% on SWE-bench Verified in Vals AI's independent harness (GPT-5.6 Sol: 96.2%, Claude Fable 5: 95.0%) and #1 on Arena.ai's Frontend Code Arena at 1679 points, ahead of Fable 5
- 93.5% on GPQA Diamond (between Fable 5's 92.6% and GPT-5.6's 94.1%), 96.1% on AIME 2025, and Elo 1668 on GDPval-AA v2 agentic work, second only to Fable 5
- Strong agentic profile: #1 on AutomationBench-AA SaaS workflows (53%), long-horizon terminal and repo navigation via Kimi Code, and ~21% fewer output tokens than K2 for more intelligence
- $3/$15 per 1M tokens (input drops to $0.30 on cache hit), flat across the full 1M context: still well under US closed-frontier pricing, with an OpenAI-compatible API and OpenRouter availability
- Native multimodal input (text, image, video) and open weights announced for 2026-07-27 under an expected Modified-MIT-style license, positioning it as the first 3T-class open-weights frontier model
Cons
- 3x price jump over Kimi K2.6 ($0.95/$4 to $3/$15) makes it the most expensive Chinese model ever shipped, at Claude Sonnet 5 list price: the '10x cheaper than US models' era is over (Simon Willison documented the hike)
- Slow: 33 output tokens/s, ranked #145 of 190 models on Artificial Analysis, a poor fit for interactive use
- Hallucination rate climbed from 39% (K2.6) to 51% on AA-Omniscience as the model now attempts answers it would previously refuse (Fable 5 sits at a comparable 54.9%)
- Political censorship on sensitive topics in Mandarin (deflects on Xi, CCP, Tiananmen) and Beijing jurisdiction for API data, a compliance blocker for many EU and enterprise deployments; UK AISI/CAISI also found its cyber guardrails failed to block exploit-development attempts during testing
- 'Open weights' is theoretical for most: ~594 GB in native MXFP4 requiring a multi-GPU datacenter to self-host, and the weights (plus final license) were still unpublished at review time
The arena’s verdict
Kimi K3 is the strongest argument yet that the frontier is no longer exclusively American: #3 on Artificial Analysis, top-tier GPQA and SWE-bench Verified scores, and the best frontend-code arena ranking in the business, at roughly half to a third of US closed-frontier prices. Take it for agentic coding, frontend work and SaaS automation where its benchmarks are strongest, or if your roadmap depends on self-hosting a frontier-class model once the weights land. Skip it for interactive products (33 tokens/s is slow), for anything touching politically sensitive content or strict EU data-residency requirements (Mandarin-language censorship, Beijing jurisdiction), and for high-accuracy retrieval where its 51% hallucination rate on AA-Omniscience demands a verification layer. Cost-obsessed teams should note DeepSeek V4 still delivers vastly more tokens per dollar; K3's pitch is peak capability per dollar, not cheapest tokens.
Thumbs up or thumbs down
Cast your verdict
Would you recommend Kimi K3, or warn the crowd away?
Top Kimi K3 alternatives
All alternativesAnthropic's fastest model: about 90% of Sonnet 4.5's coding skill at $1/$5 per 1M tokens, 200K context.
Anthropic's July 2026 Opus refresh: near-Fable 5 intelligence at half the price, same $5/$25 as Opus 4.8
Anthropic's June 2026 Mythos-class flagship: 80.3% on SWE-bench Pro, 11 points clear of every other frontier model
Compare Kimi K3 head-to-head
What the crowd says
“Tried it for our knowledge-base assistant and the hallucination rate is real: it confidently invented two API endpoints in one afternoon. 33 tokens per second on top of that. Back to waiting for the open weights to fine-tune.”
“Our Zapier-style automation stack runs on it now. Tool orchestration that GPT-5.5 fumbled works reliably, and the cache-hit pricing makes repeated workflows genuinely cheap.”
“The Frontend Arena ranking is deserved. Gave it the same dashboard spec I gave Fable 5 and K3's React came out cleaner, with fewer invented props. At $3/$15 through OpenRouter it was a drop-in swap.”
Kimi K3: frequently asked questions
How much does Kimi K3 cost per 1M tokens?
Kimi K3 costs $3/1M in per 1M input tokens and $15/1M out per 1M output tokens. Official Moonshot API list price for kimi-k3: $3/1M input (cache miss), $0.30/1M on cache hit, $15/1M output including reasoning trace, flat across the 1M context (no length tiers). Same $3/$15 on OpenRouter (no cache pricing exposed there). A 3x increase over Kimi K2.6's $0.95/$4. Verified against platform.kimi.ai/docs/pricing/chat-k3, 2026-07.