The arena · AI model review

Kimi K3

by Moonshot AI

Moonshot's 2.8T-parameter MoE that sparked the 'new DeepSeek moment': frontier-adjacent scores at $3/$15

Arena score 3.7/557% recommended · 3 votes
Reasoning4.5
Coding4.5
Writing4.0
Speed2.0
Value3.5
Visit Kimi K3frontier-level coding and agents outside US cloudsfrontend generation (top of Arena.ai's leaderboard)SaaS workflow automation and tool orchestrationteams planning to self-host 3T-class open weights
Price

$15/1M out

Official Moonshot API list price for kimi-k3: $3/1M input (cache miss), $0.30/1M on cache hit, $15/1M output including reasoning trace, flat across the 1M context (no length tiers). Same $3/$15 on OpenRouter (no cache pricing exposed there). A 3x increase over Kimi K2.6's $0.95/$4. Verified against platform.kimi.ai/docs/pricing/chat-k3, 2026-07.

Provider

Moonshot AI

Context window

1M tokens

Input price

$3/1M in

Output price

$15/1M out

Modalities

text, vision, video input

Open weights

No

What is Kimi K3?

Released 2026-07-16 by Beijing-based Moonshot AI and unveiled at the World AI Conference in Shanghai. Kimi K3 is a 2.8-trillion-parameter Mixture-of-Experts model (896 experts, 16 active per token) with a 1M-token context window and native text, image and video input. API ID kimi-k3 at $3/1M input ($0.30 on cache hit) and $15/1M output, flat across the whole context window, also on OpenRouter. It reached #3 on the Artificial Analysis Intelligence Index at launch, comparable to Claude Opus 4.8 and GPT-5.5, triggering 'new DeepSeek moment' headlines and a semiconductor sell-off. Open weights (~594 GB, MXFP4) were announced for 2026-07-27; not yet published at the time of this fiche, with a Modified-MIT-style license expected but unconfirmed.

Kimi K3 pros & cons

Pros

  • #3 on the Artificial Analysis Intelligence Index (57) at launch, comparable to Claude Opus 4.8 and GPT-5.5: the closest a Chinese lab has come to the closed US frontier
  • 93.4% on SWE-bench Verified in Vals AI's independent harness (GPT-5.6 Sol: 96.2%, Claude Fable 5: 95.0%) and #1 on Arena.ai's Frontend Code Arena at 1679 points, ahead of Fable 5
  • 93.5% on GPQA Diamond (between Fable 5's 92.6% and GPT-5.6's 94.1%), 96.1% on AIME 2025, and Elo 1668 on GDPval-AA v2 agentic work, second only to Fable 5
  • Strong agentic profile: #1 on AutomationBench-AA SaaS workflows (53%), long-horizon terminal and repo navigation via Kimi Code, and ~21% fewer output tokens than K2 for more intelligence
  • $3/$15 per 1M tokens (input drops to $0.30 on cache hit), flat across the full 1M context: still well under US closed-frontier pricing, with an OpenAI-compatible API and OpenRouter availability
  • Native multimodal input (text, image, video) and open weights announced for 2026-07-27 under an expected Modified-MIT-style license, positioning it as the first 3T-class open-weights frontier model

Cons

  • 3x price jump over Kimi K2.6 ($0.95/$4 to $3/$15) makes it the most expensive Chinese model ever shipped, at Claude Sonnet 5 list price: the '10x cheaper than US models' era is over (Simon Willison documented the hike)
  • Slow: 33 output tokens/s, ranked #145 of 190 models on Artificial Analysis, a poor fit for interactive use
  • Hallucination rate climbed from 39% (K2.6) to 51% on AA-Omniscience as the model now attempts answers it would previously refuse (Fable 5 sits at a comparable 54.9%)
  • Political censorship on sensitive topics in Mandarin (deflects on Xi, CCP, Tiananmen) and Beijing jurisdiction for API data, a compliance blocker for many EU and enterprise deployments; UK AISI/CAISI also found its cyber guardrails failed to block exploit-development attempts during testing
  • 'Open weights' is theoretical for most: ~594 GB in native MXFP4 requiring a multi-GPU datacenter to self-host, and the weights (plus final license) were still unpublished at review time

The arena’s verdict

Kimi K3 is the strongest argument yet that the frontier is no longer exclusively American: #3 on Artificial Analysis, top-tier GPQA and SWE-bench Verified scores, and the best frontend-code arena ranking in the business, at roughly half to a third of US closed-frontier prices. Take it for agentic coding, frontend work and SaaS automation where its benchmarks are strongest, or if your roadmap depends on self-hosting a frontier-class model once the weights land. Skip it for interactive products (33 tokens/s is slow), for anything touching politically sensitive content or strict EU data-residency requirements (Mandarin-language censorship, Beijing jurisdiction), and for high-accuracy retrieval where its 51% hallucination rate on AA-Omniscience demands a verification layer. Cost-obsessed teams should note DeepSeek V4 still delivers vastly more tokens per dollar; K3's pitch is peak capability per dollar, not cheapest tokens.

Thumbs up or thumbs down

Cast your verdict

Would you recommend Kimi K3, or warn the crowd away?

57%crowd score · 3

Top Kimi K3 alternatives

All alternatives
1
Claude Haiku 4.5

Anthropic's fastest model: about 90% of Sonnet 4.5's coding skill at $1/$5 per 1M tokens, 200K context.

$5/1M out67%(2)
Visit
2
Claude Opus 5

Anthropic's July 2026 Opus refresh: near-Fable 5 intelligence at half the price, same $5/$25 as Opus 4.8

$25/1M out63%(4)
Visit
3
Claude Fable 5

Anthropic's June 2026 Mythos-class flagship: 80.3% on SWE-bench Pro, 11 points clear of every other frontier model

$50/1M out63%(4)
Visit

Compare Kimi K3 head-to-head

What the crowd says

Captain Churn

Tried it for our knowledge-base assistant and the hallucination rate is real: it confidently invented two API endpoints in one afternoon. 33 tokens per second on top of that. Back to waiting for the open weights to fine-tune.

Golden Thumbicus

Our Zapier-style automation stack runs on it now. Tool orchestration that GPT-5.5 fumbled works reliably, and the cache-hit pricing makes repeated workflows genuinely cheap.

Saint Deployus

The Frontend Arena ranking is deserved. Gave it the same dashboard spec I gave Fable 5 and K3's React came out cleaner, with fewer invented props. At $3/$15 through OpenRouter it was a drop-in swap.

Kimi K3: frequently asked questions

How much does Kimi K3 cost per 1M tokens?

Kimi K3 costs $3/1M in per 1M input tokens and $15/1M out per 1M output tokens. Official Moonshot API list price for kimi-k3: $3/1M input (cache miss), $0.30/1M on cache hit, $15/1M output including reasoning trace, flat across the 1M context (no length tiers). Same $3/$15 on OpenRouter (no cache pricing exposed there). A 3x increase over Kimi K2.6's $0.95/$4. Verified against platform.kimi.ai/docs/pricing/chat-k3, 2026-07.