Head-to-head

Claude Opus 5 logovsKimi K3 logo

Claude Opus 5 vs Kimi K3: which AI model wins in 2026?

Claude Opus 5 ($25/1M out) and Kimi K3 ($15/1M out) are two of the most-used AI models in 2026. Across 7 community votes, Claude Opus 5 leads with 63% approval.

Quick verdict

On Reasoning, pick Claude Opus 5: the arena rates it 5/5 against 4.5/5 for Kimi K3. On budget, Kimi K3 wins: it starts at $15/1M out versus $25/1M out for Claude Opus 5.

Line-by-line comparison

From
$25/1M outOfficial Anthropic API list price for claude-opus-5: $5/1M input, $25/1M output, unchanged from Opus 4.8, single tier with 1M context as default and maximum, 128K max output, prompt caching from 512 tokens. Research-preview Fast mode at $10/$50 (~2.5x faster) is Claude API only. Verified against platform.claude.com (What's new in Opus 5) 2026-07.
$15/1M outOfficial Moonshot API list price for kimi-k3: $3/1M input (cache miss), $0.30/1M on cache hit, $15/1M output including reasoning trace, flat across the 1M context (no length tiers). Same $3/$15 on OpenRouter (no cache pricing exposed there). A 3x increase over Kimi K2.6's $0.95/$4. Verified against platform.kimi.ai/docs/pricing/chat-k3, 2026-07.
Provider
Anthropic
Moonshot AI
Context window
1M tokens
1M tokens
Input price
$5/1M in
$3/1M in
Output price
$25/1M out
$15/1M out
Modalities
text, vision
text, vision, video input
Open weights
No
No
Crowd score
63%(4)
57%(3)
Arena ratings (1-5)
Reasoning
5.0
4.5
Coding
5.0
4.5
Writing
4.0
4.0
Speed
2.0
2.0
Value
4.0
3.5

Strengths and weaknesses

Claude Opus 5

  • Ranked #1 in composite intelligence across 190 models on Artificial Analysis at launch, and 43.3% on Frontier-Bench v0.1 vs 34.4% for GPT-5.6 Sol and 33.7% for Claude Fable 5
  • 3.9x better than GPT-5.6 Sol on ARC-AGI-3 novel reasoning (30.2% vs 7.8%), and Elo 1861 on GDPval-AA v2 economic knowledge work, ahead of Fable 5 (1747)
  • 79.2% on SWE-bench Pro, within a point of Fable 5 (80.0%) and 10 points above Opus 4.8 (69.2%), at half Fable's price; Cursor's co-founder calls it 'near Fable 5 intelligence at Opus speed and cost'
  • Same $5/$25 pricing as Opus 4.8 with a bigger window: 1M context is now the default and only tier, with 128K max output and prompt caching from 512 tokens
  • Dual-use safety classifiers trigger 85% less often than on Fable 5, and the new default fallback mode avoids the silent mid-session refusals that plagued Fable's launch
  • Self-verifies its work without being told, handles mid-conversation tool changes (beta) without busting the prompt cache, and ships day one on Claude.ai, the API, Bedrock, Vertex and Microsoft Foundry
  • Notably slow and very verbose: 52.6 output tokens/s and 68 seconds to first token on Artificial Analysis, and it consumed ~100M output tokens during their eval vs a 63M median
  • The verbosity is a real-world cost problem: CodeRabbit measured it reading ~50% more and writing ~65% more than reference frontier models per code-review call
  • No actual price cut despite the 'cost-efficient' narrative: identical to Opus 4.8 ($5/$25), nearly GPT-5.6 money ($5/$30), and more than 2x comparable Gemini or Grok tiers
  • Reviewers describe a 'brilliant but annoying' personality: over-verification, hedging, and occasional refusals of mundane tasks like resolving a merge conflict (Lenny's Newsletter field review)
  • Breaking API change for Opus 4.8 migrants: thinking is on by default and cannot be disabled at xhigh or max effort (returns a 400); Fast mode ($10/$50, ~2.5x faster) is API-only, not on Bedrock or Vertex

Kimi K3

  • #3 on the Artificial Analysis Intelligence Index (57) at launch, comparable to Claude Opus 4.8 and GPT-5.5: the closest a Chinese lab has come to the closed US frontier
  • 93.4% on SWE-bench Verified in Vals AI's independent harness (GPT-5.6 Sol: 96.2%, Claude Fable 5: 95.0%) and #1 on Arena.ai's Frontend Code Arena at 1679 points, ahead of Fable 5
  • 93.5% on GPQA Diamond (between Fable 5's 92.6% and GPT-5.6's 94.1%), 96.1% on AIME 2025, and Elo 1668 on GDPval-AA v2 agentic work, second only to Fable 5
  • Strong agentic profile: #1 on AutomationBench-AA SaaS workflows (53%), long-horizon terminal and repo navigation via Kimi Code, and ~21% fewer output tokens than K2 for more intelligence
  • $3/$15 per 1M tokens (input drops to $0.30 on cache hit), flat across the full 1M context: still well under US closed-frontier pricing, with an OpenAI-compatible API and OpenRouter availability
  • Native multimodal input (text, image, video) and open weights announced for 2026-07-27 under an expected Modified-MIT-style license, positioning it as the first 3T-class open-weights frontier model
  • 3x price jump over Kimi K2.6 ($0.95/$4 to $3/$15) makes it the most expensive Chinese model ever shipped, at Claude Sonnet 5 list price: the '10x cheaper than US models' era is over (Simon Willison documented the hike)
  • Slow: 33 output tokens/s, ranked #145 of 190 models on Artificial Analysis, a poor fit for interactive use
  • Hallucination rate climbed from 39% (K2.6) to 51% on AA-Omniscience as the model now attempts answers it would previously refuse (Fable 5 sits at a comparable 54.9%)
  • Political censorship on sensitive topics in Mandarin (deflects on Xi, CCP, Tiananmen) and Beijing jurisdiction for API data, a compliance blocker for many EU and enterprise deployments; UK AISI/CAISI also found its cyber guardrails failed to block exploit-development attempts during testing
  • 'Open weights' is theoretical for most: ~594 GB in native MXFP4 requiring a multi-GPU datacenter to self-host, and the weights (plus final license) were still unpublished at review time

Cast your verdict

One recommendation per tool per gladiator. It reshapes the crowd score everyone sees.

Claude Opus 5$25/1M out
63%crowd score · 4
Kimi K3$15/1M out
57%crowd score · 3

The arena’s verdict on Claude Opus 5

Claude Opus 5 is the sane default of the Series 5 range: most of Fable 5's intelligence (and more than Fable on Frontier-Bench and GDPval) at exactly half the token price, with classifiers that trigger 85% less often. If you migrated workloads to Fable 5 for capability but resent the bill or the false-positive refusals, move them here; if you are still on Opus 4.8, the upgrade is 10 SWE-bench Pro points for free. The two honest reasons to look elsewhere: latency and verbosity. At 52.6 tokens/s with 68s to first token it is a poor fit for interactive UX, and its token appetite quietly inflates real costs beyond the sticker price, so budget-sensitive high-volume pipelines still belong on Sonnet 5, Gemini or DeepSeek. Keep Fable 5 only for the longest autonomous runs where its slight SWE-bench Pro edge compounds.

The arena’s verdict on Kimi K3

Kimi K3 is the strongest argument yet that the frontier is no longer exclusively American: #3 on Artificial Analysis, top-tier GPQA and SWE-bench Verified scores, and the best frontend-code arena ranking in the business, at roughly half to a third of US closed-frontier prices. Take it for agentic coding, frontend work and SaaS automation where its benchmarks are strongest, or if your roadmap depends on self-hosting a frontier-class model once the weights land. Skip it for interactive products (33 tokens/s is slow), for anything touching politically sensitive content or strict EU data-residency requirements (Mandarin-language censorship, Beijing jurisdiction), and for high-accuracy retrieval where its 51% hallucination rate on AA-Omniscience demands a verification layer. Cost-obsessed teams should note DeepSeek V4 still delivers vastly more tokens per dollar; K3's pitch is peak capability per dollar, not cheapest tokens.

What the crowd says

On Claude Opus 5

Thumbs Downicus

The sticker price is unchanged but my invoice is not: it writes essays where Opus 4.8 wrote answers. 68 seconds to first token killed it for our support chat, we went back to Sonnet 5.

The Fair Reviewer

Fed it a 40-tab financial model with cross-sheet formulas and asked for a scenario deck. It got the edge cases the analysts missed. For document-heavy enterprise work this is the best model I have used.

Sir Ships-A-Lot

We moved off Fable 5 because the bio classifier kept flagging our genomics tooling. Opus 5 does the same work at half the price and I have not seen a single silent reroute since.

Guardian of the Repo

Migrated our agents from Opus 4.8 the day it dropped. Same bill, and tasks that used to stall at the planning stage now just finish. The 10-point SWE-bench jump is not marketing, our merge queue feels it.

On Kimi K3

Captain Churn

Tried it for our knowledge-base assistant and the hallucination rate is real: it confidently invented two API endpoints in one afternoon. 33 tokens per second on top of that. Back to waiting for the open weights to fine-tune.

Golden Thumbicus

Our Zapier-style automation stack runs on it now. Tool orchestration that GPT-5.5 fumbled works reliably, and the cache-hit pricing makes repeated workflows genuinely cheap.

Saint Deployus

The Frontend Arena ranking is deserved. Gave it the same dashboard spec I gave Fable 5 and K3's React came out cleaner, with fewer invented props. At $3/$15 through OpenRouter it was a drop-in swap.

Frequently asked questions

Is Claude Opus 5 better than Kimi K3?

The crowd currently sides with Claude Opus 5: 63% recommend it, versus 57% for Kimi K3 (7 votes). On Reasoning, Claude Opus 5 rates higher (5/5 vs 4.5/5). The right pick depends on your use case. The line-by-line comparison on this page breaks down pricing, key specs and arena ratings.

Which is cheaper, Claude Opus 5 or Kimi K3?

Kimi K3 is cheaper: it starts at $15/1M out, while Claude Opus 5 starts at $25/1M out.

How much do Claude Opus 5 and Kimi K3 cost per 1M tokens?

Claude Opus 5: $5/1M in per 1M input tokens, $25/1M out per 1M output tokens. Kimi K3: $3/1M in per 1M input tokens, $15/1M out per 1M output tokens.