Head-to-head
Claude Opus 4.8 vs Kimi K3: which AI model wins in 2026?
Claude Opus 4.8 ($25/1M out) and Kimi K3 ($15/1M out) are two of the most-used AI models in 2026. Across 6 community votes, Claude Opus 4.8 leads with 57% approval.
Quick verdict
On Reasoning, Claude Opus 4.8 and Kimi K3 are tied at 4.5/5. On budget, Kimi K3 wins: it starts at $15/1M out versus $25/1M out for Claude Opus 4.8.
Line-by-line comparison
Strengths and weaknesses
Claude Opus 4.8
- SWE-Bench Pro 69.2% (vs 64.3% for Opus 4.7) and beats prior Opus models on CursorBench at every effort level; strong real-world reports on large refactors and multi-file bug hunts
- About 4x less likely than Opus 4.7 to let flaws in its own generated code pass unflagged; big jump on math reasoning (USAMO 2026: 96.7% vs 69.3%)
- 1M-token context and 128K output at unchanged $5/$25 pricing, with no long-context premium; batch API at 50% off ($2.50/$12.50)
- Fast mode (research preview) delivers up to 2.5x output speed at $10/$50, 3x cheaper than Opus 4.7's fast tier ($30/$150)
- Unique API features for agents: mid-conversation system messages that preserve the prompt cache, and Dynamic Workflows spawning parallel subagents in Claude Code
- 84% on Online-Mind2Web browser automation and record score on Legal Agent Benchmark (first model past 10% all-pass); strong enterprise knowledge work (Box reports 87% vs 77% internally)
- Turn-by-turn regressions reported: missed obvious instructions in planning docs, answering a narrow slice of the goal, and worse one-shot simple UI generation than 4.7
- Writing style criticized by heavy users: excessive hedging, over-cautious editing that 'cuts anything bold or funny' (Steve Yegge), and pushback loops even against well-evidenced theses
- Language-mixing quirk: users report random Chinese, Cyrillic, or Greek insertions in long research threads
- Visible quality degradation past ~200K tokens in hands-on use despite the advertised 1M window
- Vending-Bench regression: fell for scam suppliers about 30x more than 4.7 and negotiates worse (a side effect of stricter honesty alignment)
Kimi K3
- #3 on the Artificial Analysis Intelligence Index (57) at launch, comparable to Claude Opus 4.8 and GPT-5.5: the closest a Chinese lab has come to the closed US frontier
- 93.4% on SWE-bench Verified in Vals AI's independent harness (GPT-5.6 Sol: 96.2%, Claude Fable 5: 95.0%) and #1 on Arena.ai's Frontend Code Arena at 1679 points, ahead of Fable 5
- 93.5% on GPQA Diamond (between Fable 5's 92.6% and GPT-5.6's 94.1%), 96.1% on AIME 2025, and Elo 1668 on GDPval-AA v2 agentic work, second only to Fable 5
- Strong agentic profile: #1 on AutomationBench-AA SaaS workflows (53%), long-horizon terminal and repo navigation via Kimi Code, and ~21% fewer output tokens than K2 for more intelligence
- $3/$15 per 1M tokens (input drops to $0.30 on cache hit), flat across the full 1M context: still well under US closed-frontier pricing, with an OpenAI-compatible API and OpenRouter availability
- Native multimodal input (text, image, video) and open weights announced for 2026-07-27 under an expected Modified-MIT-style license, positioning it as the first 3T-class open-weights frontier model
- 3x price jump over Kimi K2.6 ($0.95/$4 to $3/$15) makes it the most expensive Chinese model ever shipped, at Claude Sonnet 5 list price: the '10x cheaper than US models' era is over (Simon Willison documented the hike)
- Slow: 33 output tokens/s, ranked #145 of 190 models on Artificial Analysis, a poor fit for interactive use
- Hallucination rate climbed from 39% (K2.6) to 51% on AA-Omniscience as the model now attempts answers it would previously refuse (Fable 5 sits at a comparable 54.9%)
- Political censorship on sensitive topics in Mandarin (deflects on Xi, CCP, Tiananmen) and Beijing jurisdiction for API data, a compliance blocker for many EU and enterprise deployments; UK AISI/CAISI also found its cyber guardrails failed to block exploit-development attempts during testing
- 'Open weights' is theoretical for most: ~594 GB in native MXFP4 requiring a multi-GPU datacenter to self-host, and the weights (plus final license) were still unpublished at review time
Cast your verdict
One recommendation per tool per gladiator. It reshapes the crowd score everyone sees.
The arena’s verdict on Claude Opus 4.8
A drop-in upgrade for Opus 4.7 users: identical API surface and $5/$25 pricing with real gains on long-horizon agentic coding, code review, and enterprise analysis. Choose it if you run Claude Code, multi-file migrations, security audits, or agent pipelines that inspect, act, and verify over many steps. Skip it for quick one-shot UI snippets or prompts tightly tuned to 4.7 behavior, where users report regressions, and pick Sonnet 5 ($3/$15, intro $2/$10 through Aug 2026) if cost matters more than ceiling capability. Writers sensitive to hedging and over-cautious editing may find its style frustrating.
The arena’s verdict on Kimi K3
Kimi K3 is the strongest argument yet that the frontier is no longer exclusively American: #3 on Artificial Analysis, top-tier GPQA and SWE-bench Verified scores, and the best frontend-code arena ranking in the business, at roughly half to a third of US closed-frontier prices. Take it for agentic coding, frontend work and SaaS automation where its benchmarks are strongest, or if your roadmap depends on self-hosting a frontier-class model once the weights land. Skip it for interactive products (33 tokens/s is slow), for anything touching politically sensitive content or strict EU data-residency requirements (Mandarin-language censorship, Beijing jurisdiction), and for high-accuracy retrieval where its 51% hallucination rate on AA-Omniscience demands a verification layer. Cost-obsessed teams should note DeepSeek V4 still delivers vastly more tokens per dollar; K3's pitch is peak capability per dollar, not cheapest tokens.
What the crowd says
On Claude Opus 4.8
“Writing took a hit. It hedges everything and edits any bold or funny line out of my drafts. Also caught it answering a narrow slice of my planning doc and calling it done.”
“Threw USAMO-level math at it for a lark and it just grinds through. 96.7 vs 69 for 4.7 tracks with what I see. Same $5/$25, 1M context, no excuse not to switch.”
“Upgraded from 4.7 for a monorepo refactor and the difference is real. It actually flags its own sketchy code instead of shipping it. Multi-file bug hunts feel way less babysat.”
On Kimi K3
“Tried it for our knowledge-base assistant and the hallucination rate is real: it confidently invented two API endpoints in one afternoon. 33 tokens per second on top of that. Back to waiting for the open weights to fine-tune.”
“Our Zapier-style automation stack runs on it now. Tool orchestration that GPT-5.5 fumbled works reliably, and the cache-hit pricing makes repeated workflows genuinely cheap.”
“The Frontend Arena ranking is deserved. Gave it the same dashboard spec I gave Fable 5 and K3's React came out cleaner, with fewer invented props. At $3/$15 through OpenRouter it was a drop-in swap.”
Keep comparing
Frequently asked questions
Is Claude Opus 4.8 better than Kimi K3?
The right pick depends on your use case. The line-by-line comparison on this page breaks down pricing, key specs and arena ratings.
Which is cheaper, Claude Opus 4.8 or Kimi K3?
Kimi K3 is cheaper: it starts at $15/1M out, while Claude Opus 4.8 starts at $25/1M out.
How much do Claude Opus 4.8 and Kimi K3 cost per 1M tokens?
Claude Opus 4.8: $5/1M in per 1M input tokens, $25/1M out per 1M output tokens. Kimi K3: $3/1M in per 1M input tokens, $15/1M out per 1M output tokens.