# Claude Opus 5 vs Kimi K3 (2026): side-by-side comparison Source: [GLAD-AI-TOR](https://glad-ia-tor.com) · Full page: https://glad-ia-tor.com/vs/claude-opus-5-vs-kimi-k3 Arena: llm-models · Crowd scores are live visitor verdicts (one per person per tool, never paid, Bayesian-smoothed). ## At a glance | | Claude Opus 5 | Kimi K3 | |---|---|---| | Price | $25/1M out | $15/1M out | | Crowd score | 63% (4 votes) | 57% (3 votes) | | provider | Anthropic | Moonshot AI | | contextWindow | 1M tokens | 1M tokens | | priceIn | $5/1M in | $3/1M in | | priceOut | $25/1M out | $15/1M out | | modalities | text, vision | text, vision, video input | | openWeights | no | no | | reasoning (1-5) | 5 | 4.5 | | coding (1-5) | 5 | 4.5 | | writing (1-5) | 4 | 4 | | speed (1-5) | 2 | 2 | | valueForMoney (1-5) | 4 | 3.5 | ### Claude Opus 5 > Anthropic's July 2026 Opus refresh: near-Fable 5 intelligence at half the price, same $5/$25 as Opus 4.8 Strengths: - Ranked #1 in composite intelligence across 190 models on Artificial Analysis at launch, and 43.3% on Frontier-Bench v0.1 vs 34.4% for GPT-5.6 Sol and 33.7% for Claude Fable 5 - 3.9x better than GPT-5.6 Sol on ARC-AGI-3 novel reasoning (30.2% vs 7.8%), and Elo 1861 on GDPval-AA v2 economic knowledge work, ahead of Fable 5 (1747) - 79.2% on SWE-bench Pro, within a point of Fable 5 (80.0%) and 10 points above Opus 4.8 (69.2%), at half Fable's price; Cursor's co-founder calls it 'near Fable 5 intelligence at Opus speed and cost' - Same $5/$25 pricing as Opus 4.8 with a bigger window: 1M context is now the default and only tier, with 128K max output and prompt caching from 512 tokens Weaknesses: - Notably slow and very verbose: 52.6 output tokens/s and 68 seconds to first token on Artificial Analysis, and it consumed ~100M output tokens during their eval vs a 63M median - The verbosity is a real-world cost problem: CodeRabbit measured it reading ~50% more and writing ~65% more than reference frontier models per code-review call - No actual price cut despite the 'cost-efficient' narrative: identical to Opus 4.8 ($5/$25), nearly GPT-5.6 money ($5/$30), and more than 2x comparable Gemini or Grok tiers Verdict: Claude Opus 5 is the sane default of the Series 5 range: most of Fable 5's intelligence (and more than Fable on Frontier-Bench and GDPval) at exactly half the token price, with classifiers that trigger 85% less often. If you migrated workloads to Fable 5 for capability but resent the bill or the false-positive refusals, move them here; if you are still on Opus 4.8, the upgrade is 10 SWE-bench Pro points for free. The two honest reasons to look elsewhere: latency and verbosity. At 52.6 tokens/s with 68s to first token it is a poor fit for interactive UX, and its token appetite quietly inflates real costs beyond the sticker price, so budget-sensitive high-volume pipelines still belong on Sonnet 5, Gemini or DeepSeek. Keep Fable 5 only for the longest autonomous runs where its slight SWE-bench Pro edge compounds. Full review: https://glad-ia-tor.com/tool/claude-opus-5 · Markdown: https://glad-ia-tor.com/tool/claude-opus-5.md ### Kimi K3 > Moonshot's 2.8T-parameter MoE that sparked the 'new DeepSeek moment': frontier-adjacent scores at $3/$15 Strengths: - #3 on the Artificial Analysis Intelligence Index (57) at launch, comparable to Claude Opus 4.8 and GPT-5.5: the closest a Chinese lab has come to the closed US frontier - 93.4% on SWE-bench Verified in Vals AI's independent harness (GPT-5.6 Sol: 96.2%, Claude Fable 5: 95.0%) and #1 on Arena.ai's Frontend Code Arena at 1679 points, ahead of Fable 5 - 93.5% on GPQA Diamond (between Fable 5's 92.6% and GPT-5.6's 94.1%), 96.1% on AIME 2025, and Elo 1668 on GDPval-AA v2 agentic work, second only to Fable 5 - Strong agentic profile: #1 on AutomationBench-AA SaaS workflows (53%), long-horizon terminal and repo navigation via Kimi Code, and ~21% fewer output tokens than K2 for more intelligence Weaknesses: - 3x price jump over Kimi K2.6 ($0.95/$4 to $3/$15) makes it the most expensive Chinese model ever shipped, at Claude Sonnet 5 list price: the '10x cheaper than US models' era is over (Simon Willison documented the hike) - Slow: 33 output tokens/s, ranked #145 of 190 models on Artificial Analysis, a poor fit for interactive use - Hallucination rate climbed from 39% (K2.6) to 51% on AA-Omniscience as the model now attempts answers it would previously refuse (Fable 5 sits at a comparable 54.9%) Verdict: Kimi K3 is the strongest argument yet that the frontier is no longer exclusively American: #3 on Artificial Analysis, top-tier GPQA and SWE-bench Verified scores, and the best frontend-code arena ranking in the business, at roughly half to a third of US closed-frontier prices. Take it for agentic coding, frontend work and SaaS automation where its benchmarks are strongest, or if your roadmap depends on self-hosting a frontier-class model once the weights land. Skip it for interactive products (33 tokens/s is slow), for anything touching politically sensitive content or strict EU data-residency requirements (Mandarin-language censorship, Beijing jurisdiction), and for high-accuracy retrieval where its 51% hallucination rate on AA-Omniscience demands a verification layer. Cost-obsessed teams should note DeepSeek V4 still delivers vastly more tokens per dollar; K3's pitch is peak capability per dollar, not cheapest tokens. Full review: https://glad-ia-tor.com/tool/kimi-k3 · Markdown: https://glad-ia-tor.com/tool/kimi-k3.md ## More Full llm-models ranking: https://glad-ia-tor.com/hall-of-fame/llm-models --- This markdown version exists for AI assistants; the canonical page is https://glad-ia-tor.com/vs/claude-opus-5-vs-kimi-k3