# Kimi K3 vs Claude Sonnet 5 (2026): side-by-side comparison Source: [GLAD-AI-TOR](https://glad-ia-tor.com) · Full page: https://glad-ia-tor.com/vs/kimi-k3-vs-claude-sonnet-5 Arena: llm-models · Crowd scores are live visitor verdicts (one per person per tool, never paid, Bayesian-smoothed). ## At a glance | | Kimi K3 | Claude Sonnet 5 | |---|---|---| | Price | $15/1M out | $15/1M out ($10 intro until 2026-08-31) | | Crowd score | 57% (3 votes) | 57% (3 votes) | | provider | Moonshot AI | Anthropic | | contextWindow | 1M tokens | 1M tokens (128K max output) | | priceIn | $3/1M in | $3/1M in ($2 intro until 2026-08-31) | | priceOut | $15/1M out | $15/1M out ($10 intro until 2026-08-31) | | modalities | text, vision, video input | text, vision (image input, text output) | | openWeights | no | no | | reasoning (1-5) | 4.5 | 4.5 | | coding (1-5) | 4.5 | 4.5 | | writing (1-5) | 4 | 4.5 | | speed (1-5) | 2 | 4 | | valueForMoney (1-5) | 3.5 | 4.5 | ### Kimi K3 > Moonshot's 2.8T-parameter MoE that sparked the 'new DeepSeek moment': frontier-adjacent scores at $3/$15 Strengths: - #3 on the Artificial Analysis Intelligence Index (57) at launch, comparable to Claude Opus 4.8 and GPT-5.5: the closest a Chinese lab has come to the closed US frontier - 93.4% on SWE-bench Verified in Vals AI's independent harness (GPT-5.6 Sol: 96.2%, Claude Fable 5: 95.0%) and #1 on Arena.ai's Frontend Code Arena at 1679 points, ahead of Fable 5 - 93.5% on GPQA Diamond (between Fable 5's 92.6% and GPT-5.6's 94.1%), 96.1% on AIME 2025, and Elo 1668 on GDPval-AA v2 agentic work, second only to Fable 5 - Strong agentic profile: #1 on AutomationBench-AA SaaS workflows (53%), long-horizon terminal and repo navigation via Kimi Code, and ~21% fewer output tokens than K2 for more intelligence Weaknesses: - 3x price jump over Kimi K2.6 ($0.95/$4 to $3/$15) makes it the most expensive Chinese model ever shipped, at Claude Sonnet 5 list price: the '10x cheaper than US models' era is over (Simon Willison documented the hike) - Slow: 33 output tokens/s, ranked #145 of 190 models on Artificial Analysis, a poor fit for interactive use - Hallucination rate climbed from 39% (K2.6) to 51% on AA-Omniscience as the model now attempts answers it would previously refuse (Fable 5 sits at a comparable 54.9%) Verdict: Kimi K3 is the strongest argument yet that the frontier is no longer exclusively American: #3 on Artificial Analysis, top-tier GPQA and SWE-bench Verified scores, and the best frontend-code arena ranking in the business, at roughly half to a third of US closed-frontier prices. Take it for agentic coding, frontend work and SaaS automation where its benchmarks are strongest, or if your roadmap depends on self-hosting a frontier-class model once the weights land. Skip it for interactive products (33 tokens/s is slow), for anything touching politically sensitive content or strict EU data-residency requirements (Mandarin-language censorship, Beijing jurisdiction), and for high-accuracy retrieval where its 51% hallucination rate on AA-Omniscience demands a verification layer. Cost-obsessed teams should note DeepSeek V4 still delivers vastly more tokens per dollar; K3's pitch is peak capability per dollar, not cheapest tokens. Full review: https://glad-ia-tor.com/tool/kimi-k3 · Markdown: https://glad-ia-tor.com/tool/kimi-k3.md ### Claude Sonnet 5 > Anthropic's most agentic Sonnet: near Opus 4.8 quality on coding and agents at $3/$15 with 1M context Strengths: - Large agentic gains over Sonnet 4.6: Terminal-Bench 2.1 80.4% vs 67.0%, OSWorld-Verified 81.2% vs 78.5%, SWE-bench Pro 63.2% vs 58.1% - Matches Opus 4.8 on knowledge work (GDPval-AA v2: 1,618 vs 1,615) and nearly ties it on Humanity's Last Exam with tools (57.4% vs 57.9%) at 60% of Opus 4.8 pricing (40% during the intro window) - 1M token context window and 128K max output; introductory pricing of $2/$10 per 1M tokens through Aug 31, 2026 - Persistent self-verifying agent behavior: hands-on reviews note it tests its own code and iterates on hard problems until solved, unlike Sonnet 4.6 Weaknesses: - New tokenizer inflates token counts roughly 30% for the same text (1.0-1.35x per Anthropic; ~1.4x English, ~1.28x Python measured by Simon Willison), raising effective cost despite the unchanged sticker price - Verbose and token-hungry: ~$2.29 per task vs ~$1.20 for Sonnet 4.6 in independent tests (ranked 101st of 161 for cost efficiency); at high effort cost-per-task can exceed Opus 4.8 - Measurably slower than Sonnet 4.6 on small routine edits and prone to over-engineering simple tasks (CodeRabbit hands-on review) Verdict: Choose Sonnet 5 if you run coding, terminal or computer-use agents and want near Opus 4.8 quality at Sonnet prices, especially during the $2/$10 intro window; it is a strict upgrade over Sonnet 4.6 at low and medium effort. Budget for the new tokenizer and its verbosity: real per-task costs run well above Sonnet 4.6, and at the highest effort levels Opus 4.8 can be the better deal per solved task. Avoid it for latency-sensitive small edits or pipelines that rely on temperature and top_p, which now error. Sonnet 4.6 remains the pragmatic pick for high-volume tiny-diff workloads. Full review: https://glad-ia-tor.com/tool/claude-sonnet-5 · Markdown: https://glad-ia-tor.com/tool/claude-sonnet-5.md ## More Full llm-models ranking: https://glad-ia-tor.com/hall-of-fame/llm-models --- This markdown version exists for AI assistants; the canonical page is https://glad-ia-tor.com/vs/kimi-k3-vs-claude-sonnet-5