Head-to-head

Claude Haiku 4.5 logovsKimi K3 logo

Claude Haiku 4.5 vs Kimi K3: which AI model wins in 2026?

Claude Haiku 4.5 ($5/1M out) and Kimi K3 ($15/1M out) are two of the most-used AI models in 2026. Across 5 community votes, Claude Haiku 4.5 leads with 67% approval.

Quick verdict

On Reasoning, pick Kimi K3: the arena rates it 4.5/5 against 3/5 for Claude Haiku 4.5. On budget, Claude Haiku 4.5 wins: it starts at $5/1M out versus $15/1M out for Kimi K3.

Line-by-line comparison

From
$5/1M outSingle tier: $1/1M input, $5/1M output; prompt cache reads $0.10/1M (5m writes $1.25/1M) and Batch API cuts 50% ($0.50/$2.50); no long-context surcharge (200K max).
$15/1M outOfficial Moonshot API list price for kimi-k3: $3/1M input (cache miss), $0.30/1M on cache hit, $15/1M output including reasoning trace, flat across the 1M context (no length tiers). Same $3/$15 on OpenRouter (no cache pricing exposed there). A 3x increase over Kimi K2.6's $0.95/$4. Verified against platform.kimi.ai/docs/pricing/chat-k3, 2026-07.
Provider
Anthropic
Moonshot AI
Context window
200K tokens
1M tokens
Input price
$1/1M in
$3/1M in
Output price
$5/1M out
$15/1M out
Modalities
text, vision (input); text output
text, vision, video input
Open weights
No
No
Crowd score
67%(2)
57%(3)
Arena ratings (1-5)
Reasoning
3.0
4.5
Coding
3.5
4.5
Writing
3.0
4.0
Speed
4.5
2.0
Value
4.0
3.5

Strengths and weaknesses

Claude Haiku 4.5

  • 73.3% on SWE-bench Verified, about 90% of Sonnet 4.5's agentic coding at one third of the price
  • Fast: more than 2x Sonnet 4 speed per Anthropic, with launch customers reporting 4-5x faster than Sonnet 4.5; ~92-110 output tok/s measured by Artificial Analysis
  • Devs report precise, localized code edits that avoid touching irrelevant code, better than GPT-5 mini class in early testing
  • Supports both vision input and extended thinking, rare at this price tier at launch
  • Well suited as worker model in multi-agent setups (Sonnet/Opus plans, parallel Haiku sub-agents execute)
  • Prompt caching reads at $0.10/1M and 50% Batch API discount cut effective cost further
  • $5/1M output is pricey for a small model: Gemini Flash and GPT mini tiers undercut it several-fold on output-heavy tasks
  • 200K context (vs 1M for Sonnet 5/Opus siblings) and 64K max output limit large-codebase and long-output work
  • Mediocre cross-domain reasoning: users report weak results on GPQA, MedQA, MMMU style knowledge tasks
  • Throughput varies widely in practice (82-208 tok/s reported) and quality degrades on long 7-8+ minute agentic sessions
  • Knowledge cutoff (reliable to Feb 2025) is dated by mid-2026 standards

Kimi K3

  • #3 on the Artificial Analysis Intelligence Index (57) at launch, comparable to Claude Opus 4.8 and GPT-5.5: the closest a Chinese lab has come to the closed US frontier
  • 93.4% on SWE-bench Verified in Vals AI's independent harness (GPT-5.6 Sol: 96.2%, Claude Fable 5: 95.0%) and #1 on Arena.ai's Frontend Code Arena at 1679 points, ahead of Fable 5
  • 93.5% on GPQA Diamond (between Fable 5's 92.6% and GPT-5.6's 94.1%), 96.1% on AIME 2025, and Elo 1668 on GDPval-AA v2 agentic work, second only to Fable 5
  • Strong agentic profile: #1 on AutomationBench-AA SaaS workflows (53%), long-horizon terminal and repo navigation via Kimi Code, and ~21% fewer output tokens than K2 for more intelligence
  • $3/$15 per 1M tokens (input drops to $0.30 on cache hit), flat across the full 1M context: still well under US closed-frontier pricing, with an OpenAI-compatible API and OpenRouter availability
  • Native multimodal input (text, image, video) and open weights announced for 2026-07-27 under an expected Modified-MIT-style license, positioning it as the first 3T-class open-weights frontier model
  • 3x price jump over Kimi K2.6 ($0.95/$4 to $3/$15) makes it the most expensive Chinese model ever shipped, at Claude Sonnet 5 list price: the '10x cheaper than US models' era is over (Simon Willison documented the hike)
  • Slow: 33 output tokens/s, ranked #145 of 190 models on Artificial Analysis, a poor fit for interactive use
  • Hallucination rate climbed from 39% (K2.6) to 51% on AA-Omniscience as the model now attempts answers it would previously refuse (Fable 5 sits at a comparable 54.9%)
  • Political censorship on sensitive topics in Mandarin (deflects on Xi, CCP, Tiananmen) and Beijing jurisdiction for API data, a compliance blocker for many EU and enterprise deployments; UK AISI/CAISI also found its cyber guardrails failed to block exploit-development attempts during testing
  • 'Open weights' is theoretical for most: ~594 GB in native MXFP4 requiring a multi-GPU datacenter to self-host, and the weights (plus final license) were still unpublished at review time

Cast your verdict

One recommendation per tool per gladiator. It reshapes the crowd score everyone sees.

67%crowd score · 2
Kimi K3$15/1M out
57%crowd score · 3

The arena’s verdict on Claude Haiku 4.5

Pick Haiku 4.5 if you are on the Anthropic stack and need near-Sonnet coding quality at low latency and a third of the price: it is a massive step up from Haiku 3.5 and excels as the worker model in multi-agent pipelines. It remains Anthropic's current small model as of July 2026, so it is the default cheap tier for Claude-based products. Avoid it for deep cross-domain reasoning, very large codebases (200K context cap), or pure cost-per-token shopping, where Gemini Flash and GPT mini tiers are now cheaper, and step up to Sonnet 5 when quality matters more than speed.

The arena’s verdict on Kimi K3

Kimi K3 is the strongest argument yet that the frontier is no longer exclusively American: #3 on Artificial Analysis, top-tier GPQA and SWE-bench Verified scores, and the best frontend-code arena ranking in the business, at roughly half to a third of US closed-frontier prices. Take it for agentic coding, frontend work and SaaS automation where its benchmarks are strongest, or if your roadmap depends on self-hosting a frontier-class model once the weights land. Skip it for interactive products (33 tokens/s is slow), for anything touching politically sensitive content or strict EU data-residency requirements (Mandarin-language censorship, Beijing jurisdiction), and for high-accuracy retrieval where its 51% hallucination rate on AA-Omniscience demands a verification layer. Cost-obsessed teams should note DeepSeek V4 still delivers vastly more tokens per dollar; K3's pitch is peak capability per dollar, not cheapest tokens.

What the crowd says

On Claude Haiku 4.5

Guardian of the Repo

The precise localized edits are the underrated feature. It fixes the line that needs fixing and leaves the rest alone. GPT mini class models keep rewriting half my file.

Champion of Vibes

Haiku 4.5 gives me about 90% of Sonnet agentic coding at a third of the price, and it is fast enough that edit loops feel instant. My default for quick fixes now.

On Kimi K3

Captain Churn

Tried it for our knowledge-base assistant and the hallucination rate is real: it confidently invented two API endpoints in one afternoon. 33 tokens per second on top of that. Back to waiting for the open weights to fine-tune.

Golden Thumbicus

Our Zapier-style automation stack runs on it now. Tool orchestration that GPT-5.5 fumbled works reliably, and the cache-hit pricing makes repeated workflows genuinely cheap.

Saint Deployus

The Frontend Arena ranking is deserved. Gave it the same dashboard spec I gave Fable 5 and K3's React came out cleaner, with fewer invented props. At $3/$15 through OpenRouter it was a drop-in swap.

Frequently asked questions

Is Claude Haiku 4.5 better than Kimi K3?

The crowd currently sides with Claude Haiku 4.5: 67% recommend it, versus 57% for Kimi K3 (5 votes). On Reasoning, Kimi K3 rates higher (4.5/5 vs 3/5). The right pick depends on your use case. The line-by-line comparison on this page breaks down pricing, key specs and arena ratings.

Which is cheaper, Claude Haiku 4.5 or Kimi K3?

Claude Haiku 4.5 is cheaper: it starts at $5/1M out, while Kimi K3 starts at $15/1M out.

How much do Claude Haiku 4.5 and Kimi K3 cost per 1M tokens?

Claude Haiku 4.5: $1/1M in per 1M input tokens, $5/1M out per 1M output tokens. Kimi K3: $3/1M in per 1M input tokens, $15/1M out per 1M output tokens.