Head-to-head
Kimi K3 vs DeepSeek-V4: which AI model wins in 2026?
Kimi K3 ($15/1M out) and DeepSeek-V4 ($0.87/1M out) are two of the most-used AI models in 2026. Across 6 community votes, Kimi K3 leads with 57% approval.
Quick verdict
On Reasoning, Kimi K3 and DeepSeek-V4 are tied at 4.5/5. On budget, DeepSeek-V4 wins: it starts at $0.87/1M out versus $15/1M out for Kimi K3.
Line-by-line comparison
Strengths and weaknesses
Kimi K3
- #3 on the Artificial Analysis Intelligence Index (57) at launch, comparable to Claude Opus 4.8 and GPT-5.5: the closest a Chinese lab has come to the closed US frontier
- 93.4% on SWE-bench Verified in Vals AI's independent harness (GPT-5.6 Sol: 96.2%, Claude Fable 5: 95.0%) and #1 on Arena.ai's Frontend Code Arena at 1679 points, ahead of Fable 5
- 93.5% on GPQA Diamond (between Fable 5's 92.6% and GPT-5.6's 94.1%), 96.1% on AIME 2025, and Elo 1668 on GDPval-AA v2 agentic work, second only to Fable 5
- Strong agentic profile: #1 on AutomationBench-AA SaaS workflows (53%), long-horizon terminal and repo navigation via Kimi Code, and ~21% fewer output tokens than K2 for more intelligence
- $3/$15 per 1M tokens (input drops to $0.30 on cache hit), flat across the full 1M context: still well under US closed-frontier pricing, with an OpenAI-compatible API and OpenRouter availability
- Native multimodal input (text, image, video) and open weights announced for 2026-07-27 under an expected Modified-MIT-style license, positioning it as the first 3T-class open-weights frontier model
- 3x price jump over Kimi K2.6 ($0.95/$4 to $3/$15) makes it the most expensive Chinese model ever shipped, at Claude Sonnet 5 list price: the '10x cheaper than US models' era is over (Simon Willison documented the hike)
- Slow: 33 output tokens/s, ranked #145 of 190 models on Artificial Analysis, a poor fit for interactive use
- Hallucination rate climbed from 39% (K2.6) to 51% on AA-Omniscience as the model now attempts answers it would previously refuse (Fable 5 sits at a comparable 54.9%)
- Political censorship on sensitive topics in Mandarin (deflects on Xi, CCP, Tiananmen) and Beijing jurisdiction for API data, a compliance blocker for many EU and enterprise deployments; UK AISI/CAISI also found its cyber guardrails failed to block exploit-development attempts during testing
- 'Open weights' is theoretical for most: ~594 GB in native MXFP4 requiring a multi-GPU datacenter to self-host, and the weights (plus final license) were still unpublished at review time
DeepSeek-V4
- 1M-token context window (8x the 128K of V3.2) with up to 384K output tokens, standard on the official API
- Aggressive pricing: $0.435/$0.87 per 1M tokens (V4-Pro), roughly 28.7x cheaper per output token than Claude Opus 4.8; cache-hit input drops to $0.003625/1M (over 99% discount)
- MIT-licensed open weights for both V4-Pro and V4-Flash on Hugging Face: commercial use, fine-tuning and redistribution allowed
- Open-source SOTA on agentic coding: 80.6 on SWE-bench Verified (Think Max config), tied with Gemini 3.1 Pro, plus Codeforces rating 3206 (~rank 23 vs humans)
- Ranks #3 of 93 on the Artificial Analysis Intelligence Index (score 44), well above the 25 average
- Sparse-attention stack cuts 1M-context inference to 27% of V3.2's FLOPs and 10% of its KV cache
- Intermittent malformed tool calls: function calls sometimes emitted as plain text in content instead of the tool_calls field (GitHub issue deepseek-ai #1244)
- Thinking mode breaks long multi-turn tool-call chains with 400 errors in agent frameworks (OpenClaw issue #72044, fix still incomplete)
- Developers report it fabricating nonexistent APIs in custom codebases and acting on hallucinated user input in agent loops
- Very verbose (180M eval output tokens vs 95M median) and mid-pack speed at 54.6 tok/s (#39/93), which erodes the low per-token price in practice
- Text only (no vision or audio) and still a preview: the official release planned for mid-July 2026 adds peak-hour pricing that doubles listed API rates during Beijing business hours
Cast your verdict
One recommendation per tool per gladiator. It reshapes the crowd score everyone sees.
The arena’s verdict on Kimi K3
Kimi K3 is the strongest argument yet that the frontier is no longer exclusively American: #3 on Artificial Analysis, top-tier GPQA and SWE-bench Verified scores, and the best frontend-code arena ranking in the business, at roughly half to a third of US closed-frontier prices. Take it for agentic coding, frontend work and SaaS automation where its benchmarks are strongest, or if your roadmap depends on self-hosting a frontier-class model once the weights land. Skip it for interactive products (33 tokens/s is slow), for anything touching politically sensitive content or strict EU data-residency requirements (Mandarin-language censorship, Beijing jurisdiction), and for high-accuracy retrieval where its 51% hallucination rate on AA-Omniscience demands a verification layer. Cost-obsessed teams should note DeepSeek V4 still delivers vastly more tokens per dollar; K3's pitch is peak capability per dollar, not cheapest tokens.
The arena’s verdict on DeepSeek-V4
Choose DeepSeek-V4 if you want near-frontier reasoning and agentic coding at 3x to nearly 30x below Claude Opus or GPT-5.5 pricing, or if MIT-licensed weights for self-hosting and fine-tuning matter to you. It is a decisive upgrade over V3.2: 8x longer context, far cheaper long-context inference and stronger coding, and the legacy deepseek-chat/reasoner endpoints are deprecated on July 24, 2026 anyway. Avoid it for production agents that depend on rock-solid multi-turn tool calling, where users still report malformed tool calls and fabricated APIs, and for any vision or audio work since it is text only. Latency-sensitive apps should also test first, as its verbosity and mid-pack 54.6 tok/s output speed offset some of the cost advantage, and budget for the peak-hour price doubling arriving with the official mid-July release.
What the crowd says
On Kimi K3
“Tried it for our knowledge-base assistant and the hallucination rate is real: it confidently invented two API endpoints in one afternoon. 33 tokens per second on top of that. Back to waiting for the open weights to fine-tune.”
“Our Zapier-style automation stack runs on it now. Tool orchestration that GPT-5.5 fumbled works reliably, and the cache-hit pricing makes repeated workflows genuinely cheap.”
“The Frontend Arena ranking is deserved. Gave it the same dashboard spec I gave Fable 5 and K3's React came out cleaner, with fewer invented props. At $3/$15 through OpenRouter it was a drop-in swap.”
On DeepSeek-V4
“Tool calling is flaky. Function calls sometimes land as plain text instead of the tool_calls field, and thinking mode 400s on long multi-turn chains. Not agent-ready yet.”
“MIT license on both Pro and Flash weights is the real story. Fine-tune, redistribute, ship commercially, no lawyer needed. Plus 384K output tokens for long-doc generation.”
“MIT weights, 1M context, and output tokens roughly 29x cheaper than Opus 4.8. Cache hits make input basically free. Moved my bulk pipelines over and the bill collapsed.”
Keep comparing
Frequently asked questions
Is Kimi K3 better than DeepSeek-V4?
The right pick depends on your use case. The line-by-line comparison on this page breaks down pricing, key specs and arena ratings.
Which is cheaper, Kimi K3 or DeepSeek-V4?
DeepSeek-V4 is cheaper: it starts at $0.87/1M out, while Kimi K3 starts at $15/1M out.
How much do Kimi K3 and DeepSeek-V4 cost per 1M tokens?
Kimi K3: $3/1M in per 1M input tokens, $15/1M out per 1M output tokens. DeepSeek-V4: $0.435/1M in (cache hit $0.003625) per 1M input tokens, $0.87/1M out per 1M output tokens.