Head-to-head

Llama 4 (Scout / Maverick) logovsClaude Fable 5 logo

Llama 4 (Scout / Maverick) vs Claude Fable 5: which AI model wins in 2026?

Llama 4 (Scout / Maverick) ($0.60/1M out (Maverick, hosted)) and Claude Fable 5 ($50/1M out) are two of the most-used AI models in 2026. Across 4 community votes, Claude Fable 5 leads with 63% approval.

Quick verdict

On Reasoning, pick Claude Fable 5: the arena rates it 5/5 against 2.5/5 for Llama 4 (Scout / Maverick). On budget, Llama 4 (Scout / Maverick) wins: it starts at $0.60/1M out (Maverick, hosted) versus $50/1M out for Claude Fable 5.

Line-by-line comparison

From
$0.60/1M out (Maverick, hosted)No first-party paid Meta API; typical third-party hosted rates shown for Maverick (Scout from ~$0.10/$0.30 per 1M on DeepInfra, $0.11/$0.34 on Groq). Hosted Scout context is capped well below the 10M nominal spec (e.g. ~320K on DeepInfra).
$50/1M outOfficial Anthropic API list price for claude-fable-5: $10/1M input, $50/1M output, single tier with 1M context by default (no long-context premium), 128K max output; requests refused before any output are not billed. Verified against platform.claude.com (Introducing Claude Fable 5) 2026-07.
Provider
Meta
Anthropic
Context window
10M tokens (Scout) / 1M (Maverick)
1M tokens
Input price
$0.15/1M in (Maverick, hosted)
$10/1M in
Output price
$0.60/1M out (Maverick, hosted)
$50/1M out
Modalities
text, vision (image input)
text, vision
Open weights
Yes
No
Crowd score
50%(0)
63%(4)
Arena ratings (1-5)
Reasoning
2.5
5.0
Coding
2.0
5.0
Writing
2.5
4.5
Speed
4.5
2.0
Value
3.5
3.0

Strengths and weaknesses

Llama 4 (Scout / Maverick)

  • MoE efficiency: only 17B active parameters per token (Scout 109B/16 experts, Maverick 400B/128 experts), giving near GPT-4o-class chat at a fraction of the compute
  • Very cheap hosted inference: Maverick from about $0.15/1M in and $0.60/1M out; Scout from about $0.10/$0.30 on DeepInfra or $0.11/$0.34 on Groq
  • Native early-fusion multimodality (text plus images, tested up to 8 images) in an open-weight model
  • Largest nominal context of any open-weight model at release: 10M tokens on Scout, 1M on Maverick
  • Scout fits on a single H100 GPU with Int4 quantization; pretrained on 200 languages
  • High throughput: 17B active params reach 500+ tokens/s on fast providers (Groq lists Scout at 594 TPS)
  • Benchmark trust damaged: Meta submitted an unreleased chat-optimized Maverick variant to LMArena (ELO 1417), which devs called misleading since public weights score lower
  • Long-context claims collapse in practice: Scout scored ~15.6% at 128K on Fiction.LiveBench vs 90.6% for Gemini 2.5 Pro, and hosted providers cap Scout far below 10M (e.g. ~320K on DeepInfra)
  • Coding widely panned on r/LocalLlama and HN, losing to the similarly-priced DeepSeek V3 on real dev tasks
  • Not OSI open source: license excludes EU-domiciled companies, requires a special license above 700M MAU, and mandates Llama branding
  • 109B/400B total params are too big for consumer GPUs, and the lineage stalled: no reasoning variant shipped and Behemoth was never released

Claude Fable 5

  • 80.3% on SWE-bench Pro vs 69.2% for Opus 4.8, 58.6% for GPT-5.5 and 54.2% for Gemini 3.1 Pro, roughly 11 points ahead of the next frontier model
  • 95.0% on SWE-bench Verified (Opus 4.8: 88.6%, GPT-5.5: 82.6%) and 29.3% on Cognition's FrontierCode Diamond split, more than double Opus 4.8's 13.4%
  • Long-horizon autonomy is the real story: Stripe reported a 50-million-line Ruby codebase migration done in one day instead of 2+ months, and Cursor's CEO calls it state of the art on CursorBench
  • Field reports match the benchmarks: HN engineers describe it working 'like an actual engineer' (CRDTs with minimal hand-holding, writing its own fuzzers, one 46x allocation reduction), Simon Willison measured 'several days' worth of work' in a single session
  • 1M token context window by default plus 128K output, and state-of-the-art vision on dense documents (29.8% on GDP.pdf vs 24.9% for GPT-5.5 and 22.5% for Opus 4.8)
  • Refused-before-output requests are not billed, and server-side fallback to Opus 4.8 with fallback credit is built into the API
  • Double the price of Opus 4.8 ($10/$50 vs $5/$25) and slow: single requests on hard tasks routinely run many minutes, Simon Willison bluntly calls it 'slow, expensive'
  • Dual-use safety classifiers misfire on legitimate work: a medical physicist reported fluid dynamics problems and MRI segmentation code refused as biosecurity risks, with requests silently rerouted to Opus 4.8 (the viral HN thread was titled 'If Claude Fable stops helping you, you'll never know'; Anthropic says under 5% of sessions)
  • Rocky launch: US export controls forced Anthropic to suspend access worldwide from June 12 to June 30, 2026, three days after release, with full restoration only on July 1
  • Requires 30-day data retention and is not available under zero data retention, a hard blocker for strict-compliance orgs; also no thinking-off mode, raw chain of thought never returned, assistant prefill returns a 400
  • Not universally state of the art: GPT-5.5 still leads ARC-AGI-2 (85.0% vs 77.1%), and Andon Labs found unblocked Mythos 5 underperformed both Opus 4.7 and GPT-5.5 on Vending-Bench, with reasoning that optimized for detectability rather than actual harm

Cast your verdict

One recommendation per tool per gladiator. It reshapes the crowd score everyone sees.

Llama 4 (Scout / Maverick)$0.60/1M out (Maverick, hosted)
50%crowd score · 0
Claude Fable 5$50/1M out
63%crowd score · 4

The arena’s verdict on Llama 4 (Scout / Maverick)

Pick Llama 4 if you need a cheap, fast, self-hostable multimodal model for high-volume chat, extraction, or multilingual workloads outside the EU; Maverick lands near GPT-4o quality at roughly a tenth of the price. Avoid it for coding, hard reasoning, or genuine million-token retrieval, where DeepSeek, Qwen 3, and Gemini clearly outperform it. Versus Llama 3.3 70B it adds native vision and a longer window, but many developers found the older dense model or Qwen more reliable for pure text quality, and the 10M context is mostly a paper number.

The arena’s verdict on Claude Fable 5

Take Claude Fable 5 if your workload is genuinely long-horizon: overnight agentic runs, monster migrations, tasks where one multi-hour session replaces days of supervised work. There, the 2x premium over Opus 4.8 pays for itself in task compression, and the benchmarks (80.3% SWE-bench Pro, 11 points clear of the field) are backed by real deployments at Stripe and Cursor. For interactive coding and everyday work, stay on Opus 4.8: 88.6% on SWE-bench Verified at half the price, no classifier misfires, faster turns. Cost-sensitive teams get near-Opus coding from Sonnet 5 at $3/$15 (intro $2/$10 through August 2026). Avoid Fable 5 entirely if your org requires zero data retention or if you work anywhere near biology, medical imaging or security tooling, where the dual-use classifiers still produce false positives and silently swap in Opus 4.8 mid-session.

What the crowd says

On Llama 4 (Scout / Maverick)

No verdicts yet. Be the first to speak.

On Claude Fable 5

Thumbs Downicus

I do medical imaging research and the bio classifier keeps flagging my MRI segmentation prompts, then it silently falls back to Opus 4.8 mid-session. At $50 per million output tokens I expect to at least know which model actually answered me.

Glorius Maximus

Yes it's 2x the price of Opus and yes the turns are slow. But one overnight Fable run replaced what used to be a week of supervising shorter runs. On a per-task basis it's actually the cheapest model we use.

Golden Thumbicus

The 1M context is real, not marketing. I fed it our entire service mesh config plus six months of incident postmortems and it traced a flaky timeout to a retry policy nobody remembered writing. Opus 4.8 never connected those dots.

Saint Deployus

Gave it a monorepo migration that Opus 4.8 kept stalling on. It ran for about 40 minutes, came back with the whole thing done plus a test harness it wrote for itself. Felt like reviewing a senior engineer's PR, not babysitting a chatbot.

Frequently asked questions

Is Llama 4 (Scout / Maverick) better than Claude Fable 5?

The crowd currently sides with Claude Fable 5: 63% recommend it, versus 50% for Llama 4 (Scout / Maverick) (4 votes). On Reasoning, Claude Fable 5 rates higher (5/5 vs 2.5/5). The right pick depends on your use case. The line-by-line comparison on this page breaks down pricing, key specs and arena ratings.

Which is cheaper, Llama 4 (Scout / Maverick) or Claude Fable 5?

Llama 4 (Scout / Maverick) is cheaper: it starts at $0.60/1M out (Maverick, hosted), while Claude Fable 5 starts at $50/1M out.

How much do Llama 4 (Scout / Maverick) and Claude Fable 5 cost per 1M tokens?

Llama 4 (Scout / Maverick): $0.15/1M in (Maverick, hosted) per 1M input tokens, $0.60/1M out (Maverick, hosted) per 1M output tokens. Claude Fable 5: $10/1M in per 1M input tokens, $50/1M out per 1M output tokens.