Opus 5 vs Opus 4.8: Benchmarks, Token Usage and Which to Use
Published 30 September 2026 · Last updated 30 September 2026 · Data verified 30 September 2026
Verdict: Opus 5 is clearly better than Opus 4.8 at the same price ($5 / $25 per million tokens). In Anthropic’s launch table it beats Opus 4.8 on every benchmark listed: 43.3% vs 21.1% on Frontier-Bench and 30.2% vs 1.5% on ARC-AGI-3, for example.
Independent testing (Artificial Analysis) puts Opus 5 at 51 vs 42 on the Intelligence Index, using about 18% fewer output tokens (140M vs 170M). But both are now Legacy models on Anthropic’s model pages: Opus 5.5 is current, cheaper ($4 / $20), and faster. New work should start on 5.5.
| Question | Winner | Why |
|---|---|---|
| Smarter? | Opus 5 | Leads on all 14 rows of Anthropic’s head-to-head table; AA Index 51 vs 42. |
| Better at coding? | Opus 5 | DeepSWE 68.8% vs 59.0%; Frontier-Bench 43.3% vs 21.1% (more than double). |
| Uses fewer tokens? | Opus 5 | ~18% fewer output tokens on AA’s Index (140M vs 170M); customers report up to 26% fewer. |
| Cheaper per task? | Opus 4.8 | AA measured $4.08/task for 4.8 vs $5.86 for Opus 5 — fewer tokens, pricier task mix. |
| Faster? | Opus 4.8 (barely) | 58.9 vs 54.6 tok/s on AA; but both are volatile day to day — see the speed section. |
| Which to use today? | Opus 5.5 | Current model, 20% cheaper per token, 60% cheaper cache reads. Don’t start new work on either Legacy model. |
Opus 4.8 vs 5: benchmark results
Opus 5 wins every row. Anthropic’s Opus 5 launch post (24 July 2026) puts the two models head to head on 14 benchmark rows, and Opus 5 is ahead on all of them. The gaps are biggest where agentic work is involved: ARC-AGI-3 (30.2% vs 1.5%), Frontier-Bench (43.3% vs 21.1%) and OSWorld 2.0 (70.6% vs 55.7%).
The headline number is Frontier-Bench v0.1, a test of agentic terminal coding: Opus 5 scores 43.3% against 21.1% for Opus 4.8, which Anthropic describes as “more than doubles Opus 4.8’s performance at a lower cost per task”. On OSWorld 2.0 (computer use), Opus 5’s 70.6% also beats Fable 5’s best score, at what Anthropic says is just over a third of the cost.
One caveat: Opus 5 does not beat every model on every row. Fable 5 leads it on HLE without tools (56.5% vs 56.3%), DeepSWE v1.1 (69.7% vs 68.8%), FrontierCode v1.1 (53.5% vs 53.4%) and Legal (13.3% vs 11.7%) — a different comparison, covered briefly below. The BioMysteryBench “human solved” row (90.1% vs 88.5%) is a sub-score of the same eval.
All figures below are self-reported by Anthropic and were read from the launch-post table image on 30 September 2026. For more on how we read benchmark tables, see our benchmarks hub.
| Benchmark ↕ | What it tests | Opus 5 | Opus 4.8 | Gap |
|---|---|---|---|---|
| Frontier-Bench v0.1 | Agentic terminal coding | 43.3% | 21.1% | +22.2 |
| GDPval-AA v2 | Knowledge work (Elo) | 1861 | 1593* | +268 |
| ARC-AGI-3 | Novel problem-solving | 30.2% | 1.5% | +28.7 |
| BrowseComp | Agentic search | 90.8% | 84.3% | +6.5 |
| Humanity’s Last Exam (no tools) | Reasoning | 56.3% | 49.8% | +6.5 |
| Humanity’s Last Exam (with tools) | Reasoning | 64.7% | 57.9% | +6.8 |
| OSWorld 2.0 | Computer use | 70.6% | 55.7% | +14.9 |
| DeepSWE v1.1 | Agentic coding | 68.8% | 59.0% | +9.8 |
| FrontierCode v1.1 (Main) | Agentic coding | 53.4% | 46.5% | +6.9 |
| AutomationBench | Business workflows | 26.0% | 17.0% | +9.0 |
| Legal Agent Benchmark (held-out) | Legal | 11.7% | 10.4% | +1.3 |
| HealthBench Professional | Health | 59.8% | 57.4% | +2.4 |
| BioMysteryBench (hard) | Biology | 49.4% | 42.4% | +7.0 |
| BioMysteryBench (human-solved) | Biology (sub-score) | 90.1% | 88.5% | +1.6 |
Data-integrity notes. *Opus 4.8’s GDPval-AA v2 score is 1593 in the Opus 5 post but 1615 in the Sonnet 5 post (30 Jun 2026). We use 1593 here because it comes from the same table as Opus 5. Frontier-Bench methodology: Anthropic ran it on the mini-SWE-agent harness, averaging 5 attempts per task, and Opus 4.8 served as the fallback whenever safety classifiers refused Opus 5. GDPval-AA v2 is an Elo score, not a percentage — do not compare its “gap” with percentage gaps.
For coding work specifically, the numbers that matter most are Frontier-Bench, DeepSWE and FrontierCode — Opus 5 leads on all three. If you are choosing a model mainly for code, also read our guide to the best Claude model for coding.
What independent tests show
Artificial Analysis agrees: Opus 5 is smarter, but its tasks cost more. On AA’s Intelligence Index v4.3.2 (max effort, checked 30 September 2026), Opus 5 scores 51 against 42 for Opus 4.8. The catch: Opus 5’s cost per Index task is $5.86 vs $4.08 for Opus 4.8, at the same per-token price. The token-usage section below digs into why.
Two caveats before you quote these numbers. AA marks both models deprecated and now only re-benchmarks the default workload for them. And these numbers move: Opus 5’s time-to-first-token moved by roughly 10 s in two days on the same AA page. Treat AA’s speed figures as weather, not climate.
| Metric | Opus 5 | Opus 4.8 | Opus 5.5 |
|---|---|---|---|
| AA Intelligence Index (v4.3.2) | 51 (rank #15/222) | 42 (rank #47/222) | 58 (rank #1/222) |
| Output tokens used on the Index | 140M | 170M | not published |
| Cost per Index task | $5.86 (rank #101) | $4.08 (rank #97) | $5.98 |
| Output speed | 54.6 tok/s (rank #145) | 58.9 tok/s (rank #126) | not published |
| Time to first token (includes thinking) | 65.74 s | 62.22 s | not published |
Last checked 30 September 2026. Artificial Analysis marks both Opus 5 and Opus 4.8 as deprecated: it now only benchmarks the default 10k-input-token workload for them and points users at Opus 5.5 and Opus 5 respectively. Speed and TTFT figures are volatile — Opus 5’s TTFT moved by roughly 10 s between our 28 and 30 September checks.
Opus 5 token usage vs Opus 4.8
Here is the paradox: Opus 5 uses fewer output tokens but costs more per task. On Artificial Analysis’ Intelligence Index, Opus 5 used 140M output tokens against 170M for Opus 4.8 — about 18% fewer. Yet its cost per Index task was $5.86 vs $4.08 for Opus 4.8, roughly 44% more, at the same $5/$25 per-token price. Fewer output tokens cannot explain a higher bill at identical prices, so the difference has to come from other token types: input tokens, cache writes, or cache reads.
AA’s cost breakdown shows where the money goes. Opus 5’s $5.86 is mostly answer tokens (~$3.16) plus cache writes (~$1.07); Opus 4.8’s $4.08 is mostly cache hits (~$1.60), with answer tokens at just ~$0.47. Opus 5’s tasks generated far more billable output per task, while 4.8’s leaned on cheap cached context.
One honest caveat: the per-task cost mix and the index-wide token totals measure different things, and AA now only re-benchmarks the default 10k-input-token workload for these deprecated models. So AA’s data shows higher cost per task despite fewer output tokens, but the breakdown is not public at the per-task level for your workload. The likely mechanism is mundane: longer agentic runs re-send more context each turn, and cache writes ($6.25/MTok on both models) add up fast.
Anthropic’s own customer reports point the same way — Opus 5 does more per token:
- A legal AI firm got similar quality at lower effort with 26% fewer tokens on average than Opus 4.8 at max effort.
- A trading firm reported about one-seventh of the reasoning tokens and under half the latency of Opus 4.8.
- A financial-modelling eval came out 9 points more accurate across effort levels, with a third fewer turns and tool calls and 60% less time.
The practical takeaway: Opus 5 is the more token-efficient model, but your bill depends on how much context your agent re-sends each turn. Measure with your own usage logs, use prompt caching — cache reads cost $0.50/MTok on both models, and the minimum cacheable prompt on Opus 5 is just 512 tokens — and try a lower effort level before assuming max is needed. Full price list: our Claude pricing plans guide.
| Model | Answer | Reasoning | Cache Write | Cache Hit | Input | Total |
|---|---|---|---|---|---|---|
| Opus 5 | ~$3.16 | ~$0.76 | ~$1.07 | ~$0.74 | ~$0.13 | $5.86 |
| Opus 4.8 | $0.47 | $1.30 | $0.58 | $1.60 | $0.13 | $4.08 |
Opus 5 vs Opus 4.8 token usage and cost calculator
Plug in your own workload. Prices are Anthropic’s list prices from 30 September 2026 — no estimates. The static example below (30K input tokens, 5K output, 60% cached, 300 tasks/month) is correct as printed; the inputs make it live.
| Model | Price (in/out per MTok) | Monthly cost |
|---|---|---|
| Opus 4.8 | $5 / $25 | $58.20 |
| Opus 5 | $5 / $25 | $51.45 |
| Opus 5.5 | $4 / $20 | $45.48 |
Speed: which answers faster?
Opus 4.8 is slightly faster on paper, but neither is fast and the numbers move daily. Artificial Analysis measured 58.9 tok/s for Opus 4.8 vs 54.6 tok/s for Opus 5, with time-to-first-token of 62.22 s vs 65.74 s (thinking time included). That is a ~8% speed edge for 4.8 — real, but small next to the 2× benchmark gaps.
Do not plan capacity around these figures. Opus 5’s TTFT moved by roughly 10 s in two days on the same AA page, and both models are now deprecated there, so the numbers will only get staler. What matters more than the model choice is the effort setting: both default to high effort, which is where the 60-second waits come from. If latency hurts, drop to medium effort before switching models — customers in Anthropic’s launch post reported under half the latency at lower effort with similar quality. Note that Anthropic claims Opus 5.5’s output is more than 30% faster than Opus 5’s, which is another reason new work belongs on 5.5.
Opus 4.5 vs Opus 4.8
Same price, much bigger envelope. Opus 4.8 costs exactly what Opus 4.5 costs — $5/$25 per million tokens — but gives you 5× the context window (1M vs 200K), 2× the max output (128K vs 64K), adaptive thinking instead of extended-only, a knowledge cutoff eight months newer, and fast mode. There is no benchmark table that puts both models side by side at the same benchmark version, so this is a specs-and-price comparison, not a scores one.
| Spec | Opus 4.5 | Opus 4.8 | Opus 5 | Opus 5.5 |
|---|---|---|---|---|
| API ID | claude-opus-4-5-20251101 | claude-opus-4-8 | claude-opus-5 | claude-opus-5-5 |
| Released | 24 Nov 2025 | 28 May 2026 | 24 Jul 2026 | 22 Sep 2026 |
| Model-page label* | Legacy | Legacy | Legacy | Latest |
| Input / output per MTok | $5 / $25 | $5 / $25 | $5 / $25 | $4 / $20 |
| Cache read per MTok | $0.50 | $0.50 | $0.50 | $0.20 |
| Batch (50% off) | $2.50 / $12.50 | $2.50 / $12.50 | $2.50 / $12.50 | $2 / $10 |
| Fast mode | — | $10 / $50 | $10 / $50 | $8 / $40 |
| Context window | 200K | 1M | 1M | 1M |
| Max output | 64K | 128K | 128K (300K on Batch, beta) | 128K (300K on Batch, beta) |
| Thinking | Extended | Adaptive | Adaptive (on by default) | Adaptive (always on) |
| Default effort | high | high | high | medium |
| Reliable knowledge cutoff | May 2025 | Jan 2026 | May 2026 | Jun 2026 |
| Earliest retirement | 24 Nov 2026 | 28 May 2027 | 24 Jul 2027 | 22 Sep 2027 |
*”Legacy” is Anthropic’s own label on each model’s docs page. On the deprecations page all four models still show lifecycle status Active (still served) with the tentative retirement dates above — “Legacy” means superseded, not shut off.
Should you use Opus 5, Opus 4.8 or Opus 5.5 today?
Start new work on Opus 5.5; stay put if you are pinned to Opus 5. Opus 5.5 is the current model, costs 20% less per token, charges 60% less for cache reads, and Anthropic says its output is over 30% faster than Opus 5’s. There is no technical reason to start a new project on a Legacy model. If your prompts and evals are pinned to Opus 5, that is fine too — it stays served until at least 24 July 2027. Anthropic’s current lineup also includes Claude Sonnet 5.5 at $2/$10 per million tokens for lighter workloads.
| Your situation | Pick | Why |
|---|---|---|
| New project, or free to switch | Opus 5.5 | Current model; cheapest per token ($4/$20); cheapest cache reads ($0.20). |
| Pinned to Opus 5 (prompts, evals) | Opus 5 | Better than 4.8 on every Table A row; served until at least 24 Jul 2027. |
| Still on Opus 4.8 | Opus 5.5 | Opus 5 beats 4.8 everywhere at the same price — and 5.5 undercuts both. |
| Still on Opus 4.5 | Migrate now | Retirement not sooner than 24 Nov 2026; 5× context and 2× output await on 4.8+. |
Choosing between model families rather than Opus generations? See our Claude Sonnet vs Opus comparison for when the cheaper Sonnet tier is enough.
Fable 5 vs Opus 4.8 (quick answer)
Mostly Fable 5, at twice the price. Where Anthropic’s launch table includes a Fable 5 column, Fable 5 beats Opus 4.8 on every row shown: Frontier-Bench 33.7% vs 21.1%, GDPval-AA v2 1747 vs 1593, BrowseComp 87.4% vs 84.3%, OSWorld 2.0 66.1% vs 55.7%, HLE 56.5%/63.9% (no tools / with tools) vs 49.8%/57.9%, DeepSWE 69.7% vs 59.0%, FrontierCode 53.5% vs 46.5% and Legal 13.3% vs 10.4%. Against Opus 5, the HLE picture is split: Fable 5 leads only on HLE without tools (56.5% vs 56.3%); with tools, Opus 5 leads 64.7% vs 63.9%. Fable 5 is $10/$50 per million tokens — double Opus pricing. For agentic coding it is a credible alternative to Opus 4.8; for everything else, Opus 5 or 5.5 win on value. We will publish a full Fable 5 comparison separately.
Label correction: the HealthBench row’s 66.0% in the launch-post table is labeled Mythos 5, not Fable 5 — we do not present it as a Fable 5 score. Fable 5’s own model page confirms its $10/$50 pricing (checked 30 Sep 2026).
How to switch models
Switching is a one-line change in each surface:
- API: change the
modelparameter toclaude-opus-5-5,claude-opus-5orclaude-opus-4-8. Note the dateless IDs — only pre-4.6 generations like Opus 4.5 use dated snapshot IDs (claude-opus-4-5-20251101). - Claude Code: run
/modeland pick from the list. If you have not set it up yet, see how to install Claude Code. - claude.ai: use the model picker. Heads-up: on claude.ai, Claude Code and Cowork, requests that Opus 5’s cyber safety classifiers flag fall back to Opus 4.8 by default — so an occasional 4.8-flavoured answer while on Opus 5 is the safety fallback, not a bug. On the API, automatic fallbacks are a beta option.
Switching models is also the standard fix when one model is having a bad day: if the API keeps timing out, see Claude API error 529 (overloaded) before retrying.
How we gathered the data
Three kinds of sources, kept separate. (1) Official: Anthropic’s Opus 5 launch post for Table A; platform.claude.com docs for every spec and price in Table C. (2) Independent: Artificial Analysis’ Intelligence Index and cost breakdown for Table B and Charts 2–4. (3) Practical: documented customer reports from the launch post — paraphrased and labelled as customer claims.
Self-reported and independent data never share a chart. Every number carries its source and a “checked 30 September 2026” note. Where sources conflicted we showed both: Opus 4.8’s GDPval-AA v2 is 1593 in the Opus 5 post and 1615 in the Sonnet 5 post; we use 1593 with a footnote because it comes from the same table as Opus 5.
Limitations. We did not run our own API tests for this page, so there is no “our test” section — nothing here is estimated or invented. AA’s speed and cost figures move between checks and both legacy models are now deprecated there, so treat them as dated snapshots. Benchmark versions were never mixed: GDPval-AA v2 and v2.1 are different versions and are not compared.
Frequently asked questions
Is Opus 5 better than Opus 4.8?
Yes. In this opus 5 vs 4.8 comparison, Anthropic’s launch table shows Opus 5 beating Opus 4.8 on all 14 benchmark rows, including 43.3% vs 21.1% on Frontier-Bench and 30.2% vs 1.5% on ARC-AGI-3. Independent testing by Artificial Analysis scores Opus 5 at 51 vs 42 on the Intelligence Index. Both cost the same per token ($5/$25 per million).
Does Opus 5 cost more than Opus 4.8?
Per token, no: both are $5 per million input and $25 per million output tokens. Per task, it depends. Artificial Analysis measured $5.86 per Index task for Opus 5 vs $4.08 for Opus 4.8, because Opus 5’s tasks used more answer tokens and cache writes. With prompt caching and medium effort, Opus 5 can cost less per completed task.
Does Opus 5 use fewer tokens?
On Artificial Analysis’ Intelligence Index, Opus 5 used 140M output tokens vs 170M for Opus 4.8, about 18% fewer. Anthropic customers reported similar savings: one legal AI firm saw 26% fewer tokens at equal quality, and a trading firm reported roughly one-seventh the reasoning tokens. Your mileage depends on how much context your agent re-sends each turn.
Why did Claude answer with Opus 4.8 when I picked Opus 5?
Opus 4.8 is the silent fallback. On claude.ai, Claude Code and Cowork, requests that Opus 5’s cyber safety classifiers flag fall back to Opus 4.8 by default. On the API, automatic fallbacks are a beta option you can switch on or off. If an answer feels like 4.8 while you picked Opus 5, the safety classifier probably rerouted your request.
Is Opus 4.5 being retired?
Not yet, but it is first in line. Anthropic’s deprecations page lists Opus 4.5 as still served with retirement not sooner than 24 November 2026, under two months away. Its model page already carries Anthropic’s Legacy label. Anyone still on Opus 4.5 should plan a migration now using Anthropic’s migration guide.
Opus 4.5 vs 4.8: what changed?
Same price ($5/$25 per million tokens), much bigger envelope. Opus 4.8 has a 1M-token context window vs 200K on Opus 4.5, 128K max output vs 64K, adaptive thinking instead of extended-only, and a knowledge cutoff eight months newer (Jan 2026 vs May 2025). Opus 4.8 also supports fast mode ($10/$50); Opus 4.5 does not.
Is Opus 5 still the newest Opus?
No. Claude Opus 5.5 launched on 22 September 2026 and is now the current model: $4/$20 per million tokens (20% cheaper), cache reads at $0.20 instead of $0.50, and over 30% faster output than Opus 5’s according to Anthropic. Opus 5, 4.8 and 4.5 all carry Anthropic’s Legacy label on their model pages, though all three are still served.
Is Fable 5 better than Opus 4.8?
On the benchmarks Anthropic published, mostly yes but not everywhere. Fable 5 beats Opus 4.8 on Frontier-Bench (33.7% vs 21.1%), GDPval-AA v2 (1747 vs 1593), DeepSWE (69.7% vs 59.0%) and Legal (13.3% vs 10.4%). It costs twice as much per token ($10/$50). For agentic coding it is a credible alternative; for most work Opus 5 or 5.5 win on value.
Sources
- Anthropic, “Claude Opus 5” launch post — anthropic.com/news/claude-opus-5 — checked 30 Sep 2026 (Table A, customer claims, safety score, fallback behaviour, beta API features)
- Anthropic, “Claude Opus 5.5” announcement — anthropic.com/news/claude-opus-5-5 — checked 30 Sep 2026 (5.5 release date, pricing, speed claim)
- Anthropic, “Claude Opus 4.8” launch post — anthropic.com/news/claude-opus-4-8 — checked 30 Sep 2026 (4.8 release date, API ID, fast mode)
- Artificial Analysis, Claude Opus 5 model page — artificialanalysis.ai/models/claude-opus-5 — checked 30 Sep 2026 (Table B, Charts 2–4)
- Artificial Analysis, Claude Opus 4.8 model page — artificialanalysis.ai/models/claude-opus-4-8 — checked 30 Sep 2026 (Table B, Charts 2–4)
- Anthropic platform docs, Pricing — platform.claude.com/docs/en/about-claude/pricing — checked 30 Sep 2026 (Table C prices)
- Anthropic platform docs, Model deprecations — platform.claude.com/docs/en/about-claude/model-deprecations — checked 30 Sep 2026 (retirement dates, API IDs)
- Anthropic platform docs, Opus 4.5 / 4.8 / 5 / 5.5 model overviews — platform.claude.com/docs/en/models/ — checked 30 Sep 2026 (specs, Legacy labels, cutoffs)
- Anthropic platform docs, Fable 5 overview — platform.claude.com/docs/en/models/fable-5/overview — checked 30 Sep 2026 (Fable 5 pricing, specs)
- Anthropic platform docs, Prompt caching — platform.claude.com/docs/en/build-with-claude/prompt-caching — checked 30 Sep 2026 (512-token minimum)
- Anthropic platform docs, Migration guide — platform.claude.com/docs/en/about-claude/models/migration-guide — checked 30 Sep 2026