LIVE TELEMETRY • FRONTIER LLM BENCHMARKS

AI Coding Benchmarks & Rankings Leaderboard

Empirical evaluations measuring coding proficiency, reasoning benchmarks (GPQA, SWE-bench Verified, HLE), and real-time value-for-money ($/point) metrics.

Highlights

Updated Daily • Empirical Telemetry

Artificial Analysis Intelligence Index

↗

Intelligence Index v4.3 incorporates evaluations: SWE-bench Verified (40%), Terminal-Bench 2.0 (25%), Aider Polyglot (20%), LiveCodeBench (15%). Every figure is verified against primary source logs without prompt leakage.

Filters:
Anthropic OpenAI Google DeepSeek Alibaba

Intelligence Index vs. Cost per 1M Tokens

Code Intelligence Index vs. Weighted average cost (USD, Log Scale)

Most attractive quadrant Pareto frontier

Frontier Language Model Intelligence, Over Time

Progressive evolution across Anthropic, OpenAI, DeepSeek, Google, and Meta from 2023 to 2026.

Canonical Benchmark Suite

Breakdown across isolated execution harnesses matching Artificial Analysis methodologies.

Resolved Rate %
CLI & SHELL USE

Terminal-Bench 2.0 →

Task Success %
POLYGLOT EDITING

Aider Polyglot →

Pass Rate %
ALGORITHMS & MATH

LiveCodeBench →

Pass@1 %
PhD REASONING

GPQA Diamond →

Accuracy %
SCIENTIFIC CODE

SciCode →

Problem Solving %
Leads on Reasoning GPT-6 Astra (max) HLE 0.0%
Best Value — $0.0000/pt
Fastest Output Gemini 3.5 Flash-Lite 358 t/s
Longest Context Kimi K3 (max) 1.1M tokens
Best Open-Weights — Stats 0.0%
Customize Weights:

Value-for-Money Intelligence vs. Cost Matrix

BEST VALUE ZONE 0% 25% 50% 75% 100% STATS SCORE (%) Free $7.50 $15.00 $22.50 $30.00+ BLENDED COST / 1M TOKENS (USD) Proprietary API Open Weights GPT-6 Astra (max) Stats Score: 53.2% Blended Cost: $12.0000 Claude Fable 5.1 (max with fallback) Stats Score: 53.1% Blended Cost: $30.0000 Claude Opus 5 (max) Stats Score: 51.0% Blended Cost: $20.0000 Muse Spark 1.3 (max) Stats Score: 48.0% Blended Cost: $2.2500 GPT-5.6 Sol (max) Stats Score: 47.0% Blended Cost: $6.1200 GLM-5.3 (max) Stats Score: 45.1% Blended Cost: $1.5000 Qwen3.8 Max (0902) Stats Score: 45.0% Blended Cost: $3.0000 Step 5 Preview Stats Score: 44.1% Blended Cost: $1.2000 Grok 4.6 (high) Stats Score: 44.0% Blended Cost: $4.0000 Kimi K3 (max) Stats Score: 44.0% Blended Cost: $1.8000 GPT-5.6 Terra (max) Stats Score: 42.1% Blended Cost: $3.1500 GLM-5.3-Flash Stats Score: 42.1% Blended Cost: $0.3500 Gemini 3.8 Flash (high) Stats Score: 41.0% Blended Cost: $1.3100 DeepSeek V4.1 Flash (max) Stats Score: 40.0% Blended Cost: $0.3700 Qwen3.8-Flash-Next Stats Score: 40.0% Blended Cost: $0.4500 Qwen3.8 2.4T A95B Stats Score: 40.0% Blended Cost: $2.2500 Claude Sonnet 5 (max) Stats Score: 38.1% Blended Cost: $8.0000 GPT-5.6 Luna (max) Stats Score: 38.0% Blended Cost: $0.5200 DeepSeek V4 Pro 0813 (max) Stats Score: 36.0% Blended Cost: $0.9600 DeepSeek V4 Flash Vision (max) Stats Score: 35.0% Blended Cost: $0.4700 Qwen3.8 27B (xhigh) Stats Score: 34.0% Blended Cost: $0.9000 K2 Horizon 375B A23B Stats Score: 31.0% Blended Cost: $1.2000 MiniMax-M3 Stats Score: 30.0% Blended Cost: $0.6000 Gemini 3.1 Pro Preview Stats Score: 29.9% Blended Cost: $1.4000 Inkling Stats Score: 26.1% Blended Cost: $0.7500 Nemotron 3 Ultra Stats Score: 23.0% Blended Cost: $0.7500 Gemini 3.5 Flash-Lite Stats Score: 23.0% Blended Cost: $0.1700 Muse Glimmer (high) Stats Score: 18.1% Blended Cost: $0.1200 Mistral Medium 3.5 Stats Score: 15.1% Blended Cost: $0.6000 gpt-oss-120b (high) Stats Score: 12.0% Blended Cost: $0.1800
RANK MODEL & CREATOR
01 GPT-6 Astra (max) OpenAI • API • ▲ +29
53.2%*
—% —% 1M 57 t/s TTFT: 283660ms $12.0000 —
02 Claude Fable 5.1 (max with fallback) Anthropic • API • ▼ -1
53.1%*
—% —% 1M 68 t/s TTFT: 262210ms $30.0000 —
03 Claude Opus 5 (max) Anthropic • API • ▲ +27
51.0%*
—% —% 1M 51 t/s TTFT: 57030ms $20.0000 —
04 Muse Spark 1.3 (max) Meta • API • ▲ +26
48.0%*
—% —% 1M 214 t/s TTFT: 24160ms $2.2500 —
05 GPT-5.6 Sol (max) OpenAI • API • ▲ +25
47.0%*
—% —% 1M 61 t/s TTFT: 129500ms $6.1200 —
06 GLM-5.3 (max) Z AI • API • ▲ +24
45.1%*
—% —% 1M 72 t/s TTFT: 2960ms $1.5000 —
07 Qwen3.8 Max (0902) Alibaba • API • ▼ -1
45.0%*
—% —% 984K 37 t/s TTFT: 2720ms $3.0000 —
08 Step 5 Preview StepFun • API • ▲ +22
44.1%*
—% —% 1M 100 t/s TTFT: 2960ms $1.2000 —
09 Grok 4.6 (high) SpaceXAI • API • ▼ -1
44.0%*
—% —% 500K 60 t/s TTFT: 44250ms $4.0000 —
10 Kimi K3 (max) Kimi • API • ▼ -1
44.0%*
—% —% 1.05M 38 t/s TTFT: 3990ms $1.8000 —
11 GPT-5.6 Terra (max) OpenAI • API • ▲ +19
42.1%*
—% —% 1M 84 t/s TTFT: 220860ms $3.1500 —
12 GLM-5.3-Flash Z AI • API • ▲ +18
42.1%*
—% —% 1M 95 t/s TTFT: 2500ms $0.3500 —
13 Gemini 3.8 Flash (high) Google • API • ▲ +17
41.0%*
—% —% 1M 298 t/s TTFT: 15470ms $1.3100 —
14 DeepSeek V4.1 Flash (max) DeepSeek • API • ▲ +16
40.0%*
—% —% 1M 208 t/s TTFT: 1140ms $0.3700 —
15 Qwen3.8-Flash-Next Alibaba • API • ▼ -1
40.0%*
—% —% 256K 55 t/s TTFT: 2890ms $0.4500 —
16 Qwen3.8 2.4T A95B Alibaba • API • ▼ -1
40.0%*
—% —% 984K 38 t/s TTFT: 2690ms $2.2500 —
17 Claude Sonnet 5 (max) Anthropic • API • ▲ +13
38.1%*
—% —% 1M 74 t/s TTFT: 133400ms $8.0000 —
18 GPT-5.6 Luna (max) OpenAI • API • ▼ -1
38.0%*
—% —% 1M 133 t/s TTFT: 113240ms $0.5200 —
19 DeepSeek V4 Pro 0813 (max) DeepSeek • API • ▲ +11
36.0%*
—% —% 1M 88 t/s TTFT: 1700ms $0.9600 —
20 DeepSeek V4 Flash Vision (max) DeepSeek • API • ▲ +10
35.0%*
—% —% 1M 214 t/s TTFT: 1050ms $0.4700 —
21 Qwen3.8 27B (xhigh) Alibaba • API • ▲ +9
34.0%*
—% —% 256K 43 t/s TTFT: 3840ms $0.9000 —
22 K2 Horizon 375B A23B Institute of Foundation Models • API • ▲ +8
31.0%*
—% —% 524K 80 t/s TTFT: 3500ms $1.2000 —
23 MiniMax-M3 MiniMax • API • ▲ +7
30.0%*
—% —% 1M 186 t/s TTFT: 1000ms $0.6000 —
24 Gemini 3.1 Pro Preview Google • API • ▲ +6
29.9%*
—% —% 1M 119 t/s TTFT: 33160ms $1.4000 —
25 Inkling Thinking Machines • API • ▲ +5
26.1%*
—% —% 1M 81 t/s TTFT: 2480ms $0.7500 —
26 Nemotron 3 Ultra NVIDIA • API • ▲ +4
23.0%*
—% —% 262K 164 t/s TTFT: 2400ms $0.7500 —
27 Gemini 3.5 Flash-Lite Google • API • ▲ +3
23.0%*
—% —% 1M 358 t/s TTFT: 9780ms $0.1700 —
28 Muse Glimmer (high) Meta • API • ▲ +2
18.1%*
—% —% 131K 93 t/s TTFT: 990ms $0.1200 —
29 Mistral Medium 3.5 Mistral • API • ▲ +1
15.1%*
—% —% 256K 137 t/s TTFT: 2290ms $0.6000 —
30 gpt-oss-120b (high) OpenAI • API • ▼ -1
12.0%*
—% —% 131K 181 t/s TTFT: 850ms $0.1800 —
* Measured in-house using Vibecoder Journal's test harness † Vendor-reported or compiled from public benchmarks
STATS SCORE: 53.2%
GPQA: —%
SWE-BENCH: —%
CONTEXT: 1M
SPEED: 57 t/s (283660ms)
BLENDED $/M: $12.0000
$ / POINT: —
OpenAI • API • ▲ +29
STATS SCORE: 53.1%
GPQA: —%
SWE-BENCH: —%
CONTEXT: 1M
SPEED: 68 t/s (262210ms)
BLENDED $/M: $30.0000
$ / POINT: —
Anthropic • API • ▼ -1
STATS SCORE: 51.0%
GPQA: —%
SWE-BENCH: —%
CONTEXT: 1M
SPEED: 51 t/s (57030ms)
BLENDED $/M: $20.0000
$ / POINT: —
Anthropic • API • ▲ +27
STATS SCORE: 48.0%
GPQA: —%
SWE-BENCH: —%
CONTEXT: 1M
SPEED: 214 t/s (24160ms)
BLENDED $/M: $2.2500
$ / POINT: —
Meta • API • ▲ +26
STATS SCORE: 47.0%
GPQA: —%
SWE-BENCH: —%
CONTEXT: 1M
SPEED: 61 t/s (129500ms)
BLENDED $/M: $6.1200
$ / POINT: —
OpenAI • API • ▲ +25
4.5
STATS SCORE: 45.1%
GPQA: —%
SWE-BENCH: —%
CONTEXT: 1M
SPEED: 72 t/s (2960ms)
BLENDED $/M: $1.5000
$ / POINT: —
Z AI • API • ▲ +24
STATS SCORE: 45.0%
GPQA: —%
SWE-BENCH: —%
CONTEXT: 984K
SPEED: 37 t/s (2720ms)
BLENDED $/M: $3.0000
$ / POINT: —
Alibaba • API • ▼ -1
4.4
STATS SCORE: 44.1%
GPQA: —%
SWE-BENCH: —%
CONTEXT: 1M
SPEED: 100 t/s (2960ms)
BLENDED $/M: $1.2000
$ / POINT: —
StepFun • API • ▲ +22
4.4
STATS SCORE: 44.0%
GPQA: —%
SWE-BENCH: —%
CONTEXT: 500K
SPEED: 60 t/s (44250ms)
BLENDED $/M: $4.0000
$ / POINT: —
SpaceXAI • API • ▼ -1
4.4
STATS SCORE: 44.0%
GPQA: —%
SWE-BENCH: —%
CONTEXT: 1.05M
SPEED: 38 t/s (3990ms)
BLENDED $/M: $1.8000
$ / POINT: —
Kimi • API • ▼ -1
STATS SCORE: 42.1%
GPQA: —%
SWE-BENCH: —%
CONTEXT: 1M
SPEED: 84 t/s (220860ms)
BLENDED $/M: $3.1500
$ / POINT: —
OpenAI • API • ▲ +19
4.2
STATS SCORE: 42.1%
GPQA: —%
SWE-BENCH: —%
CONTEXT: 1M
SPEED: 95 t/s (2500ms)
BLENDED $/M: $0.3500
$ / POINT: —
Z AI • API • ▲ +18
STATS SCORE: 41.0%
GPQA: —%
SWE-BENCH: —%
CONTEXT: 1M
SPEED: 298 t/s (15470ms)
BLENDED $/M: $1.3100
$ / POINT: —
Google • API • ▲ +17
STATS SCORE: 40.0%
GPQA: —%
SWE-BENCH: —%
CONTEXT: 1M
SPEED: 208 t/s (1140ms)
BLENDED $/M: $0.3700
$ / POINT: —
DeepSeek • API • ▲ +16
STATS SCORE: 40.0%
GPQA: —%
SWE-BENCH: —%
CONTEXT: 256K
SPEED: 55 t/s (2890ms)
BLENDED $/M: $0.4500
$ / POINT: —
Alibaba • API • ▼ -1
STATS SCORE: 40.0%
GPQA: —%
SWE-BENCH: —%
CONTEXT: 984K
SPEED: 38 t/s (2690ms)
BLENDED $/M: $2.2500
$ / POINT: —
Alibaba • API • ▼ -1
STATS SCORE: 38.1%
GPQA: —%
SWE-BENCH: —%
CONTEXT: 1M
SPEED: 74 t/s (133400ms)
BLENDED $/M: $8.0000
$ / POINT: —
Anthropic • API • ▲ +13
STATS SCORE: 38.0%
GPQA: —%
SWE-BENCH: —%
CONTEXT: 1M
SPEED: 133 t/s (113240ms)
BLENDED $/M: $0.5200
$ / POINT: —
OpenAI • API • ▼ -1
STATS SCORE: 36.0%
GPQA: —%
SWE-BENCH: —%
CONTEXT: 1M
SPEED: 88 t/s (1700ms)
BLENDED $/M: $0.9600
$ / POINT: —
DeepSeek • API • ▲ +11
STATS SCORE: 35.0%
GPQA: —%
SWE-BENCH: —%
CONTEXT: 1M
SPEED: 214 t/s (1050ms)
BLENDED $/M: $0.4700
$ / POINT: —
DeepSeek • API • ▲ +10
STATS SCORE: 34.0%
GPQA: —%
SWE-BENCH: —%
CONTEXT: 256K
SPEED: 43 t/s (3840ms)
BLENDED $/M: $0.9000
$ / POINT: —
Alibaba • API • ▲ +9
STATS SCORE: 31.0%
GPQA: —%
SWE-BENCH: —%
CONTEXT: 524K
SPEED: 80 t/s (3500ms)
BLENDED $/M: $1.2000
$ / POINT: —
Institute of Foundation Models • API • ▲ +8
3
STATS SCORE: 30.0%
GPQA: —%
SWE-BENCH: —%
CONTEXT: 1M
SPEED: 186 t/s (1000ms)
BLENDED $/M: $0.6000
$ / POINT: —
MiniMax • API • ▲ +7
STATS SCORE: 29.9%
GPQA: —%
SWE-BENCH: —%
CONTEXT: 1M
SPEED: 119 t/s (33160ms)
BLENDED $/M: $1.4000
$ / POINT: —
Google • API • ▲ +6
#25 Inkling
2.6
STATS SCORE: 26.1%
GPQA: —%
SWE-BENCH: —%
CONTEXT: 1M
SPEED: 81 t/s (2480ms)
BLENDED $/M: $0.7500
$ / POINT: —
Thinking Machines • API • ▲ +5
STATS SCORE: 23.0%
GPQA: —%
SWE-BENCH: —%
CONTEXT: 262K
SPEED: 164 t/s (2400ms)
BLENDED $/M: $0.7500
$ / POINT: —
NVIDIA • API • ▲ +4
STATS SCORE: 23.0%
GPQA: —%
SWE-BENCH: —%
CONTEXT: 1M
SPEED: 358 t/s (9780ms)
BLENDED $/M: $0.1700
$ / POINT: —
Google • API • ▲ +3
STATS SCORE: 18.1%
GPQA: —%
SWE-BENCH: —%
CONTEXT: 131K
SPEED: 93 t/s (990ms)
BLENDED $/M: $0.1200
$ / POINT: —
Meta • API • ▲ +2
STATS SCORE: 15.1%
GPQA: —%
SWE-BENCH: —%
CONTEXT: 256K
SPEED: 137 t/s (2290ms)
BLENDED $/M: $0.6000
$ / POINT: —
Mistral • API • ▲ +1
STATS SCORE: 12.0%
GPQA: —%
SWE-BENCH: —%
CONTEXT: 131K
SPEED: 181 t/s (850ms)
BLENDED $/M: $0.1800
$ / POINT: —
OpenAI • API • ▼ -1