EMPIRICAL BENCHMARK • INDEPENDENT EVALUATION

LMArena Explained: How Chatbot Arena Really Ranks Models

Live 2026 Chatbot Arena rankings, category charts, confidence intervals, and complete methodology breakdown of Bradley-Terry, AutoEval, and Style Control.

Leaderboard Overview

See how leading AI models stack up across text, image, vision, and agent workflows. Sourced directly from live Chatbot Arena blind battle telemetry. Explore dedicated tabs below for in-depth category ratings.

Vote on Arena.ai
Agent Leaderboard
Multi-turn Coding Agent • Win Rate %
1
Claude Fable 5.1 (Max)
▲ 13.71%±1.72%
2
GPT 6 Astra (Max)
▲ 11.54%±2.10%
3
Claude Opus 5 (High)
▲ 10.25%±1.41%
4
Claude Opus 5 (Max)
▲ 10.16%±1.55%
5
Claude Fable 5 (High)
▲ 8.81%±1.25%
6
Claude Opus 4.8 (High)
▲ 8.19%±1.27%
7
GPT 5.6 Sol (xHigh)
▲ 7.1%±1.28%
8
Kimi K3 (Max)
▲ 6.22%±0.62%
9
Claude Sonnet 5 (High)
▲ 5.97%±0.85%
Text
🏆 Overall ▾
1
claude-fable-5-high
1506±5
2
claude-opus-4-6-high
1505±4
3
claude-opus-4-7-high
1502±4
4
muse-spark-1.2 (xHigh)
1500±11
5
claude-fable-5.1-max
1498±8
6
claude-opus-4-6
1497±3
7
claude-opus-4-7
1494±4
8
muse-spark-1.3-max
1493±9
9
gemini-3.8-flash-high
1493ⓘ±9
10
claude-opus-5-high
1493±4
WebDev
🏆 Overall ▾
1
gpt-6-astra-max
1800+16/-16
2
claude-fable-5.1-max
1758+14/-14
3
claude-opus-5-max
1687+7/-7
4
qwen3.8-max-0902
1681ⓘ+15/-15
5
kimi-k3-max
1674+11/-11
6
qwen3.8-max
1671ⓘ+12/-12
7
claude-opus-5-high
1660+7/-7
8
muse-spark-1.3-max
1652+12/-12
9
1635ⓘ+13/-13
10
claude-fable-5-high
1628+7/-7
Vision
🏆 Overall ▾
1
claude-fable-5-high
1310±8
2
qwen3.8-max
1302±8
3
claude-opus-4-7-high
1301±7
4
claude-opus-4-7
1300±7
5
claude-opus-4-6-high
1299±7
6
muse-spark-1.3-max
1294±15
7
muse-spark
1294±9
8
claude-opus-4-6
1293±7
9
muse-spark-1.2 (xHigh)
1292±15
10
claude-opus-5-high
1289±8
Document
PDF & Tables ▾
1
claude-opus-5-high
1516±7
2
claude-fable-5.1-max
1513±15
3
claude-opus-4-6-high
1507±6
4
claude-opus-4-6
1507±6
5
claude-fable-5-high
1496±8
6
claude-opus-4-7
1495±6
7
claude-opus-4-7-high
1495±7
8
gpt-5.5
1486±6
9
gpt-5.5-high
1484±4
10
gpt-5.6-sol-xhigh
1483±9
Text-to-Image
🏆 Overall ▾
1
gpt-image-2.5-sunburst
1421ⓘ±13
2
gpt-image-2.5-flare
1399ⓘ±13
3
gpt-image-2 (medium)
1381±4
4
mai-image-2.6
1331±7
5
grok-imagine-image-2.0 (low)
1315ⓘ±12
6
reve-2.1
1301±8
7
muse-image
1277±6
8
reve-2.0
1270±6
9
gemini-3.1-flash-image (nano-banana-2)
1261±5
10
seedream-5.0-pro
1257±4
Image Edit
🏆 Single Image Edit ▾
1
gpt-image-2.5-sunburst
1520ⓘ±9
2
gpt-image-2.5-flare
1491ⓘ±9
3
gpt-image-2 (medium)
1461±8
4
grok-imagine-image-2.0 (low)
1439ⓘ±8
5
mai-image-2.6
1434±8
6
muse-image
1403±5
7
mai-image-2.5
1400±4
8
seedream-5.0-pro
1394±4
9
gemini-3-pro-image-2k (nano-banana-pro)
1390±3
10
grok-imagine-image-quality (20260519)
1390±6
Image-to-WebDev
Vision to Code ▾
1
gpt-6-astra-max
1733+21/-21
2
claude-fable-5.1-max
1710+21/-21
3
claude-opus-5-max
1665+13/-13
4
muse-spark-1.3-max
1645+18/-18
5
qwen3.8-max-0902
1639+19/-19
6
claude-fable-5-high
1623+11/-11
7
qwen3.8-max
1618ⓘ+13/-13
8
gpt-5.6-sol-xhigh (codex-harness)
1604+13/-13
9
grok-4.6-high
1596+18/-18
10
1588+15/-15
Search Grounding
Web Retrieval ▾
1
gpt-5.6-sol-xhigh
1257±7
2
claude-opus-4-6-search
1253±5
3
gpt-5.5-search
1242±5
4
claude-opus-4-7
1233±5
5
claude-fable-5-high
1230±8
6
ernie-5.1
1227±10
7
claude-sonnet-4-6-search
1221±6
8
grok-4.5
1213±7
9
gemini-3.1-pro-grounding
1210±5
10
gemini-3-pro-grounding
1207±6
Autonomous Agent Benchmark (Hugging Face Telemetry)
Multi-turn Win Rate %
1
Claude Fable 5.1 (Max)
▲ 13.71%±1.72%
2
GPT 6 Astra (Max)
▲ 11.54%±2.10%
3
Claude Opus 5 (High)
▲ 10.25%±1.41%
4
Claude Opus 5 (Max)
▲ 10.16%±1.55%
5
Claude Fable 5 (High)
▲ 8.81%±1.25%
6
Claude Opus 4.8 (High)
▲ 8.19%±1.27%
7
GPT 5.6 Sol (xHigh)
▲ 7.1%±1.28%
8
Kimi K3 (Max)
▲ 6.22%±0.62%
9
Claude Sonnet 5 (High)
▲ 5.97%±0.85%
WebDev & Programming Leaderboard
Bradley-Terry Elo Ratings
1
gpt-6-astra-max
1800+16/-16
2
claude-fable-5.1-max
1758+14/-14
3
claude-opus-5-max
1687+7/-7
4
qwen3.8-max-0902
1681ⓘ+15/-15
5
kimi-k3-max
1674+11/-11
6
qwen3.8-max
1671ⓘ+12/-12
7
claude-opus-5-high
1660+7/-7
8
muse-spark-1.3-max
1652+12/-12
9
qwen3.8-flash-next
1635ⓘ+13/-13
10
claude-fable-5-high
1628+7/-7
Conversational & General Text Leaderboard
Overall Text Elo
1
claude-fable-5-high
1506±5
2
claude-opus-4-6-high
1505±4
3
claude-opus-4-7-high
1502±4
4
muse-spark-1.2 (xHigh)
1500±11
5
claude-fable-5.1-max
1498±8
6
claude-opus-4-6
1497±3
7
claude-opus-4-7
1494±4
8
muse-spark-1.3-max
1493±9
9
gemini-3.8-flash-high
1493ⓘ±9
10
claude-opus-5-high
1493±4
Vision (Multimodal)
Overall Vision Elo
1
claude-fable-5-high
1310±8
2
qwen3.8-max
1302±8
3
claude-opus-4-7-high
1301±7
4
claude-opus-4-7
1300±7
5
claude-opus-4-6-high
1299±7
6
muse-spark-1.3-max
1294±15
7
muse-spark
1294±9
8
claude-opus-4-6
1293±7
9
muse-spark-1.2 (xHigh)
1292±15
10
claude-opus-5-high
1289±8
Document Understanding
PDF & Charts Elo
1
claude-opus-5-high
1516±7
2
claude-fable-5.1-max
1513±15
3
claude-opus-4-6-high
1507±6
4
claude-opus-4-6
1507±6
5
claude-fable-5-high
1496±8
6
claude-opus-4-7
1495±6
7
claude-opus-4-7-high
1495±7
8
gpt-5.5
1486±6
9
gpt-5.5-high
1484±4
10
gpt-5.6-sol-xhigh
1483±9
Text-to-Image Generation
Image Elo
1
gpt-image-2.5-sunburst
1421ⓘ±13
2
gpt-image-2.5-flare
1399ⓘ±13
3
gpt-image-2 (medium)
1381±4
4
mai-image-2.6
1331±7
5
grok-imagine-image-2.0 (low)
1315ⓘ±12
6
reve-2.1
1301±8
7
muse-image
1277±6
8
reve-2.0
1270±6
9
gemini-3.1-flash-image (nano-banana-2)
1261±5
10
seedream-5.0-pro
1257±4
Image Editing
Single Image Edit Elo
1
gpt-image-2.5-sunburst
1520ⓘ±9
2
gpt-image-2.5-flare
1491ⓘ±9
3
gpt-image-2 (medium)
1461±8
4
grok-imagine-image-2.0 (low)
1439ⓘ±8
5
mai-image-2.6
1434±8
6
muse-image
1403±5
7
mai-image-2.5
1400±4
8
seedream-5.0-pro
1394±4
9
gemini-3-pro-image-2k (nano-banana-pro)
1390±3
10
grok-imagine-image-quality (20260519)
1390±6
Master Arena Leaderboard (52 Frontier Models)
First Place Second Place Third Place
🔍 595 / 595
Rank ▴▾ Model ▴▾ Provider ▴▾ Overall ▴▾ Expert ▴▾ Hard Prompts ▴▾ Coding ▴▾ Math ▴▾ Creative Writing ▴▾ Instruction Following ▴▾ Longer Query ▴▾ Access ▴▾
1
claude-fable-5.1-max
Anthropic11122111Proprietary
2
gpt-6-astra-max
OpenAI22211222Proprietary
3
claude-opus-5-max
Anthropic33333333Proprietary
4
claude-fable-5-high
Anthropic444104444Proprietary
5
claude-opus-4-6-high
Anthropic55565555Proprietary
6
claude-opus-4-7-high
Anthropic66676666Proprietary
7
muse-spark-1.2-xhigh
Meta78788878Open
8
qwen3.8-max-0902
Alibaba Qwen878471087Proprietary
9
kimi-k3-max
Moonshot999591299Proprietary
10
qwen3.8-max
Alibaba Qwen101010610111010Proprietary
11
claude-opus-4-6
Anthropic111211111291111Proprietary
12
claude-opus-4-7
Anthropic1213121213131212Proprietary
13
muse-spark-1.3-max
Meta131413814141313Open
14
gemini-3.8-flash-high
Google1415141415151415Proprietary
15
claude-opus-5-high
Anthropic15111571171510Proprietary
16
qwen3.8-flash-next
Alibaba Qwen161616916161616Open
17
gpt-5.6-sol-xhigh
OpenAI1717171717181717Proprietary
18
claude-sonnet-5-high
Anthropic1820181820191818Proprietary
19
gpt-5.5
OpenAI1922191922201919Proprietary
20
gpt-5.5-high
OpenAI2023202023222020Proprietary
21
grok-4.6-high
xAI2124212124252121Proprietary
22
glm-5.3-flash
Zhipu AI2225222225272222Open
23
claude-opus-4-8-high
Anthropic2327232327282323Proprietary
24
ernie-5.1-pro
Baidu2429242429302424Proprietary
25
gemini-3.1-pro
Google2530252530312525Proprietary
26
deepseek-v4-pro
DeepSeek2631262631322626Open
27
mini-v2.5-max
MiniMax2733272733332727Proprietary
28
gpt-6.0
OpenAI2821305018392429Proprietary
29
glm-6.3-flash
Zhipu AI291828207602944Open
30
grok-4.20-beta1
xAI3072495265235764Proprietary
31
gemini-3.5-flash-medium
Google3152386328173842Proprietary
32
gpt-5.5-instant
OpenAI3266413852284846Proprietary
33
gemini-3-flash
Google3338365337264643Proprietary
34
qwen3.7-max-preview
Alibaba Qwen3425352121493620Proprietary
35
claude-opus-4-8
Anthropic3514211539212114Proprietary
36
claude-opus-4-5-20251101
Anthropic3634261443161418Proprietary
37
claude-sonnet-4-6
Anthropic3723201658372219Proprietary
38
glm-5.2-max
Zhipu AI3840374826353231Open
39
grok-4.20-beta-0909
xAI3959454747446262Proprietary
40
grok-4.20-multi-agent
xAI4065525163406669Proprietary
41
claude-opus-4-5-base
Anthropic4129272350242321Proprietary
42
grok-4.5
xAI4243393332363333Proprietary
43
ernie-5.1
Baidu4341474035675259Proprietary
44
mini-v2.5-pro
MiniMax4424332533593027Proprietary
45
gpt-5.6-terra-xhigh
OpenAI4526423538694155Proprietary
46
gpt-5.4
OpenAI4645464155564241Proprietary
47
glm-5.1
Zhipu AI4739444534434440Open
48
grok-4.1-thinking
xAI4883616973688488Proprietary
49
qwen3.5-max-preview
Alibaba Qwen4935404344423136Proprietary
50
deepseek-v4-pro-high
DeepSeek5052535951514849Open
51
claude-sonnet-5-high
Anthropic5120432736553435Proprietary
52
kimi-k2.6
Moonshot5232503927655145Proprietary

If you arrived here searching for “chatbot arena ranking methodology glicko-2 blind battles,” that query contains a foundational factual error: Chatbot Arena has never used Glicko-2. The platform moved from an early online Elo prototype directly to a joint Bradley-Terry Maximum Likelihood Estimation (MLE) regression, as confirmed in Arena’s official FAQ (arena.ai/faq). Furthermore, LMArena is now officially Arena, operated by Arena Intelligence Inc. at arena.ai. Below is how votes become ratings, why Style Control alters the leaderboard, what the 2026 AutoEval additions mean for human evaluation, and the exact boundaries of what an Arena score measures.

What is Chatbot Arena, and Who Created It?

Chatbot Arena was launched in May 2023 by LMSYS Org (Large Model Systems Organization), a research collective founded by researchers at UC Berkeley SkyLab with UC San Diego (UCSD) and Carnegie Mellon University (CMU). Its goal was replacing contaminated, static multiple-choice benchmarks (like MMLU) with real-world conversational evaluation.

The project’s academic foundation was published in “Chatbot Arena: An Open Platform for Evaluating LLMs by Human Preference” (arXiv:2403.04132, March 2024), authored by Wei-Lin Chiang, Lianmin Zheng, Ying Sheng, Anastasios Nikolas Angelopoulos, Tianle Li, Dacheng Li, Hao Zhang, Banghua Zhu, Michael Jordan, Joseph E. Gonzalez, and Ion Stoica. By publication, the platform had collected over 240,000 blind pairwise human comparisons.

As the platform grew into the industry’s default consumer sentiment benchmark, its branding evolved across three distinct eras:

Era Official Name Primary URL Governance & Infrastructure
May 2023 – Mid 2024 Chatbot Arena chat.lmsys.org UC Berkeley SkyLab academic open-source research project
Mid 2024 – Early 2026 LMArena lmarena.ai LMSYS Org consortium hosted on Hugging Face Spaces
2026 – Present Arena arena.ai Commercial spin-out under Arena Intelligence Inc. ($100M ARR commercial enterprise)

Understanding the purpose requires understanding what standard academic benchmarks failed to deliver. Static benchmarks suffer from rapid test-set leakage, unrepresentative prompt distributions, and metric saturation. Arena addressed this by crowdsourcing natural language queries from millions of real humans globally.

How Blind Battles Work: From User Prompt to Vote

The platform operates on a single core mechanism: the crowdsourced blind pairwise comparison. A user enters a prompt, two anonymous models generate answers side-by-side under identical system constraints, and the user selects which response was better.

Architecture diagram of the Chatbot Arena blind battle voting pipeline
Figure 1 • End-to-End Chatbot Arena Blind Battle Voting Pipeline

The sequence enforces strict blind conditions:

  1. Prompt Submission: The human user provides an open-ended input. Prompts range from Python bug debugging to multi-paragraph creative essays.
  2. Model Routing & Pairing: Arena’s matchmaking backend routes the prompt simultaneously to Model A and Model B. Match pairings are chosen dynamically based on information-gain algorithms that prioritize models with close ratings or wide confidence intervals.
  3. Generation: Both models generate responses in parallel. Brand-identifying pre-prompts and self-identification headers are stripped.
  4. Human Adjudication: The user reviews both answers and picks one of four buttons: Model A is better, Model B is better, Tie, or Both are bad.
  5. De-anonymization: Only after the vote is recorded in the immutable database are the true identities of Model A and Model B revealed to the user.

The Voting Threshold: To prevent statistical manipulation, a model must accumulate a minimum of 1,000 verified blind battle votes before its Bradley-Terry score stabilizes on the public public leaderboard.

The Bradley-Terry Model: Why Arena Does Not Use Glicko-2 (or Elo)

The persistent misconception that Chatbot Arena uses Glicko-2 stems from its origins in competitive gaming. While chess and video games rely on Glicko-2 to incorporate rating volatility, Arena has never deployed Glicko-2.

At launch in May 2023, Arena used a standard sequential Elo calculation. While computationally lightweight, online Elo suffered from three major flaws:

  • Order Dependency: In sequential Elo, the order in which matches are processed changes the final rating. If a new model faces top models early, its rating trajectory diverges compared to facing them later.
  • Time Drift: Ratings from June 2023 were not directly comparable to ratings calculated in December 2023 due to non-stationary opponent pools.
  • Tie Inefficiency: Classical Elo handles draws arbitrarily by awarding 0.5 points, distorting log-odds calibrations.

To resolve this, LMSYS migrated in early 2024 to Bradley-Terry Maximum Likelihood Estimation (MLE).

Comparison diagram contrasting sequential Elo updating against joint Bradley-Terry MLE regression
Figure 2 • Sequential Online Elo vs Joint Bradley-Terry MLE

The Bradley-Terry Mathematical Formulation

The Bradley-Terry model assumes that each model i possesses an latent quality parameter θi. When Model i faces Model j, the probability that i is preferred over j is modeled via the logistic sigmoid function:

P(i beats j) = 1 / [ 1 + 10(θj − θi) / 400 ]

Instead of updating ratings game-by-game, Arena aggregates the entire matrix of pairwise battles across all models simultaneously. The joint log-likelihood function across all recorded matchups is maximized:

L(θ) = ∑i < j [ Wij · ln P(i beats j) + Wji · ln P(j beats i) ]

Where Wij is the empirical number of times Model i defeated Model j. This optimization is solved using convex logistic regression (L-BFGS). The output θ vector provides a globally optimal rating scale anchored to an arbitrarily chosen baseline (typically 1,000 or an established open model).

Confidence Intervals via Non-Parametric Bootstrapping

Unlike sequential Elo which produces single-point estimates, Arena reports 95% bootstrap confidence intervals. Arena repeatedly resamples the battle dataset 1,000 times with replacement, computes the Bradley-Terry MLE for each sample, and records the 2.5th and 97.5th percentiles. When two models have overlapping error whiskers on the leaderboard, their performance difference is statistically indistinguishable.

Style Control: The Battle Against Length and Formatting Bias

By late 2023, researchers discovered a serious vulnerability: crowdsourced human evaluators exhibit systemic superficial biases. In blind tests, human judges consistently prefer:

  • Verbosity: Longer, wordier responses beat concise, technically accurate answers up to 70% of the time.
  • Markdown Formatting: Extensive bolding, numbered lists, and bullet points artificially inflate preference scores.
  • Polite Conversational Tone: Apologetic, friendly sycophancy ranks higher than blunt correctness.

This phenomenon created an incentive for model creators to tune models for excessive length rather than reasoning ability.

Bar chart showing rating shifts before and after Style Control normalization
Figure 3 • Impact of Style Control on Frontier Model Ratings

To neutralize this, LMSYS implemented Style Control (June 2024). Style Control extends the Bradley-Terry log-linear model by adding explicit covariate coefficients for response length and markdown complexity:

logit(P(i beats j)) = (θi − θj) + βlen · ΔLengthij + βmd · ΔMarkdownij

By conditioning the regression on length differential (ΔLength) and markdown usage differential (ΔMarkdown), the resulting θ parameters measure preference independent of verbosity. When Style Control is applied, overly verbose models lose 20 to 50 rating points, while concise reasoning models move up significantly.

2026 Additions: AutoEval, Factuality, and Multi-Turn Benchmarking

As Arena scaled into 2026, LMSYS introduced three structural architectural upgrades to address modern frontier LLM capabilities:

Evolution timeline of Chatbot Arena milestones from 2023 launch to 2026 AutoEval
Figure 4 • Architectural Evolution of Chatbot Arena: 2023 – 2026

1. AutoEval (July 2026)

Accumulating 1,000 human votes for every open-weights release took weeks. In July 2026, Arena launched AutoEval (arena.ai/blog/autoeval), a calibrated synthetic evaluation pipeline. AutoEval deploys a panel of frontier LLM judges using prompt templates rigorously calibrated against historical human consensus. Preliminary ratings appear within hours of model weights dropping, marked with an automated badge until 1,000 human votes confirm the rating.

2. The Factuality Benchmark

Human judges are notoriously poor at detecting subtle hallucinations in complex domains (e.g. quantum physics, advanced tax law, specialized medical coding). A model that hallucinates authoritatively often wins a human vote over one that admits uncertainty. The Factuality track pairs human prompts with automated fact-checking pipelines to penalize confident inaccuracies.

3. Multi-Turn Conversations and Agent Benchmarks

Single-turn prompts cannot evaluate context retention or agentic error correction. The 2026 Agent Leaderboard evaluates models across multi-turn interactions where models execute code in sandboxes, navigate terminals, and debug complex multi-file repositories.

Commercialization and Governance: UC Berkeley to Arena Intelligence Inc.

In early 2026, LMSYS transitioned its operational infrastructure from an academic research project to a commercial entity: Arena Intelligence Inc., operating at arena.ai. The commercial transition was driven by staggering computational costs: serving millions of free blind battle inferences across hundreds of frontier models requires millions of dollars in monthly GPU compute.

By late 2026, Arena reached an estimated $100M ARR through enterprise API benchmarking, private model pre-testing, and customized evaluation suites for AI labs. While the core public leaderboard remains openly accessible, this commercialization introduces critical governance considerations.

The Selective Disclosure Loophole: AI labs can submit unreleased checkpoint weights to Arena for private evaluation before public launch. If a checkpoint scores at the top of the leaderboard, the lab publishes the result as marketing material. If the model underperforms, the lab simply suppresses the score and discards the checkpoint. Because negative results remain private, public rankings reflect an inherent survivor bias.

What an Arena Score Measures — and What It Does Not

Because Chatbot Arena is widely cited in earnings calls and research announcements, practitioners frequently treat it as a universal proxy for model intelligence. It is not.

Evaluation boundaries matrix comparing controlled variables against uncontrolled biases in LMArena
Figure 5 • Controlled Variables vs Uncontrolled Biases in Arena Evaluation

The core principle: an Arena score measures conversational preference, not technical correctness. In creative writing, preference is useful; in programming, polite models with broken code routinely defeat working concise solutions. Arena acknowledged this in July 2026 by adding its factuality tier (arena.ai/blog/factuality-in-arena).

How Arena Compares to Canonical Coding Benchmarks

At Vibecoder Journal, we benchmark models against reproducible test suites. For software engineering, Arena should be evaluated alongside deterministic benchmarks:

Benchmark Evaluation Harness Primary Metric Core Strengths & Vulnerabilities
Arena (Coding) Crowdsourced blind pairwise web comparisons Bradley-Terry Elo (Preference) Conversational alignment; vulnerable to verbosity and unverified execution.
SWE-bench Verified Docker unit test execution (N=500) Resolve Rate (% binary PASS) Gold standard for GitHub bug resolution. Deterministic; expensive to evaluate.
Terminal-Bench 2.0 Linux sandbox terminal execution CLI Completion Rate (%) Evaluates bash scripting, git operations, and tool calling in an isolated container.
LiveCodeBench Competitive programming contests Pass@1 on fresh problems Tests algorithmic correctness on post-training cutoff coding contests (LeetCode, Codeforces).
Aider Polyglot Multi-language git diff search/replace harness Polyglot Resolve Rate (%) Tests real-world git patch editing across 6 languages under zero-shot conditions.
OpenRouter Free Uptime telemetry (N=100) + 25 tasks Availability Rate & Unit Tests Measures cloud availability, rate limits, and task pass rates for free AI models.

Frequently Asked Questions (FAQ)

Does Chatbot Arena use Elo or Glicko-2?
Neither. Chatbot Arena has never used Glicko-2. While it launched in 2023 with an online sequential Elo engine, it migrated entirely to Bradley-Terry Maximum Likelihood Estimation (MLE) in early 2024. Bradley-Terry jointly solves the full pairwise battle matrix simultaneously, eliminating match-sequence bias and time drift.
Who created Chatbot Arena?
Chatbot Arena was created by LMSYS Org, a research consortium founded at UC Berkeley SkyLab with UC San Diego (UCSD) and Carnegie Mellon University (CMU). Key founding researchers include Wei-Lin Chiang, Lianmin Zheng, Ying Sheng, and advisors Michael Jordan, Joseph E. Gonzalez, and Ion Stoica.
Is LMArena the same as Arena.ai?
Yes. The platform launched at chat.lmsys.org in 2023, transitioned to lmarena.ai in 2024, and rebranded to arena.ai under Arena Intelligence Inc. in 2026.
How many votes does a model need to appear on the leaderboard?
Models require ≥1,000 blind battle votes before ratings stabilize on the public leaderboard. With AutoEval (July 2026), Arena also displays preliminary calibrated synthetic ratings for early-stage models before 1,000 human votes accumulate.
What is Style Control?
Style Control is a mathematical covariate inside Arena’s Bradley-Terry regression that strips out length and formatting bias, preventing models from boosting ratings purely through verbosity.
Are Arena rankings reliable for programming and software engineering?
Arena rankings reflect conversational preference, which does not predict code correctness. A model outputting broken code politely can easily win an Arena battle. Developers should reference deterministic test suites like SWE-bench Verified, Terminal-Bench, and Aider Polyglot.
Abdullah Zulfiqar
Abdullah Zulfiqar Founder & Technical Editor

Abdullah Zulfiqar is the founder and editor of Vibe Coder Journal, an independent publication that benchmarks AI coding tools. He verifies every figure against primary sources — official documentation, real release files and live leaderboards — rather than repeating secondary reporting. His work has corrected widely-circulated errors in Terminal-Bench scores, Ollama's official uninstall instructions and Anthropic's documented install commands. Vibe Coder Journal accepts no sponsorships or affiliate commissions.

Related AI Coding Benchmarks