| Rank ▴▾ | Model ▴▾ | Provider ▴▾ | Overall ▴▾ | Expert ▴▾ | Hard Prompts ▴▾ | Coding ▴▾ | Math ▴▾ | Creative Writing ▴▾ | Instruction Following ▴▾ | Longer Query ▴▾ | Access ▴▾ |
|---|---|---|---|---|---|---|---|---|---|---|---|
| 1 | claude-fable-5.1-max | Anthropic | 1 | 1 | 1 | 2 | 2 | 1 | 1 | 1 | Proprietary |
| 2 | gpt-6-astra-max | OpenAI | 2 | 2 | 2 | 1 | 1 | 2 | 2 | 2 | Proprietary |
| 3 | claude-opus-5-max | Anthropic | 3 | 3 | 3 | 3 | 3 | 3 | 3 | 3 | Proprietary |
| 4 | claude-fable-5-high | Anthropic | 4 | 4 | 4 | 10 | 4 | 4 | 4 | 4 | Proprietary |
| 5 | claude-opus-4-6-high | Anthropic | 5 | 5 | 5 | 6 | 5 | 5 | 5 | 5 | Proprietary |
| 6 | claude-opus-4-7-high | Anthropic | 6 | 6 | 6 | 7 | 6 | 6 | 6 | 6 | Proprietary |
| 7 | muse-spark-1.2-xhigh | Meta | 7 | 8 | 7 | 8 | 8 | 8 | 7 | 8 | Open |
| 8 | qwen3.8-max-0902 | Alibaba Qwen | 8 | 7 | 8 | 4 | 7 | 10 | 8 | 7 | Proprietary |
| 9 | kimi-k3-max | Moonshot | 9 | 9 | 9 | 5 | 9 | 12 | 9 | 9 | Proprietary |
| 10 | qwen3.8-max | Alibaba Qwen | 10 | 10 | 10 | 6 | 10 | 11 | 10 | 10 | Proprietary |
| 11 | claude-opus-4-6 | Anthropic | 11 | 12 | 11 | 11 | 12 | 9 | 11 | 11 | Proprietary |
| 12 | claude-opus-4-7 | Anthropic | 12 | 13 | 12 | 12 | 13 | 13 | 12 | 12 | Proprietary |
| 13 | muse-spark-1.3-max | Meta | 13 | 14 | 13 | 8 | 14 | 14 | 13 | 13 | Open |
| 14 | gemini-3.8-flash-high | 14 | 15 | 14 | 14 | 15 | 15 | 14 | 15 | Proprietary | |
| 15 | claude-opus-5-high | Anthropic | 15 | 11 | 15 | 7 | 11 | 7 | 15 | 10 | Proprietary |
| 16 | qwen3.8-flash-next | Alibaba Qwen | 16 | 16 | 16 | 9 | 16 | 16 | 16 | 16 | Open |
| 17 | gpt-5.6-sol-xhigh | OpenAI | 17 | 17 | 17 | 17 | 17 | 18 | 17 | 17 | Proprietary |
| 18 | claude-sonnet-5-high | Anthropic | 18 | 20 | 18 | 18 | 20 | 19 | 18 | 18 | Proprietary |
| 19 | gpt-5.5 | OpenAI | 19 | 22 | 19 | 19 | 22 | 20 | 19 | 19 | Proprietary |
| 20 | gpt-5.5-high | OpenAI | 20 | 23 | 20 | 20 | 23 | 22 | 20 | 20 | Proprietary |
| 21 | grok-4.6-high | xAI | 21 | 24 | 21 | 21 | 24 | 25 | 21 | 21 | Proprietary |
| 22 | glm-5.3-flash | Zhipu AI | 22 | 25 | 22 | 22 | 25 | 27 | 22 | 22 | Open |
| 23 | claude-opus-4-8-high | Anthropic | 23 | 27 | 23 | 23 | 27 | 28 | 23 | 23 | Proprietary |
| 24 | ernie-5.1-pro | Baidu | 24 | 29 | 24 | 24 | 29 | 30 | 24 | 24 | Proprietary |
| 25 | gemini-3.1-pro | 25 | 30 | 25 | 25 | 30 | 31 | 25 | 25 | Proprietary | |
| 26 | deepseek-v4-pro | DeepSeek | 26 | 31 | 26 | 26 | 31 | 32 | 26 | 26 | Open |
| 27 | mini-v2.5-max | MiniMax | 27 | 33 | 27 | 27 | 33 | 33 | 27 | 27 | Proprietary |
| 28 | gpt-6.0 | OpenAI | 28 | 21 | 30 | 50 | 18 | 39 | 24 | 29 | Proprietary |
| 29 | glm-6.3-flash | Zhipu AI | 29 | 18 | 28 | 20 | 7 | 60 | 29 | 44 | Open |
| 30 | grok-4.20-beta1 | xAI | 30 | 72 | 49 | 52 | 65 | 23 | 57 | 64 | Proprietary |
| 31 | gemini-3.5-flash-medium | 31 | 52 | 38 | 63 | 28 | 17 | 38 | 42 | Proprietary | |
| 32 | gpt-5.5-instant | OpenAI | 32 | 66 | 41 | 38 | 52 | 28 | 48 | 46 | Proprietary |
| 33 | gemini-3-flash | 33 | 38 | 36 | 53 | 37 | 26 | 46 | 43 | Proprietary | |
| 34 | qwen3.7-max-preview | Alibaba Qwen | 34 | 25 | 35 | 21 | 21 | 49 | 36 | 20 | Proprietary |
| 35 | claude-opus-4-8 | Anthropic | 35 | 14 | 21 | 15 | 39 | 21 | 21 | 14 | Proprietary |
| 36 | claude-opus-4-5-20251101 | Anthropic | 36 | 34 | 26 | 14 | 43 | 16 | 14 | 18 | Proprietary |
| 37 | claude-sonnet-4-6 | Anthropic | 37 | 23 | 20 | 16 | 58 | 37 | 22 | 19 | Proprietary |
| 38 | glm-5.2-max | Zhipu AI | 38 | 40 | 37 | 48 | 26 | 35 | 32 | 31 | Open |
| 39 | grok-4.20-beta-0909 | xAI | 39 | 59 | 45 | 47 | 47 | 44 | 62 | 62 | Proprietary |
| 40 | grok-4.20-multi-agent | xAI | 40 | 65 | 52 | 51 | 63 | 40 | 66 | 69 | Proprietary |
| 41 | claude-opus-4-5-base | Anthropic | 41 | 29 | 27 | 23 | 50 | 24 | 23 | 21 | Proprietary |
| 42 | grok-4.5 | xAI | 42 | 43 | 39 | 33 | 32 | 36 | 33 | 33 | Proprietary |
| 43 | ernie-5.1 | Baidu | 43 | 41 | 47 | 40 | 35 | 67 | 52 | 59 | Proprietary |
| 44 | mini-v2.5-pro | MiniMax | 44 | 24 | 33 | 25 | 33 | 59 | 30 | 27 | Proprietary |
| 45 | gpt-5.6-terra-xhigh | OpenAI | 45 | 26 | 42 | 35 | 38 | 69 | 41 | 55 | Proprietary |
| 46 | gpt-5.4 | OpenAI | 46 | 45 | 46 | 41 | 55 | 56 | 42 | 41 | Proprietary |
| 47 | glm-5.1 | Zhipu AI | 47 | 39 | 44 | 45 | 34 | 43 | 44 | 40 | Open |
| 48 | grok-4.1-thinking | xAI | 48 | 83 | 61 | 69 | 73 | 68 | 84 | 88 | Proprietary |
| 49 | qwen3.5-max-preview | Alibaba Qwen | 49 | 35 | 40 | 43 | 44 | 42 | 31 | 36 | Proprietary |
| 50 | deepseek-v4-pro-high | DeepSeek | 50 | 52 | 53 | 59 | 51 | 51 | 48 | 49 | Open |
| 51 | claude-sonnet-5-high | Anthropic | 51 | 20 | 43 | 27 | 36 | 55 | 34 | 35 | Proprietary |
| 52 | kimi-k2.6 | Moonshot | 52 | 32 | 50 | 39 | 27 | 65 | 51 | 45 | Proprietary |
If you arrived here searching for “chatbot arena ranking methodology glicko-2 blind battles,” that query contains a foundational factual error: Chatbot Arena has never used Glicko-2. The platform moved from an early online Elo prototype directly to a joint Bradley-Terry Maximum Likelihood Estimation (MLE) regression, as confirmed in Arena’s official FAQ (arena.ai/faq). Furthermore, LMArena is now officially Arena, operated by Arena Intelligence Inc. at arena.ai. Below is how votes become ratings, why Style Control alters the leaderboard, what the 2026 AutoEval additions mean for human evaluation, and the exact boundaries of what an Arena score measures.
What is Chatbot Arena, and Who Created It?
Chatbot Arena was launched in May 2023 by LMSYS Org (Large Model Systems Organization), a research collective founded by researchers at UC Berkeley SkyLab with UC San Diego (UCSD) and Carnegie Mellon University (CMU). Its goal was replacing contaminated, static multiple-choice benchmarks (like MMLU) with real-world conversational evaluation.
The project’s academic foundation was published in “Chatbot Arena: An Open Platform for Evaluating LLMs by Human Preference” (arXiv:2403.04132, March 2024), authored by Wei-Lin Chiang, Lianmin Zheng, Ying Sheng, Anastasios Nikolas Angelopoulos, Tianle Li, Dacheng Li, Hao Zhang, Banghua Zhu, Michael Jordan, Joseph E. Gonzalez, and Ion Stoica. By publication, the platform had collected over 240,000 blind pairwise human comparisons.
As the platform grew into the industry’s default consumer sentiment benchmark, its branding evolved across three distinct eras:
| Era | Official Name | Primary URL | Governance & Infrastructure |
|---|---|---|---|
| May 2023 – Mid 2024 | Chatbot Arena | chat.lmsys.org |
UC Berkeley SkyLab academic open-source research project |
| Mid 2024 – Early 2026 | LMArena | lmarena.ai |
LMSYS Org consortium hosted on Hugging Face Spaces |
| 2026 – Present | Arena | arena.ai |
Commercial spin-out under Arena Intelligence Inc. ($100M ARR commercial enterprise) |
Understanding the purpose requires understanding what standard academic benchmarks failed to deliver. Static benchmarks suffer from rapid test-set leakage, unrepresentative prompt distributions, and metric saturation. Arena addressed this by crowdsourcing natural language queries from millions of real humans globally.
How Blind Battles Work: From User Prompt to Vote
The platform operates on a single core mechanism: the crowdsourced blind pairwise comparison. A user enters a prompt, two anonymous models generate answers side-by-side under identical system constraints, and the user selects which response was better.
The sequence enforces strict blind conditions:
- Prompt Submission: The human user provides an open-ended input. Prompts range from Python bug debugging to multi-paragraph creative essays.
- Model Routing & Pairing: Arena’s matchmaking backend routes the prompt simultaneously to Model A and Model B. Match pairings are chosen dynamically based on information-gain algorithms that prioritize models with close ratings or wide confidence intervals.
- Generation: Both models generate responses in parallel. Brand-identifying pre-prompts and self-identification headers are stripped.
- Human Adjudication: The user reviews both answers and picks one of four buttons: Model A is better, Model B is better, Tie, or Both are bad.
- De-anonymization: Only after the vote is recorded in the immutable database are the true identities of Model A and Model B revealed to the user.
The Voting Threshold: To prevent statistical manipulation, a model must accumulate a minimum of 1,000 verified blind battle votes before its Bradley-Terry score stabilizes on the public public leaderboard.
The Bradley-Terry Model: Why Arena Does Not Use Glicko-2 (or Elo)
The persistent misconception that Chatbot Arena uses Glicko-2 stems from its origins in competitive gaming. While chess and video games rely on Glicko-2 to incorporate rating volatility, Arena has never deployed Glicko-2.
At launch in May 2023, Arena used a standard sequential Elo calculation. While computationally lightweight, online Elo suffered from three major flaws:
- Order Dependency: In sequential Elo, the order in which matches are processed changes the final rating. If a new model faces top models early, its rating trajectory diverges compared to facing them later.
- Time Drift: Ratings from June 2023 were not directly comparable to ratings calculated in December 2023 due to non-stationary opponent pools.
- Tie Inefficiency: Classical Elo handles draws arbitrarily by awarding 0.5 points, distorting log-odds calibrations.
To resolve this, LMSYS migrated in early 2024 to Bradley-Terry Maximum Likelihood Estimation (MLE).
The Bradley-Terry Mathematical Formulation
The Bradley-Terry model assumes that each model i possesses an latent quality parameter θi. When Model i faces Model j, the probability that i is preferred over j is modeled via the logistic sigmoid function:
P(i beats j) = 1 / [ 1 + 10(θj − θi) / 400 ]
Instead of updating ratings game-by-game, Arena aggregates the entire matrix of pairwise battles across all models simultaneously. The joint log-likelihood function across all recorded matchups is maximized:
L(θ) = ∑i < j [ Wij · ln P(i beats j) + Wji · ln P(j beats i) ]
Where Wij is the empirical number of times Model i defeated Model j. This optimization is solved using convex logistic regression (L-BFGS). The output θ vector provides a globally optimal rating scale anchored to an arbitrarily chosen baseline (typically 1,000 or an established open model).
Confidence Intervals via Non-Parametric Bootstrapping
Unlike sequential Elo which produces single-point estimates, Arena reports 95% bootstrap confidence intervals. Arena repeatedly resamples the battle dataset 1,000 times with replacement, computes the Bradley-Terry MLE for each sample, and records the 2.5th and 97.5th percentiles. When two models have overlapping error whiskers on the leaderboard, their performance difference is statistically indistinguishable.
Style Control: The Battle Against Length and Formatting Bias
By late 2023, researchers discovered a serious vulnerability: crowdsourced human evaluators exhibit systemic superficial biases. In blind tests, human judges consistently prefer:
- Verbosity: Longer, wordier responses beat concise, technically accurate answers up to 70% of the time.
- Markdown Formatting: Extensive bolding, numbered lists, and bullet points artificially inflate preference scores.
- Polite Conversational Tone: Apologetic, friendly sycophancy ranks higher than blunt correctness.
This phenomenon created an incentive for model creators to tune models for excessive length rather than reasoning ability.
To neutralize this, LMSYS implemented Style Control (June 2024). Style Control extends the Bradley-Terry log-linear model by adding explicit covariate coefficients for response length and markdown complexity:
logit(P(i beats j)) = (θi − θj) + βlen · ΔLengthij + βmd · ΔMarkdownij
By conditioning the regression on length differential (ΔLength) and markdown usage differential (ΔMarkdown), the resulting θ parameters measure preference independent of verbosity. When Style Control is applied, overly verbose models lose 20 to 50 rating points, while concise reasoning models move up significantly.
2026 Additions: AutoEval, Factuality, and Multi-Turn Benchmarking
As Arena scaled into 2026, LMSYS introduced three structural architectural upgrades to address modern frontier LLM capabilities:
1. AutoEval (July 2026)
Accumulating 1,000 human votes for every open-weights release took weeks. In July 2026, Arena launched AutoEval (arena.ai/blog/autoeval), a calibrated synthetic evaluation pipeline. AutoEval deploys a panel of frontier LLM judges using prompt templates rigorously calibrated against historical human consensus. Preliminary ratings appear within hours of model weights dropping, marked with an automated badge until 1,000 human votes confirm the rating.
2. The Factuality Benchmark
Human judges are notoriously poor at detecting subtle hallucinations in complex domains (e.g. quantum physics, advanced tax law, specialized medical coding). A model that hallucinates authoritatively often wins a human vote over one that admits uncertainty. The Factuality track pairs human prompts with automated fact-checking pipelines to penalize confident inaccuracies.
3. Multi-Turn Conversations and Agent Benchmarks
Single-turn prompts cannot evaluate context retention or agentic error correction. The 2026 Agent Leaderboard evaluates models across multi-turn interactions where models execute code in sandboxes, navigate terminals, and debug complex multi-file repositories.
Commercialization and Governance: UC Berkeley to Arena Intelligence Inc.
In early 2026, LMSYS transitioned its operational infrastructure from an academic research project to a commercial entity: Arena Intelligence Inc., operating at arena.ai. The commercial transition was driven by staggering computational costs: serving millions of free blind battle inferences across hundreds of frontier models requires millions of dollars in monthly GPU compute.
By late 2026, Arena reached an estimated $100M ARR through enterprise API benchmarking, private model pre-testing, and customized evaluation suites for AI labs. While the core public leaderboard remains openly accessible, this commercialization introduces critical governance considerations.
The Selective Disclosure Loophole: AI labs can submit unreleased checkpoint weights to Arena for private evaluation before public launch. If a checkpoint scores at the top of the leaderboard, the lab publishes the result as marketing material. If the model underperforms, the lab simply suppresses the score and discards the checkpoint. Because negative results remain private, public rankings reflect an inherent survivor bias.
What an Arena Score Measures — and What It Does Not
Because Chatbot Arena is widely cited in earnings calls and research announcements, practitioners frequently treat it as a universal proxy for model intelligence. It is not.
The core principle: an Arena score measures conversational preference, not technical correctness. In creative writing, preference is useful; in programming, polite models with broken code routinely defeat working concise solutions. Arena acknowledged this in July 2026 by adding its factuality tier (arena.ai/blog/factuality-in-arena).
How Arena Compares to Canonical Coding Benchmarks
At Vibecoder Journal, we benchmark models against reproducible test suites. For software engineering, Arena should be evaluated alongside deterministic benchmarks:
| Benchmark | Evaluation Harness | Primary Metric | Core Strengths & Vulnerabilities |
|---|---|---|---|
| Arena (Coding) | Crowdsourced blind pairwise web comparisons | Bradley-Terry Elo (Preference) | Conversational alignment; vulnerable to verbosity and unverified execution. |
| SWE-bench Verified | Docker unit test execution (N=500) | Resolve Rate (% binary PASS) | Gold standard for GitHub bug resolution. Deterministic; expensive to evaluate. |
| Terminal-Bench 2.0 | Linux sandbox terminal execution | CLI Completion Rate (%) | Evaluates bash scripting, git operations, and tool calling in an isolated container. |
| LiveCodeBench | Competitive programming contests | Pass@1 on fresh problems | Tests algorithmic correctness on post-training cutoff coding contests (LeetCode, Codeforces). |
| Aider Polyglot | Multi-language git diff search/replace harness | Polyglot Resolve Rate (%) | Tests real-world git patch editing across 6 languages under zero-shot conditions. |
| OpenRouter Free | Uptime telemetry (N=100) + 25 tasks | Availability Rate & Unit Tests | Measures cloud availability, rate limits, and task pass rates for free AI models. |
Frequently Asked Questions (FAQ)
Does Chatbot Arena use Elo or Glicko-2?
Who created Chatbot Arena?
Is LMArena the same as Arena.ai?
chat.lmsys.org in 2023, transitioned to lmarena.ai in 2024, and rebranded to arena.ai under Arena Intelligence Inc. in 2026.