Chatbot Arena ELO measures relative user preference from blind head-to-head votes in LMArena. It is a live human-judgment signal for conversational quality under real prompts.
Large-scale blinded pairwise voting reduces single-rater bias; continuous updates surface quality shifts quickly after model releases.
Voter population is self-selected and may not match enterprise or domain users; style, verbosity, and safety tone can influence votes independently of factual correctness.
Higher is better. Scores are ELO-like ratings (typically 900-1400). Compare relative ranking — absolute values shift as new models enter.
Voter population is self-selected and may not match enterprise or domain users.
| # | Model | Score |
|---|---|---|
| 1 | 4203.0 | |
| 2 | 3412.0 | |
| 3 | 2391.0 | |
| 4 | Claude Fable 5 (Adaptive Reasoning, Max Effort, Opus 4.8 Fallback) source chatbot_arena Evidence observed Jul 31, 2026 | 1508.0 |
| 5 | 1502.0 | |
| 6 | 1502.0 | |
| 7 | 1497.0 | |
| 8 | Claude Opus 5 (Adaptive Reasoning, Max Effort) source chatbot_arena Preliminary Evidence observed Jul 31, 2026 | 1495.0 |
| 9 | 1493.0 | |
| 10 | 1491.0 | |
| 11 | 1488.0 | |
| 12 | 1486.0 | |
| 13 | 1486.0 | |
| 14 | 1486.0 | |
| 15 | 1485.0 | |
| 16 | 1485.0 | |
| 17 | 1482.0 | |
| 18 | 1482.0 | |
| 19 | 1482.0 | |
| 20 | 1482.0 | |
| 21 | 1481.0 | |
| 22 | 1476.0 | |
| 23 | 1476.0 | |
| 24 | 1475.0 | |
| 25 | 1474.0 | |
| 26 | 1474.0 | |
| 27 | 1473.0 | |
| 28 | 1473.0 | |
| 29 | 1473.0 | |
| 30 | 1473.0 | |
| 31 | 1473.0 | |
| 32 | 1472.0 | |
| 33 | 1472.0 | |
| 34 | 1472.0 | |
| 35 | 1469.0 | |
| 36 | 1469.0 | |
| 37 | 1468.0 | |
| 38 | 1468.0 | |
| 39 | 1466.0 | |
| 40 | 1465.0 | |
| 41 | 1465.0 | |
| 42 | 1462.0 | |
| 43 | 1461.0 | |
| 44 | 1460.0 | |
| 45 | 1460.0 | |
| 46 | 1460.0 | |
| 47 | 1457.0 | |
| 48 | 1457.0 | |
| 49 | 1457.0 | |
| 50 | 1457.0 |