LiveBench Overall tracks broad model performance on recently refreshed questions with verifiable answers. It is used as a lower-contamination snapshot across major capability categories.
Monthly refresh with verifiable answers reduces contamination relative to static sets; aggregate blends math, coding, reasoning, language, and data analysis for a balanced top-line view.
Monthly test-set changes make cross-month deltas imperfectly comparable; aggregate score can hide category-specific weaknesses that matter for deployment fit.
Higher is better. Global average across all categories. Top models score 60-80%.
Max score: 100
Monthly test-set changes make cross-month deltas imperfectly comparable.
| # | Model | Score |
|---|---|---|
| 1 | 98.5 | |
| 2 | 97.3 | |
| 3 | 96.2 | |
| 4 | 86.8 | |
| 5 | 86.8 | |
| 6 | 84.6 | |
| 7 | 84.4 | |
| 8 | 83.1 | |
| 9 | 82.6 | |
| 10 | 82.3 | |
| 11 | 82.0 | |
| 12 | 81.7 | |
| 13 | 81.4 | |
| 14 | 80.8 | |
| 15 | 80.7 | |
| 16 | 80.3 | |
| 17 | 80.3 | |
| 18 | 79.9 | |
| 19 | 79.5 | |
| 20 | 79.4 | |
| 21 | 78.3 | |
| 22 | 78.3 | |
| 23 | 78.2 | |
| 24 | 78.1 | |
| 25 | 77.9 | |
| 26 | 77.7 | |
| 27 | 77.0 | |
| 28 | 77.0 | |
| 29 | 76.8 | |
| 30 | 76.6 | |
| 31 | 76.3 | |
| 32 | 76.2 | |
| 33 | 76.0 | |
| 34 | 76.0 | |
| 35 | 76.0 | |
| 36 | 75.6 | |
| 37 | 75.3 | |
| 38 | 75.2 | |
| 39 | 75.1 | |
| 40 | 75.0 | |
| 41 | 75.0 | |
| 42 | 74.7 | |
| 43 | 74.6 | |
| 44 | 74.4 | |
| 45 | 74.3 | |
| 46 | 74.2 | |
| 47 | 74.2 | |
| 48 | 74.0 | |
| 49 | 73.9 | |
| 50 | 73.6 |