AgMoDB
ModelsAgentsEvalsCompositesVisualizeIndustry
AgMoDB by @mistakeknot
Methodology workbench · live calculation

Composite Benchmark Builder

Combine evidence, define the population, and inspect how every model earns its score.

Loading local draft

01 · Inputs

Source Benchmarks

1/225

02 · Method

Formula

Intelligence Index100%
Population filters

Agent support · any selected

Include or exclude specific models

Selecting any Include entry creates an allowlist. Exclude entries always win.

gpt-oss-20b (low)
gpt-oss-120b (low)
GPT-5.1 Codex mini (high)
gpt-oss-120b (high)
gpt-oss-20b (high)
GPT-5.2 (medium)
GPT-5 nano (high)
Grok-1
GPT-5.2 Codex (xhigh)
o3
GPT-5.2 (Non-reasoning)
GPT-5.2 (xhigh)
GPT-5.1 Codex (high)
GPT-5 mini (high)
Llama 3.3 Instruct 70B
Llama 3.1 Instruct 405B
Llama 3.2 Instruct 90B (Vision)
Llama 3.2 Instruct 11B (Vision)
Llama 4 Maverick
Llama 4 Scout
Gemma 3 1B Instruct
Gemini 3 Pro Preview (low)
Gemma 3 4B Instruct
Gemma 3n E2B Instruct
Gemini 2.5 Flash-Lite Preview (Sep '25) (Non-reasoning)
Gemini 2.5 Flash-Lite Preview (Sep '25) (Reasoning)
Gemini 3 Flash Preview (Non-reasoning)
Gemma 3 270M
Gemini 2.5 Pro
Gemma 3 27B Instruct
Gemma 3 12B Instruct
Gemini 3 Flash Preview (Reasoning)
Gemini 3 Pro Preview (high)
Gemma 3n E4B Instruct
Claude 4.5 Haiku (Reasoning)
Claude 4.5 Sonnet (Non-reasoning)
Claude Opus 4.6 (Adaptive Reasoning, Max Effort)
Claude 4.5 Sonnet (Reasoning)
Claude 4.5 Haiku (Non-reasoning)
Claude Opus 4.6 (Non-reasoning, High Effort)
Claude Opus 4.5 (Non-reasoning)
Claude Opus 4.5 (Reasoning)
Magistral Small 1.2
Magistral Medium 1.2
Mistral Large 3
Mistral Small 3.2
Ministral 3 8B
Ministral 3 3B
Devstral Small 2
Mistral Medium 3.1
Ministral 3 14B
Devstral 2
DeepSeek R1 Distill Llama 70B
DeepSeek V3.2 Speciale
DeepSeek V3.2 (Reasoning)
DeepSeek R1 0528 (May '25)
DeepSeek V3.2 (Non-reasoning)
DeepSeek R1 0528 Qwen3 8B
DeepSeek-OCR
R1 1776
Falcon-H1R-7B
Grok 4.1 Fast (Reasoning)
Grok 3 mini Reasoning (high)
Grok 4
Grok 4.1 Fast (Non-reasoning)
Grok Voice Agent
Grok Code Fast 1
Nova Micro
Nova Premier
Nova 2.0 Omni (Non-reasoning)
Nova 2.0 Omni (medium)
Nova 2.0 Lite (medium)
Nova 2.0 Lite (Non-reasoning)
Nova 2.0 Pro Preview (medium)
Nova 2.0 Omni (low)
Nova 2.0 Pro Preview (low)
Nova 2.0 Pro Preview (Non-reasoning)
Nova 2.0 Lite (low)
Phi-4
Phi-4 Mini Instruct
Phi-4 Multimodal Instruct
LFM2.5-VL-1.6B
LFM2.5-1.2B-Thinking
LFM2.5-1.2B-Instruct
LFM2 2.6B
LFM2 8B A1B
Solar Open 100B (Reasoning)
Solar Pro 2 (Reasoning)
Solar Pro 2 (Non-reasoning)
MiniMax-M2.1
Llama 3.1 Nemotron Instruct 70B
NVIDIA Nemotron Nano 12B v2 VL (Non-reasoning)
NVIDIA Nemotron 3 Nano 30B A3B (Non-reasoning)
NVIDIA Nemotron 3 Nano 30B A3B (Reasoning)
NVIDIA Nemotron Nano 9B V2 (Reasoning)
Llama 3.3 Nemotron Super 49B v1 (Non-reasoning)
NVIDIA Nemotron Nano 12B v2 VL (Reasoning)
Llama Nemotron Super 49B v1.5 (Reasoning)
Llama 3.3 Nemotron Super 49B v1 (Reasoning)
Llama Nemotron Super 49B v1.5 (Non-reasoning)
Llama 3.1 Nemotron Ultra 253B v1 (Reasoning)
NVIDIA Nemotron Nano 9B V2 (Non-reasoning)
Llama 3.1 Nemotron Nano 4B v1.1 (Reasoning)
Kimi K2.5 (Reasoning)
Kimi K2 Thinking
Kimi K2.5 (Non-reasoning)
Kimi K2 0905
Kimi Linear 48B A3B Instruct
Step3 VL 10B
Olmo 3 7B Instruct
Olmo 3.1 32B Think
Olmo 3 7B Think
Olmo 3.1 32B Instruct
Molmo 7B-D
Molmo2-8B
Granite 4.0 H Small
Granite 4.0 H 350M
Granite 4.0 H 1B
Granite 4.0 1B
Granite 4.0 Micro
Granite 4.0 350M
Reka Flash 3
Hermes 4 - Llama-3.1 405B (Reasoning)
DeepHermes 3 - Llama-3.1 8B Preview (Non-reasoning)
Hermes 4 - Llama-3.1 70B (Reasoning)
Hermes 4 - Llama-3.1 70B (Non-reasoning)
Hermes 4 - Llama-3.1 405B (Non-reasoning)
DeepHermes 3 - Mistral 24B Preview (Non-reasoning)
K-EXAONE (Non-reasoning)
EXAONE 4.0 32B (Non-reasoning)
EXAONE 4.0 32B (Reasoning)
Exaone 4.0 1.2B (Non-reasoning)
Exaone 4.0 1.2B (Reasoning)
K-EXAONE (Reasoning)
MiMo-V2-Flash (Reasoning)
MiMo-V2-Flash (Non-reasoning)
ERNIE 4.5 300B A47B
ERNIE 5.0 Thinking Preview
Llama 65B
Cogito v2.1 (Reasoning)
KAT-Coder-Pro V1
INTELLECT-3
Motif-2-12.7B-Reasoning
K2-V2 (low)
K2-V2 (high)
K2 Think V2
K2-V2 (medium)
Mi:dm K 2.5 Pro
Mi:dm K 2.5 Pro Preview
HyperCLOVA X SEED Think (32B)
GLM-4.7 (Non-reasoning)
GLM-4.5-Air
GLM-4.7 (Reasoning)
GLM-4.6V (Reasoning)
GLM-4.6V (Non-reasoning)
GLM-4.7-Flash (Reasoning)
GLM-4.7-Flash (Non-reasoning)
Command A
Apriel-v1.6-15B-Thinker
Jamba 1.7 Mini
Jamba Reasoning 3B
Jamba 1.7 Large
Qwen3 Omni 30B A3B Instruct
Qwen3 Coder 30B A3B Instruct
Qwen3 30B A3B 2507 Instruct
Qwen3 235B A22B 2507 Instruct
Qwen3 VL 30B A3B Instruct
Qwen3 235B A22B 2507 (Reasoning)
Qwen Chat 14B
Qwen3 30B A3B 2507 (Reasoning)
Qwen3 VL 4B (Reasoning)
Qwen3 VL 8B Instruct
Qwen3 VL 4B Instruct
Qwen3 Max Thinking
Qwen3 Next 80B A3B (Reasoning)
Qwen3 Coder 480B A35B Instruct
Qwen3 Max Thinking (Preview)
Qwen3 Next 80B A3B Instruct
Qwen3 1.7B (Reasoning)
Qwen3 VL 235B A22B Instruct
Qwen3 VL 30B A3B (Reasoning)
Qwen3 VL 32B Instruct
Qwen3 4B 2507 (Reasoning)
Qwen3 4B 2507 Instruct
Qwen3 VL 235B A22B (Reasoning)
Qwen3 Coder Next
Qwen3 Omni 30B A3B (Reasoning)
Qwen3 0.6B (Non-reasoning)
Qwen3 0.6B (Reasoning)
Qwen3 1.7B (Non-reasoning)
Qwen3 VL 8B (Reasoning)
Qwen3 VL 32B (Reasoning)
Qwen3 Max
Ling-mini-2.0
Ring-1T
Ling-1T
Ring-flash-2.0
Ling-flash-2.0
Doubao Seed Code
Doubao-Seed-1.8
o1
o1-preview
o1-mini
GPT-4o (Aug '24)
GPT-4o (May '24)
GPT-4 Turbo
GPT-4o (Nov '24)
GPT-4o mini
GPT-3.5 Turbo
GPT-4.1 nano
GPT-5.1 (high)
GPT-5 (minimal)
o4-mini (high)
GPT-4.1
GPT-5 Codex (high)
o3-pro
GPT-5.1 (Non-reasoning)
GPT-5 (high)
GPT-5 nano (medium)
GPT-5 (medium)
GPT-4.1 mini
GPT-4
GPT-5 (low)
GPT-5 mini (medium)
GPT-5 nano (minimal)
o3-mini
GPT-4o mini Realtime (Dec '24)
o3-mini (high)
GPT-5 mini (minimal)
GPT-4.5 (Preview)
GPT-4o (ChatGPT)
GPT-4o Realtime (Dec '24)
GPT-3.5 Turbo (0613)
o1-pro
GPT-5 (ChatGPT)
GPT-4o (March 2025, chatgpt-4o-latest)
Llama 3.1 Instruct 70B
Llama 3.1 Instruct 8B
Llama 3.2 Instruct 3B
Llama 3 Instruct 70B
Llama 3 Instruct 8B
Llama 3.2 Instruct 1B
Llama 2 Chat 7B
Llama 2 Chat 70B
Llama 2 Chat 13B
Gemini 2.0 Pro Experimental (Feb '25)
Gemini 2.0 Flash (experimental)
Gemini 1.5 Pro (Sep '24)
Gemini 2.0 Flash-Lite (Preview)
Gemini 2.0 Flash (Feb '25)
Gemini 1.5 Flash (Sep '24)
Gemini 1.5 Flash-8B
Gemini 1.0 Ultra
Gemma 3n E4B Instruct Preview (May '25)
Gemini 2.5 Flash (Non-reasoning)
Gemini 2.5 Flash Preview (Non-reasoning)
PALM-2
Gemini 2.5 Flash Preview (Sep '25) (Non-reasoning)
Gemini 2.5 Flash (Reasoning)
Gemini 2.5 Flash Preview (Reasoning)
Gemini 2.0 Flash Thinking Experimental (Jan '25)
Gemini 1.5 Pro (May '24)
Gemini 1.5 Flash (May '24)
Gemini 2.5 Flash-Lite (Non-reasoning)
Gemini 1.0 Pro
Gemini 2.5 Pro Preview (May' 25)
Gemini 2.5 Flash Preview (Sep '25) (Reasoning)
Gemini 2.0 Flash-Lite (Feb '25)
Gemini 2.0 Flash Thinking Experimental (Dec '24)
Gemini 2.5 Pro Preview (Mar' 25)
Gemini 2.5 Flash-Lite (Reasoning)
Claude 3.5 Sonnet (Oct '24)
Claude 3.5 Sonnet (June '24)
Claude 3 Opus
Claude 3.5 Haiku
Claude 3 Sonnet
Claude 3 Haiku
Claude Instant
Claude 3.7 Sonnet (Non-reasoning)
Claude 4.1 Opus (Reasoning)
Claude 4.1 Opus (Non-reasoning)
Claude 4 Sonnet (Non-reasoning)
Claude 3.7 Sonnet (Reasoning)
Claude 4 Opus (Non-reasoning)
Claude 4 Sonnet (Reasoning)
Claude 4 Opus (Reasoning)
Claude 2.1
Claude 2.0
Mistral Large 2 (Nov '24)
Mistral Large 2 (Jul '24)
Pixtral Large
Mistral Small 3
Mistral Small (Sep '24)
Mixtral 8x22B Instruct
Mistral Small (Feb '24)
Mistral Large (Feb '24)
Mixtral 8x7B Instruct
Mistral 7B Instruct
Mistral Small 3.1
Mistral Saba
Devstral Medium
Mistral Medium 3
Devstral Small (Jul '25)
Devstral Small (May '25)
Magistral Small 1
Mistral Medium
Magistral Medium 1
DeepSeek R1 Distill Qwen 32B
DeepSeek V3 (Dec '24)
DeepSeek R1 Distill Qwen 14B
DeepSeek-V2.5 (Dec '24)
DeepSeek-Coder-V2
DeepSeek R1 Distill Llama 8B
DeepSeek LLM 67B Chat (V1)
DeepSeek R1 Distill Qwen 1.5B
DeepSeek V3.1 Terminus (Reasoning)
DeepSeek Coder V2 Lite Instruct
DeepSeek R1 (Jan '25)
DeepSeek V3.1 Terminus (Non-reasoning)
DeepSeek V3 0324
DeepSeek V3.2 Exp (Non-reasoning)
DeepSeek-V2.5
DeepSeek V3.1 (Non-reasoning)
DeepSeek V3.1 (Reasoning)
DeepSeek-V2-Chat
DeepSeek V3.2 Exp (Reasoning)
Sonar Pro
Sonar Reasoning Pro
Sonar Reasoning
Sonar
Grok Beta
Grok 3
Grok 4 Fast (Reasoning)
Grok 4 Fast (Non-reasoning)
Grok 3 Reasoning Beta
Grok 2 (Dec '24)
OpenChat 3.5 (1210)
Nova Pro
Nova Lite
Phi-3 Mini Instruct 3.8B
LFM 40B
LFM2 1.2B
Solar Mini
Solar Pro 2 (Preview) (Reasoning)
Solar Pro 2 (Preview) (Non-reasoning)
DBRX Instruct
MiniMax-M2
MiniMax M1 80k
MiniMax M1 40k
Kimi K2
Llama 3.1 Tulu3 405B
OLMo 2 32B
OLMo 2 7B
Olmo 3 32B Think
Granite 3.3 8B (Non-reasoning)
Reka Flash (Sep '24)
Hermes 3 - Llama-3.1 70B
GLM-4.5 (Reasoning)
GLM-4.5V (Reasoning)
GLM-4.6 (Non-reasoning)
GLM-4.6 (Reasoning)
GLM-4.5V (Non-reasoning)
Command-R+ (Apr '24)
Command-R (Mar '24)
Apriel-v1.5-15B-Thinker
Jamba 1.5 Mini
Jamba 1.5 Large
Jamba 1.6 Large
Jamba 1.6 Mini
Arctic Instruct
Qwen2.5 Max
Qwen2.5 Instruct 72B
Qwen2.5 Coder Instruct 32B
Qwen2.5 Turbo
Qwen2 Instruct 72B
Qwen3 8B (Reasoning)
Qwen3 8B (Non-reasoning)
Qwen3 4B (Reasoning)
QwQ 32B-Preview
Qwen2.5 Coder Instruct 7B
QwQ 32B
Qwen2.5 Instruct 32B
Qwen Chat 72B
Qwen1.5 Chat 110B
Qwen3 235B A22B (Reasoning)
Qwen3 32B (Non-reasoning)
Qwen3 30B A3B (Reasoning)
Qwen3 32B (Reasoning)
Qwen3 30B A3B (Non-reasoning)
Qwen3 235B A22B (Non-reasoning)
Qwen3 14B (Reasoning)
Qwen3 4B (Non-reasoning)
Qwen3 14B (Non-reasoning)
Qwen3 Max (Preview)
Seed-OSS-36B-Instruct
MiMo-V2-Flash (Feb 2026)
GLM-5 (Reasoning)
Aurora Alpha
Free Models Router
StepFun: Step 3.5 Flash (free)
Arcee AI: Trinity Large Preview (free)
Upstage: Solar Pro 3 (free)
MiniMax: MiniMax M2-her
Writer: Palmyra X5
LiquidAI: LFM2.5-1.2B-Thinking (free)
LiquidAI: LFM2.5-1.2B-Instruct (free)
OpenAI: GPT Audio
OpenAI: GPT Audio Mini
AllenAI: Molmo2 8B
ByteDance Seed: Seed 1.6 Flash
ByteDance Seed: Seed 1.6
Google: Gemini 3 Flash Preview
Mistral: Mistral Small Creative
OpenAI: GPT-5.2 Chat
OpenAI: GPT-5.2 Pro
Mistral: Devstral 2 2512
Relace: Relace Search
Nex AGI: DeepSeek V3.1 Nex N1
EssentialAI: Rnj 1 Instruct
Body Builder (beta)
OpenAI: GPT-5.1-Codex-Max
Mistral: Ministral 3 14B 2512
Mistral: Ministral 3 8B 2512
Mistral: Ministral 3 3B 2512
Mistral: Mistral Large 3 2512
Arcee AI: Trinity Mini (free)
TNG: R1T Chimera (free)
Google: Nano Banana Pro (Gemini 3 Pro Image Preview)
Google: Gemini 3 Pro Preview
Deep Cogito: Cogito v2.1 671B
OpenAI: GPT-5.1 Chat
Kwaipilot: KAT-Coder-Pro V1
Perplexity: Sonar Pro Search
Mistral: Voxtral Small 24B 2507
OpenAI: gpt-oss-safeguard-20b
LiquidAI: LFM2-2.6B
IBM: Granite 4.0 Micro
OpenAI: GPT-5 Image Mini
Anthropic: Claude Haiku 4.5
Qwen: Qwen3 VL 8B Thinking
OpenAI: GPT-5 Image
OpenAI: o3 Deep Research
OpenAI: o4 Mini Deep Research
NVIDIA: Llama 3.3 Nemotron Super 49B V1.5
Baidu: ERNIE 4.5 21B A3B Thinking
Google: Gemini 2.5 Flash Image (Nano Banana)
Qwen: Qwen3 VL 30B A3B Thinking
OpenAI: GPT-5 Pro
Anthropic: Claude Sonnet 4.5
DeepSeek: DeepSeek V3.2 Exp
TheDrummer: Cydonia 24B V4.1
Relace: Relace Apply 3
Qwen: Qwen3 VL 235B A22B Thinking
Qwen: Qwen3 Coder Plus
Tongyi DeepResearch 30B A3B
Qwen: Qwen3 Coder Flash
OpenGVLab: InternVL3 78B
Qwen: Qwen3 Next 80B A3B Thinking
Meituan: LongCat Flash Chat
Qwen: Qwen Plus 0728
Qwen: Qwen3 30B A3B Thinking 2507
Nous: Hermes 4 70B
Nous: Hermes 4 405B
DeepSeek: DeepSeek V3.1
OpenAI: GPT-4o Audio
Baidu: ERNIE 4.5 21B A3B
Baidu: ERNIE 4.5 VL 28B A3B
AI21: Jamba Large 1.7
OpenAI: GPT-5 Chat
Anthropic: Claude Opus 4.1
Mistral: Codestral 2508
Qwen: Qwen3 30B A3B Instruct 2507
Z.ai: GLM 4.5
Qwen: Qwen3 235B A22B Thinking 2507
Z.ai: GLM 4 32B
Qwen: Qwen3 Coder 480B A35B (free)
ByteDance: UI-TARS 7B
Qwen: Qwen3 235B A22B Instruct 2507
Switchpoint Router
Venice: Uncensored (free)
Google: Gemma 3n 2B (free)
Tencent: Hunyuan A13B Instruct
TNG: DeepSeek R1T2 Chimera (free)
Morph: Morph V3 Large
Morph: Morph V3 Fast
Baidu: ERNIE 4.5 VL 424B A47B
Inception: Mercury
Mistral: Mistral Small 3.2 24B
MiniMax: MiniMax M1
xAI: Grok 3 Mini
Google: Gemini 2.5 Pro Preview 06-05
DeepSeek: R1 0528 (free)
Anthropic: Claude Opus 4
Anthropic: Claude Sonnet 4
Google: Gemma 3n 4B (free)
Google: Gemini 2.5 Pro Preview 05-06
Arcee AI: Spotlight
Arcee AI: Maestro Reasoning
Arcee AI: Virtuoso Large
Arcee AI: Coder Large
Inception: Mercury Coder
Qwen: Qwen3 4B (free)
Meta: Llama Guard 4 12B
Qwen: Qwen3 30B A3B
Qwen: Qwen3 8B
Qwen: Qwen3 14B
Qwen: Qwen3 32B
Qwen: Qwen3 235B A22B
TNG: DeepSeek R1T Chimera (free)
OpenAI: o4 Mini High
EleutherAI: Llemma 7b
AlfredPros: CodeLLaMa 7B Instruct Solidity
xAI: Grok 3 Mini Beta
xAI: Grok 3 Beta
NVIDIA: Llama 3.1 Nemotron Ultra 253B v1
Qwen: Qwen2.5 VL 32B Instruct
DeepSeek: DeepSeek V3 0324
Mistral: Mistral Small 3.1 24B (free)
AllenAI: Olmo 2 32B Instruct
Google: Gemma 3 4B (free)
Google: Gemma 3 12B (free)
OpenAI: GPT-4o-mini Search Preview
OpenAI: GPT-4o Search Preview
Google: Gemma 3 27B (free)
TheDrummer: Skyfall 36B V2
Perplexity: Sonar Deep Research
Llama Guard 3 8B
Google: Gemini 2.0 Flash
Qwen: Qwen VL Plus
AionLabs: Aion-1.0
AionLabs: Aion-1.0-Mini
AionLabs: Aion-RP 1.0 (8B)
Qwen: Qwen VL Max
Qwen: Qwen2.5 VL 72B Instruct
Qwen: Qwen-Plus
Qwen: Qwen-Max
Mistral: Mistral Small 3
MiniMax: MiniMax-01
Sao10K: Llama 3.1 70B Hanami x1
DeepSeek: DeepSeek V3
Sao10K: Llama 3.3 Euryale 70B
Cohere: Command R7B (12-2024)
Meta: Llama 3.3 70B Instruct (free)
OpenAI: GPT-4o (2024-11-20)
Mistral Large 2411
Qwen2.5 Coder 32B Instruct
SorcererLM 8x22B
TheDrummer: UnslopNemo 12B
Magnum v4 72B
Anthropic: Claude 3.5 Sonnet
Qwen: Qwen2.5 7B Instruct
NVIDIA: Llama 3.1 Nemotron 70B Instruct
Inflection: Inflection 3 Pi
Inflection: Inflection 3 Productivity
TheDrummer: Rocinante 12B
Meta: Llama 3.2 3B Instruct (free)
Meta: Llama 3.2 1B Instruct
Meta: Llama 3.2 11B Vision Instruct
Qwen2.5 72B Instruct
NeverSleep: Lumimaid v0.2 8B
Mistral: Pixtral 12B
Cohere: Command R (08-2024)
Cohere: Command R+ (08-2024)
Sao10K: Llama 3.1 Euryale 70B v2.2
Qwen: Qwen2.5-VL 7B Instruct
Nous: Hermes 3 405B Instruct (free)
OpenAI: ChatGPT-4o
Sao10K: Llama 3 8B Lunaris
Meta: Llama 3.1 405B (base)
Meta: Llama 3.1 8B Instruct
Meta: Llama 3.1 405B Instruct
Meta: Llama 3.1 70B Instruct
Mistral: Mistral Nemo
OpenAI: GPT-4o-mini (2024-07-18)
Google: Gemma 2 27B
Google: Gemma 2 9B
Sao10k: Llama 3 Euryale 70B v2.1
NousResearch: Hermes 2 Pro - Llama-3 8B
Mistral: Mistral 7B Instruct v0.3
Meta: LlamaGuard 2 8B
Meta: Llama 3 70B Instruct
Meta: Llama 3 8B Instruct
Mistral: Mixtral 8x22B Instruct
WizardLM-2 8x22B
OpenAI: GPT-4 Turbo Preview
Mistral: Mistral 7B Instruct v0.2
Noromaid 20B
Goliath 120B
Auto Router
OpenAI: GPT-4 Turbo (older v1106)
OpenAI: GPT-3.5 Turbo Instruct
Mistral: Mistral 7B Instruct v0.1
OpenAI: GPT-3.5 Turbo 16k
Mancer: Weaver (alpha)
ReMM SLERP 13B
MythoMax 13B
OpenAI: GPT-4 (older v0314)
OpenAI: GPT-3.5 Turbo
MiniMax-M2.5
GLM-5 (Non-reasoning)
Qwen: Qwen Plus 0728 (thinking)
Qwen: Qwen3 Coder 480B A35B
Qwen: Qwen3 4B
Mistral: Mistral Small 3.1 24B
Meta: Llama 3.3 70B Instruct
Meta: Llama 3.2 3B Instruct
Nous: Hermes 3 405B Instruct
Qwen: Qwen3.5 Plus 2026-02-15
Qwen: Qwen3.5 397B A17B
Qwen3.5 397B A17B (Reasoning)
Claude Sonnet 4.6 (Adaptive Reasoning, Max Effort)
Claude Sonnet 4.6 (Non-reasoning, High Effort)
Amazon: Nova 2 Lite
Amazon: Nova Premier 1.0
Amazon: Nova Lite 1.0
Amazon: Nova Micro 1.0
Amazon: Nova Pro 1.0
Qwen3.5 397B A17B (Non-reasoning)
Gemini 3.1 Pro Preview
Tri-21B-Think
Tri-21B-think Preview
Claude Sonnet 4.6 (Non-reasoning, Low Effort)
Tiny Aya Global
Doubao Seed 2.0 lite (Reasoning)
OpenAI: GPT-5.3-Codex
AionLabs: Aion-2.0
Mercury 2
Google: Gemini 3.1 Pro Preview Custom Tools
Qwen3.5 35B A3B (Reasoning)
Qwen3.5 122B A10B (Reasoning)
Qwen3.5 27B (Reasoning)
Qwen: Qwen3.5-Flash
LiquidAI: LFM2-24B-A2B
ByteDance Seed: Seed-2.0-Mini
Google: Nano Banana 2 (Gemini 3.1 Flash Image Preview)
GPT-5.3 Codex (xhigh)
LFM2 24B A2B
Qwen3.5 122B A10B (Non-reasoning)
Qwen3.5 35B A3B (Non-reasoning)
Qwen3.5 27B (Non-reasoning)
Gemini 3.1 Flash-Lite
Qwen3.5 9B (Reasoning)
Qwen3.5 4B (Reasoning)
Qwen3.5 0.8B (Reasoning)
Qwen3.5 2B (Reasoning)
OpenAI: GPT-5.4 Pro
OpenAI: GPT-5.4
OpenAI: GPT-5.3 Chat
Step 3.5 Flash
GPT-5.4 (xhigh)
GPT-5.4 Pro (xhigh)
Qwen3.5 2B (Non-reasoning)
Qwen3.5 4B (Non-reasoning)
Qwen3.5 9B (Non-reasoning)
LongCat Flash Lite
NVIDIA Nemotron 3 Super 120B A12B (Reasoning)
Qwen3.5 0.8B (Non-reasoning)
Sarvam M (Reasoning)
Grok 4.20 0309 (Non-reasoning)
Grok 4.20 0309 (Reasoning)
Z.ai: GLM 5 Turbo
xAI: Grok 4.20 Multi-Agent Beta
xAI: Grok 4.20 Beta
Hunter Alpha
Healer Alpha
ByteDance Seed: Seed-2.0-Lite
Mistral: Mistral Small 4
OpenAI: GPT-5.4 Nano
OpenAI: GPT-5.4 Mini
MiMo-V2-Pro
Mistral Small 4 (Reasoning)
Mistral Small 4 (Non-reasoning)
MiniMax: MiniMax M2.7
GPT-5.4 (Non-reasoning)
MiniMax-M2.7
Sarvam 30B (high)
Sarvam 105B (high)
Xiaomi: MiMo-V2-Omni
GPT-5.4 mini (xhigh)
GPT-5.4 mini (Non-Reasoning)
GPT-5.4 nano (xhigh)
GPT-5.4 mini (medium)
GPT-5.4 nano (medium)
GPT-5.4 nano (Non-Reasoning)
MiMo-V2-Omni
Nanbeige4.1-3B
Apertus 70B Instruct
Apertus 8B Instruct
GLM-5-Turbo
NVIDIA Nemotron 3 Nano 4B
Reka Edge
Gemini 3 Deep Think
Kwaipilot: KAT-Coder-Pro V2
KAT Coder Pro V2
Nemotron Cascade 2 30B A3B
Google: Lyria 3 Pro Preview
Google: Lyria 3 Clip Preview
Qwen: Qwen3.6 Plus Preview (free)
xAI: Grok 4.20 Multi-Agent
Reka Edge
Z.ai: GLM 5V Turbo
Arcee AI: Trinity Large Thinking
GLM 5V Turbo (Reasoning)
Qwen: Qwen3.6 Plus (free)
Google: Gemma 4 31B
Gemma 4 31B (Reasoning)
Gemma 4 26B A4B (Reasoning)
Qwen3.5 Omni Plus
Gemma 4 E2B (Reasoning)
Gemma 4 E4B (Reasoning)
Gemma 4 E2B (Non-reasoning)
Gemma 4 E4B (Non-reasoning)
Gemma 4 31B (Non-reasoning)
Gemma 4 26B A4B (Non-reasoning)
Solar Pro 3
Nova 2.0 Lite (high)
Muse Spark
MiMo-V2-Omni-0327
Trinity Large Thinking
GLM-5.1 (Non-reasoning)
GLM-5.1 (Reasoning)
Qwen3.6 Plus
Qwen3.5 Omni Flash
Grok 4.20 0309 v2 (Non-reasoning)
Grok 4.20 0309 v2 (Reasoning)
Step 3.5 Flash 2603
JT-MINI
Anthropic: Claude Opus 4.7
Elephant
Anthropic: Claude Opus 4.6 (Fast)
Google: Gemma 4 26B A4B (free)
Meta: Llama Guard 4 12B (free)
Claude Opus 4.7
Claude Mythos Preview
Nanonets OCR-3
Nanonets OCR2+
GLM-OCR
GPT-5.5 (xhigh)
GPT-5.5 (high)
GPT-5.5 (low)
GPT-5.5 (Non-reasoning)
GPT-5.5 (medium)
Claude Opus 4.7 (Non-reasoning, High Effort)
DeepSeek V4 Flash (Reasoning, High Effort)
DeepSeek V4 Pro (Reasoning, High Effort)
DeepSeek V4 Pro (Reasoning, Max Effort)
DeepSeek V4 Flash (Reasoning, Max Effort)
Kimi K2.6
MiMo-V2.5-Pro
MiMo-V2.5
Qwen3.6 27B (Reasoning)
Qwen3.6 35B A3B (Reasoning)
Qwen3.6 35B A3B (Non-reasoning)
Qwen3.6 27B (Non-reasoning)
Qwen3.6 Max Preview
Ling-2.6-1T
Ling 2.6 Flash
Anthropic Claude Haiku Latest
OpenAI GPT Mini Latest
Google Gemini Pro Latest
MoonshotAI Kimi Latest
Google Gemini Flash Latest
Anthropic Claude Sonnet Latest
OpenAI GPT Latest
Qwen: Qwen3.5 Plus 2026-04-20
Qwen: Qwen3.6 Flash
Qwen: Qwen3.6 Max Preview
OpenAI: GPT-5.5 Pro
Tencent: Hy3 preview (free)
Xiaomi: MiMo-V2.5
OpenAI: GPT-5.4 Image 2
Anthropic: Claude Opus Latest
Pareto Code Router
Baidu: Qianfan-OCR-Fast (free)
EXAONE 4.5 33B (Non-reasoning)
EXAONE 4.5 33B
Hy3-preview (Reasoning)
NVIDIA: Nemotron 3 Nano Omni (free)
Poolside: Laguna XS.2 (free)
Poolside: Laguna M.1 (free)
Granite 4.1 8B
Granite 4.1 30B
Granite 4.1 3B
inclusionAI: Ling-2.6-flash
MiniMax: MiniMax M2.5 (free)
Z.ai: GLM 4.6
Qwen: Qwen3 Next 80B A3B Instruct (free)
MoonshotAI: Kimi K2 0905
OpenAI: gpt-oss-120b (free)
OpenAI: gpt-oss-20b (free)
Z.ai: GLM 4.5 Air (free)
Anthropic: Claude 3.7 Sonnet (thinking)
OpenAI: GPT-4o
GPT-5.4 (low)
DeepSeek V4 Flash (Non-reasoning)
DeepSeek V4 Pro (Non-reasoning)
Kimi K2.6 (Non-reasoning)
MiMo-V2.5-Pro (Non-reasoning)
Owl Alpha
GPT-5.5 Pro (xhigh)
Mistral Medium 3.5
Grok 4.3 (high)
Hy3-preview (Non-reasoning)
Nemotron 3 Nano Omni 30B A3B Reasoning
OpenAI: GPT Chat Latest
Microsoft: Phi 4 Mini Instruct
Baidu Qianfan: CoBuddy (free)
Google: Gemini 3.1 Flash Lite
inclusionAI: Ling-2.6-1T
Grok 4.3 (Non-reasoning)
inclusionAI: Ring-2.6-1T (free)
MiniCPM-V 4.6 1.3B
ZAYA1-8B
Anthropic: Claude Opus 4.7 (Fast)
Perceptron: Perceptron Mk1
JT-35B-Flash
DeepSeek: DeepSeek V4 Flash (free)
Ring-2.6-1T
Gemini 3.5 Flash (high)
Qwen3.7 Max
Command A+
xAI: Grok Build 0.1
GPT-5.5 Instant (May 2026)
Grok 4.3 (low)
Gemini 3.5 Flash (minimal)
Grok 4.3 (medium)
MiniCPM5-1B (Non-reasoning)
Gemini 3.5 Flash (medium)
MoonshotAI: Kimi K2.6 (free)
Claude Opus 4.8 (Adaptive Reasoning, Max Effort)
Anthropic: Claude Opus 4.8 (Fast)
StepFun: Step 3.7 Flash
MiniMax: MiniMax M3
Step 3.7 Flash
OpenRouter: Fusion
Qwen3.7 Plus
MiniMax-M3
Nemotron 3 Ultra 550B A55B (Reasoning)
MiniCPM5-1B (Reasoning)
NVIDIA: Nemotron 3.5 Content Safety (free)
Gemma 4 12B (Reasoning)
LFM2.5-8B-A1B
Nex AGI: Nex-N2-Pro (free)
Claude Fable 5 (Adaptive Reasoning, Max Effort, Opus 4.8 Fallback)
North Mini Code
Anthropic: Claude Fable Latest
Gemma 4 12B (Non-reasoning)
HyperNova 60B 2605
MoonshotAI: Kimi K2.7 Code
Z.ai: GLM 5.2
Kimi K2.7 Code
GLM-5.2 (max)
Google: Nano Banana 2 (Gemini 3.1 Flash Image)
Google: Nano Banana Pro (Gemini 3 Pro Image)
Grok Build 0.1 0616
Sakana: Fugu Ultra
Nex-N2-Pro
GPT-5.5 Instant (June 2026)
DiffusionGemma 26B A4B
Claude Sonnet 5 (Adaptive Reasoning, Max Effort)
Claude Sonnet 5 (Adaptive Reasoning, High Effort)
Claude Sonnet 5 (Adaptive Reasoning, Medium Effort)
Claude Sonnet 5 (Adaptive Reasoning, Low Effort)
Claude Sonnet 5 (Adaptive Reasoning, Xhigh Effort)
Google: Nano Banana 2 Lite (Gemini 3.1 Flash Lite Image)
Claude Sonnet 5 (Non-reasoning, High Effort)
Poolside: Laguna XS 2.1 (free)
Tencent: Hy3
Nex AGI: Nex-N2-Mini
AionLabs: Aion-3.0-Mini
AionLabs: Aion-3.0
Grok 4.5 (high)
xAI: Grok Latest
GPT-5.6 Terra (medium)
GPT-5.6 Sol (low)
GPT-5.6 Luna (low)
GPT-5.6 Sol (xhigh)
GPT-5.6 Sol (max)
GPT-5.6 Terra (low)
GPT-5.6 Terra (max)
GPT-5.6 Luna (high)
GPT-5.6 Terra (xhigh)
GPT-5.6 Terra (Non-reasoning)
GPT-5.6 Sol (high)
GPT-5.6 Luna (xhigh)
GPT-5.6 Terra (high)
GPT-5.6 Sol (Non-reasoning)
GPT-5.6 Sol (medium)
GPT-5.6 Luna (Non-reasoning)
GPT-5.6 Luna (medium)
GPT-5.6 Luna (max)
OpenAI: GPT-5.6 Luna Pro
OpenAI: GPT-5.6 Terra Pro
OpenAI: GPT-5.6 Sol Pro
JT-4.1 Flash 236B A21B
Muse Spark 1.1 (xhigh)
Gemini 3.6 Flash (high)
Gemini 3.5 Flash-Lite
Claude Opus 5 (Adaptive Reasoning, Max Effort)
Claude Opus 5 (Adaptive Reasoning, High Effort)
Claude Opus 5 (Adaptive Reasoning, Xhigh Effort)
Claude Opus 5 (Adaptive Reasoning, Low Effort)
Claude Opus 5 (Adaptive Reasoning, Medium Effort)
Kimi K3 (max)
Motif 3 (Beta)
LongCat 2.0
Inkling (xhigh)
G9v3-3B
GLM-5.2 (Non-reasoning)
Agnes 2.5 Pro Alpha
Claude Opus 5 (Fast)
Ling-3.0-flash (free)
Poolside: Laguna S 2.1
Meituan: LongCat 2.0
Auto Router (Beta)
MoonshotAI: Kimi K3
Kwaipilot: KAT-Coder-Air V2.5
Kwaipilot: KAT-Coder-Pro V2.5
Hy3
Qwen: Qwen3.7 Flash
Claude Opus 5 (batch)
Anthropic: Claude Sonnet 5 (batch)
Anthropic: Claude Fable 5 (batch)
Anthropic: Claude Opus 4.8 (batch)
OpenAI: GPT-5.5 (batch)
Anthropic: Claude Opus 4.6 (batch)
OpenAI: GPT-5.2 (batch)
Anthropic: Claude Opus 4.5 (batch)
OpenAI: GPT-5.1 (batch)
OpenAI: GPT-5 (batch)
OpenAI: GPT-5 Mini (batch)
OpenAI: GPT-5 Nano (batch)
Google: Gemini 3.6 Flash (batch)
Google: Gemini 3.5 Flash Lite (batch)
Google: Gemini 3.5 Flash (batch)
Google: Gemini 3.1 Pro Preview (batch)
Google: Gemini 2.5 Flash Lite (batch)
Google: Gemini 2.5 Flash (batch)
Google: Gemini 2.5 Pro (batch)
Claude Opus 4.7 (Adaptive Reasoning, Max Effort)
Inkling Small
Kimi K3 (low)
DeepSeek V4 Flash 0731 (Reasoning, Max Effort)
Celeris-1
DeepSeek: DeepSeek V4 Flash 0731
Constraints & coverage

Live methodology

1 Source Benchmark

Snapshot 7/31/2026, 11:33:21 PM
Intelligence Index × 10–100 Composite Score

Result · updates immediately

Leaderboard

614 ranked · 331 excluded

Selected claude-opus-5, rank 1, Composite Score 100.0.

614 ranked models and 331 excluded models. Select a model name to inspect its score evidence.
ChangeContextDecision
1—Anthropic100.0100%0%1056.21Mpareto
2—xhigh reasoningAnthropic99.8100%0%1054.7—dominated
3—Anthropic99.7100%0%2062.71Mdominated
4—high reasoningAnthropic99.4100%0%1052.3—dominated
4—max reasoningOpenAI99.4100%0%11.369.31.1Mdominated
6—xhigh reasoningOpenAI99.2100%0%11.362.8—dominated
7—max reasoningKimi99.0100%0%633.1—pareto
8—medium reasoningAnthropic98.9100%0%1055.2—dominated
9—high reasoningOpenAI98.7100%0%11.359.2—dominated
10—Anthropic98.5100%0%10—1Mdominated
11—max reasoningOpenAI98.4100%0%4.5139.11.1Mpareto
12—xhigh reasoningOpenAI98.2100%0%11.3—1.1Mdominated
13—high reasoningSpaceXAI98.0100%0%355500Kpareto
14—medium reasoningOpenAI97.8100%0%11.356.3—dominated
14—OpenAI97.8100%0%4.890.5400Kdominated
16—Anthropic97.6100%0%10——dominated
17—Anthropic97.4100%0%479.21Mdominated
18—high reasoningOpenAI97.2100%0%11.3——dominated
19—Alibaba97.1100%0%2.933.7262.1Kpareto
20—xhigh reasoningOpenAI96.9100%0%4.5109.7—dominated
21—xhigh reasoningOpenAI96.7100%0%5.6—1.1Mdominated
22—max reasoningOpenAI96.6100%0%0.5174.21.1Mpareto
23—max reasoningZ AI96.4100%0%2.2107.3—dominated
24—low reasoningAnthropic96.2100%0%1051.9—dominated
24—xhigh reasoningMeta96.2100%0%2146.41.0Mdominated
26—medium reasoningOpenAI95.9100%0%11.3——dominated
27—high reasoningGoogle95.8100%0%3.4220.61.0Mdominated
28—high reasoningGoogle95.6100%0%3220.41.0Mdominated
29—DeepSeek95.4100%0%0.2——pareto
30—MiniMax95.3100%0%0.546204.8Kdominated
31—low reasoningOpenAI95.1100%0%11.357.7—dominated
32—xhigh reasoningOpenAI94.9100%0%0.5174—dominated
33—high reasoningOpenAI94.7100%0%4.5116.3—dominated
33—Xiaomi94.7100%0%0.8—1.0Mdominated
35—OpenAI94.5100%0%1.7174.4400Kdominated
36—Anthropic94.3100%0%6——dominated
37—Z AI94.1100%0%1.5—202.8Kdominated
38—low reasoningKimi94.0100%0%633.7—dominated
39—Google93.8100%0%4.5135.61.0Mdominated
40—high reasoningOpenAI93.6100%0%0.5172.1—dominated
41—max reasoningAlibaba93.5100%0%3.8200.21Mdominated
42—medium reasoningOpenAI93.3100%0%4.598.4—dominated
43—medium reasoningGoogle93.1100%0%3.4240.9—dominated
44—Alibaba93.0100%0%1.350.9262.1Kdominated
45—MiniMax92.8100%0%0.574.1524.3Kdominated
46—DeepSeek92.6100%0%0.561.41.0Mdominated
46—xhigh reasoningOpenAI92.6100%0%4.8109.1—dominated
48—Kimi92.3100%0%1.7—256Kdominated
49—Motif Technologies92.2100%0%0——pareto
50—OpenAI92.0100%0%0.5163.2400Kdominated
51—KwaiKAT91.8100%0%0.5111.5256Kdominated
52—Anthropic91.7100%0%10——dominated
53—low reasoningOpenAI91.5100%0%11.3——dominated
54—Xiaomi91.4100%0%0.8—262.1Kdominated
55—high reasoningDeepSeek91.1100%0%0.555.3—dominated
55—Meta91.1100%0%0——dominated
57—Z AI90.9100%0%1.9—202.8Kdominated
58—non-reasoning reasoningAnthropic90.7100%0%10——dominated
59—xhigh reasoningOpenAI90.5100%0%4.8—400Kdominated
59—Xiaomi90.5100%0%0.546.51.0Mdominated
61—Kimi90.1100%0%1.739.6—dominated
61—Tencent90.1100%0%086.4262.1Kdominated
63—non-reasoning reasoningAnthropic89.9100%0%465—dominated
64—non-reasoning reasoningOpenAI89.6100%0%11.360.2—dominated
64—Tencent89.6100%0%0.272.2—dominated
66—Nex AGI89.4100%0%1130.7—dominated
67—Anthropic89.2100%0%10——dominated
68—xhigh reasoningThinking Machines89.1100%0%2.679.61.0Mdominated
69—low reasoningOpenAI88.9100%0%4.5109.6—dominated
70—DeepSeek88.7100%0%0.2106.71.0Mdominated
70—Xiaomi88.7100%0%0—1.0Mdominated
72—Z AI88.3100%0%2.1—202.8Kdominated
72—Thinking Machines88.3100%0%0.595.2524.3Kdominated
74—xhigh reasoningOpenAI88.1100%0%4.8—400Kdominated
75—xhigh reasoningOpenAI87.8100%0%1.7—400Kdominated
75—max reasoningAlibaba87.8100%0%2.9——dominated
77—SpaceXAI87.6100%0%1.3——dominated
78—high reasoningGoogle87.4100%0%4.5——dominated
78—Alibaba87.4100%0%1.1551Mdominated
80—Z AI87.1100%0%1.6—204.8Kdominated
81—low reasoningOpenAI86.9100%0%5.6——dominated
82—Alibaba86.8100%0%0.751.71Mdominated
83—Sapiens AI86.5100%0%0.6127.2—dominated
83—China Mobile86.5100%0%0——dominated
85—xhigh reasoningOpenAI86.3100%0%0.5—400Kdominated
86—Z AI86.0100%0%0——dominated
86—medium reasoningOpenAI86.0100%0%0.5158.4—dominated
86—MiniMax86.0100%0%0.5——dominated
89—medium reasoningOpenAI85.6100%0%4.8——dominated
90—Anthropic85.3100%0%10—1Mdominated
90—Google85.3100%0%1.1——dominated
90—NVIDIA85.3100%0%1.2222.31Mdominated
93—high reasoningSpaceXAI85.0100%0%1.6—1Mdominated
94—high reasoningDeepSeek84.8100%0%0.2——dominated
95—Xiaomi84.7100%0%0.276.4—dominated
96—Anthropic84.4100%0%649.21Mdominated
96—Alibaba84.4100%0%1.458262.1Kdominated
98—SpaceXAI84.2100%0%3——dominated
99—high reasoningOpenAI84.0100%0%3.4—400Kdominated
100—Google83.8100%0%0.9351.21.0Mdominated
100—SpaceXAI83.8100%0%3—2Mdominated
102—Anthropic83.4100%0%6——dominated
102—Xiaomi83.4100%0%0——dominated
104—ByteDance Seed83.2100%0%0——dominated
105—high reasoningOpenAI83.0100%0%3.4—400Kdominated
106—Anthropic82.8100%0%3038.1200Kdominated
106—medium reasoningSpaceXAI82.8100%0%1.6129.8—dominated
108—Anthropic82.5100%0%6—1Mdominated
109—non-reasoning reasoningZ AI82.1100%0%2.1——dominated
109—non-reasoning reasoningOpenAI82.1100%0%11.3——dominated
109—low reasoningSpaceXAI82.1100%0%1.6137.5—dominated
109—Kimi82.1100%0%1.2—262.1Kdominated
113—Google81.6100%0%1.1179.31.0Mdominated
113—Xiaomi81.6100%0%0——dominated
115—minimal reasoningGoogle81.4100%0%3.4223.3—dominated
116—non-reasoning reasoningAnthropic81.1100%0%10—200Kdominated
116—high reasoningOpenAI81.1100%0%3.4—400Kdominated
116—high reasoningOpenAI81.1100%0%3.4—400Kdominated
119—non-reasoning reasoningKimi80.8100%0%1.737.8—dominated
120—Z AI80.6100%0%0——dominated
121—Anthropic80.4100%0%648.1—dominated
122—non-reasoning reasoningZ AI80.3100%0%2.164.3—dominated
123—non-reasoning reasoningOpenAI80.1100%0%4.5102.9—dominated
124—Alibaba79.9100%0%0.8—262.1Kdominated
125—Anthropic79.4100%0%30——dominated
125—Z AI79.4100%0%1—202.8Kdominated
125—medium reasoningOpenAI79.4100%0%3.4——dominated
125—KwaiKAT79.4100%0%0.5104.6—dominated
125—MiniMax79.4100%0%0.5—196.6Kdominated
125—Alibaba79.4100%0%1.464.9—dominated
131—Tencent78.8100%0%0.1163.5262.1Kdominated
132—OpenAI78.5100%0%11.3——dominated
132—LongCat78.5100%0%1.345.5—dominated
134—low reasoningOpenAI78.2100%0%0.5162.1—dominated
134—SpaceXAI78.2100%0%6—256Kdominated
136—Xiaomi78.0100%0%0——dominated
137—low reasoningGoogle77.8100%0%4.5——dominated
138—Anthropic77.6100%0%3037.9200Kdominated
138—Anthropic77.6100%0%649.41Mdominated
140—Kimi77.3100%0%1.1—262.1Kdominated
141—OpenAI77.2100%0%35—200Kdominated
142—non-reasoning reasoningZ AI77.0100%0%1.6——dominated
143—Alibaba76.8100%0%1.1137.9262.1Kdominated
144—DeepSeek76.6100%0%0.3——dominated
144—non-reasoning reasoningAlibaba76.6100%0%1.464.3—dominated
146—Alibaba76.3100%0%1.6—262.1Kdominated
147—Alibaba76.2100%0%0.6131.1262.1Kdominated
148—MiniMax76.0100%0%0.5—196.6Kdominated
149—non-reasoning reasoningDeepSeek75.7100%0%0.563.8—dominated
149—low reasoningOpenAI75.7100%0%3.4——dominated
149—Xiaomi75.7100%0%0.2——dominated
152—Anthropic75.4100%0%2109.9200Kdominated
153—Anthropic75.2100%0%30——dominated
154—medium reasoningOpenAI75.0100%0%0.7——dominated
155—high reasoningOpenAI74.6100%0%0.7—400Kdominated
155—SpaceXAI74.6100%0%0——dominated
155—Alibaba74.6100%0%1.553—dominated
155—InclusionAI74.6100%0%0.9125.7262.1Kdominated
159—non-reasoning reasoningAlibaba74.2100%0%1.457.6—dominated
160—DeepSeek74.0100%0%1.9——dominated
160—OpenAI74.0100%0%3.5142.3200Kdominated
162—StepFun73.7100%0%0.4394.3—dominated
163—medium reasoningOpenAI73.6100%0%0.5——dominated
164—Mistral73.4100%0%372.9262.1Kdominated
165—medium reasoningOpenAI73.2100%0%1.7——dominated
166—Anthropic73.1100%0%2140.1—dominated
167—Google72.8100%0%035.4—dominated
167—non-reasoning reasoningKimi72.8100%0%1.2——dominated
169—non-reasoning reasoningAnthropic72.4100%0%6——dominated
169—non-reasoning reasoningAlibaba72.4100%0%0.8——dominated
169—Alibaba72.4100%0%0.7—262.1Kdominated
172—Anthropic72.0100%0%6——dominated
172—OpenAI72.0100%0%11.3——dominated
174—non-reasoning reasoningDeepSeek71.7100%0%0.2106.4—dominated
174—Z AI71.7100%0%1——dominated
176—China Mobile71.5100%0%0——dominated
177—KwaiKAT71.2100%0%0——dominated
177—MiniMax71.2100%0%0.5—196.6Kdominated
179—non-reasoning reasoningAnthropic71.0100%0%30——dominated
180—non-reasoning reasoningXiaomi70.8100%0%0.546.3—dominated
181—non-reasoning reasoningOpenAI70.6100%0%5.6——dominated
182—non-reasoning reasoningAlibaba70.5100%0%1.1153.2—dominated
183—non-reasoning reasoningGoogle70.2100%0%1.1——dominated
183—SpaceXAI70.2100%0%0.3——dominated
185—Anthropic70.0100%0%0——dominated
186—non-reasoning reasoningZ AI69.7100%0%1——dominated
186—non-reasoning reasoningOpenAI69.7100%0%0.5165.6—dominated
188—Z AI69.5100%0%0.648.9131.1Kdominated
189—non-reasoning reasoningTencent69.2100%0%0.1163.6—dominated
189—InclusionAI69.2100%0%0.9—262.1Kdominated
191—ByteDance Seed68.8100%0%0——dominated
191—non-reasoning reasoningOpenAI68.8100%0%4.8——dominated
191—StepFun68.8100%0%0.2——dominated
194—Google68.5100%0%3.4131.31.0Mdominated
195—Google68.4100%0%0.2—262.1Kdominated
196—high reasoningOpenAI68.2100%0%1.9—200Kdominated
197—non-reasoning reasoningAnthropic67.9100%0%30——dominated
197—non-reasoning reasoningAnthropic67.9100%0%6——dominated
197—StepFun67.9100%0%0.2—256Kdominated
200—DeepSeek67.5100%0%0.3——dominated
200—NVIDIA67.5100%0%0.4146.6262.1Kdominated
202—high reasoningOpenAI67.2100%0%0.7—400Kdominated
203—Google67.0100%0%0.6302.41.0Mdominated
203—Alibaba67.0100%0%2.4——dominated
205—non-reasoning reasoningSpaceXAI66.7100%0%1.6119.8—dominated
206—non-reasoning reasoningDeepSeek66.5100%0%0.3—163.8Kdominated
206—non-reasoning reasoningXiaomi66.5100%0%0—262.1Kdominated
208—non-reasoning reasoningAlibaba66.2100%0%0.8101.2—dominated
209—non-reasoning reasoningAlibaba66.0100%0%0.7175.9—dominated
209—max reasoningAlibaba66.0100%0%2.4—262.1Kdominated
211—Google65.7100%0%0——dominated
211—high reasoningOpenAI65.7100%0%0.3193.6131.1Kdominated
213—non-reasoning reasoningAnthropic65.4100%0%2103.8—dominated
214—non-reasoning reasoningAnthropic65.2100%0%6—200Kdominated
214—Kimi65.2100%0%1.1—131.1Kdominated
216—OpenAI64.9100%0%26.3—200Kdominated
217—Google64.7100%0%0——dominated
217—non-reasoning reasoningZ AI64.7100%0%1—202.8Kdominated
219—Z AI64.4100%0%0.2—202.8Kdominated
220—Cohere64.1100%0%0195.6—dominated
220—high reasoningSpaceXAI64.1100%0%0.4——dominated
220—non-reasoning reasoningSpaceXAI64.1100%0%3——dominated
223—Google63.8100%0%3.4——dominated
224—DeepSeek63.6100%0%0—163.8Kdominated
225—LG AI Research63.5100%0%0——dominated
226—Baidu63.3100%0%0——dominated
227—Google62.9100%0%0.229.5—dominated
227—non-reasoning reasoningGoogle62.9100%0%0.275.2—dominated
227—non-reasoning reasoningSpaceXAI62.9100%0%3——dominated
227—medium reasoningAmazon62.9100%0%3.4138.1—dominated
231—SpaceXAI62.5100%0%0—256Kdominated
232—non-reasoning reasoningDeepSeek62.2100%0%0.5—163.8Kdominated
232—Inception62.2100%0%0.4904.6128Kdominated
232—Alibaba62.2100%0%0.266.9256Kdominated
235—non-reasoning reasoningDeepSeek61.8100%0%0.3——dominated
236—ServiceNow61.7100%0%0——dominated
237—Alibaba61.5100%0%0.6126.8262.1Kdominated
238—non-reasoning reasoningDeepSeek61.3100%0%0.8——dominated
239—medium reasoningAmazon61.2100%0%0.9——dominated
240—DeepSeek61.0100%0%0.9——dominated
241—Alibaba60.8100%0%2.6——dominated
242—ServiceNow60.7100%0%0——dominated
243—non-reasoning reasoningOpenAI60.5100%0%3.4——dominated
244—non-reasoning reasoningAlibaba60.4100%0%0——dominated
245—LG AI Research60.2100%0%0——dominated
246—DeepSeek59.8100%0%2.1—64Kdominated
246—Google59.8100%0%0.9——dominated
246—non-reasoning reasoningGoogle59.8100%0%0.255.9—dominated
246—Alibaba59.8100%0%0.121.7—dominated
250—high reasoningOpenAI59.4100%0%0.1—400Kdominated
251—Cohere59.2100%0%021.9256Kdominated
252—Mistral58.9100%0%0.3149.9—dominated
252—low reasoningAmazon58.9100%0%3.4150.4—dominated
252—Alibaba58.9100%0%2.6——dominated
255—Z AI58.6100%0%0——dominated
256—OpenAI58.3100%0%3.5—1.0Mdominated
256—Kimi58.3100%0%1—131.1Kdominated
258—Mistral58.0100%0%053.2—dominated
258—Alibaba58.0100%0%2.4——dominated
260—medium reasoningOpenAI57.5100%0%0.1——dominated
260—medium reasoningAmazon57.5100%0%0.9188.4—dominated
260—OpenAI57.5100%0%1.9—200Kdominated
260—Alibaba57.5100%0%0.3246.6—dominated
264—OpenAI57.1100%0%262.5—200Kdominated
265—non-reasoning reasoningGoogle56.9100%0%0—1.0Mdominated
266—DeepSeek56.7100%0%2.4——dominated
266—China Mobile56.7100%0%0——dominated
268—SpaceXAI56.4100%0%8—131.1Kdominated
269—ByteDance Seed56.3100%0%0.3——dominated
270—high reasoningAmazon56.0100%0%0.9190.3—dominated
270—Alibaba56.0100%0%1.2——dominated
270—Arcee AI56.0100%0%0.4209.5—dominated
273—Alibaba55.6100%0%3——dominated
274—Mistral55.4100%0%2.845.7—dominated
274—Alibaba55.4100%0%2.6——dominated
276—Multiverse Computing55.0100%0%0.1418—dominated
276—low reasoningAmazon55.0100%0%0.9184.1—dominated
276—Perplexity55.0100%0%3.5—128Kdominated
279—MiniMax54.6100%0%1——dominated
280—non-reasoning reasoningOpenAI54.4100%0%0.5——dominated
280—NVIDIA54.4100%0%0——dominated
282—Google54.2100%0%0——dominated
283—Mistral54.0100%0%049.4—dominated
284—MBZUAI Institute of Foundation Models53.8100%0%0——dominated
285—minimal reasoningOpenAI53.6100%0%3.4——dominated
285—LongCat53.6100%0%0——dominated
287—Naver53.3100%0%0——dominated
287—OpenAI53.3100%0%28.9——dominated
289—non-reasoning reasoningSpaceXAI53.0100%0%0—2Mdominated
290—Z AI52.9100%0%0.5——dominated
291—non-reasoning reasoningLG AI Research52.6100%0%0——dominated
291—Alibaba52.6100%0%1.9202.9—dominated
293—non-reasoning reasoningOpenAI52.3100%0%1.7——dominated
293—low reasoningAmazon52.3100%0%0.9——dominated
295—DeepSeek51.9100%0%0.5—163.8Kdominated
295—Z AI51.9100%0%0.4—131.1Kdominated
295—non-reasoning reasoningSpaceXAI51.9100%0%0.3—2Mdominated
298—Korea Telecom51.5100%0%0——dominated
299—InclusionAI51.4100%0%0——dominated
300—AI9Stars51.2100%0%0——dominated
301—non-reasoning reasoningAlibaba51.1100%0%0.117.2—dominated
302—Mistral50.9100%0%0.863.1—dominated
303—Prime Intellect50.7100%0%0—131.1Kdominated
303—high reasoningOpenAI50.7100%0%1.9—200Kdominated
305—non-reasoning reasoningZ AI50.4100%0%0.2——dominated
306—DeepSeek50.2100%0%0.5——dominated
307—OpenAI50.1100%0%0——dominated
308—Google49.8100%0%0.2——dominated
308—SpaceXAI49.8100%0%0——dominated
308—Upstage49.8100%0%0——dominated
311—Alibaba49.4100%0%0.171262.1Kdominated
312—low reasoningOpenAI49.1100%0%0.3225—dominated
312—high reasoningOpenAI49.1100%0%0.1182.5131.1Kdominated
312—NVIDIA49.1100%0%0.1319.3—dominated
315—OpenAI48.8100%0%0.7—1.0Mdominated
316—Mistral48.5100%0%0.8—131.1Kdominated
316—Mistral48.5100%0%0.2——dominated
318—Meta48.2100%0%0.290.3131.1Kdominated
318—Meta48.2100%0%090.3128Kdominated
320—MiniMax47.8100%0%0——dominated
320—non-reasoning reasoningAmazon47.8100%0%3.4124.2—dominated
320—Alibaba47.8100%0%0.8——dominated
323—minimal reasoningOpenAI47.2100%0%0.7——dominated
323—low reasoningOpenAI47.2100%0%0.1198.5—dominated
323—Meta47.2100%0%0.4104.81.0Mdominated
323—Alibaba47.2100%0%1.2—262.1Kdominated
327—DeepSeek46.7100%0%0.5——dominated
327—high reasoningMBZUAI Institute of Foundation Models46.7100%0%0——dominated
327—NVIDIA46.7100%0%0.1259.4—dominated
330—non-reasoning reasoningGoogle46.1100%0%0.9—1.0Mdominated
330—InclusionAI46.1100%0%0.282.8262.1Kdominated
330—Upstage46.1100%0%0.315128Kdominated
330—Zyphra46.1100%0%0——dominated
334—OpenAI45.7100%0%0——dominated
335—Alibaba45.5100%0%0.9194.4262.1Kdominated
336—OpenAI45.2100%0%0——dominated
336—Alibaba45.2100%0%0.9—160Kdominated
336—Trillion Labs45.2100%0%0——dominated
339—Google44.9100%0%0——dominated
340—NVIDIA44.5100%0%1.236.6131.1Kdominated
340—Alibaba44.5100%0%2.6——dominated
340—Alibaba44.5100%0%0.7—32.8Kdominated
343—Google44.1100%0%0——dominated
343—Alibaba44.1100%0%0.8——dominated
345—non-reasoning reasoningGoogle43.9100%0%0.229.7—dominated
346—non-reasoning reasoningGoogle43.7100%0%0.2—1.0Mdominated
347—InclusionAI43.5100%0%0——dominated
347—Motif Technologies43.5100%0%0——dominated
349—Amazon43.2100%0%565.3—dominated
350—medium reasoningMistral42.8100%0%0——dominated
350—Meta42.8100%0%0.431.2131.1Kdominated
350—Mistral42.8100%0%0.8—131.1Kdominated
350—Upstage42.8100%0%0——dominated
354—medium reasoningMistral42.1100%0%0—131.1Kdominated
354—medium reasoningMBZUAI Institute of Foundation Models42.1100%0%0——dominated
354—NVIDIA42.1100%0%0.493.4—dominated
354—non-reasoning reasoningMistral42.1100%0%0.3145.3—dominated
354—Trillion Labs42.1100%0%0——dominated
359—Anthropic41.4100%0%0—200Kdominated
359—Google41.4100%0%0——dominated
359—OpenAI41.4100%0%0——dominated
362—Google41.0100%0%0105.1—dominated
362—NVIDIA41.0100%0%0——dominated
364—OpenBMB40.7100%0%0——dominated
364—Alibaba40.7100%0%0——dominated
366—high reasoningSarvam40.5100%0%0.1——dominated
367—Celeris39.9100%0%32,116.3—dominated
367—Anthropic39.9100%0%30——dominated
367—Mistral39.9100%0%0——dominated
367—Google39.9100%0%0——dominated
367—Meta39.9100%0%0169.216.4Kdominated
367—non-reasoning reasoningAmazon39.9100%0%0.9214.7—dominated
373—non-reasoning reasoningGoogle39.2100%0%0——dominated
373—non-reasoning reasoningOpenBMB39.2100%0%0——dominated
373—Perplexity39.2100%0%0——dominated
376—Alibaba38.8100%0%2.6——dominated
377—Google38.7100%0%0.2——dominated
378—Mistral38.5100%0%0.894.6—dominated
379—OpenAI38.3100%0%4.4—128Kdominated
380—Mistral38.0100%0%0.281.3—dominated
380—Nanbeige38.0100%0%0——dominated
380—Alibaba38.0100%0%1.2—131.1Kdominated
383—DeepSeek37.6100%0%0—32.8Kdominated
383—non-reasoning reasoningZ AI37.6100%0%0.4—131.1Kdominated
385—AI21 Labs37.3100%0%3.557.4256Kdominated
385—non-reasoning reasoningAlibaba37.3100%0%1.2——dominated
387—Google36.9100%0%0——dominated
387—Mistral36.9100%0%0——dominated
389—LG AI Research36.5100%0%0——dominated
389—Mistral36.5100%0%0.2——dominated
389—Alibaba36.5100%0%0.7——dominated
392—non-reasoning reasoningAmazon36.2100%0%0.9——dominated
393—DeepSeek36.0100%0%0——dominated
393—Alibaba36.0100%0%1.3——dominated
395—max reasoningAlibaba35.7100%0%0——dominated
396—Google35.2100%0%0——dominated
396—Nous Research35.2100%0%0.291.2—dominated
396—Meta35.2100%0%0.382327.7Kdominated
396—Alibaba35.2100%0%0.4—131.1Kdominated
396—non-reasoning reasoningUpstage35.2100%0%0——dominated
401—Anthropic34.7100%0%6——dominated
401—DeepSeek34.7100%0%0.8—131.1Kdominated
403—DeepSeek34.3100%0%0——dominated
403—TII UAE34.3100%0%0——dominated
403—Mistral34.3100%0%3—65.5Kdominated
406—Google33.7100%0%0——dominated
406—InclusionAI33.7100%0%0.2——dominated
406—Meta33.7100%0%052.1131.1Kdominated
406—Meta33.7100%0%052.1131.1Kdominated
410—OpenAI33.0100%0%0.2—1.0Mdominated
410—OpenAI33.0100%0%4.4—128Kdominated
410—Alibaba33.0100%0%0.5——dominated
410—Alibaba33.0100%0%0.4107.5—dominated
414—Perplexity32.5100%0%1—127.1Kdominated
414—StepFun32.5100%0%0——dominated
416—Meta32.3100%0%0.675.5—dominated
417—Mistral32.0100%0%0—131.1Kdominated
417—Alibaba32.0100%0%0.8——dominated
417—Perplexity32.0100%0%6—200Kdominated
420—Alibaba31.6100%0%0——dominated
421—Z AI31.2100%0%0.9——dominated
421—NVIDIA31.2100%0%0.951.4—dominated
421—Mistral31.2100%0%0——dominated
421—Alibaba31.2100%0%0.4——dominated
425—Baidu30.4100%0%0.5—123Kdominated
425—Nous Research30.4100%0%1.536.3—dominated
425—Mistral30.4100%0%0.291.5—dominated
425—NVIDIA30.4100%0%0.379.4—dominated
425—OpenAI30.4100%0%0.8103.616.4Kdominated
425—Upstage30.4100%0%0——dominated
431—non-reasoning reasoningGoogle29.7100%0%0102.9—dominated
431—IBM29.7100%0%0——dominated
431—Meta29.7100%0%0.6468.2Kdominated
434—Google29.2100%0%0—1.0Mdominated
434—non-reasoning reasoningNous Research29.2100%0%1.536.2—dominated
434—NVIDIA29.2100%0%0.1245.7—dominated
437—non-reasoning reasoningNVIDIA28.7100%0%0.499.9—dominated
437—Meta28.7100%0%077.4131.1Kdominated
437—NVIDIA28.7100%0%0——dominated
440—Google28.1100%0%0——dominated
440—OpenAI28.1100%0%7.5—128Kdominated
440—low reasoningMBZUAI Institute of Foundation Models28.1100%0%0——dominated
440—non-reasoning reasoningAlibaba28.1100%0%1.2——dominated
444—Kimi27.5100%0%0——dominated
444—Meta27.5100%0%4.4——dominated
444—NVIDIA27.5100%0%0——dominated
444—non-reasoning reasoningNVIDIA27.5100%0%0——dominated
448—Alibaba27.0100%0%0——dominated
448—Alibaba27.0100%0%0.3—131.1Kdominated
450—Anthropic26.5100%0%6——dominated
450—Liquid AI26.5100%0%0342—dominated
450—Alibaba26.5100%0%0.7——dominated
450—Allen Institute for AI26.5100%0%0——dominated
454—OpenAI26.0100%0%0——dominated
454—InclusionAI26.0100%0%0.2——dominated
456—Allen Institute for AI25.7100%0%0—65.5Kdominated
456—Mistral25.7100%0%0—131.1Kdominated
458—Google25.3100%0%0——dominated
458—minimal reasoningOpenAI25.3100%0%0.1——dominated
458—SpaceXAI25.3100%0%0——dominated
461—OpenAI24.9100%0%15—128Kdominated
461—Alibaba24.9100%0%0——dominated
463—non-reasoning reasoningUpstage24.6100%0%0——dominated
464—Cohere24.4100%0%4.453.6256Kdominated
464—Amazon24.4100%0%1.4——dominated
466—Meta24.0100%0%0.1——dominated
466—NVIDIA24.0100%0%1.278.3—dominated
466—Alibaba24.0100%0%0——dominated
469—SpaceXAI23.6100%0%0——dominated
469—Alibaba23.6100%0%0——dominated
471—Google23.2100%0%0——dominated
471—non-reasoning reasoningNVIDIA23.2100%0%0.1218256Kdominated
471—non-reasoning reasoningNVIDIA23.2100%0%0.1176.5128Kdominated
474—Mistral22.8100%0%3—131.1Kdominated
475—Alibaba22.6100%0%0——dominated
475—Alibaba22.6100%0%0——dominated
477—non-reasoning reasoningZ AI22.2100%0%0.9—65.5Kdominated
477—OpenAI22.2100%0%37.5—8.2Kdominated
477—non-reasoning reasoningAlibaba22.2100%0%0.6——dominated
480—non-reasoning reasoningGoogle21.5100%0%0.2—1.0Mdominated
480—OpenAI21.5100%0%0.3—128Kdominated
480—non-reasoning reasoningNous Research21.5100%0%0.286—dominated
480—Mistral21.5100%0%0.2——dominated
480—Amazon21.5100%0%0.1——dominated
485—DeepSeek20.7100%0%0——dominated
485—Meta20.7100%0%0.6——dominated
485—Mistral20.7100%0%0.1295.3—dominated
485—non-reasoning reasoningAlibaba20.7100%0%0.4——dominated
485—non-reasoning reasoningAlibaba20.7100%0%0——dominated
490—IBM20.2100%0%0.1116.2131.1Kdominated
491—DeepSeek19.9100%0%0——dominated
491—Google19.9100%0%0——dominated
491—high reasoningSarvam19.9100%0%0——dominated
494—Allen Institute for AI19.6100%0%0—65.5Kdominated
495—DeepSeek19.1100%0%0——dominated
495—non-reasoning reasoningGoogle19.1100%0%0——dominated
495—Meta19.1100%0%082.28.2Kdominated
495—Mistral19.1100%0%0.3—32.8Kdominated
495—Allen Institute for AI19.1100%0%0.2—65.5Kdominated
500—Google18.3100%0%0——dominated
500—Meta18.3100%0%0.1139.860Kdominated
500—Alibaba18.3100%0%0.1—131.1Kdominated
500—Perplexity18.3100%0%0——dominated
500—Reka AI18.3100%0%0.4——dominated
505—Meta17.7100%0%0——dominated
505—Upstage17.7100%0%0.2——dominated
507—non-reasoning reasoningLG AI Research17.2100%0%0——dominated
507—SpaceXAI17.2100%0%0——dominated
507—Microsoft17.2100%0%044.9—dominated
507—Alibaba17.2100%0%0——dominated
511—non-reasoning reasoningAlibaba16.8100%0%0——dominated
512—Google16.5100%0%0——dominated
512—Google16.5100%0%0——dominated
512—Alibaba16.5100%0%0——dominated
515—non-reasoning reasoningNous Research16.1100%0%0—32.8Kdominated
515—AI21 Labs16.1100%0%3.556.2—dominated
517—IBM15.8100%0%0.1448.9—dominated
518—DeepSeek15.3100%0%0——dominated
518—Nous Research15.3100%0%0.7—65.5Kdominated
518—AI21 Labs15.3100%0%3.5——dominated
518—non-reasoning reasoningAlibaba15.3100%0%0.3——dominated
518—Alibaba15.3100%0%0.4107.9—dominated
523—AI21 Labs14.8100%0%3.5——dominated
523—Allen Institute for AI14.8100%0%0——dominated
525—Google14.4100%0%0——dominated
525—Liquid AI14.4100%0%0——dominated
525—Microsoft14.4100%0%0.241.916.4Kdominated
528—Anthropic13.8100%0%6——dominated
528—IBM13.8100%0%0——dominated
528—Mistral13.8100%0%0.3——dominated
528—Amazon13.8100%0%0.1265.4—dominated
532—Google13.1100%0%0——dominated
532—Google13.1100%0%0——dominated
532—non-reasoning reasoningNVIDIA13.1100%0%0.3159.7128Kdominated
532—Microsoft13.1100%0%0——dominated
536—Microsoft12.6100%0%026.6—dominated
536—Alibaba12.6100%0%0—32.8Kdominated
538—Mistral12.3100%0%0——dominated
538—Mistral12.3100%0%6—128Kdominated
540—Meta12.1100%0%0.1——dominated
541—Meta11.8100%0%0——dominated
541—OpenBMB11.8100%0%0——dominated
543—AI21 Labs11.3100%0%0——dominated
543—Alibaba11.3100%0%0——dominated
543—Alibaba11.3100%0%0——dominated
543—Reka AI11.3100%0%0.496.965.5Kdominated
547—Allen Institute for AI10.9100%0%0—65.5Kdominated
548—Anthropic10.6100%0%0——dominated
548—Anthropic10.6100%0%0.5—200Kdominated
548—Allen Institute for AI10.6100%0%0——dominated
551—InclusionAI10.2100%0%0——dominated
551—Allen Institute for AI10.2100%0%0——dominated
553—DeepSeek10.0100%0%0——dominated
554—Anthropic9.5100%0%0——dominated
554—DeepSeek9.5100%0%0——dominated
554—OpenAI9.5100%0%0.8——dominated
554—medium reasoningMistral9.5100%0%3——dominated
554—Mistral9.5100%0%0.3——dominated
559—Meta9.0100%0%1.2——dominated
560—Snowflake8.6100%0%0——dominated
560—Liquid AI8.6100%0%0——dominated
560—Alibaba8.6100%0%0——dominated
563—Meta8.2100%0%0.32.5—dominated
563—non-reasoning reasoningAlibaba8.2100%0%0——dominated
565—Google8.0100%0%0——dominated
566—DeepSeek7.7100%0%0——dominated
566—Google7.7100%0%0——dominated
568—Cohere7.0100%0%6——dominated
568—Databricks7.0100%0%0——dominated
568—DeepSeek7.0100%0%0——dominated
568—Meta7.0100%0%0——dominated
568—Meta7.0100%0%0——dominated
568—OpenChat7.0100%0%0——dominated
568—Sarvam7.0100%0%0——dominated
575—LG AI Research6.4100%0%0——dominated
576—non-reasoning reasoningLG AI Research6.1100%0%0——dominated
576—Allen Institute for AI6.1100%0%0.1—65.5Kdominated
578—IBM5.5100%0%0——dominated
578—AI21 Labs5.5100%0%0.3——dominated
578—AI21 Labs5.5100%0%0——dominated
578—Liquid AI5.5100%0%0——dominated
578—Liquid AI5.5100%0%0——dominated
578—Liquid AI5.5100%0%0——dominated
584—AI21 Labs4.8100%0%0.3——dominated
584—Alibaba4.8100%0%0——dominated
586—Swiss AI Initiative4.3100%0%1.3——dominated
586—Google4.3100%0%0——dominated
586—IBM4.3100%0%0——dominated
586—Mistral4.3100%0%0.5—32.8Kdominated
590—non-reasoning reasoningNous Research3.9100%0%0——dominated
591—Anthropic3.3100%0%0——dominated
591—Cohere3.3100%0%0.8——dominated
591—IBM3.3100%0%0——dominated
591—Meta3.3100%0%0——dominated
591—Mistral3.3100%0%0.3—32.8Kdominated
591—Alibaba3.3100%0%0——dominated
597—Allen Institute for AI2.8100%0%0——dominated
598—non-reasoning reasoningIBM2.5100%0%0.1——dominated
598—Liquid AI2.5100%0%0—32.8Kdominated
600—non-reasoning reasoningAlibaba2.3100%0%0——dominated
601—Alibaba2.1100%0%0——dominated
602—Google1.9100%0%0.1——dominated
602—Meta1.9100%0%0.1——dominated
604—Google1.5100%0%0——dominated
604—Liquid AI1.5100%0%0——dominated
604—Meta1.5100%0%0——dominated
607—Swiss AI Initiative0.6100%0%0.1——dominated
607—Google0.6100%0%0——dominated
607—Google0.6100%0%0——dominated
607—IBM0.6100%0%0——dominated
607—IBM0.6100%0%0——dominated
607—Liquid AI0.6100%0%0373.9—dominated
607—non-reasoning reasoningAlibaba0.6100%0%0——dominated
607—Cohere0.6100%0%0130—dominated
331 excludedby population, evidence, or constraints
  • DeepSeek-OCR
    • Too few Source Benchmarks contribute to this score.
    • Too few contributions are directly observed.
  • Grok Voice Agent
    • Too few Source Benchmarks contribute to this score.
    • Too few contributions are directly observed.
  • Cogito v2.1 (Reasoning)
    • Too few Source Benchmarks contribute to this score.
    • Too few contributions are directly observed.
  • Mi:dm K 2.5 Pro Preview
    • Too few Source Benchmarks contribute to this score.
    • Too few contributions are directly observed.
  • Doubao-Seed-1.8
    • Too few Source Benchmarks contribute to this score.
    • Too few contributions are directly observed.
  • GPT-4o mini Realtime (Dec '24)
    • Too few Source Benchmarks contribute to this score.
    • Too few contributions are directly observed.
  • GPT-4o Realtime (Dec '24)
    • Too few Source Benchmarks contribute to this score.
    • Too few contributions are directly observed.
  • GPT-3.5 Turbo (0613)
    • Too few Source Benchmarks contribute to this score.
    • Too few contributions are directly observed.
  • Aurora Alpha
    • Too few Source Benchmarks contribute to this score.
    • Too few contributions are directly observed.
  • Free Models Router
    • Too few Source Benchmarks contribute to this score.
    • Too few contributions are directly observed.
  • StepFun: Step 3.5 Flash (free)
    • Too few Source Benchmarks contribute to this score.
    • Too few contributions are directly observed.
  • Arcee AI: Trinity Large Preview (free)
    • Too few Source Benchmarks contribute to this score.
    • Too few contributions are directly observed.
  • Upstage: Solar Pro 3 (free)
    • Too few Source Benchmarks contribute to this score.
    • Too few contributions are directly observed.
  • MiniMax: MiniMax M2-her
    • Too few Source Benchmarks contribute to this score.
    • Too few contributions are directly observed.
  • Writer: Palmyra X5
    • Too few Source Benchmarks contribute to this score.
    • Too few contributions are directly observed.
  • LiquidAI: LFM2.5-1.2B-Thinking (free)
    • Too few Source Benchmarks contribute to this score.
    • Too few contributions are directly observed.
  • LiquidAI: LFM2.5-1.2B-Instruct (free)
    • Too few Source Benchmarks contribute to this score.
    • Too few contributions are directly observed.
  • OpenAI: GPT Audio
    • Too few Source Benchmarks contribute to this score.
    • Too few contributions are directly observed.
  • OpenAI: GPT Audio Mini
    • Too few Source Benchmarks contribute to this score.
    • Too few contributions are directly observed.
  • AllenAI: Molmo2 8B
    • Too few Source Benchmarks contribute to this score.
    • Too few contributions are directly observed.
  • ByteDance Seed: Seed 1.6 Flash
    • Too few Source Benchmarks contribute to this score.
    • Too few contributions are directly observed.
  • ByteDance Seed: Seed 1.6
    • Too few Source Benchmarks contribute to this score.
    • Too few contributions are directly observed.
  • Mistral: Mistral Small Creative
    • Too few Source Benchmarks contribute to this score.
    • Too few contributions are directly observed.
  • OpenAI: GPT-5.2 Chat
    • Too few Source Benchmarks contribute to this score.
    • Too few contributions are directly observed.
  • OpenAI: GPT-5.2 Pro
    • Too few Source Benchmarks contribute to this score.
    • Too few contributions are directly observed.
  • Mistral: Devstral 2 2512
    • Too few Source Benchmarks contribute to this score.
    • Too few contributions are directly observed.
  • Relace: Relace Search
    • Too few Source Benchmarks contribute to this score.
    • Too few contributions are directly observed.
  • Nex AGI: DeepSeek V3.1 Nex N1
    • Too few Source Benchmarks contribute to this score.
    • Too few contributions are directly observed.
  • EssentialAI: Rnj 1 Instruct
    • Too few Source Benchmarks contribute to this score.
    • Too few contributions are directly observed.
  • Body Builder (beta)
    • Too few Source Benchmarks contribute to this score.
    • Too few contributions are directly observed.
  • OpenAI: GPT-5.1-Codex-Max
    • Too few Source Benchmarks contribute to this score.
    • Too few contributions are directly observed.
  • Mistral: Ministral 3 14B 2512
    • Too few Source Benchmarks contribute to this score.
    • Too few contributions are directly observed.
  • Mistral: Ministral 3 8B 2512
    • Too few Source Benchmarks contribute to this score.
    • Too few contributions are directly observed.
  • Mistral: Ministral 3 3B 2512
    • Too few Source Benchmarks contribute to this score.
    • Too few contributions are directly observed.
  • Mistral: Mistral Large 3 2512
    • Too few Source Benchmarks contribute to this score.
    • Too few contributions are directly observed.
  • Arcee AI: Trinity Mini (free)
    • Too few Source Benchmarks contribute to this score.
    • Too few contributions are directly observed.
  • TNG: R1T Chimera (free)
    • Too few Source Benchmarks contribute to this score.
    • Too few contributions are directly observed.
  • Google: Nano Banana Pro (Gemini 3 Pro Image Preview)
    • Too few Source Benchmarks contribute to this score.
    • Too few contributions are directly observed.
  • Google: Gemini 3 Pro Preview
    • Too few Source Benchmarks contribute to this score.
    • Too few contributions are directly observed.
  • Deep Cogito: Cogito v2.1 671B
    • Too few Source Benchmarks contribute to this score.
    • Too few contributions are directly observed.
  • OpenAI: GPT-5.1 Chat
    • Too few Source Benchmarks contribute to this score.
    • Too few contributions are directly observed.
  • Kwaipilot: KAT-Coder-Pro V1
    • Too few Source Benchmarks contribute to this score.
    • Too few contributions are directly observed.
  • Perplexity: Sonar Pro Search
    • Too few Source Benchmarks contribute to this score.
    • Too few contributions are directly observed.
  • Mistral: Voxtral Small 24B 2507
    • Too few Source Benchmarks contribute to this score.
    • Too few contributions are directly observed.
  • OpenAI: gpt-oss-safeguard-20b
    • Too few Source Benchmarks contribute to this score.
    • Too few contributions are directly observed.
  • LiquidAI: LFM2-2.6B
    • Too few Source Benchmarks contribute to this score.
    • Too few contributions are directly observed.
  • IBM: Granite 4.0 Micro
    • Too few Source Benchmarks contribute to this score.
    • Too few contributions are directly observed.
  • OpenAI: GPT-5 Image Mini
    • Too few Source Benchmarks contribute to this score.
    • Too few contributions are directly observed.
  • Qwen: Qwen3 VL 8B Thinking
    • Too few Source Benchmarks contribute to this score.
    • Too few contributions are directly observed.
  • OpenAI: GPT-5 Image
    • Too few Source Benchmarks contribute to this score.
    • Too few contributions are directly observed.
  • OpenAI: o3 Deep Research
    • Too few Source Benchmarks contribute to this score.
    • Too few contributions are directly observed.
  • OpenAI: o4 Mini Deep Research
    • Too few Source Benchmarks contribute to this score.
    • Too few contributions are directly observed.
  • NVIDIA: Llama 3.3 Nemotron Super 49B V1.5
    • Too few Source Benchmarks contribute to this score.
    • Too few contributions are directly observed.
  • Baidu: ERNIE 4.5 21B A3B Thinking
    • Too few Source Benchmarks contribute to this score.
    • Too few contributions are directly observed.
  • Google: Gemini 2.5 Flash Image (Nano Banana)
    • Too few Source Benchmarks contribute to this score.
    • Too few contributions are directly observed.
  • Qwen: Qwen3 VL 30B A3B Thinking
    • Too few Source Benchmarks contribute to this score.
    • Too few contributions are directly observed.
  • OpenAI: GPT-5 Pro
    • Too few Source Benchmarks contribute to this score.
    • Too few contributions are directly observed.
  • DeepSeek: DeepSeek V3.2 Exp
    • Too few Source Benchmarks contribute to this score.
    • Too few contributions are directly observed.
  • TheDrummer: Cydonia 24B V4.1
    • Too few Source Benchmarks contribute to this score.
    • Too few contributions are directly observed.
  • Relace: Relace Apply 3
    • Too few Source Benchmarks contribute to this score.
    • Too few contributions are directly observed.
  • Qwen: Qwen3 VL 235B A22B Thinking
    • Too few Source Benchmarks contribute to this score.
    • Too few contributions are directly observed.
  • Qwen: Qwen3 Coder Plus
    • Too few Source Benchmarks contribute to this score.
    • Too few contributions are directly observed.
  • Tongyi DeepResearch 30B A3B
    • Too few Source Benchmarks contribute to this score.
    • Too few contributions are directly observed.
  • Qwen: Qwen3 Coder Flash
    • Too few Source Benchmarks contribute to this score.
    • Too few contributions are directly observed.
  • OpenGVLab: InternVL3 78B
    • Too few Source Benchmarks contribute to this score.
    • Too few contributions are directly observed.
  • Qwen: Qwen3 Next 80B A3B Thinking
    • Too few Source Benchmarks contribute to this score.
    • Too few contributions are directly observed.
  • Meituan: LongCat Flash Chat
    • Too few Source Benchmarks contribute to this score.
    • Too few contributions are directly observed.
  • Qwen: Qwen Plus 0728
    • Too few Source Benchmarks contribute to this score.
    • Too few contributions are directly observed.
  • Qwen: Qwen3 30B A3B Thinking 2507
    • Too few Source Benchmarks contribute to this score.
    • Too few contributions are directly observed.
  • Nous: Hermes 4 70B
    • Too few Source Benchmarks contribute to this score.
    • Too few contributions are directly observed.
  • Nous: Hermes 4 405B
    • Too few Source Benchmarks contribute to this score.
    • Too few contributions are directly observed.
  • DeepSeek: DeepSeek V3.1
    • Too few Source Benchmarks contribute to this score.
    • Too few contributions are directly observed.
  • OpenAI: GPT-4o Audio
    • Too few Source Benchmarks contribute to this score.
    • Too few contributions are directly observed.
  • Baidu: ERNIE 4.5 21B A3B
    • Too few Source Benchmarks contribute to this score.
    • Too few contributions are directly observed.
  • Baidu: ERNIE 4.5 VL 28B A3B
    • Too few Source Benchmarks contribute to this score.
    • Too few contributions are directly observed.
  • OpenAI: GPT-5 Chat
    • Too few Source Benchmarks contribute to this score.
    • Too few contributions are directly observed.
  • Mistral: Codestral 2508
    • Too few Source Benchmarks contribute to this score.
    • Too few contributions are directly observed.
  • Qwen: Qwen3 235B A22B Thinking 2507
    • Too few Source Benchmarks contribute to this score.
    • Too few contributions are directly observed.
  • Z.ai: GLM 4 32B
    • Too few Source Benchmarks contribute to this score.
    • Too few contributions are directly observed.
  • Qwen: Qwen3 Coder 480B A35B (free)
    • Too few Source Benchmarks contribute to this score.
    • Too few contributions are directly observed.
  • ByteDance: UI-TARS 7B
    • Too few Source Benchmarks contribute to this score.
    • Too few contributions are directly observed.
  • Qwen: Qwen3 235B A22B Instruct 2507
    • Too few Source Benchmarks contribute to this score.
    • Too few contributions are directly observed.
  • Switchpoint Router
    • Too few Source Benchmarks contribute to this score.
    • Too few contributions are directly observed.
  • Venice: Uncensored (free)
    • Too few Source Benchmarks contribute to this score.
    • Too few contributions are directly observed.
  • Google: Gemma 3n 2B (free)
    • Too few Source Benchmarks contribute to this score.
    • Too few contributions are directly observed.
  • Tencent: Hunyuan A13B Instruct
    • Too few Source Benchmarks contribute to this score.
    • Too few contributions are directly observed.
  • TNG: DeepSeek R1T2 Chimera (free)
    • Too few Source Benchmarks contribute to this score.
    • Too few contributions are directly observed.
  • Morph: Morph V3 Large
    • Too few Source Benchmarks contribute to this score.
    • Too few contributions are directly observed.
  • Morph: Morph V3 Fast
    • Too few Source Benchmarks contribute to this score.
    • Too few contributions are directly observed.
  • Baidu: ERNIE 4.5 VL 424B A47B
    • Too few Source Benchmarks contribute to this score.
    • Too few contributions are directly observed.
  • Inception: Mercury
    • Too few Source Benchmarks contribute to this score.
    • Too few contributions are directly observed.
  • Mistral: Mistral Small 3.2 24B
    • Too few Source Benchmarks contribute to this score.
    • Too few contributions are directly observed.
  • MiniMax: MiniMax M1
    • Too few Source Benchmarks contribute to this score.
    • Too few contributions are directly observed.
  • xAI: Grok 3 Mini
    • Too few Source Benchmarks contribute to this score.
    • Too few contributions are directly observed.
  • Google: Gemini 2.5 Pro Preview 06-05
    • Too few Source Benchmarks contribute to this score.
    • Too few contributions are directly observed.
  • DeepSeek: R1 0528 (free)
    • Too few Source Benchmarks contribute to this score.
    • Too few contributions are directly observed.
  • Google: Gemma 3n 4B (free)
    • Too few Source Benchmarks contribute to this score.
    • Too few contributions are directly observed.
  • Google: Gemini 2.5 Pro Preview 05-06
    • Too few Source Benchmarks contribute to this score.
    • Too few contributions are directly observed.
  • Arcee AI: Spotlight
    • Too few Source Benchmarks contribute to this score.
    • Too few contributions are directly observed.
  • Arcee AI: Maestro Reasoning
    • Too few Source Benchmarks contribute to this score.
    • Too few contributions are directly observed.
  • Arcee AI: Virtuoso Large
    • Too few Source Benchmarks contribute to this score.
    • Too few contributions are directly observed.
  • Arcee AI: Coder Large
    • Too few Source Benchmarks contribute to this score.
    • Too few contributions are directly observed.
  • Inception: Mercury Coder
    • Too few Source Benchmarks contribute to this score.
    • Too few contributions are directly observed.
  • Qwen: Qwen3 4B (free)
    • Too few Source Benchmarks contribute to this score.
    • Too few contributions are directly observed.
  • Meta: Llama Guard 4 12B
    • Too few Source Benchmarks contribute to this score.
    • Too few contributions are directly observed.
  • Qwen: Qwen3 30B A3B
    • Too few Source Benchmarks contribute to this score.
    • Too few contributions are directly observed.
  • Qwen: Qwen3 8B
    • Too few Source Benchmarks contribute to this score.
    • Too few contributions are directly observed.
  • Qwen: Qwen3 14B
    • Too few Source Benchmarks contribute to this score.
    • Too few contributions are directly observed.
  • Qwen: Qwen3 32B
    • Too few Source Benchmarks contribute to this score.
    • Too few contributions are directly observed.
  • Qwen: Qwen3 235B A22B
    • Too few Source Benchmarks contribute to this score.
    • Too few contributions are directly observed.
  • TNG: DeepSeek R1T Chimera (free)
    • Too few Source Benchmarks contribute to this score.
    • Too few contributions are directly observed.
  • OpenAI: o4 Mini High
    • Too few Source Benchmarks contribute to this score.
    • Too few contributions are directly observed.
  • EleutherAI: Llemma 7b
    • Too few Source Benchmarks contribute to this score.
    • Too few contributions are directly observed.
  • AlfredPros: CodeLLaMa 7B Instruct Solidity
    • Too few Source Benchmarks contribute to this score.
    • Too few contributions are directly observed.
  • xAI: Grok 3 Mini Beta
    • Too few Source Benchmarks contribute to this score.
    • Too few contributions are directly observed.
  • xAI: Grok 3 Beta
    • Too few Source Benchmarks contribute to this score.
    • Too few contributions are directly observed.
  • NVIDIA: Llama 3.1 Nemotron Ultra 253B v1
    • Too few Source Benchmarks contribute to this score.
    • Too few contributions are directly observed.
  • Qwen: Qwen2.5 VL 32B Instruct
    • Too few Source Benchmarks contribute to this score.
    • Too few contributions are directly observed.
  • DeepSeek: DeepSeek V3 0324
    • Too few Source Benchmarks contribute to this score.
    • Too few contributions are directly observed.
  • Mistral: Mistral Small 3.1 24B (free)
    • Too few Source Benchmarks contribute to this score.
    • Too few contributions are directly observed.
  • AllenAI: Olmo 2 32B Instruct
    • Too few Source Benchmarks contribute to this score.
    • Too few contributions are directly observed.
  • Google: Gemma 3 4B (free)
    • Too few Source Benchmarks contribute to this score.
    • Too few contributions are directly observed.
  • Google: Gemma 3 12B (free)
    • Too few Source Benchmarks contribute to this score.
    • Too few contributions are directly observed.
  • OpenAI: GPT-4o-mini Search Preview
    • Too few Source Benchmarks contribute to this score.
    • Too few contributions are directly observed.
  • OpenAI: GPT-4o Search Preview
    • Too few Source Benchmarks contribute to this score.
    • Too few contributions are directly observed.
  • Google: Gemma 3 27B (free)
    • Too few Source Benchmarks contribute to this score.
    • Too few contributions are directly observed.
  • TheDrummer: Skyfall 36B V2
    • Too few Source Benchmarks contribute to this score.
    • Too few contributions are directly observed.
  • Perplexity: Sonar Deep Research
    • Too few Source Benchmarks contribute to this score.
    • Too few contributions are directly observed.
  • Llama Guard 3 8B
    • Too few Source Benchmarks contribute to this score.
    • Too few contributions are directly observed.
  • Google: Gemini 2.0 Flash
    • Too few Source Benchmarks contribute to this score.
    • Too few contributions are directly observed.
  • Qwen: Qwen VL Plus
    • Too few Source Benchmarks contribute to this score.
    • Too few contributions are directly observed.
  • AionLabs: Aion-1.0
    • Too few Source Benchmarks contribute to this score.
    • Too few contributions are directly observed.
  • AionLabs: Aion-1.0-Mini
    • Too few Source Benchmarks contribute to this score.
    • Too few contributions are directly observed.
  • AionLabs: Aion-RP 1.0 (8B)
    • Too few Source Benchmarks contribute to this score.
    • Too few contributions are directly observed.
  • Qwen: Qwen VL Max
    • Too few Source Benchmarks contribute to this score.
    • Too few contributions are directly observed.
  • Qwen: Qwen2.5 VL 72B Instruct
    • Too few Source Benchmarks contribute to this score.
    • Too few contributions are directly observed.
  • Qwen: Qwen-Plus
    • Too few Source Benchmarks contribute to this score.
    • Too few contributions are directly observed.
  • Qwen: Qwen-Max
    • Too few Source Benchmarks contribute to this score.
    • Too few contributions are directly observed.
  • Mistral: Mistral Small 3
    • Too few Source Benchmarks contribute to this score.
    • Too few contributions are directly observed.
  • MiniMax: MiniMax-01
    • Too few Source Benchmarks contribute to this score.
    • Too few contributions are directly observed.
  • Sao10K: Llama 3.1 70B Hanami x1
    • Too few Source Benchmarks contribute to this score.
    • Too few contributions are directly observed.
  • Sao10K: Llama 3.3 Euryale 70B
    • Too few Source Benchmarks contribute to this score.
    • Too few contributions are directly observed.
  • Cohere: Command R7B (12-2024)
    • Too few Source Benchmarks contribute to this score.
    • Too few contributions are directly observed.
  • OpenAI: GPT-4o (2024-11-20)
    • Too few Source Benchmarks contribute to this score.
    • Too few contributions are directly observed.
  • Mistral Large 2411
    • Too few Source Benchmarks contribute to this score.
    • Too few contributions are directly observed.
  • Qwen2.5 Coder 32B Instruct
    • Too few Source Benchmarks contribute to this score.
    • Too few contributions are directly observed.
  • SorcererLM 8x22B
    • Too few Source Benchmarks contribute to this score.
    • Too few contributions are directly observed.
  • TheDrummer: UnslopNemo 12B
    • Too few Source Benchmarks contribute to this score.
    • Too few contributions are directly observed.
  • Magnum v4 72B
    • Too few Source Benchmarks contribute to this score.
    • Too few contributions are directly observed.
  • Anthropic: Claude 3.5 Sonnet
    • Too few Source Benchmarks contribute to this score.
    • Too few contributions are directly observed.
  • Qwen: Qwen2.5 7B Instruct
    • Too few Source Benchmarks contribute to this score.
    • Too few contributions are directly observed.
  • Inflection: Inflection 3 Pi
    • Too few Source Benchmarks contribute to this score.
    • Too few contributions are directly observed.
  • Inflection: Inflection 3 Productivity
    • Too few Source Benchmarks contribute to this score.
    • Too few contributions are directly observed.
  • TheDrummer: Rocinante 12B
    • Too few Source Benchmarks contribute to this score.
    • Too few contributions are directly observed.
  • Qwen2.5 72B Instruct
    • Too few Source Benchmarks contribute to this score.
    • Too few contributions are directly observed.
  • NeverSleep: Lumimaid v0.2 8B
    • Too few Source Benchmarks contribute to this score.
    • Too few contributions are directly observed.
  • Mistral: Pixtral 12B
    • Too few Source Benchmarks contribute to this score.
    • Too few contributions are directly observed.
  • Cohere: Command R (08-2024)
    • Too few Source Benchmarks contribute to this score.
    • Too few contributions are directly observed.
  • Cohere: Command R+ (08-2024)
    • Too few Source Benchmarks contribute to this score.
    • Too few contributions are directly observed.
  • Sao10K: Llama 3.1 Euryale 70B v2.2
    • Too few Source Benchmarks contribute to this score.
    • Too few contributions are directly observed.
  • Qwen: Qwen2.5-VL 7B Instruct
    • Too few Source Benchmarks contribute to this score.
    • Too few contributions are directly observed.
  • Nous: Hermes 3 405B Instruct (free)
    • Too few Source Benchmarks contribute to this score.
    • Too few contributions are directly observed.
  • OpenAI: ChatGPT-4o
    • Too few Source Benchmarks contribute to this score.
    • Too few contributions are directly observed.
  • Sao10K: Llama 3 8B Lunaris
    • Too few Source Benchmarks contribute to this score.
    • Too few contributions are directly observed.
  • Meta: Llama 3.1 405B (base)
    • Too few Source Benchmarks contribute to this score.
    • Too few contributions are directly observed.
  • Meta: Llama 3.1 405B Instruct
    • Too few Source Benchmarks contribute to this score.
    • Too few contributions are directly observed.
  • Mistral: Mistral Nemo
    • Too few Source Benchmarks contribute to this score.
    • Too few contributions are directly observed.
  • OpenAI: GPT-4o-mini (2024-07-18)
    • Too few Source Benchmarks contribute to this score.
    • Too few contributions are directly observed.
  • Google: Gemma 2 27B
    • Too few Source Benchmarks contribute to this score.
    • Too few contributions are directly observed.
  • Google: Gemma 2 9B
    • Too few Source Benchmarks contribute to this score.
    • Too few contributions are directly observed.
  • Sao10k: Llama 3 Euryale 70B v2.1
    • Too few Source Benchmarks contribute to this score.
    • Too few contributions are directly observed.
  • NousResearch: Hermes 2 Pro - Llama-3 8B
    • Too few Source Benchmarks contribute to this score.
    • Too few contributions are directly observed.
  • Mistral: Mistral 7B Instruct v0.3
    • Too few Source Benchmarks contribute to this score.
    • Too few contributions are directly observed.
  • Meta: LlamaGuard 2 8B
    • Too few Source Benchmarks contribute to this score.
    • Too few contributions are directly observed.
  • WizardLM-2 8x22B
    • Too few Source Benchmarks contribute to this score.
    • Too few contributions are directly observed.
  • OpenAI: GPT-4 Turbo Preview
    • Too few Source Benchmarks contribute to this score.
    • Too few contributions are directly observed.
  • Mistral: Mistral 7B Instruct v0.2
    • Too few Source Benchmarks contribute to this score.
    • Too few contributions are directly observed.
  • Noromaid 20B
    • Too few Source Benchmarks contribute to this score.
    • Too few contributions are directly observed.
  • Goliath 120B
    • Too few Source Benchmarks contribute to this score.
    • Too few contributions are directly observed.
  • Auto Router
    • Too few Source Benchmarks contribute to this score.
    • Too few contributions are directly observed.
  • OpenAI: GPT-4 Turbo (older v1106)
    • Too few Source Benchmarks contribute to this score.
    • Too few contributions are directly observed.
  • OpenAI: GPT-3.5 Turbo Instruct
    • Too few Source Benchmarks contribute to this score.
    • Too few contributions are directly observed.
  • Mistral: Mistral 7B Instruct v0.1
    • Too few Source Benchmarks contribute to this score.
    • Too few contributions are directly observed.
  • OpenAI: GPT-3.5 Turbo 16k
    • Too few Source Benchmarks contribute to this score.
    • Too few contributions are directly observed.
  • Mancer: Weaver (alpha)
    • Too few Source Benchmarks contribute to this score.
    • Too few contributions are directly observed.
  • ReMM SLERP 13B
    • Too few Source Benchmarks contribute to this score.
    • Too few contributions are directly observed.
  • MythoMax 13B
    • Too few Source Benchmarks contribute to this score.
    • Too few contributions are directly observed.
  • OpenAI: GPT-4 (older v0314)
    • Too few Source Benchmarks contribute to this score.
    • Too few contributions are directly observed.
  • Qwen: Qwen Plus 0728 (thinking)
    • Too few Source Benchmarks contribute to this score.
    • Too few contributions are directly observed.
  • Qwen: Qwen3 Coder 480B A35B
    • Too few Source Benchmarks contribute to this score.
    • Too few contributions are directly observed.
  • Qwen: Qwen3 4B
    • Too few Source Benchmarks contribute to this score.
    • Too few contributions are directly observed.
  • Mistral: Mistral Small 3.1 24B
    • Too few Source Benchmarks contribute to this score.
    • Too few contributions are directly observed.
  • Nous: Hermes 3 405B Instruct
    • Too few Source Benchmarks contribute to this score.
    • Too few contributions are directly observed.
  • Qwen: Qwen3.5 Plus 2026-02-15
    • Too few Source Benchmarks contribute to this score.
    • Too few contributions are directly observed.
  • Amazon: Nova 2 Lite
    • Too few Source Benchmarks contribute to this score.
    • Too few contributions are directly observed.
  • Amazon: Nova Premier 1.0
    • Too few Source Benchmarks contribute to this score.
    • Too few contributions are directly observed.
  • Amazon: Nova Lite 1.0
    • Too few Source Benchmarks contribute to this score.
    • Too few contributions are directly observed.
  • Amazon: Nova Micro 1.0
    • Too few Source Benchmarks contribute to this score.
    • Too few contributions are directly observed.
  • Amazon: Nova Pro 1.0
    • Too few Source Benchmarks contribute to this score.
    • Too few contributions are directly observed.
  • AionLabs: Aion-2.0
    • Too few Source Benchmarks contribute to this score.
    • Too few contributions are directly observed.
  • Google: Gemini 3.1 Pro Preview Custom Tools
    • Too few Source Benchmarks contribute to this score.
    • Too few contributions are directly observed.
  • Qwen: Qwen3.5-Flash
    • Too few Source Benchmarks contribute to this score.
    • Too few contributions are directly observed.
  • LiquidAI: LFM2-24B-A2B
    • Too few Source Benchmarks contribute to this score.
    • Too few contributions are directly observed.
  • ByteDance Seed: Seed-2.0-Mini
    • Too few Source Benchmarks contribute to this score.
    • Too few contributions are directly observed.
  • Google: Nano Banana 2 (Gemini 3.1 Flash Image Preview)
    • Too few Source Benchmarks contribute to this score.
    • Too few contributions are directly observed.
  • OpenAI: GPT-5.4 Pro
    • Too few Source Benchmarks contribute to this score.
    • Too few contributions are directly observed.
  • OpenAI: GPT-5.4
    • Too few Source Benchmarks contribute to this score.
    • Too few contributions are directly observed.
  • OpenAI: GPT-5.3 Chat
    • Too few Source Benchmarks contribute to this score.
    • Too few contributions are directly observed.
  • GPT-5.4 Pro (xhigh)
    • Too few Source Benchmarks contribute to this score.
    • Too few contributions are directly observed.
  • xAI: Grok 4.20 Multi-Agent Beta
    • Too few Source Benchmarks contribute to this score.
    • Too few contributions are directly observed.
  • xAI: Grok 4.20 Beta
    • Too few Source Benchmarks contribute to this score.
    • Too few contributions are directly observed.
  • Hunter Alpha
    • Too few Source Benchmarks contribute to this score.
    • Too few contributions are directly observed.
  • Healer Alpha
    • Too few Source Benchmarks contribute to this score.
    • Too few contributions are directly observed.
  • ByteDance Seed: Seed-2.0-Lite
    • Too few Source Benchmarks contribute to this score.
    • Too few contributions are directly observed.
  • Mistral: Mistral Small 4
    • Too few Source Benchmarks contribute to this score.
    • Too few contributions are directly observed.
  • Reka Edge
    • Too few Source Benchmarks contribute to this score.
    • Too few contributions are directly observed.
  • Gemini 3 Deep Think
    • Too few Source Benchmarks contribute to this score.
    • Too few contributions are directly observed.
  • Google: Lyria 3 Pro Preview
    • Too few Source Benchmarks contribute to this score.
    • Too few contributions are directly observed.
  • Google: Lyria 3 Clip Preview
    • Too few Source Benchmarks contribute to this score.
    • Too few contributions are directly observed.
  • Qwen: Qwen3.6 Plus Preview (free)
    • Too few Source Benchmarks contribute to this score.
    • Too few contributions are directly observed.
  • xAI: Grok 4.20 Multi-Agent
    • Too few Source Benchmarks contribute to this score.
    • Too few contributions are directly observed.
  • Reka Edge
    • Too few Source Benchmarks contribute to this score.
    • Too few contributions are directly observed.
  • Arcee AI: Trinity Large Thinking
    • Too few Source Benchmarks contribute to this score.
    • Too few contributions are directly observed.
  • Qwen: Qwen3.6 Plus (free)
    • Too few Source Benchmarks contribute to this score.
    • Too few contributions are directly observed.
  • Google: Gemma 4 31B
    • Too few Source Benchmarks contribute to this score.
    • Too few contributions are directly observed.
  • Anthropic: Claude Opus 4.7
    • Too few Source Benchmarks contribute to this score.
    • Too few contributions are directly observed.
  • Elephant
    • Too few Source Benchmarks contribute to this score.
    • Too few contributions are directly observed.
  • Anthropic: Claude Opus 4.6 (Fast)
    • Too few Source Benchmarks contribute to this score.
    • Too few contributions are directly observed.
  • Google: Gemma 4 26B A4B (free)
    • Too few Source Benchmarks contribute to this score.
    • Too few contributions are directly observed.
  • Meta: Llama Guard 4 12B (free)
    • Too few Source Benchmarks contribute to this score.
    • Too few contributions are directly observed.
  • Claude Opus 4.7
    • Too few Source Benchmarks contribute to this score.
    • Too few contributions are directly observed.
  • Claude Mythos Preview
    • Too few Source Benchmarks contribute to this score.
    • Too few contributions are directly observed.
  • Nanonets OCR-3
    • Too few Source Benchmarks contribute to this score.
    • Too few contributions are directly observed.
  • Nanonets OCR2+
    • Too few Source Benchmarks contribute to this score.
    • Too few contributions are directly observed.
  • GLM-OCR
    • Too few Source Benchmarks contribute to this score.
    • Too few contributions are directly observed.
  • Anthropic Claude Haiku Latest
    • Too few Source Benchmarks contribute to this score.
    • Too few contributions are directly observed.
  • OpenAI GPT Mini Latest
    • Too few Source Benchmarks contribute to this score.
    • Too few contributions are directly observed.
  • Google Gemini Pro Latest
    • Too few Source Benchmarks contribute to this score.
    • Too few contributions are directly observed.
  • MoonshotAI Kimi Latest
    • Too few Source Benchmarks contribute to this score.
    • Too few contributions are directly observed.
  • Google Gemini Flash Latest
    • Too few Source Benchmarks contribute to this score.
    • Too few contributions are directly observed.
  • Anthropic Claude Sonnet Latest
    • Too few Source Benchmarks contribute to this score.
    • Too few contributions are directly observed.
  • OpenAI GPT Latest
    • Too few Source Benchmarks contribute to this score.
    • Too few contributions are directly observed.
  • Qwen: Qwen3.5 Plus 2026-04-20
    • Too few Source Benchmarks contribute to this score.
    • Too few contributions are directly observed.
  • Qwen: Qwen3.6 Flash
    • Too few Source Benchmarks contribute to this score.
    • Too few contributions are directly observed.
  • OpenAI: GPT-5.5 Pro
    • Too few Source Benchmarks contribute to this score.
    • Too few contributions are directly observed.
  • OpenAI: GPT-5.4 Image 2
    • Too few Source Benchmarks contribute to this score.
    • Too few contributions are directly observed.
  • Anthropic: Claude Opus Latest
    • Too few Source Benchmarks contribute to this score.
    • Too few contributions are directly observed.
  • Pareto Code Router
    • Too few Source Benchmarks contribute to this score.
    • Too few contributions are directly observed.
  • Baidu: Qianfan-OCR-Fast (free)
    • Too few Source Benchmarks contribute to this score.
    • Too few contributions are directly observed.
  • EXAONE 4.5 33B (Non-reasoning)
    • Too few Source Benchmarks contribute to this score.
    • Too few contributions are directly observed.
  • NVIDIA: Nemotron 3 Nano Omni (free)
    • Too few Source Benchmarks contribute to this score.
    • Too few contributions are directly observed.
  • Poolside: Laguna XS.2 (free)
    • Too few Source Benchmarks contribute to this score.
    • Too few contributions are directly observed.
  • Poolside: Laguna M.1 (free)
    • Too few Source Benchmarks contribute to this score.
    • Too few contributions are directly observed.
  • inclusionAI: Ling-2.6-flash
    • Too few Source Benchmarks contribute to this score.
    • Too few contributions are directly observed.
  • MiniMax: MiniMax M2.5 (free)
    • Too few Source Benchmarks contribute to this score.
    • Too few contributions are directly observed.
  • Z.ai: GLM 4.6
    • Too few Source Benchmarks contribute to this score.
    • Too few contributions are directly observed.
  • Qwen: Qwen3 Next 80B A3B Instruct (free)
    • Too few Source Benchmarks contribute to this score.
    • Too few contributions are directly observed.
  • MoonshotAI: Kimi K2 0905
    • Too few Source Benchmarks contribute to this score.
    • Too few contributions are directly observed.
  • OpenAI: gpt-oss-120b (free)
    • Too few Source Benchmarks contribute to this score.
    • Too few contributions are directly observed.
  • OpenAI: gpt-oss-20b (free)
    • Too few Source Benchmarks contribute to this score.
    • Too few contributions are directly observed.
  • Z.ai: GLM 4.5 Air (free)
    • Too few Source Benchmarks contribute to this score.
    • Too few contributions are directly observed.
  • Anthropic: Claude 3.7 Sonnet (thinking)
    • Too few Source Benchmarks contribute to this score.
    • Too few contributions are directly observed.
  • OpenAI: GPT-4o
    • Too few Source Benchmarks contribute to this score.
    • Too few contributions are directly observed.
  • Owl Alpha
    • Too few Source Benchmarks contribute to this score.
    • Too few contributions are directly observed.
  • GPT-5.5 Pro (xhigh)
    • Too few Source Benchmarks contribute to this score.
    • Too few contributions are directly observed.
  • OpenAI: GPT Chat Latest
    • Too few Source Benchmarks contribute to this score.
    • Too few contributions are directly observed.
  • Microsoft: Phi 4 Mini Instruct
    • Too few Source Benchmarks contribute to this score.
    • Too few contributions are directly observed.
  • Baidu Qianfan: CoBuddy (free)
    • Too few Source Benchmarks contribute to this score.
    • Too few contributions are directly observed.
  • Google: Gemini 3.1 Flash Lite
    • Too few Source Benchmarks contribute to this score.
    • Too few contributions are directly observed.
  • inclusionAI: Ling-2.6-1T
    • Too few Source Benchmarks contribute to this score.
    • Too few contributions are directly observed.
  • inclusionAI: Ring-2.6-1T (free)
    • Too few Source Benchmarks contribute to this score.
    • Too few contributions are directly observed.
  • Anthropic: Claude Opus 4.7 (Fast)
    • Too few Source Benchmarks contribute to this score.
    • Too few contributions are directly observed.
  • Perceptron: Perceptron Mk1
    • Too few Source Benchmarks contribute to this score.
    • Too few contributions are directly observed.
  • DeepSeek: DeepSeek V4 Flash (free)
    • Too few Source Benchmarks contribute to this score.
    • Too few contributions are directly observed.
  • xAI: Grok Build 0.1
    • Too few Source Benchmarks contribute to this score.
    • Too few contributions are directly observed.
  • MoonshotAI: Kimi K2.6 (free)
    • Too few Source Benchmarks contribute to this score.
    • Too few contributions are directly observed.
  • Anthropic: Claude Opus 4.8 (Fast)
    • Too few Source Benchmarks contribute to this score.
    • Too few contributions are directly observed.
  • StepFun: Step 3.7 Flash
    • Too few Source Benchmarks contribute to this score.
    • Too few contributions are directly observed.
  • MiniMax: MiniMax M3
    • Too few Source Benchmarks contribute to this score.
    • Too few contributions are directly observed.
  • OpenRouter: Fusion
    • Too few Source Benchmarks contribute to this score.
    • Too few contributions are directly observed.
  • NVIDIA: Nemotron 3.5 Content Safety (free)
    • Too few Source Benchmarks contribute to this score.
    • Too few contributions are directly observed.
  • Nex AGI: Nex-N2-Pro (free)
    • Too few Source Benchmarks contribute to this score.
    • Too few contributions are directly observed.
  • Anthropic: Claude Fable Latest
    • Too few Source Benchmarks contribute to this score.
    • Too few contributions are directly observed.
  • MoonshotAI: Kimi K2.7 Code
    • Too few Source Benchmarks contribute to this score.
    • Too few contributions are directly observed.
  • Z.ai: GLM 5.2
    • Too few Source Benchmarks contribute to this score.
    • Too few contributions are directly observed.
  • Google: Nano Banana 2 (Gemini 3.1 Flash Image)
    • Too few Source Benchmarks contribute to this score.
    • Too few contributions are directly observed.
  • Google: Nano Banana Pro (Gemini 3 Pro Image)
    • Too few Source Benchmarks contribute to this score.
    • Too few contributions are directly observed.
  • Sakana: Fugu Ultra
    • Too few Source Benchmarks contribute to this score.
    • Too few contributions are directly observed.
  • Claude Sonnet 5 (Adaptive Reasoning, High Effort)
    • Too few Source Benchmarks contribute to this score.
    • Too few contributions are directly observed.
  • Claude Sonnet 5 (Adaptive Reasoning, Medium Effort)
    • Too few Source Benchmarks contribute to this score.
    • Too few contributions are directly observed.
  • Claude Sonnet 5 (Adaptive Reasoning, Low Effort)
    • Too few Source Benchmarks contribute to this score.
    • Too few contributions are directly observed.
  • Claude Sonnet 5 (Adaptive Reasoning, Xhigh Effort)
    • Too few Source Benchmarks contribute to this score.
    • Too few contributions are directly observed.
  • Google: Nano Banana 2 Lite (Gemini 3.1 Flash Lite Image)
    • Too few Source Benchmarks contribute to this score.
    • Too few contributions are directly observed.
  • Poolside: Laguna XS 2.1 (free)
    • Too few Source Benchmarks contribute to this score.
    • Too few contributions are directly observed.
  • Tencent: Hy3
    • Too few Source Benchmarks contribute to this score.
    • Too few contributions are directly observed.
  • Nex AGI: Nex-N2-Mini
    • Too few Source Benchmarks contribute to this score.
    • Too few contributions are directly observed.
  • AionLabs: Aion-3.0-Mini
    • Too few Source Benchmarks contribute to this score.
    • Too few contributions are directly observed.
  • AionLabs: Aion-3.0
    • Too few Source Benchmarks contribute to this score.
    • Too few contributions are directly observed.
  • xAI: Grok Latest
    • Too few Source Benchmarks contribute to this score.
    • Too few contributions are directly observed.
  • OpenAI: GPT-5.6 Luna Pro
    • Too few Source Benchmarks contribute to this score.
    • Too few contributions are directly observed.
  • OpenAI: GPT-5.6 Terra Pro
    • Too few Source Benchmarks contribute to this score.
    • Too few contributions are directly observed.
  • OpenAI: GPT-5.6 Sol Pro
    • Too few Source Benchmarks contribute to this score.
    • Too few contributions are directly observed.
  • Claude Opus 5 (Fast)
    • Too few Source Benchmarks contribute to this score.
    • Too few contributions are directly observed.
  • Ling-3.0-flash (free)
    • Too few Source Benchmarks contribute to this score.
    • Too few contributions are directly observed.
  • Poolside: Laguna S 2.1
    • Too few Source Benchmarks contribute to this score.
    • Too few contributions are directly observed.
  • Meituan: LongCat 2.0
    • Too few Source Benchmarks contribute to this score.
    • Too few contributions are directly observed.
  • Auto Router (Beta)
    • Too few Source Benchmarks contribute to this score.
    • Too few contributions are directly observed.
  • MoonshotAI: Kimi K3
    • Too few Source Benchmarks contribute to this score.
    • Too few contributions are directly observed.
  • Kwaipilot: KAT-Coder-Air V2.5
    • Too few Source Benchmarks contribute to this score.
    • Too few contributions are directly observed.
  • Kwaipilot: KAT-Coder-Pro V2.5
    • Too few Source Benchmarks contribute to this score.
    • Too few contributions are directly observed.
  • Qwen: Qwen3.7 Flash
    • Too few Source Benchmarks contribute to this score.
    • Too few contributions are directly observed.
  • Claude Opus 5 (batch)
    • Too few Source Benchmarks contribute to this score.
    • Too few contributions are directly observed.
  • Anthropic: Claude Sonnet 5 (batch)
    • Too few Source Benchmarks contribute to this score.
    • Too few contributions are directly observed.
  • Anthropic: Claude Fable 5 (batch)
    • Too few Source Benchmarks contribute to this score.
    • Too few contributions are directly observed.
  • Anthropic: Claude Opus 4.8 (batch)
    • Too few Source Benchmarks contribute to this score.
    • Too few contributions are directly observed.
  • OpenAI: GPT-5.5 (batch)
    • Too few Source Benchmarks contribute to this score.
    • Too few contributions are directly observed.
  • Anthropic: Claude Opus 4.6 (batch)
    • Too few Source Benchmarks contribute to this score.
    • Too few contributions are directly observed.
  • OpenAI: GPT-5.2 (batch)
    • Too few Source Benchmarks contribute to this score.
    • Too few contributions are directly observed.
  • Anthropic: Claude Opus 4.5 (batch)
    • Too few Source Benchmarks contribute to this score.
    • Too few contributions are directly observed.
  • OpenAI: GPT-5.1 (batch)
    • Too few Source Benchmarks contribute to this score.
    • Too few contributions are directly observed.
  • OpenAI: GPT-5 (batch)
    • Too few Source Benchmarks contribute to this score.
    • Too few contributions are directly observed.
  • OpenAI: GPT-5 Mini (batch)
    • Too few Source Benchmarks contribute to this score.
    • Too few contributions are directly observed.
  • OpenAI: GPT-5 Nano (batch)
    • Too few Source Benchmarks contribute to this score.
    • Too few contributions are directly observed.
  • Google: Gemini 3.6 Flash (batch)
    • Too few Source Benchmarks contribute to this score.
    • Too few contributions are directly observed.
  • Google: Gemini 3.5 Flash Lite (batch)
    • Too few Source Benchmarks contribute to this score.
    • Too few contributions are directly observed.
  • Google: Gemini 3.5 Flash (batch)
    • Too few Source Benchmarks contribute to this score.
    • Too few contributions are directly observed.
  • Google: Gemini 3.1 Pro Preview (batch)
    • Too few Source Benchmarks contribute to this score.
    • Too few contributions are directly observed.
  • Google: Gemini 2.5 Flash Lite (batch)
    • Too few Source Benchmarks contribute to this score.
    • Too few contributions are directly observed.
  • Google: Gemini 2.5 Flash (batch)
    • Too few Source Benchmarks contribute to this score.
    • Too few contributions are directly observed.
  • Google: Gemini 2.5 Pro (batch)
    • Too few Source Benchmarks contribute to this score.
    • Too few contributions are directly observed.
  • DeepSeek: DeepSeek V4 Flash 0731
    • Too few Source Benchmarks contribute to this score.
    • Too few contributions are directly observed.

Why this score?

Why Claude Opus 5 (Adaptive Reasoning, Max Effort) scored 100.0

#1Rank
100.0%Coverage
1Observed
0Predicted

Intelligence Index

Observed
+100.0score points
Raw
60.70
Normalized
100.00
Configured weight
100.0%
Effective weight
100.0%
artificial_analysis
Source date 2026-07-31 · recorded 2026-07-31
Methodology & snapshot

Formula

  1. Intelligence Index × 100.0%higher · percentile
Missing data
reweight available
Evidence
observed only
Population
Population: all models
Coverage gate
0% · 1 observed · 1 contributing
Snapshot
Captured 2026-07-31 · 945 models · 83992 receipts
Integrity warnings1
  • Intelligence Index has moderate contamination riskinfo

    Weighting is set by Artificial Analysis and encodes subjective design choices even when documented; strong performance on one high-value workload can be hidden by weaker results in other categories.

    Remove it, lower its weight, or document why its contamination risk is acceptable.

Sensitivity

Perturb every configured weight by ±20% to test whether small judgment changes reorder the result.