AgMoDB
ModelsAgentsEvalsCompositesVisualizeIndustry
AgMoDB by @mistakeknot

AI Evals & Benchmarks

What each benchmark measures, why it matters, and how to interpret scores.

Heatmap

Your methodology, inspectable

Build a Composite Benchmark

Combine any Source Benchmarks with explicit weights, normalization, evidence policy, population, and constraints. See the ranking move live, then publish an immutable version others can inspect and fork.

Open workbench

AgMoBench

Transparent composite indices computed by AgMoDB using percentile-rank normalization across core domain benchmarks, with both full and confidence-oriented variants.

AgMoBench Overall
AgMoDB

Weighted blend of Reasoning (22%), Coding (22%), Math (14%), Agentic (14%), Robustness (18%), and Document Intelligence (10%) domain indices using observed benchmark data only (BenchPress predictions excluded). Uses percentile-rank normalization across the full model population.

Components

AgMoBench ReasoningAgMoBench CodingAgMoBench MathAgMoBench AgenticAgMoBench RobustnessAgMoBench Document Intelligence
AgMoBench Trust
AgMoDB

Trust-weighted blend of domain scores that emphasizes reasoning and coding signal. Uses prediction-inclusive data. Weights: Reasoning (27%), Coding (27%), Math (14%), Agentic (9%), Robustness (13%), and Document Intelligence (10%).

Components

AgMoBench ReasoningAgMoBench CodingAgMoBench MathAgMoBench AgenticAgMoBench RobustnessAgMoBench Document Intelligence
AgMoBench Predicted
AgMoDB

Prediction-inclusive variant of AgMoBench that includes BenchPress ML-predicted benchmark cells alongside observed data. Uses confidence adjustment, shrinking each domain's weight toward a 25% floor when that domain is prediction-heavy.

Components

AgMoBench ReasoningAgMoBench CodingAgMoBench MathAgMoBench AgenticAgMoBench RobustnessAgMoBench Document Intelligence
AgMoBench Reasoning
AgMoDB

Composite reasoning score averaging percentile ranks across knowledge and reasoning benchmarks.

Components

MMLU ProGPQA DiamondHLELiveBench OverallIFBenchARC-AGI-2LCRArena ELO
AgMoBench Coding
AgMoDB

Composite coding score averaging percentile ranks across code generation and software engineering benchmarks.

Components

LiveCodeBenchSciCodeTerminal-Bench HardAider PolyglotBigCodeBench CompleteArena ELO: Coding
AgMoBench Math
AgMoDB

Composite math score averaging percentile ranks across mathematical reasoning benchmarks.

Components

AIME 2025MATH-500FrontierMath
AgMoBench Agentic
AgMoDB

Composite agentic score averaging percentile ranks across tool use, web browsing, and autonomous task benchmarks.

Components

GAIATauBench AirlineWebArenaMLEBenchBFCLBrowseComp
AgMoBench Robustness
AgMoDB

Composite factual accuracy score averaging percentile ranks across factuality and critical thinking benchmarks — how well models resist nonsense, avoid fabrication, and know what they don't know.

Components

BullshitBenchTruthfulQASimpleQAAA-Omniscience Index
AgMoBench Document Intelligence
AgMoDB

Composite document processing score averaging percentile ranks across OCR, parsing, table understanding, and document retrieval benchmarks — how well models handle enterprise document workflows.

Components

IDP OverallIDP OlmOCRIDP OmniDocIDP CoreOmniDocBenchOCRBench v2MMTUViDoRe v3

Artificial Analysis Indices

Proprietary composite scores from Artificial Analysis. Methodology is not publicly disclosed.

Intelligence Index
Artificial Analysis

Artificial Analysis's proprietary composite intelligence score aggregating multiple benchmarks. Methodology is not publicly disclosed.

Coding Index
Artificial Analysis

Artificial Analysis's proprietary composite coding ability score.

Math Index
Artificial Analysis

Artificial Analysis's proprietary composite mathematical reasoning score.

Other Aggregate Benchmarks

Third-party composite scores and leaderboard aggregates.

AA Intelligence Index (Matrix)
benchmark_matrixMetadata incomplete

Artificial Analysis's composite Intelligence Index as reported in the LLM Benchmark Matrix. Aggregates multiple evals into a single capability score.

Ale Bench
epoch_aiMetadata incomplete

Ale Bench benchmark scores from Epoch AI's benchmark data collection.

Algotune
epoch_aiMetadata incomplete

Algotune benchmark scores from Epoch AI's benchmark data collection.

Apex Agents
epoch_aiMetadata incomplete

Apex Agents benchmark scores from Epoch AI's benchmark data collection.

Arc Agi 2
epoch_aiMetadata incomplete

Arc Agi 2 benchmark scores from Epoch AI's benchmark data collection.

Blueprint Bench 2
epoch_aiMetadata incomplete

Blueprint Bench 2 benchmark scores from Epoch AI's benchmark data collection.

Btf3
epoch_aiMetadata incomplete

Btf3 benchmark scores from Epoch AI's benchmark data collection.

Chatbot Arena ELO
chatbot_arena

Chatbot Arena ELO measures relative user preference from blind head-to-head votes in LMArena. It is a live human-judgment signal for conversational quality under real prompts.

Cl Bench
epoch_aiMetadata incomplete

Cl Bench benchmark scores from Epoch AI's benchmark data collection.

Cl Bench Life
epoch_aiMetadata incomplete

Cl Bench Life benchmark scores from Epoch AI's benchmark data collection.

Critpt
epoch_aiMetadata incomplete

Critpt benchmark scores from Epoch AI's benchmark data collection.

Cursorbench
epoch_aiMetadata incomplete

Cursorbench benchmark scores from Epoch AI's benchmark data collection.

Enigma Eval
epoch_aiMetadata incomplete

Enigma Eval benchmark scores from Epoch AI's benchmark data collection.

Epoch Capabilities Index
epoch_aiMetadata incomplete

Epoch Capabilities Index benchmark scores from Epoch AI's benchmark data collection.

Exploitbench
epoch_aiMetadata incomplete

Exploitbench benchmark scores from Epoch AI's benchmark data collection.

Forecastbench
epoch_aiMetadata incomplete

Forecastbench benchmark scores from Epoch AI's benchmark data collection.

Frontiercode
epoch_aiMetadata incomplete

Frontiercode benchmark scores from Epoch AI's benchmark data collection.

Frontierswe
epoch_aiMetadata incomplete

Frontierswe benchmark scores from Epoch AI's benchmark data collection.

Gbaeval
epoch_aiMetadata incomplete

Gbaeval benchmark scores from Epoch AI's benchmark data collection.

Gdp Pdf
epoch_aiMetadata incomplete

Gdp Pdf benchmark scores from Epoch AI's benchmark data collection.

Gdpval
epoch_aiMetadata incomplete

Gdpval benchmark scores from Epoch AI's benchmark data collection.

Hle
epoch_aiMetadata incomplete

Hle benchmark scores from Epoch AI's benchmark data collection.

LiveBench Overall
livebench

LiveBench Overall tracks broad model performance on recently refreshed questions with verifiable answers. It is used as a lower-contamination snapshot across major capability categories.

Mindcube
epoch_aiMetadata incomplete

Mindcube benchmark scores from Epoch AI's benchmark data collection.

Mystery Game Puzzles
epoch_aiMetadata incomplete

Mystery Game Puzzles benchmark scores from Epoch AI's benchmark data collection.

Open LLM Average
open_llm_leaderboard

Average score across HuggingFace's Open LLM Leaderboard benchmarks (IFEval, BBH, MATH Level 5, GPQA, MUSR, MMLU-PRO) — the community standard for open-source model evaluation.

Parameter Count
epoch_ai

Total number of trainable parameters in the model, as catalogued by Epoch AI's frontier model database.

Posttrainbench
epoch_aiMetadata incomplete

Posttrainbench benchmark scores from Epoch AI's benchmark data collection.

Proofbench
epoch_aiMetadata incomplete

Proofbench benchmark scores from Epoch AI's benchmark data collection.

Rli
epoch_aiMetadata incomplete

Rli benchmark scores from Epoch AI's benchmark data collection.

Scicode
epoch_aiMetadata incomplete

Scicode benchmark scores from Epoch AI's benchmark data collection.

Spatialviz Bench
epoch_aiMetadata incomplete

Spatialviz Bench benchmark scores from Epoch AI's benchmark data collection.

Surface Evolver Bench
epoch_aiMetadata incomplete

Surface Evolver Bench benchmark scores from Epoch AI's benchmark data collection.

Training Compute
epoch_ai

Total floating-point operations (FLOP) used during model training, as estimated by Epoch AI from public disclosures, hardware counts, and training duration.

Training Cost (USD)
epoch_ai

Estimated total cost to train the model in 2024 US dollars, including compute hardware rental or amortization.

Vending Bench 2
epoch_aiMetadata incomplete

Vending Bench 2 benchmark scores from Epoch AI's benchmark data collection.