AI Model Benchmark Comparator
Side-by-side comparison of 50+ AI models across 15 benchmarks — MMLU, HumanEval, GPQA, math, coding, and reasoning.
Key Highlights
- ✓ 50+ models compared
- ✓ 15 benchmarks (MMLU, HumanEval, GPQA, MATH, etc.)
- ✓ Filter by model size and provider
- ✓ Cost per million tokens
- ✓ Context window comparison
Overview
An interactive tool comparing 50+ AI language models across 15 standardized benchmarks. Filter by benchmark, model size, and provider to find the best model for your use case.
What's Inside
Benchmark Coverage
We track MMLU, MMLU-Pro, HumanEval, GPQA, GSM8K, MATH, BBH, HellaSwag, ARC, TruthfulQA, MT-Bench, AlpacaEval, and LMSYS Chatbot Arena scores across all major models.
Cost-Performance Analysis
Compare cost per million tokens against benchmark performance to identify the best value models. Includes API pricing from OpenAI, Anthropic, Google, Meta, and Mistral.
Ready to dive in?
Explore this resource and discover more across our 12 technology frontiers.