OpenAI o3 vs DeepSeek R1 vs Claude 3.5 Thinking vs Gemini 2.0 Thinking
Frontier Reasoning Model Comparison
Comparison of the top reasoning-optimized LLMs — chain-of-thought depth, math benchmarks, coding, cost, and latency.
o3DeepSeek R1Claude 3.5 ThinkingGemini 2.0 Thinking
Full Specification Comparison
| Specification | o3 | DeepSeek R1 | Claude 3.5 Thinking | Gemini 2.0 Thinking |
|---|---|---|---|---|
| AIME 2024 Score | 96.7% | 79.8% | 78.3% | 75.7% |
| ARC-AGI Score | 87.5% | 32.0% | 27.2% | 47.6% |
| Open Weights | No | Yes (MIT) | No | No |
| API Cost / 1M tokens | $15 / $60 | $0.55 / $2.19 | $3 / $15 | $1.25 / $5 |
| Coding (SWE-bench) | 71.7% | 49.2% | 64.9% | 53.6% |
| Reasoning Tokens | Adaptive | Explicit CoT | Extended CoT | Adaptive |
| Speed (tokens/sec) | ~50 | ~60 | ~40 | ~80 |
Expert Verdict
o3 leads on AIME and ARC-AGI; DeepSeek R1 offers the best open-weight reasoning at 10x lower cost; Claude excels at coding reasoning; Gemini provides the longest reasoning context.
Subject Breakdown
o3
3 Category Wins
- ★ AIME 2024 Score: 96.7%
- ★ ARC-AGI Score: 87.5%
- ★ Coding (SWE-bench): 71.7%
DeepSeek R1
2 Category Wins
- ★ Open Weights: Yes (MIT)
- ★ API Cost / 1M tokens: $0.55 / $2.19
Claude 3.5 Thinking
0 Category Wins
Gemini 2.0 Thinking
1 Category Wins
- ★ Speed (tokens/sec): ~80