← Atlas home Benchmark view 20 models

MMLU-Pro (knowledge)

Unaided means: answer without seeing the multiple-choice optionsProbes: Prospective + Counterfactual

Operating-point board

MMLU-Pro (knowledge) · Prospective · Fβ β
#ModelFβ ?Prec ?Rec ?Task acc ?
1 GPT-5.20.8610.820.910.73
2 Gemini 2.5 Pro0.8220.730.940.67
3 Claude Sonnet 4.50.8190.780.870.70
4 Gemini 3 Pro0.8140.750.890.66
5 DeepSeek R10.7990.720.900.65
6 Gemini 3 Flash0.7780.650.980.61
7 Claude Haiku 4.50.7440.730.760.65
8 Claude 3.5 Sonnet0.7050.580.890.54
9 DeepSeek Chat0.6930.550.950.51
10 GPT-4o0.6590.510.940.47
11 Mistral Medium 3.10.6480.520.870.48
12 Gemini 2.5 Flash0.6360.480.960.46
13 Gemini 2.0 Flash0.6230.451.000.45
14 GPT-4o mini0.5950.430.950.40
15 Mistral Small 3.20.5730.410.950.39
16 Qwen 2.5 72B0.5700.460.760.41
17 Llama 3.1 70B0.5490.390.950.36
18 Llama 3.30.5250.360.940.34
19 Claude 3 Haiku0.5120.350.930.35
20 Qwen 2.5 Coder 32B0.4050.251.000.25

Pairwise signal on MMLU-Pro

Observed correlation
0.041
mean τ-b · 0 = no signal
Baseline correlation
0.001
permutation null · ≈ 0
Model pairs
190
unit = pair, not model
Significant pairs
35.8%
p<0.05 after FDR · not effect size
-1.0-0.50.00.51.0Pair signal: do confidence gaps rank performance gaps? (Kendall tau-b)
Observed model pairs Base-rate-matched null Calibration-preserving null 5%-95% observed: -0.030 to 0.130

Confidence structure: the shared-difficulty factor

PC1 = 48.3% of variance?Eigenvalue rank →Variance share20 models