← Atlas home Benchmark view 18 models

MathBench (competition math)

Unaided means: solve the problem without hints or worked solutionProbes: Prospective + Counterfactual

Operating-point board

MathBench (competition math) · Prospective · Fβ β
#ModelFβ ?Prec ?Rec ?Task acc ?
1 Gemini 3 Pro1.0001.001.001.00
2 GPT-5.20.9981.001.001.00
3 Gemini 3 Flash0.9910.981.000.98
4 DeepSeek R10.9880.981.000.98
5 Gemini 2.5 Flash0.9730.951.000.95
6 Claude Haiku 4.50.9720.951.000.95
7 Claude Sonnet 4.50.9670.941.000.93
8 DeepSeek Chat0.9620.950.970.95
9 Gemini 2.5 Pro0.9550.911.000.91
10 Gemini 2.0 Flash0.9450.901.000.89
11 Mistral Small 3.20.9230.870.980.87
12 Mistral Medium 3.10.9160.950.880.94
13 Qwen 2.5 72B0.9100.870.950.83
14 GPT-4o0.8940.811.000.80
15 GPT-4o mini0.8780.781.000.78
16 Llama 3.30.8530.750.990.74
17 Llama 3.1 70B0.7820.650.990.63
18 Claude 3 Haiku0.5950.430.980.41

Pairwise signal on MathBench

Observed correlation
0.072
mean τ-b · 0 = no signal
Baseline correlation
0.000
permutation null · ≈ 0
Model pairs
135
unit = pair, not model
Significant pairs
28.9%
p<0.05 after FDR · not effect size
-1.0-0.50.00.51.0Pair signal: do confidence gaps rank performance gaps? (Kendall tau-b)
Observed model pairs Base-rate-matched null Calibration-preserving null 5%-95% observed: -0.188 to 0.280

Confidence structure: the shared-difficulty factor

PC1 = 25.3% of variance?Eigenvalue rank →Variance share19 models