Operating-point board
MathBench (competition math) · Prospective · Fβ β
| # | Model | Fβ ? | Prec ? | Rec ? | Task acc ? |
|---|---|---|---|---|---|
| 1 | Gemini 3 Pro | 1.000 | 1.00 | 1.00 | 1.00 |
| 2 | GPT-5.2 | 0.998 | 1.00 | 1.00 | 1.00 |
| 3 | Gemini 3 Flash | 0.991 | 0.98 | 1.00 | 0.98 |
| 4 | DeepSeek R1 | 0.988 | 0.98 | 1.00 | 0.98 |
| 5 | Gemini 2.5 Flash | 0.973 | 0.95 | 1.00 | 0.95 |
| 6 | Claude Haiku 4.5 | 0.972 | 0.95 | 1.00 | 0.95 |
| 7 | Claude Sonnet 4.5 | 0.967 | 0.94 | 1.00 | 0.93 |
| 8 | DeepSeek Chat | 0.962 | 0.95 | 0.97 | 0.95 |
| 9 | Gemini 2.5 Pro | 0.955 | 0.91 | 1.00 | 0.91 |
| 10 | Gemini 2.0 Flash | 0.945 | 0.90 | 1.00 | 0.89 |
| 11 | Mistral Small 3.2 | 0.923 | 0.87 | 0.98 | 0.87 |
| 12 | Mistral Medium 3.1 | 0.916 | 0.95 | 0.88 | 0.94 |
| 13 | Qwen 2.5 72B | 0.910 | 0.87 | 0.95 | 0.83 |
| 14 | GPT-4o | 0.894 | 0.81 | 1.00 | 0.80 |
| 15 | GPT-4o mini | 0.878 | 0.78 | 1.00 | 0.78 |
| 16 | Llama 3.3 | 0.853 | 0.75 | 0.99 | 0.74 |
| 17 | Llama 3.1 70B | 0.782 | 0.65 | 0.99 | 0.63 |
| 18 | Claude 3 Haiku | 0.595 | 0.43 | 0.98 | 0.41 |
Pairwise signal on MathBench
Observed correlation
0.072
mean τ-b · 0 = no signal
Baseline correlation
0.000
permutation null · ≈ 0
Model pairs
135
unit = pair, not model
Significant pairs
28.9%
p<0.05 after FDR · not effect size
Observed model pairs Base-rate-matched null Calibration-preserving null 5%-95% observed: -0.188 to 0.280