Positioning spread
| Benchmark | Task acc | Confidence | F₁ | Leans |
|---|---|---|---|---|
| SQuAD (factual recall) | 0.58 | 0.68 | 0.76 | +0.10 overconfident |
| MMLU-Pro (knowledge) | 0.67 | 0.87 | 0.82 | +0.20 overconfident |
| LegalBench (legal reasoning) | 0.87 | 0.79 | 0.82 | -0.09 cautious |
| MathBench (competition math) | 0.91 | 1.00 | 0.95 | +0.09 overconfident |
| OmniMath (advanced math) | 0.70 | 0.99 | 0.83 | +0.29 overconfident |
| SciCode (scientific code) | 0.61 | 0.38 | 0.54 | -0.23 cautious |
In the full cloud
Gemini 2.5 Pro conditions all other model/condition points equal relative confidence and pass rate
Pairwise footprint
Match accuracy controls for the performance base-rate gap
Gemini 2.5 Pro pairs
18/ 171
Gemini 2.5 Pro mean tau
+0.053
All-pairs mean
+0.037
Gemini 2.5 Pro p<0.05
12(66.7%)
all model pairs (observed) base-rate-matched null calibration-preserving null Gemini 2.5 Pro pair (filled = p<0.05) Gemini 2.5 Pro mean all-pairs mean