Positioning spread
| Benchmark | Task acc | Confidence | F₁ | Leans |
|---|---|---|---|---|
| SQuAD (factual recall) | 0.54 | 0.63 | 0.74 | +0.09 overconfident |
| MMLU-Pro (knowledge) | 0.47 | 0.87 | 0.66 | +0.40 overconfident |
| LegalBench (legal reasoning) | 0.83 | 0.45 | 0.58 | -0.38 cautious |
| MathBench (competition math) | 0.80 | 0.99 | 0.89 | +0.19 overconfident |
| OmniMath (advanced math) | 0.30 | 0.58 | 0.52 | +0.28 overconfident |
| SciCode (scientific code) | 0.51 | 0.39 | 0.56 | -0.12 cautious |
In the full cloud
GPT-4o conditions all other model/condition points equal relative confidence and pass rate
Pairwise footprint
Match accuracy controls for the performance base-rate gap
GPT-4o pairs
18/ 171
GPT-4o mean tau
+0.041
All-pairs mean
+0.037
GPT-4o p<0.05
11(61.1%)
all model pairs (observed) base-rate-matched null calibration-preserving null GPT-4o pair (filled = p<0.05) GPT-4o mean all-pairs mean