Positioning spread
| Benchmark | Task acc | Confidence | F₁ | Leans |
|---|---|---|---|---|
| SQuAD (factual recall) | 0.59 | 0.55 | 0.75 | -0.04 calibrated |
| MMLU-Pro (knowledge) | 0.70 | 0.79 | 0.82 | +0.09 overconfident |
| LegalBench (legal reasoning) | 0.86 | 0.73 | 0.79 | -0.13 cautious |
| MathBench (competition math) | 0.93 | 1.00 | 0.97 | +0.06 overconfident |
| OmniMath (advanced math) | 0.57 | 0.75 | 0.72 | +0.18 overconfident |
| SciCode (scientific code) | 0.57 | 0.29 | 0.43 | -0.29 cautious |
In the full cloud
Claude Sonnet 4.5 conditions all other model/condition points equal relative confidence and pass rate
Pairwise footprint
Match accuracy controls for the performance base-rate gap
Claude Sonnet 4.5 pairs
18/ 171
Claude Sonnet 4.5 mean tau
+0.036
All-pairs mean
+0.037
Claude Sonnet 4.5 p<0.05
10(55.6%)
all model pairs (observed) base-rate-matched null calibration-preserving null Claude Sonnet 4.5 pair (filled = p<0.05) Claude Sonnet 4.5 mean all-pairs mean