Positioning spread
| Benchmark | Task acc | Confidence | F₁ | Leans |
|---|---|---|---|---|
| SQuAD (factual recall) | 0.56 | 0.39 | 0.64 | -0.18 cautious |
| MMLU-Pro (knowledge) | 0.54 | 0.83 | 0.70 | +0.28 overconfident |
| LegalBench (legal reasoning) | 0.86 | 0.73 | 0.82 | -0.13 cautious |
In the full cloud
Claude 3.5 Sonnet conditions all other model/condition points equal relative confidence and pass rate
Pairwise footprint
Match accuracy controls for the performance base-rate gap
Claude 3.5 Sonnet pairs
18/ 171
Claude 3.5 Sonnet mean tau
+0.006
All-pairs mean
+0.037
Claude 3.5 Sonnet p<0.05
4(22.2%)
all model pairs (observed) base-rate-matched null calibration-preserving null Claude 3.5 Sonnet pair (filled = p<0.05) Claude 3.5 Sonnet mean all-pairs mean