Positioning spread
| Benchmark | Task acc | Confidence | F₁ | Leans |
|---|---|---|---|---|
| SQuAD (factual recall) | 0.45 | 0.39 | 0.59 | -0.06 cautious |
| MMLU-Pro (knowledge) | 0.34 | 0.89 | 0.53 | +0.55 overconfident |
| LegalBench (legal reasoning) | 0.88 | 0.68 | 0.76 | -0.20 cautious |
| MathBench (competition math) | 0.74 | 0.98 | 0.85 | +0.24 overconfident |
| OmniMath (advanced math) | 0.33 | 0.51 | 0.54 | +0.18 overconfident |
| SciCode (scientific code) | 0.46 | 0.44 | 0.51 | -0.02 calibrated |
In the full cloud
Llama 3.3 conditions all other model/condition points equal relative confidence and pass rate
Pairwise footprint
Match accuracy controls for the performance base-rate gap
Llama 3.3 pairs
18/ 171
Llama 3.3 mean tau
+0.034
All-pairs mean
+0.037
Llama 3.3 p<0.05
5(27.8%)
all model pairs (observed) base-rate-matched null calibration-preserving null Llama 3.3 pair (filled = p<0.05) Llama 3.3 mean all-pairs mean