Positioning spread
| Benchmark | Task acc | Confidence | F₁ | Leans |
|---|---|---|---|---|
| SQuAD (factual recall) | 0.31 | 1.00 | 0.48 | +0.69 overconfident |
| MMLU-Pro (knowledge) | 0.25 | 1.00 | 0.41 | +0.75 overconfident |
In the full cloud
Qwen 2.5 Coder 32B conditions all other model/condition points equal relative confidence and pass rate
Pairwise footprint
Match accuracy controls for the performance base-rate gap
Qwen 2.5 Coder 32B pairs
19/ 190
Qwen 2.5 Coder 32B mean tau
+0.127
All-pairs mean
+0.041
Qwen 2.5 Coder 32B p<0.05
17(89.5%)
all model pairs (observed) base-rate-matched null calibration-preserving null Qwen 2.5 Coder 32B pair (filled = p<0.05) Qwen 2.5 Coder 32B mean all-pairs mean