Positioning spread
| Benchmark | Task acc | Confidence | F₁ | Leans |
|---|---|---|---|---|
| SQuAD (factual recall) | 0.44 | 0.61 | 0.67 | +0.17 overconfident |
| MMLU-Pro (knowledge) | 0.39 | 0.90 | 0.57 | +0.51 overconfident |
| LegalBench (legal reasoning) | 0.85 | 0.86 | 0.85 | +0.01 calibrated |
| MathBench (competition math) | 0.87 | 0.98 | 0.92 | +0.11 overconfident |
| OmniMath (advanced math) | 0.44 | 0.82 | 0.63 | +0.38 overconfident |
| SciCode (scientific code) | 0.41 | 0.47 | 0.56 | +0.06 overconfident |
In the full cloud
Mistral Small 3.2 conditions all other model/condition points equal relative confidence and pass rate
Pairwise footprint
Match accuracy controls for the performance base-rate gap
Mistral Small 3.2 pairs
18/ 171
Mistral Small 3.2 mean tau
+0.038
All-pairs mean
+0.037
Mistral Small 3.2 p<0.05
8(44.4%)
all model pairs (observed) base-rate-matched null calibration-preserving null Mistral Small 3.2 pair (filled = p<0.05) Mistral Small 3.2 mean all-pairs mean