← Qwen Individual model view Qwen

Qwen 2.5 72B

Appears in 6 benchmarksMean lean (confidence − pass rate): -0.013/6 benchmarks lean overconfident (prospective probe)

Positioning spread

0.000.250.500.751.00SQuAD (factual recall)MMLU-Pro (knowledge)LegalBench (legal reasoning)MathBench (competition math)OmniMath (advanced math)SciCode (scientific code)performanceconfidence (red gap = overconfident)
Performance vs. confidence for Qwen 2.5 72B, per benchmark (prospective probe).
BenchmarkTask accConfidenceF₁Leans
SQuAD (factual recall)0.470.300.57-0.18 cautious
MMLU-Pro (knowledge)0.410.690.57+0.28 overconfident
LegalBench (legal reasoning)0.840.510.63-0.33 cautious
MathBench (competition math)0.830.910.91+0.08 overconfident
OmniMath (advanced math)0.340.430.57+0.10 overconfident
SciCode (scientific code)0.430.420.56-0.01 calibrated

In the full cloud

-3-3-2-2-1-100112233Performance z-score within benchmark/probe →Confidence z-score →
Qwen 2.5 72B conditions all other model/condition points equal relative confidence and pass rate

Pairwise footprint

Match accuracy controls for the performance base-rate gap
Qwen 2.5 72B pairs
18/ 171
Qwen 2.5 72B mean tau
+0.052
All-pairs mean
+0.037
Qwen 2.5 72B p<0.05
11(61.1%)
-1.0-0.50.00.51.0Pair signal: do confidence gaps rank performance gaps? (Kendall tau-b)
all model pairs (observed) base-rate-matched null calibration-preserving null Qwen 2.5 72B pair (filled = p<0.05) Qwen 2.5 72B mean all-pairs mean
vs Qwen 2.5 Coder 32B → Compare with anything →