← Atlas home Company / family view 2 models in the bench

Qwen

Same public data as the home page, specialized to one familyParametric view: company_family

Qwen has the lowest family mean on the performance z-score axis while sitting near the field mean on confidence z-score. The family has the largest confidence-over-performance separation in this view. The highlighted cloud sits above the diagonal at the family mean.

Where Qwen sits in the cloud

-3-3-2-2-1-100112233Performance z-score within benchmark/probe →Confidence z-score →
Qwen model/condition points Qwen model/condition points mean x/y all other model/condition points equal relative confidence and pass rate

Disposition: where each model draws the line

0.000.250.500.751.00Qwen 2.5 Coder 32BQwen 2.5 72Bpass rateimplied confidence (red gap = overconfident)

Competitive context, benchmark by benchmark

SQuAD (factual recall) · Prospective · Fβ β
#ModelFβ ?Prec ?Rec ?Task acc ?
1 Gemini 3 Pro0.7850.770.800.65
2 Gemini 3 Flash0.7820.700.880.59
3 Gemini 2.5 Pro0.7640.710.830.58
4 Claude Sonnet 4.50.7530.780.720.59
5 GPT-4o0.7390.690.800.54
6 DeepSeek R10.7330.760.700.57
7 Claude Haiku 4.50.7160.790.650.48
8 Mistral Medium 3.10.7030.650.770.52
9 GPT-4o mini0.6860.560.880.45
10 Gemini 2.0 Flash0.6830.580.820.45
11 GPT-5.20.6730.840.560.62
12 Gemini 2.5 Flash0.6700.580.790.45
13 Mistral Small 3.20.6660.570.790.44
14 Llama 3.1 70B0.6430.650.640.45
15 Claude 3.5 Sonnet0.6430.790.540.56
16 Llama 3.30.5930.640.550.45
17 Qwen 2.5 72B0.5720.740.470.47
18 DeepSeek Chat0.5670.780.450.55
19 Claude 3 Haiku0.5540.550.560.40
20 Qwen 2.5 Coder 32B0.4760.311.000.31

The four metacognitive outcomes — Qwen

Competent claimed it could — and could
Qwen 2.5 72B · SQuAD

How often does the General Conference meet?

It said it could answer: “The General Conference, in the context of the United Methodist Church, meets every four years. This is a well-established fact within the structure of the church and does not require additional information to answer accurately.”

→ It and it did.It answered “The General Conference meets every four years.”; expected “every four years”.

Overconfident claimed it could — but couldn’t
Qwen 2.5 72B · SQuAD

Who was the 4th to rule the Yuan dynasty?

It said it could answer: “As a knowledgeable assistant, I am aware of the succession of Yuan dynasty emperors, and I can provide the name of the 4th ruler without additional information.”

→ It but it couldn't.It answered “The 4th ruler of the Yuan dynasty was Emperor Wenzong, also known as Tugh Temur.”; expected “Buyantu Khan”.

Underconfident declined — but could
Qwen 2.5 72B · SQuAD

How long had John Paul II been the pope in 1979?

It said it couldn't answer: “John Paul II was elected Pope in October 1978, so to determine how long he had been Pope in 1979, we would need to calculate the exact duration from his election date to a specific point in 1979. Without specifying a particular date in 1979, it is not possible to provide an exact answer.”

→ It but it could.It answered “John Paul II had been the pope for about 5 months in 1979.”; expected “less than a year”.

Well-declined declined — and couldn’t
Qwen 2.5 72B · SQuAD

Who were the announcers of Super Bowl 50?

It said it couldn't answer: “I do not have the specific information about the announcers for Super Bowl 50 memorized, and without looking up the details, I cannot provide an accurate answer.”

→ It and it couldn't.It answered “The announcers for Super Bowl 50 were Joe Buck and Troy Aikman.”; expected “Jim Nantz and Phil Simms”.

Rows: top = claimed it could answer · bottom = declined Columns: left = was actually correct · right = was actually wrong

Pairwise signal: pairs involving Qwen

Match accuracy controls for the performance base-rate gap
Qwen pairs
18/ 171
Qwen mean tau
+0.052
All-pairs mean
+0.037
Qwen p<0.05
11(61.1%)
-1.0-0.50.00.51.0Pair signal: do confidence gaps rank performance gaps? (Kendall tau-b)
all model pairs (observed) base-rate-matched null calibration-preserving null Qwen pair (filled = p<0.05) Qwen mean all-pairs mean
Qwen 2.5 72B →Qwen 2.5 Coder 32B → Compare against peers →