← Atlas home Company / family view 2 models in the bench

DeepSeek

Same public data as the home page, specialized to one familyParametric view: company_family

DeepSeek sits above the field mean on both performance and confidence z-score axes. Relative performance is higher than relative confidence at the family mean. The highlighted cloud therefore falls below the diagonal on the z-score map.

Where DeepSeek sits in the cloud

-3-3-2-2-1-100112233Performance z-score within benchmark/probe →Confidence z-score →
DeepSeek model/condition points DeepSeek model/condition points mean x/y all other model/condition points equal relative confidence and pass rate

Disposition: where each model draws the line

0.000.250.500.751.00DeepSeek R1DeepSeek Chatpass rateimplied confidence (red gap = overconfident)

Competitive context, benchmark by benchmark

SQuAD (factual recall) · Prospective · Fβ β
#ModelFβ ?Prec ?Rec ?Task acc ?
1 Gemini 3 Pro0.7850.770.800.65
2 Gemini 3 Flash0.7820.700.880.59
3 Gemini 2.5 Pro0.7640.710.830.58
4 Claude Sonnet 4.50.7530.780.720.59
5 GPT-4o0.7390.690.800.54
6 DeepSeek R10.7330.760.700.57
7 Claude Haiku 4.50.7160.790.650.48
8 Mistral Medium 3.10.7030.650.770.52
9 GPT-4o mini0.6860.560.880.45
10 Gemini 2.0 Flash0.6830.580.820.45
11 GPT-5.20.6730.840.560.62
12 Gemini 2.5 Flash0.6700.580.790.45
13 Mistral Small 3.20.6660.570.790.44
14 Llama 3.1 70B0.6430.650.640.45
15 Claude 3.5 Sonnet0.6430.790.540.56
16 Llama 3.30.5930.640.550.45
17 Qwen 2.5 72B0.5720.740.470.47
18 DeepSeek Chat0.5670.780.450.55
19 Claude 3 Haiku0.5540.550.560.40
20 Qwen 2.5 Coder 32B0.4760.311.000.31

The four metacognitive outcomes — DeepSeek

Competent claimed it could — and could
DeepSeek Chat · SQuAD

What did the Greek root pharmakos imply?

It said it could answer: “The Greek root 'pharmakos' is well-documented in linguistic and historical contexts. It originally referred to a scapegoat or a person who was ritually expelled or sacrificed to purify a community. Over time, it evolved to be associated with medicine and healing, leading to the modern term 'pharmacy.' This knowledge is based on established etymological and historical sources.”

→ It and it did.It answered “The Greek root pharmakos implied both "remedy" and "poison," reflecting its dual meaning of healing and harm. It also referred to a scapegoat or sacrificial victim in religious rituals.”; expected “sorcery or even poison”.

Overconfident claimed it could — but couldn’t
DeepSeek Chat · SQuAD

Of what is ozone a reactive part of oxygen?

It said it could answer: “Ozone is a reactive form of oxygen consisting of three oxygen atoms (O₃). It is a well-known chemical compound, and this information is fundamental in chemistry.”

→ It but it couldn't.It answered “Ozone (O₃) is a reactive form of oxygen composed of three oxygen atoms. It is distinct from the diatomic oxygen (O₂) that we breathe.”; expected “allotrope”.

Underconfident declined — but could
DeepSeek Chat · SQuAD

In what venue did Super Bowl XIX take place?

It said it couldn't answer: “The question asks for the venue of Super Bowl XIX, which is a specific historical fact. Without additional information or access to external data, I cannot provide the correct answer.”

→ It but it could.It answered “Super Bowl XIX took place at Stanford Stadium in Stanford, California.”; expected “Stanford Stadium”.

Well-declined declined — and couldn’t
DeepSeek Chat · SQuAD

What did BSkyB name their interactive service?

It said it couldn't answer: “The question asks for the name of BSkyB's interactive service, which is specific knowledge about a company's product. Without additional information or context, it is not possible to provide the correct answer.”

→ It and it couldn't.It answered “BSkyB named their interactive service Sky Interactive.”; expected “Sky Active”.

Rows: top = claimed it could answer · bottom = declined Columns: left = was actually correct · right = was actually wrong

Pairwise signal: pairs involving DeepSeek

Match accuracy controls for the performance base-rate gap
DeepSeek pairs
35/ 171
DeepSeek mean tau
+0.019
All-pairs mean
+0.037
DeepSeek p<0.05
11(31.4%)
-1.0-0.50.00.51.0Pair signal: do confidence gaps rank performance gaps? (Kendall tau-b)
all model pairs (observed) base-rate-matched null calibration-preserving null DeepSeek pair (filled = p<0.05) DeepSeek mean all-pairs mean
DeepSeek Chat →DeepSeek R1 → Compare against peers →