← Atlas home Company / family view 3 models in the bench

OpenAI · GPT

Same public data as the home page, specialized to one familyParametric view: company_family

GPT sits close to the field mean on performance z-score and slightly above the field mean on confidence z-score. The family profile spans a wider performance range than confidence range. The family mean sits slightly above the diagonal, indicating higher relative confidence than relative performance.

Where OpenAI · GPT sits in the cloud

-3-3-2-2-1-100112233Performance z-score within benchmark/probe →Confidence z-score →
OpenAI · GPT model/condition points OpenAI · GPT model/condition points mean x/y all other model/condition points equal relative confidence and pass rate

Disposition: where each model draws the line

0.000.250.500.751.00GPT-4o miniGPT-4oGPT-5.2pass rateimplied confidence (red gap = overconfident)

Competitive context, benchmark by benchmark

SQuAD (factual recall) · Prospective · Fβ β
#ModelFβ ?Prec ?Rec ?Task acc ?
1 Gemini 3 Pro0.7850.770.800.65
2 Gemini 3 Flash0.7820.700.880.59
3 Gemini 2.5 Pro0.7640.710.830.58
4 Claude Sonnet 4.50.7530.780.720.59
5 GPT-4o0.7390.690.800.54
6 DeepSeek R10.7330.760.700.57
7 Claude Haiku 4.50.7160.790.650.48
8 Mistral Medium 3.10.7030.650.770.52
9 GPT-4o mini0.6860.560.880.45
10 Gemini 2.0 Flash0.6830.580.820.45
11 GPT-5.20.6730.840.560.62
12 Gemini 2.5 Flash0.6700.580.790.45
13 Mistral Small 3.20.6660.570.790.44
14 Llama 3.1 70B0.6430.650.640.45
15 Claude 3.5 Sonnet0.6430.790.540.56
16 Llama 3.30.5930.640.550.45
17 Qwen 2.5 72B0.5720.740.470.47
18 DeepSeek Chat0.5670.780.450.55
19 Claude 3 Haiku0.5540.550.560.40
20 Qwen 2.5 Coder 32B0.4760.311.000.31

The four metacognitive outcomes — OpenAI · GPT

Competent claimed it could — and could
GPT-4o · SQuAD

How many universities does Newcastle have?

It said it could answer: “I have knowledge of the universities located in Newcastle, which include Newcastle University and Northumbria University. Therefore, I can provide an accurate count of the universities in Newcastle.”

→ It and it did.It answered “Newcastle has two universities: Newcastle University and Northumbria University.”; expected “two”.

Overconfident claimed it could — but couldn’t
GPT-4o · SQuAD

What's the nickname for Newcastle's A167?

It said it could answer: “The A167 in Newcastle is commonly known as the 'Central Motorway'. This is a well-known nickname for this road, and I can provide this information based on existing knowledge.”

→ It but it couldn't.It answered “The nickname for Newcastle's A167 is "The Central Motorway."”; expected “the old "Great North Road"”.

Underconfident declined — but could
GPT-5.2 · SQuAD

How many Catholic schools were in Victoria?

It said it couldn't answer: “The number of Catholic schools in Victoria depends on the timeframe (specific year), what counts as a Catholic school (primary/secondary, special schools, campuses), and the data source. Without that context, any number would be a guess.”

→ It but it could.It answered “There were 489 Catholic schools in Victoria.”; expected “489”.

Well-declined declined — and couldn’t
GPT-4o · MMLU-Pro

Find the last 3 digits of 2003^(2002^2001).

It said it couldn't answer: “To find the last 3 digits of a number, we need to compute it modulo 1000. This involves using modular arithmetic techniques such as Euler's theorem or the Chinese Remainder Theorem. Given the large exponents, the problem can be simplified using properties of modular arithmetic and possibly the Carmichael function to reduce the exponent. However, without performing detailed calculations, it's not immediately clear if the solution can be found without error. Therefore, while I can attempt to solve it, the complexity suggests that having multiple-choice options would help verify the solution.”

→ It and it couldn't.It answered “947”; expected “241”.

Rows: top = claimed it could answer · bottom = declined Columns: left = was actually correct · right = was actually wrong

Pairwise signal: pairs involving OpenAI · GPT

Match accuracy controls for the performance base-rate gap
OpenAI · GPT pairs
51/ 171
OpenAI · GPT mean tau
+0.025
All-pairs mean
+0.037
OpenAI · GPT p<0.05
19(37.3%)
-1.0-0.50.00.51.0Pair signal: do confidence gaps rank performance gaps? (Kendall tau-b)
all model pairs (observed) base-rate-matched null calibration-preserving null OpenAI · GPT pair (filled = p<0.05) OpenAI · GPT mean all-pairs mean
GPT-4o →GPT-4o mini →GPT-5.2 → Compare against peers →