← Atlas home Company / family view 2 models in the bench

Meta · Llama

Same public data as the home page, specialized to one familyParametric view: company_family

Llama sits below the field mean on both performance and confidence z-score axes. Confidence is less far below the field mean than performance. The family mean therefore sits above the diagonal despite both axes remaining below zero.

Where Meta · Llama sits in the cloud

-3-3-2-2-1-100112233Performance z-score within benchmark/probe →Confidence z-score →
Meta · Llama model/condition points Meta · Llama model/condition points mean x/y all other model/condition points equal relative confidence and pass rate

Disposition: where each model draws the line

0.000.250.500.751.00Llama 3.1 70BLlama 3.3pass rateimplied confidence (red gap = overconfident)

Competitive context, benchmark by benchmark

SQuAD (factual recall) · Prospective · Fβ β
#ModelFβ ?Prec ?Rec ?Task acc ?
1 Gemini 3 Pro0.7850.770.800.65
2 Gemini 3 Flash0.7820.700.880.59
3 Gemini 2.5 Pro0.7640.710.830.58
4 Claude Sonnet 4.50.7530.780.720.59
5 GPT-4o0.7390.690.800.54
6 DeepSeek R10.7330.760.700.57
7 Claude Haiku 4.50.7160.790.650.48
8 Mistral Medium 3.10.7030.650.770.52
9 GPT-4o mini0.6860.560.880.45
10 Gemini 2.0 Flash0.6830.580.820.45
11 GPT-5.20.6730.840.560.62
12 Gemini 2.5 Flash0.6700.580.790.45
13 Mistral Small 3.20.6660.570.790.44
14 Llama 3.1 70B0.6430.650.640.45
15 Claude 3.5 Sonnet0.6430.790.540.56
16 Llama 3.30.5930.640.550.45
17 Qwen 2.5 72B0.5720.740.470.47
18 DeepSeek Chat0.5670.780.450.55
19 Claude 3 Haiku0.5540.550.560.40
20 Qwen 2.5 Coder 32B0.4760.311.000.31

The four metacognitive outcomes — Meta · Llama

Competent claimed it could — and could
Llama 3.1 70B · SQuAD

What function is a teacher's role similar to?

It said it could answer: “The question is general and doesn't require specific context or information. A teacher's role can be compared to various functions based on common knowledge and general understanding of their responsibilities.”

→ It and it did.It answered “A teacher's role is similar to that of a coach, mentor, or facilitator.”; expected “a coach”.

Overconfident claimed it could — but couldn’t
Llama 3.3 · MMLU-Pro

How many kcal are there in one gram of ethanol?

It said it could answer: “The question requires knowledge of the energy content of ethanol, which is a specific and well-defined value. One gram of ethanol contains approximately 7 kcal of energy. This information is widely available and can be recalled without needing to see the multiple-choice options.”

→ It but it couldn't.It answered “29 kcal”; expected “29.7 kJ or 7.1 kcal per g”.

Underconfident declined — but could
Llama 3.1 70B · SQuAD

How much of the sun is made up of oxygen?

It said it couldn't answer: “The composition of the sun is a complex topic that requires specific knowledge of astrophysics and the sun's elemental abundance. Without additional information or context, it's unlikely to provide an accurate answer.”

→ It but it could.It answered “About 1% of the sun's mass is made up of oxygen.”; expected “0.9%”.

Well-declined declined — and couldn’t
Llama 3.3 · SQuAD

Why did al-Qaeda tell ISIL to take a hike?

It said it couldn't answer: “The question requires specific knowledge about the interactions and relationships between al-Qaeda and ISIL, which may not be widely known or may require up-to-date information on their activities and communications.”

→ It and it couldn't.It answered “Due to ideological and strategic differences.”; expected “notorious intransigence”.

Rows: top = claimed it could answer · bottom = declined Columns: left = was actually correct · right = was actually wrong

Pairwise signal: pairs involving Meta · Llama

Match accuracy controls for the performance base-rate gap
Meta · Llama pairs
35/ 171
Meta · Llama mean tau
+0.043
All-pairs mean
+0.037
Meta · Llama p<0.05
16(45.7%)
-1.0-0.50.00.51.0Pair signal: do confidence gaps rank performance gaps? (Kendall tau-b)
all model pairs (observed) base-rate-matched null calibration-preserving null Meta · Llama pair (filled = p<0.05) Meta · Llama mean all-pairs mean
Llama 3.1 70B →Llama 3.3 → Compare against peers →