← Atlas home Company / family view 2 models in the bench

Mistral

Same public data as the home page, specialized to one familyParametric view: company_family

Mistral sits below the field mean on both performance and confidence z-score axes. The family cloud spans both sides of the diagonal. The family mean is slightly below the diagonal because confidence sits lower than performance in aggregate.

Where Mistral sits in the cloud

-3-3-2-2-1-100112233Performance z-score within benchmark/probe →Confidence z-score →
Mistral model/condition points Mistral model/condition points mean x/y all other model/condition points equal relative confidence and pass rate

Disposition: where each model draws the line

0.000.250.500.751.00Mistral Small 3.2Mistral Medium 3.1pass rateimplied confidence (red gap = overconfident)

Competitive context, benchmark by benchmark

SQuAD (factual recall) · Prospective · Fβ β
#ModelFβ ?Prec ?Rec ?Task acc ?
1 Gemini 3 Pro0.7850.770.800.65
2 Gemini 3 Flash0.7820.700.880.59
3 Gemini 2.5 Pro0.7640.710.830.58
4 Claude Sonnet 4.50.7530.780.720.59
5 GPT-4o0.7390.690.800.54
6 DeepSeek R10.7330.760.700.57
7 Claude Haiku 4.50.7160.790.650.48
8 Mistral Medium 3.10.7030.650.770.52
9 GPT-4o mini0.6860.560.880.45
10 Gemini 2.0 Flash0.6830.580.820.45
11 GPT-5.20.6730.840.560.62
12 Gemini 2.5 Flash0.6700.580.790.45
13 Mistral Small 3.20.6660.570.790.44
14 Llama 3.1 70B0.6430.650.640.45
15 Claude 3.5 Sonnet0.6430.790.540.56
16 Llama 3.30.5930.640.550.45
17 Qwen 2.5 72B0.5720.740.470.47
18 DeepSeek Chat0.5670.780.450.55
19 Claude 3 Haiku0.5540.550.560.40
20 Qwen 2.5 Coder 32B0.4760.311.000.31

The four metacognitive outcomes — Mistral

Competent claimed it could — and could
Mistral Medium 3.1 · SQuAD

What are some proposals to connect campuses?

It said it could answer: “The question is general and can be answered with common knowledge about typical proposals for connecting campuses, such as shuttle services, bike paths, or technological solutions.”

→ It and it did.It answered “1. Shuttle Services: Dedicated buses between campuses. 2. Bike/Lanes & Rentals: Bike paths and shared bike programs. 3. Pedestrian Walkways: Safe, covered paths for walking. 4. Light Rail/Trams: Public transit links. 5. Ca”; expected “new and enlarged bridges, a shuttle service and/or a tram.”.

Overconfident claimed it could — but couldn’t
Mistral Small 3.2 · SQuAD

Who built a castle in Newcastle in 1080?

It said it could answer: “The castle in Newcastle was built by William the Conqueror in 1080 as part of his efforts to control the region after the Norman Conquest of England. This is a well-documented historical fact.”

→ It but it couldn't.It answered “William the Conqueror built a castle in Newcastle in 1080.”; expected “Robert Curthose”.

Underconfident declined — but could
Mistral Small 3.2 · SQuAD

What was suggested at the Symposium in 1967

It said it couldn't answer: “The question lacks specific context about the Symposium in 1967. Without knowing the event's topic, participants, or field (e.g., science, philosophy, politics), it's impossible to determine what was suggested.”

→ It but it could.It answered “The use of computers to automate the process of cataloging and retrieving information was suggested at the Symposium in 1967.”; expected “suggested it for use in the ARPANET”.

Well-declined declined — and couldn’t
Mistral Small 3.2 · SQuAD

What was William Johnson's Iroquois name?

It said it couldn't answer: “William Johnson, a prominent British Superintendent of Indian Affairs in the 18th century, was given an Iroquois name, but the exact name is not widely known or documented in general knowledge. Additional research or specific historical sources would be required to provide the correct answer.”

→ It and it couldn't.It answered “William Johnson's Iroquois name was "Warrenstaden."”; expected “Warraghiggey, meaning "He who does great things."”.

Rows: top = claimed it could answer · bottom = declined Columns: left = was actually correct · right = was actually wrong

Pairwise signal: pairs involving Mistral

Match accuracy controls for the performance base-rate gap
Mistral pairs
35/ 171
Mistral mean tau
+0.032
All-pairs mean
+0.037
Mistral p<0.05
14(40%)
-1.0-0.50.00.51.0Pair signal: do confidence gaps rank performance gaps? (Kendall tau-b)
all model pairs (observed) base-rate-matched null calibration-preserving null Mistral pair (filled = p<0.05) Mistral mean all-pairs mean
Mistral Medium 3.1 →Mistral Small 3.2 → Compare against peers →