← Google · Gemini Individual model view Google · Gemini

Gemini 2.5 Flash

Appears in 6 benchmarksMean lean (confidence − pass rate): +0.154/6 benchmarks lean overconfident (prospective probe)

Positioning spread

0.000.250.500.751.00SQuAD (factual recall)MMLU-Pro (knowledge)LegalBench (legal reasoning)MathBench (competition math)OmniMath (advanced math)SciCode (scientific code)performanceconfidence (red gap = overconfident)
Performance vs. confidence for Gemini 2.5 Flash, per benchmark (prospective probe).
BenchmarkTask accConfidenceF₁Leans
SQuAD (factual recall)0.450.610.67+0.16 overconfident
MMLU-Pro (knowledge)0.460.920.64+0.46 overconfident
LegalBench (legal reasoning)0.840.780.81-0.06 cautious
MathBench (competition math)0.951.000.97+0.05 overconfident
OmniMath (advanced math)0.640.990.79+0.35 overconfident
SciCode (scientific code)0.560.470.52-0.08 cautious

In the full cloud

-3-3-2-2-1-100112233Performance z-score within benchmark/probe →Confidence z-score →
Gemini 2.5 Flash conditions all other model/condition points equal relative confidence and pass rate

Pairwise footprint

Match accuracy controls for the performance base-rate gap
Gemini 2.5 Flash pairs
18/ 171
Gemini 2.5 Flash mean tau
+0.006
All-pairs mean
+0.037
Gemini 2.5 Flash p<0.05
6(33.3%)
-1.0-0.50.00.51.0Pair signal: do confidence gaps rank performance gaps? (Kendall tau-b)
all model pairs (observed) base-rate-matched null calibration-preserving null Gemini 2.5 Flash pair (filled = p<0.05) Gemini 2.5 Flash mean all-pairs mean
vs Gemini 2.0 Flash →vs Gemini 2.5 Pro →vs Gemini 3 Flash →vs Gemini 3 Pro → Compare with anything →