The Metacognition Bench?

We rank the risk sensitive (F-β) accuracy of Large Language Model verbalised confidence with 20 models tested across 6 benchmarks. While we uncover winners and losers within each performance band, and consistent biases towards under and overconfidence, the real story is found in the inter-model structure of our results. Models share an underlying confidence heuristic, and verbalised confidences are as likely to describe other models as themselves. We found no model where differences in confidence with other models correlated with differences in performance.

Unaided means: answer the question without the supporting context passage.
Open the SQuAD page →
Forecast probe
Prospective: asked before attempting. Counterfactual: asked after.
β<1 · weight overconfidence
precision matters
weight underconfidence · β>1
recall matters
Task-accuracy band = 0.30–0.65 Showing 20 / 20 models in band

Operating-point ranking

#Model Fβ ? Precision ? Recall ? Task acc ?
01Gemini 3 Pro deep dive ↗0.7850.7680.8030.648
02Gemini 3 Flash deep dive ↗0.7820.7050.8790.594
03Gemini 2.5 Pro deep dive ↗0.7640.7090.8290.585
04Claude Sonnet 4.5 deep dive ↗0.7530.7830.7250.594
05GPT-4o deep dive ↗0.7390.6860.8010.543
06DeepSeek R1 deep dive ↗0.7330.7640.7050.574
07Claude Haiku 4.5 deep dive ↗0.7160.7920.6530.482
08Mistral Medium 3.1 deep dive ↗0.7030.6500.7650.515
09GPT-4o mini deep dive ↗0.6860.5620.8800.452
10Gemini 2.0 Flash deep dive ↗0.6830.5830.8240.453
11GPT-5.2 deep dive ↗0.6730.8400.5610.619
12Gemini 2.5 Flash deep dive ↗0.6700.5800.7930.445
13Mistral Small 3.2 deep dive ↗0.6660.5740.7920.439
14Llama 3.1 70B deep dive ↗0.6430.6470.6390.454
15Claude 3.5 Sonnet deep dive ↗0.6430.7910.5410.565
16Llama 3.3 deep dive ↗0.5930.6400.5530.452
17Qwen 2.5 72B deep dive ↗0.5720.7410.4660.474
18DeepSeek Chat deep dive ↗0.5670.7760.4470.552
19Claude 3 Haiku deep dive ↗0.5540.5450.5630.403
20Qwen 2.5 Coder 32B deep dive ↗0.4760.3121.0000.312