The Metacognition Bench?
We rank the risk sensitive (F-β) accuracy of Large Language Model verbalised confidence with 20 models tested across 6 benchmarks. While we uncover winners and losers within each performance band, and consistent biases towards under and overconfidence, the real story is found in the inter-model structure of our results. Models share an underlying confidence heuristic, and verbalised confidences are as likely to describe other models as themselves. We found no model where differences in confidence with other models correlated with differences in performance.
See our paper for further details on the methodology and analysis
Forecast probe
Prospective: asked before attempting.
Counterfactual: asked after.
β<1 · weight overconfidence
precision matters weight underconfidence · β>1
recall matters
precision matters weight underconfidence · β>1
recall matters
Task-accuracy band = 0.30–0.65 Showing 20 / 20 models in band
Operating-point ranking
| # | Model | Fβ ? ↓ | Precision ? | Recall ? | Task acc ? |
|---|---|---|---|---|---|
| 01 | Gemini 3 Pro deep dive ↗ | 0.785 | 0.768 | 0.803 | 0.648 |
| 02 | Gemini 3 Flash deep dive ↗ | 0.782 | 0.705 | 0.879 | 0.594 |
| 03 | Gemini 2.5 Pro deep dive ↗ | 0.764 | 0.709 | 0.829 | 0.585 |
| 04 | Claude Sonnet 4.5 deep dive ↗ | 0.753 | 0.783 | 0.725 | 0.594 |
| 05 | GPT-4o deep dive ↗ | 0.739 | 0.686 | 0.801 | 0.543 |
| 06 | DeepSeek R1 deep dive ↗ | 0.733 | 0.764 | 0.705 | 0.574 |
| 07 | Claude Haiku 4.5 deep dive ↗ | 0.716 | 0.792 | 0.653 | 0.482 |
| 08 | Mistral Medium 3.1 deep dive ↗ | 0.703 | 0.650 | 0.765 | 0.515 |
| 09 | GPT-4o mini deep dive ↗ | 0.686 | 0.562 | 0.880 | 0.452 |
| 10 | Gemini 2.0 Flash deep dive ↗ | 0.683 | 0.583 | 0.824 | 0.453 |
| 11 | GPT-5.2 deep dive ↗ | 0.673 | 0.840 | 0.561 | 0.619 |
| 12 | Gemini 2.5 Flash deep dive ↗ | 0.670 | 0.580 | 0.793 | 0.445 |
| 13 | Mistral Small 3.2 deep dive ↗ | 0.666 | 0.574 | 0.792 | 0.439 |
| 14 | Llama 3.1 70B deep dive ↗ | 0.643 | 0.647 | 0.639 | 0.454 |
| 15 | Claude 3.5 Sonnet deep dive ↗ | 0.643 | 0.791 | 0.541 | 0.565 |
| 16 | Llama 3.3 deep dive ↗ | 0.593 | 0.640 | 0.553 | 0.452 |
| 17 | Qwen 2.5 72B deep dive ↗ | 0.572 | 0.741 | 0.466 | 0.474 |
| 18 | DeepSeek Chat deep dive ↗ | 0.567 | 0.776 | 0.447 | 0.552 |
| 19 | Claude 3 Haiku deep dive ↗ | 0.554 | 0.545 | 0.563 | 0.403 |
| 20 | Qwen 2.5 Coder 32B deep dive ↗ | 0.476 | 0.312 | 1.000 | 0.312 |