← Atlas home Benchmark view 18 models

SciCode (scientific code)

Unaided means: implement the function without the extra background factsProbes: Prospective

Operating-point board

SciCode (scientific code) · Prospective · Fβ β
#ModelFβ ?Prec ?Rec ?Task acc ?
1 GPT-5.20.7390.700.790.58
2 Claude Haiku 4.50.7300.680.790.56
3 DeepSeek Chat0.6840.620.770.55
4 DeepSeek R10.6770.590.800.56
5 Gemini 2.0 Flash0.6530.500.940.51
6 Gemini 3 Flash0.6350.680.600.59
7 Gemini 3 Pro0.6230.720.550.64
8 Claude 3 Haiku0.5960.421.000.40
9 Llama 3.1 70B0.5940.510.710.44
10 GPT-4o0.5650.650.500.51
11 Mistral Small 3.20.5580.520.600.41
12 Qwen 2.5 72B0.5570.560.550.43
13 GPT-4o mini0.5560.540.570.37
14 Gemini 2.5 Pro0.5420.700.440.61
15 Gemini 2.5 Flash0.5200.560.480.56
16 Llama 3.30.5120.520.500.46
17 Claude Sonnet 4.50.4280.640.320.57
18 Mistral Medium 3.10.3380.550.240.48

Pairwise signal on SciCode

Observed correlation
0.049
mean τ-b · 0 = no signal
Baseline correlation
-0.000
permutation null · ≈ 0
Model pairs
153
unit = pair, not model
Significant pairs
11.1%
p<0.05 after FDR · not effect size
-1.0-0.50.00.51.0Pair signal: do confidence gaps rank performance gaps? (Kendall tau-b)
Observed model pairs Base-rate-matched null Calibration-preserving null 5%-95% observed: -0.100 to 0.205

Confidence structure: the shared-difficulty factor

0PC1 = 37.3% of variance??Eigenvalue rank →Variance share19 models · 4 negative eigenvalues