CORE Bench

The DCLM CORE metric: 22 in-context-learning tasks, each accuracy centered against random chance and averaged. The same number nanochat speedruns chase — GPT-2 1.6B sits at 0.2565. Scores are comparable only within a tier (the per-task example budget): smoke 50, standard 500, full everything. How to hook a model up.

paste into your coding agent — it implements the one scoring endpoint your model needs
CORE over timeone measurement scale at a time —
0.00.10.2GPT-2 1.6B (nanochat target)GPT-2 124MGPT-2 124M (reference)
GPT-2 124M (reference)

Standings

Each model's newest complete attempt, per tier.

ModelCORETierCheckpointWhenHistory
GPT-2 124M (reference)0.0796smokeparity gate2026-08-20 08:36attempts