CORE Bench
The DCLM CORE metric: 22 in-context-learning tasks, each accuracy centered against random chance and averaged. The same number nanochat speedruns chase — GPT-2 1.6B sits at 0.2565. Scores are comparable only within a tier (the per-task example budget): smoke 50, standard 500, full everything. How to hook a model up.
paste into your coding agent — it implements the one scoring endpoint your model needs
CORE over timeone measurement scale at a time —
GPT-2 124M (reference)