CORE Bench
DCLM COREThe 22-task centered-accuracy metric from the DCLM paper — the number nanochat speedruns chase (GPT-2 1.6B = 0.2565, GPT-2 124M = 0.1139) — with every attempt stored forever: per-task scores, every example's outcome, config, and the exact benchmark fingerprint it ran under.
Why your chat endpoint can't do this
CORE is loglikelihood evaluation: the benchmark supplies the text and asks what probability your model assigns to it — the mean per-token loss of each given answer option, and greedy argmax-vs-actual over a given continuation. A chat/completions endpoint only reports on tokens the model generates; it cannot score text the caller supplies. The legacy OpenAI /v1/completions with echo + logprobs can, but almost no self-hosted server implements it — none of the model servers registered here do — and reconstructing token spans from echoed output is exactly the fragility this protocol avoids.
Parsing a small model's chat answers instead would be a different, incomparable benchmark: at 60–128M parameters the only reliable signal is the loss surface, which is why nanochat scores loglikelihoods too. The scoring endpoint below is the same forward pass your chat server already runs, teacher-forced — about 40 lines. If your model loads with Hugging Face, the reference server is a complete implementation: python reference-server.py --model <your-checkpoint> and you are done.
The boundary
The platform owns the benchmark: it renders every prompt (few-shot draws, delimiters, task data — byte-identical for every model, reproduced draw-for-draw from nanochat's pipeline) and computes accuracy, centering, and CORE. Your runtime owns tokenization and the forward pass, because CORE's answer spans live in token space and only your tokenizer knows where tokens fall. Your endpoint never sees gold answers — it returns per-option mean losses (the platform takes the argmin) and greedy-match verdicts. Works unchanged for byte-level vocabularies: a token is a byte, losses are per-byte.
The protocol — blah-core-http/1
GET /healthz
-> {"status":"ok","protocol":"blah-core-http/1","model_revision":"<id>"}
POST /v1/score {"examples":[...]} 1..64 examples per call
-> {"results":[...]} same order, same length
examples (one of):
{"kind":"choice_set", "prompts":[p1, p2, ...]} same context, different endings
{"kind":"suffix_set", "prompts":[p1, p2, ...]} different contexts, same ending
{"kind":"greedy_pair", "prompts":[without, with]} without = strict prefix of with
results (matching kind and index):
{"kind":"choice_set", "mean_losses":[...]} nats, mean over continuation tokens
{"kind":"suffix_set", "mean_losses":[...]}
{"kind":"greedy_pair", "match": true|false} exact greedy match over the spanAnswer spans are found in token space: the common token prefix across a choice set, the common token suffix across a schema set, the without/with prefix split for greedy pairs — string-level splitting disagrees with the benchmark whenever a token straddles the boundary. Sequences longer than your context are cropped from the left. Full semantics, batching guidance, sanity checks, and debug signatures are in the agent prompt. If a bearer token is stored for your model it arrives as Authorization: Bearer; return 503 when saturated and the platform retries with backoff.
Hand this to your coding agent
The prompt carries everything above plus the exact scoring semantics ported from nanochat's core_eval.py, verification steps, and the registration flow.
curl -L https://evals.blah.dev/core/agent-prompt.txt -o BLAH_CORE_TASK.md
Attempts and tiers
Start an attempt from your model's /models/<id>/core page or the API. The endpoint is health-checked first — a wrong URL fails in a second with the reason, not after a queued run.
curl -s -X POST https://evals.blah.dev/api/v1/models/MODEL_ID/core/attempts \
-H "Authorization: Bearer blah_YOUR_KEY" -H "Content-Type: application/json" \
-d '{"score_uri":"https://my-model.example.com","tier":"smoke","model_label":"step 1200"}'Tiers fix the per-task example budget, subsampled with nanochat's fixed seed so a tier is the same examples for everyone: smoke 50 (~2,750 forwards — quick signal, fine on CPU), standard 500 (~27,500 — the default comparison scale, wants a GPU), full everything (~350,000). Scores only compare within a tier; the chart never mixes them.
Attempts are append-only documents. Each runs in the background with live progress on its permanent page (/core/attempts/<id>), records the bundle fingerprint, scorer version, config, and endpoint it ran against, and stores every example's outcome — per-option losses, greedy verdicts, correctness — not just the totals. Task results persist as each task finishes, so a failed attempt keeps them and a new request resumes from what is missing. Attempt pages show per-task deltas against your previous same-tier attempt; benchmark every checkpoint and the chart becomes your training story, with GPT-2 124M and GPT-2 1.6B as reference rules.
Fidelity — verified, not vibes
The platform's renderer reproduces nanochat's byte-for-byte (13 fixtures generated by nanochat's own Jinja templates, checked in CI), including a bit-for-bit port of CPython's random.Random for the data shuffles and few-shot draws. Scoring follows core_eval.py: argmin mean per-token loss for multiple choice and schema, exact greedy match for language modeling, centering against the bundle's random baselines. The whole pipeline was gated on a head-to-head: GPT-2 124M through nanochat's own code and through this platform on the identical subset — 22/22 tasks exact, 0/440 examples disagreeing, CORE identical to six decimals. A frozen GPT-2 124M endpoint stays registered as the standing anchor.
Comparisons to numbers published elsewhere remain approximate — harnesses differ (nanochat itself notes its squad diverges from the DCLM reference); comparisons within this site are exact. Don't train on the eval bundle. The platform computes scores from raw losses, but it cannot detect memorization — the numbers are only worth what everyone's restraint makes them.