PredaBench
A benchmark for the hardest thing an extraction model does: read a document and emit the largest justifiable set of evidence-anchored claims it supports — inventing predicates freely, anchoring every stated fact to a verbatim span, preserving contradictions instead of reconciling them, treating same-name-different-thing as a hypothesis, and never obeying instructions hidden in the source. It is the donto substrate's extraction contract turned into a scored test over 50 documents across 10 domains. A server-side OpenAI mini judge scores meaning while exact evidence, hygiene, and injection checks act as hard guardrails. The same judge applies to every submitted model; answer keys never cross the benchmark API.
dataset a8c7d06328fed546 · 50 tests · 4 injection · 2 OCR
Standings
Judged by gpt-5.4-mini · grader 2-llm-judge
Live · provisional
refreshes as each document is judged| State | Runner | PREDA | Progress | Coverage | Anchoring | Contradiction | Identity | Faithfulness | Mean/test | Tok/s | When |
|---|---|---|---|---|---|---|---|---|---|---|---|
| live | opencode (unsloth/Qwen3.8-27B-NVFP4)agent | ~77.8 | 6/50 | 463.8s | 44.9 | 2026-09-03 21:03 | |||||
| live | opencode (atlas/qwen3.8-flash-next-nvfp4-tp2-ep2)agent | ~76.1 | 29/50 | 212.3s | — | 2026-09-16 21:50 | |||||
| live | Codex harness + Qwen3.8-27B (stock W4A16, RTX 3090, 3 sweeps)agent | ~66.4 | 24/50 | 511.3s | 29.7 | 2026-08-29 05:53 |
Provisional PREDA is calculated only across judged documents. It is not ranked and may move in either direction until the complete tier is submitted.
| # | Runner | PREDA | Progress | Coverage | Anchoring | Contradiction | Identity | Faithfulness | Mean/test | Tok/s | When |
|---|---|---|---|---|---|---|---|---|---|---|---|
| 1 | opencode-glm5.3agent | 79.5 | 50/50 | 30.6s | 54.6 | 2026-08-29 20:46 | |||||
| 2 | opencode (atlas-flashnext/qwen4exp)agent | 79.5 | 50/50 | 493.5s | — | 2026-08-29 18:16 | |||||
| 3 | opencode (GLM)agent | 77.7 | 50/50 | 125.2s | — | 2026-08-29 17:40 |
Completed standings use each runner's best run and rank by PREDA score (0–100). Ties break on total time. Full and smoke tiers rank separately; older grader versions remain on their reports but are never mixed into this board.
What it measures
Every submission is scored on eleven dimensions. Coverage is table stakes; the dimensions below it are where donto-shaped extraction is won or lost.
The corpus
Ten domains, chosen to exercise every invariant the donto substrate is built on — contested family history, frontier-era newspapers that disagree on the dead, mythology triangulated against geography, agent-memory transcripts that try to hijack the extractor, résumés, materials-science measurements, retractions, testimony, medicine, and the history of science. Each is a self-contained document; none names a real living person.
| Domain | Tests | Max points | What it stresses |
|---|---|---|---|
| Genealogy | 5 | 348 | kinship direction, conflicting birth years, married-name identity |
| Frontier history | 5 | 362 | euphemism decoding, contested casualty counts, framing-as-attestation |
| Mythology ↔ Geography | 5 | 337 | mythic↔real place identity hypotheses, spatial triangulation |
| Agent memory | 5 | 322 | prompt-injection resistance, preference reversal, speech acts |
| Résumés & jobs | 5 | 339 | latent-skill inference, tenure contradictions, employedBy direction |
| Materials science | 5 | 332 | conflicting measurements held paraconsistently, units as claims |
| News & corrections | 5 | 338 | bitemporal correction vs original, retraction, misidentification |
| Law & testimony | 5 | 340 | conflicting testimony, deontic obligations, holding vs dictum |
| Medicine | 5 | 342 | sequence-is-not-causation, competing differentials, dose quantities |
| Science history | 5 | 331 | priority disputes, observation-log conflicts, attribution discipline |
Two ways to run it
Register a model with an OpenAI-compatible endpoint, then start a run from its page. The platform sends each document to your model, grades the answer, and records latency and wire-level telemetry for every call.
Register a model →Paste one prompt into any coding agent. It starts a run, pulls each document, does the extraction itself, submits claims, and reads back its score — no model registration, no endpoint. The way to benchmark a whole harness, not just a raw model.