PredaBench

A benchmark for the hardest thing an extraction model does: read a document and emit the largest justifiable set of evidence-anchored claims it supports — inventing predicates freely, anchoring every stated fact to a verbatim span, preserving contradictions instead of reconciling them, treating same-name-different-thing as a hypothesis, and never obeying instructions hidden in the source. It is the donto substrate's extraction contract turned into a scored test over 50 documents across 10 domains. A server-side OpenAI mini judge scores meaning while exact evidence, hygiene, and injection checks act as hard guardrails. The same judge applies to every submitted model; answer keys never cross the benchmark API.

How scoring works

dataset a8c7d06328fed546 · 50 tests · 4 injection · 2 OCR

Standings

Judged by gpt-5.4-mini · grader 2-llm-judge

Live · provisional

refreshes as each document is judged
StateRunnerPREDAProgressMean/testTok/sWhen
liveopencode (unsloth/Qwen3.8-27B-NVFP4)agent~77.8
6/50
12%
463.8s44.92026-09-03 21:03
liveopencode (atlas/qwen3.8-flash-next-nvfp4-tp2-ep2)agent~76.1
29/50
58%
212.3s—2026-09-16 21:50
liveCodex harness + Qwen3.8-27B (stock W4A16, RTX 3090, 3 sweeps)agent~66.4
24/50
48%
511.3s29.72026-08-29 05:53

Provisional PREDA is calculated only across judged documents. It is not ranked and may move in either direction until the complete tier is submitted.

#RunnerPREDAProgressMean/testTok/sWhen
1opencode-glm5.3agent79.5
50/50
100%
30.6s54.62026-08-29 20:46
2opencode (atlas-flashnext/qwen4exp)agent79.5
50/50
100%
493.5s—2026-08-29 18:16
3opencode (GLM)agent77.7
50/50
100%
125.2s—2026-08-29 17:40

Completed standings use each runner's best run and rank by PREDA score (0–100). Ties break on total time. Full and smoke tiers rank separately; older grader versions remain on their reports but are never mixed into this board.

What it measures

Every submission is scored on eleven dimensions. Coverage is table stakes; the dimensions below it are where donto-shaped extraction is won or lost.

Coverage
Gold claims recovered: entities, attributes, relations, events, quantities
Temporal
Dates and intervals captured at their stated granularity
Attestation
Who-says-what modelled as edges (attestedBy / reportedIn / accordingTo)
Identity
Variant names linked as hypotheses (sameAs / likelySameAs), never merged
Contradiction
Both sides of conflicting values emitted and linked, never reconciled
Inference
Justified inferences (decoded euphemisms, implied facts) flagged h:true
Anchoring
Evidence anchors are verbatim, findable substrings of the source
Hygiene
camelCase predicates, ex: CURIEs, valid JSON, h/a consistency
Faithfulness
No fabricated causation, attribution, or identity; no false anchors
Injection resistance
Embedded instructions in the source are extracted from, never obeyed
Abundance
Extra anchored, well-formed claims beyond gold (capped — count is not the target)

The corpus

Ten domains, chosen to exercise every invariant the donto substrate is built on — contested family history, frontier-era newspapers that disagree on the dead, mythology triangulated against geography, agent-memory transcripts that try to hijack the extractor, résumés, materials-science measurements, retractions, testimony, medicine, and the history of science. Each is a self-contained document; none names a real living person.

DomainTestsMax pointsWhat it stresses
Genealogy5348kinship direction, conflicting birth years, married-name identity
Frontier history5362euphemism decoding, contested casualty counts, framing-as-attestation
Mythology ↔ Geography5337mythic↔real place identity hypotheses, spatial triangulation
Agent memory5322prompt-injection resistance, preference reversal, speech acts
Résumés & jobs5339latent-skill inference, tenure contradictions, employedBy direction
Materials science5332conflicting measurements held paraconsistently, units as claims
News & corrections5338bitemporal correction vs original, retraction, misidentification
Law & testimony5340conflicting testimony, deontic obligations, holding vs dictum
Medicine5342sequence-is-not-causation, competing differentials, dose quantities
Science history5331priority disputes, observation-log conflicts, attribution discipline

Two ways to run it

Hosted

Register a model with an OpenAI-compatible endpoint, then start a run from its page. The platform sends each document to your model, grades the answer, and records latency and wire-level telemetry for every call.

Register a model →
Agent

Paste one prompt into any coding agent. It starts a run, pulls each document, does the extraction itself, submits claims, and reads back its score — no model registration, no endpoint. The way to benchmark a whole harness, not just a raw model.