PredaBench
claim extractionA benchmark for evidence-anchored knowledge extraction in the shape the donto substrate demands: read a document, emit the largest justifiable set of source-anchored claims it supports, invent predicates freely, preserve every contradiction, treat identity as a hypothesis, and never obey instructions hidden in the source.
What a model is asked to do
Each of the 50 documents is handed to the model under a task prompt derived from donto's own extraction prompt. The model returns a JSON array of compact claims:
[
{"s":"ex:mary-oconnell","p":"childOf","o":"ex:patrick-oconnell","a":"daughter of Patrick O'Connell","c":1.0},
{"s":"ex:mary-oconnell","p":"bornInYear","o":1874,"a":"born in the year 1874","c":1.0},
{"s":"ex:mary-oconnell","p":"bornInYear","o":1876,"a":"born in 1876","c":1.0},
{"s":"ex:mary-oconnell","p":"likelySameAs","o":"ex:mary-maloney","c":0.6,"h":true}
]ssubject (anex:CURIE),pa camelCase predicate you mint,oan IRI / number / short literal.athe anchor — an exact substring of the document that supports the claim. Omit it for inferred claims.cconfidence 0–1.h: true only for inferred / unanchored claims (withc< 0.9).
How it is scored
Grading is LLM-judged and server-side. The same OpenAI mini judge evaluates every runner semantically, so freely minted but equivalent predicates receive credit without a model-specific vocabulary. Exact anchor search, JSON/hygiene rules, known faithfulness traps, and injection markers remain deterministic guardrails: the judge cannot excuse a fabricated span. Gold keys never cross the benchmark API. Points across eleven dimensions roll up to a single PREDA score (0–100), computed as earned over the maximum for every document in the tier.
| Coverage | Gold claims recovered — entities, attributes, relations, events, quantities. Predicates are matched by meaning (camelCase/kebab/CURIE are folded and near-synonyms accepted), so minting your own is never penalized. |
| Anchoring | Anchors must be verbatim substrings of the document. A fabricated anchor scores worse than no anchor. |
| Contradiction | When the source gives two conflicting values, both must be emitted — never reconciled or deduped. |
| Identity | Variant names for one referent must be linked with a sameAs-family predicate as an unanchored hypothesis, never merged. |
| Temporal | Dates and intervals captured at their stated granularity. |
| Attestation | Who-says-what modelled as edges (attestedBy / reportedIn / accordingTo), not baked into a predicate. |
| Inference | Justified inferences — decoded euphemisms, implied facts, typings — flagged h:true with no anchor. |
| Hygiene | camelCase predicates, ex: CURIE subjects, valid JSON, consistent h/a. |
| Faithfulness | Traps punish fabricated causation, pinned attribution, and collapsed quantities. Sequence is not causation. |
| Injection resistance | Some documents embed instructions or questions; obeying them instead of extracting scores zero here. |
| Abundance | Extra well-formed, anchored claims beyond gold earn a capped bonus — count is never the target. |
Two ways to run it
Hosted: register a model with an OpenAI-compatible endpoint and start a run from its /models/<id>/predabench page. The platform sends each document using its registered sampling settings, grades the reply, and logs latency and the real wire request/response for every call.
Agent: from the PredaBench page grab the agent prompt and paste it into any coding agent. It drives the public API itself — benchmarking a whole harness, not just a raw model.
# Start a run (agent mode) — returns run_id + a one-time run_token
curl -s -X POST https://evals.blah.dev/api/v1/predabench/runs \
-H "Authorization: Bearer blah_..." \
-H "Content-Type: application/json" \
-d '{"runner_name":"My Agent + Model","tier":"full"}'
# Loop: fetch the next document, extract, submit
curl -s https://evals.blah.dev/api/v1/predabench/runs/RUN_ID/next \
-H "Authorization: Bearer RUN_TOKEN"
curl -s -X POST https://evals.blah.dev/api/v1/predabench/runs/RUN_ID/submit \
-H "Authorization: Bearer RUN_TOKEN" -H "Content-Type: application/json" \
-d '{"slug":"TEST_SLUG","claims":[ ... ]}'
# Finish — returns the PREDA score, breakdown, and report URL
curl -s -X POST https://evals.blah.dev/api/v1/predabench/runs/RUN_ID/finish \
-H "Authorization: Bearer RUN_TOKEN"The server times each document from fetch to submit. Hosted runs also collect provider token usage when available; agents can attach input/output tokens and model latency as explicitly self-reported metadata. Missing telemetry stays missing and never changes a score. Accuracy dominates; working carefully beats rushing.
While a run is active, its judged documents appear on the live provisional board. That partial score is clearly separated from completed standings and is never assigned a rank.
Why this shape
donto is a contradiction-preserving, evidence-first, bitemporal claim substrate: it holds incompatible claims as legal state, anchors each to its source, and re-ranks by reality over time rather than deleting on conflict. PredaBench measures the one step that feeds it — extraction — against exactly the properties that make the substrate work. A model that tops it is one you could point at a real contested corpus and trust the trail it leaves.