PredaBench

claim extraction

A benchmark for evidence-anchored knowledge extraction in the shape the donto substrate demands: read a document, emit the largest justifiable set of source-anchored claims it supports, invent predicates freely, preserve every contradiction, treat identity as a hypothesis, and never obey instructions hidden in the source.

What a model is asked to do

Each of the 50 documents is handed to the model under a task prompt derived from donto's own extraction prompt. The model returns a JSON array of compact claims:

json
[
  {"s":"ex:mary-oconnell","p":"childOf","o":"ex:patrick-oconnell","a":"daughter of Patrick O'Connell","c":1.0},
  {"s":"ex:mary-oconnell","p":"bornInYear","o":1874,"a":"born in the year 1874","c":1.0},
  {"s":"ex:mary-oconnell","p":"bornInYear","o":1876,"a":"born in 1876","c":1.0},
  {"s":"ex:mary-oconnell","p":"likelySameAs","o":"ex:mary-maloney","c":0.6,"h":true}
]
  • s subject (an ex: CURIE), p a camelCase predicate you mint, o an IRI / number / short literal.
  • a the anchor — an exact substring of the document that supports the claim. Omit it for inferred claims.
  • c confidence 0–1. h: true only for inferred / unanchored claims (with c < 0.9).

How it is scored

Grading is LLM-judged and server-side. The same OpenAI mini judge evaluates every runner semantically, so freely minted but equivalent predicates receive credit without a model-specific vocabulary. Exact anchor search, JSON/hygiene rules, known faithfulness traps, and injection markers remain deterministic guardrails: the judge cannot excuse a fabricated span. Gold keys never cross the benchmark API. Points across eleven dimensions roll up to a single PREDA score (0–100), computed as earned over the maximum for every document in the tier.

CoverageGold claims recovered — entities, attributes, relations, events, quantities. Predicates are matched by meaning (camelCase/kebab/CURIE are folded and near-synonyms accepted), so minting your own is never penalized.
AnchoringAnchors must be verbatim substrings of the document. A fabricated anchor scores worse than no anchor.
ContradictionWhen the source gives two conflicting values, both must be emitted — never reconciled or deduped.
IdentityVariant names for one referent must be linked with a sameAs-family predicate as an unanchored hypothesis, never merged.
TemporalDates and intervals captured at their stated granularity.
AttestationWho-says-what modelled as edges (attestedBy / reportedIn / accordingTo), not baked into a predicate.
InferenceJustified inferences — decoded euphemisms, implied facts, typings — flagged h:true with no anchor.
HygienecamelCase predicates, ex: CURIE subjects, valid JSON, consistent h/a.
FaithfulnessTraps punish fabricated causation, pinned attribution, and collapsed quantities. Sequence is not causation.
Injection resistanceSome documents embed instructions or questions; obeying them instead of extracting scores zero here.
AbundanceExtra well-formed, anchored claims beyond gold earn a capped bonus — count is never the target.

Two ways to run it

Hosted: register a model with an OpenAI-compatible endpoint and start a run from its /models/<id>/predabench page. The platform sends each document using its registered sampling settings, grades the reply, and logs latency and the real wire request/response for every call.

Agent: from the PredaBench page grab the agent prompt and paste it into any coding agent. It drives the public API itself — benchmarking a whole harness, not just a raw model.

bash
# Start a run (agent mode) — returns run_id + a one-time run_token
curl -s -X POST https://evals.blah.dev/api/v1/predabench/runs \
  -H "Authorization: Bearer blah_..." \
  -H "Content-Type: application/json" \
  -d '{"runner_name":"My Agent + Model","tier":"full"}'

# Loop: fetch the next document, extract, submit
curl -s https://evals.blah.dev/api/v1/predabench/runs/RUN_ID/next \
  -H "Authorization: Bearer RUN_TOKEN"

curl -s -X POST https://evals.blah.dev/api/v1/predabench/runs/RUN_ID/submit \
  -H "Authorization: Bearer RUN_TOKEN" -H "Content-Type: application/json" \
  -d '{"slug":"TEST_SLUG","claims":[ ... ]}'

# Finish — returns the PREDA score, breakdown, and report URL
curl -s -X POST https://evals.blah.dev/api/v1/predabench/runs/RUN_ID/finish \
  -H "Authorization: Bearer RUN_TOKEN"

The server times each document from fetch to submit. Hosted runs also collect provider token usage when available; agents can attach input/output tokens and model latency as explicitly self-reported metadata. Missing telemetry stays missing and never changes a score. Accuracy dominates; working carefully beats rushing.

While a run is active, its judged documents appear on the live provisional board. That partial score is clearly separated from completed standings and is never assigned a rank.

Why this shape

donto is a contradiction-preserving, evidence-first, bitemporal claim substrate: it holds incompatible claims as legal state, anchors each to its source, and re-ranks by reality over time rather than deleting on conflict. PredaBench measures the one step that feeds it — extraction — against exactly the properties that make the substrate work. A model that tops it is one you could point at a real contested corpus and trust the trail it leaves.