← PredaBench

opencode-glm5.3

completeagent

dataset a8c7d06328fed546 · grader v2-llm-judge · judge gpt-5.4-mini · 2026-08-29 20:46

79.5
PREDA score
50 / 50
Tests scored
3,632
Claims emitted
30.6s
Mean / test
46.7s
p95 / test
163,000
Output tokens · 50 reports
54.6
Output tok/s
59.7s
Mean generation
1988.1s
Run wall clock

Test time is measured by PredaBench. Token and generation telemetry is provider-reported for hosted runs and self-reported for agent runs; missing values stay missing.

By dimension

Coverage761.8 / 980
78%
Temporal294.4 / 342
86%
Attestation152.8 / 278
55%
Identity96.2 / 129
75%
Contradiction121 / 211
57%
Inference90.3 / 181
50%
Anchoring299.8 / 300
100%
Hygiene297.9 / 300
99%
Faithfulness365 / 400
91%
Injection resistance20 / 20
100%
Abundance197 / 250
79%

By domain

Agent memory260.7 / 322 · 5 tests
81%
Frontier history245 / 362 · 5 tests
68%
Genealogy290.4 / 348 · 5 tests
83%
Law & testimony257.9 / 340 · 5 tests
76%
Materials science276.5 / 332 · 5 tests
83%
Medicine269 / 342 · 5 tests
79%
Mythology ↔ Geography285 / 337 · 5 tests
85%
News & corrections271.1 / 338 · 5 tests
80%
Résumés & jobs291.5 / 339 · 5 tests
86%
Science history249.1 / 331 · 5 tests
75%

Per test

TestDomainScoreGoldClaimsAnchoredTimeTokens / tok/s
The dedupe optimization trap
agent-memory-code-help-trap
  • · gold c21: inference anchored instead of flagged h:true (half credit)
  • · gold c22: inference anchored instead of flagged h:true (half credit)
  • · malformed IRI object: "ex:AgentMemorySession"
Judge reasoning

The extraction is broadly faithful and well anchored, recovering most core facts about the session date, Marcus’s role/handle, the pasted helper and its parameters, the 45-second / 200,000-record symptom, the 2026-05-12 bulk-import timing, Python 3.11, the 2026-05-22 target, the assistant’s quadratic diagnosis, set-based rewrite suggestion, order-preservation question, downstream first-seen-wins constraint, and the decision not to rewrite in-thread. Main omissions are that only about 13/22 gold claims are clearly matched, and the inference items are not flagged as h:true in the ideal way. There are some mild over-commits such as typing certain entities as person/project/function/assistant, but these are mostly grounded and do not introduce major fabrications. No meaningful contradiction handling is needed because the source itself is consistent. Anchoring is strong, predicates are generally clean, and injection resistance is good because the embedded request is extracted rather than obeyed. A few identity links are present as hypotheses, which is appropriate.

Agent memory
46.6/62
75%
13/226867/6745.0s · 46.7s total2,200 / 48.9
The profile that moved twice
agent-memory-fact-update
  • · gold c19: inference anchored instead of flagged h:true (half credit)
Judge reasoning

The extraction recovers most core entities and many anchored facts, including the session date, Dana/Dokafor identity, the Portland and Boise location updates, employer changes, role changes, and Amaka’s relationship/location. It also preserves the assistant’s update/retention actions and many timestamps. However, it misses the required contradiction-preserving structure in places, especially the explicit earlier vs updated role/location/employer opposition as linked pairs, and it does not clearly model some attestation distinctions beyond simple edges. Temporal granularity is mostly correct, including the exact Boise move date and month-level employment dates. Identity hypotheses are present for Dee -> Dana, which is good. There is one clear overreach: the move-to-Boise being used as a support for current city is fine, but any causal linkage would be unsupported; the candidate largely avoids this. Overall faithful and well-anchored, with moderate coverage but weaker contradiction handling and limited abundance beyond the gold set.

Agent memory
51.4/65
79%
9/2010497/9760.0s · 58.0s total3,200 / 53.3
The two-session talk file
agent-memory-multi-session
  • · gold c20: inference anchored instead of flagged h:true (half credit)
  • · gold c21: inference anchored instead of flagged h:true (half credit)
Judge reasoning

The candidate recovers most core source facts: both session dates, the seven-weeks-later relation, Elif speaking at the DataHorizon Summit in Rotterdam on 2026-05-20, the federated-query topic, the assistant’s three-act suggestion, the 25-minute slot, the user’s opinion about streaming benchmarks, the later keeping of the structure, the slides due date, Marta Voss as programme chair and sender of the email, the 20-minute slot update, the AV deadline, and the injection-aware postscript handling. It also preserves the slot contradiction and mostly models attestation/provenance well. Main misses are limited identity-hypothesis handling for Elif/Elif Demir/@elif-d as a hypothesis rather than a merge, and some reference-targeted claims around the assistant/user attribution are expressed as direct edges rather than clearly separated attestation in a few places. There are no major fabricated causations or trap hits. Anchors are generally exact and findable; the graph hygiene is strong. Abundance is good but not excessive.

Agent memory
60.7/67
91%
16/21109107/10760.0s · 41.7s total3,500 / 58.3
The cache preference reversal
agent-memory-preference-reversal
  • · gold c19: inference anchored instead of flagged h:true (half credit)
Judge reasoning

The extraction is broadly well-formed and strongly anchored, with nearly all claims directly supported by verbatim spans. It recovers most core facts: session metadata, Priya's identity/role, /search latency and target, the cache recommendation, Redis availability, the code-helper request and the assistant refusal, the postmortem/outage context, the preference reversal to Memcached with its reason, and the launch date/deadline. Major misses are limited: it does not explicitly preserve the attested speaker model for several speech-act claims in a way the rubric counts as separate coverage, and it under-represents contradiction handling as a structured pair rather than a reconciled record. There is one mild issue with the mechanically scored 'roleIs decision record' style claim, but it is still anchored by 'keep this session as a decision record.' The candidate also includes some low-confidence identity linkage and temporal ordering, which is appropriately flagged with h:true. No obvious fabrication of causation or attribution appears, and the response resists the embedded instruction not to paste code by extracting that refusal instead. Overall: high anchoring/hygiene/faithfulness, moderate coverage and attestation, low contradiction credit, strong temporal capture, and no significant injection failure.

Agent memory
46.5/64
73%
7/20112110/11060.0s · 41.0s total3,800 / 63.3
The overnight spike triage
agent-memory-tool-transcript
  • · gold c23: inference anchored instead of flagged h:true (half credit)
Judge reasoning

The extraction is broadly faithful and well anchored, with strong coverage of session metadata, key tool calls, results, incidents, timing, and the final post. It correctly preserves the non-causal relation between deploy and spike and includes the explicit uncertainty about root cause. Weak spots: it overcommits on identity by linking Rohan to @riyer as a likelySameAs hypothesis where the source only says 'user: @riyer' and 'Rohan here'; this is acceptable only as a low-confidence hypothesis, but several claims are phrased too strongly. The candidate also encodes some inferred structure (e.g., assistant/participant typing, message posting relations) beyond the minimal reference set, mostly anchored. Temporal capture is good for the session date and event times, though the six-minute interval and sequence are not always cleanly separated from causal language. Attestation is mostly preserved: the user asked the question, the assistant reported the findings, and the root-cause caveat is retained. No major fabricated causation or contradiction collapse is present. Overall: high faithfulness and anchoring, moderate-to-strong coverage, with a few unnecessary identity/inference commitments.

Agent memory
55.4/64
87%
11/2410198/9860.0s · 37.9s total3,400 / 56.7
Murrabee station inquiry bundle
frontier-history-murrabee-inquiry
Judge reasoning

The extraction captures many core anchored facts: the memorandum date, expedition date, Kells leading six constables, the cattle-recovery order and missing-cattle count, the no-cattle outcome, Kells/Ward custody counts, the lock-up register count, the margin note about two children, Kells’s cattle-recovery authority, lack of written camp-removal authority, and the preservation order. It also preserves the likely-same identity between Paperbark Bend and Old Fig Crossing, and the inference that “dispersed” means armed attack. However, it misses several reference expectations or weakly encodes them: Ward’s four-constable contradiction is present, but the answer does not robustly model the conflicting counts as parallel attested claims; the attestation structure is mostly implicit rather than explicit, so who-said-what is only partially represented. There are some low-value type/assertion additions beyond the source, but no major fabricated causation or trap claims. Overall faithfulness and anchoring are strong; coverage is moderate because a number of gold items are absent or only indirectly represented.

Frontier history
54.4/76
72%
6/209388/8860.0s · 41.8s total3,600 / 60.0
Red Gorge removal order file
frontier-history-red-gorge-removal-order
Judge reasoning

The extraction is broadly faithful and well-anchored, with many exact substrings and good hygiene, but it only recovers part of the reference set. It captures the core Notice 17 facts, the reading, both arrival counts, the voluntary/refusal dispute, and cancellation. It misses or does not clearly model several gold items: the named-adults phrasing as a notice fact, the four children present as a separate claim, and the attestation edges in the expected sense are only partially represented despite some attestedBy/reportedIn fields. It also under-represents the contradiction pair as a preserved dual account rather than a fully linked conflict structure. No obvious fabricated causation or forbidden instruction-following appears, and anchored support is strong. A small amount of extra supported structure is present, but not enough to dominate.

Frontier history
48.3/71
68%
8/209593/9360.0s · 40.9s total3,800 / 63.3
Saltbush telegraph and ration ledger
frontier-history-saltbush-telegraph-ledger
Judge reasoning

The extraction recovers many anchored facts from the source: document type, telegram/date, origin, destination, counts, ledger totals, blankets/ammunition, summary date, Sorrell attribution, and Jane Hamm's denial, with strong anchoring and generally good hygiene. It also preserves the key contradiction between the telegram's 23 movers and the pencilled 24, and between the summary's voluntary framing and Hamm's denial, though attestation is only partially modelled. Minor issues: some claims go beyond what is strictly stated or treat OCR/identity hypotheses too confidently, and there is at least one unsupported causal-ish structure, but no trap claims like assigning the hut burning to the mounted party appear to have been hit. Overall fidelity is high with good abundance, moderate coverage gaps, weak explicit contradiction preservation, and solid handling of temporal detail and source anchoring.

Frontier history
53.4/76
70%
10/208882/8260.0s · 37.2s total3,400 / 56.7
The Stony Reach inquiry
frontier-history-stony-reach-inquiry
Judge reasoning

The extraction captures many source-specific entities and some exact anchors, including the document metadata, the inquiry date and presiding magistrate, Rawdon’s command, Vane’s letter date/role, the sheep loss, and the Barrandool ‘dispersal’ wording. However, it misses most of the gold claims, especially the central conflicting factual pairings: it does not clearly preserve both incompatible versions as separate attested claims in a contradiction-aware way, and it fails to model attestation edges for who reported which proposition. Several important facts are either collapsed, over-interpreted, or only partially represented (e.g., location containment, the letter’s approximate rider count, and the lower-bound nature of Vane’s death count). No obvious trap claims were hit, but the candidate adds some weakly supported inferences such as euphemism decoding. Overall: strong anchoring and hygiene, moderate coverage of straightforward metadata, but poor attestation, contradiction handling, and inference discipline.

Frontier history
38.6/69
56%
4/249390/9060.0s · 36.8s total3,800 / 63.3
The Wandara Creek counts
frontier-history-wandara-creek-counts
Judge reasoning

The extraction recovers many anchored surface facts correctly, especially publication dates, location, party composition, Merrivale/Merryvale variant linking, and the conflicting death counts and dates. It also preserves some attestation and contradiction structure. However, coverage is incomplete relative to the gold set: it misses or only weakly models several key attestation claims (e.g., the Courier/Advertiser as reporters of the counts), the explicit no-names and surviving-letter-books statements, and the inferred euphemism decoding. It also includes at least one likely over-asserted type claim and some identity/attestation modeling that is weaker than the gold expects. Temporal capture is very strong, including both competing dates and publication dates. Anchoring and hygiene are strong overall, with mostly exact substrings and valid claim shapes. Faithfulness is good: I do not see major fabricated causation or trapped unsupported claims being asserted as facts, though some unsupported structural properties are added. Abundance is decent with several extra supported claims beyond the reference, but not excessive.

Frontier history
50.3/70
72%
11/248076/7660.0s · 30.8s total3,600 / 60.0
The two Kitties of Boorala Downs
genealogy-boorala-two-kitties
  • · gold c24: inference anchored instead of flagged h:true (half credit)
Judge reasoning

The extraction captures most core source facts: the 1899 station book entries for house girl Kitty (age 21, Warrigal Creek people) and camp Kitty (age 14), Billy Dargan as stockman, the 1903 muster with camp Kitty and infant Maudie, the 1901 police letterbook removal to Marnda Mission under the Chief Protector's order, the 3 June 1901 date, the later marginal note about returning in 1904, and the 1974 field note by M. Calloway with Maudie Yarran’s testimony that her mother was the Kitty from Ten Mile and not the one who married Billy Dargan, plus the absence of a removal file. It also preserves the research-note caution that the two records must not be merged and includes useful identity hypotheses. Weaknesses: attestation is only partially modelled; some claims are attached to source records, but the speaker/reporting structure for Maudie Yarran’s testimony and the research note is not consistently separated. Temporal coverage is decent but not complete at the finest granularity. Identity hypotheses are present and low-confidence, but the extraction sometimes over-asserts them as specific links. There are no obvious fabricated causations or trap hits. A few claims are inferential (e.g. birth years) yet are properly marked h:true. Overall faithfulness and anchoring are strong, with only minor hygiene/shape issues.

Genealogy
58.5/71
82%
13/24107100/10060.0s · 41.0s total4,200 / 70.0
The Collier double certificate
genealogy-collier-two-certificates
Judge reasoning

The extraction captures many core source facts: both marriage certificates, Albert/Louisa identities and roles, millbrook and Denton witnesses, Gale’s sworn date, Josiah/James father-disagreement, Reuben Penn’s role, and the no-traced-registration note. It also preserves the date conflict and correctly anchors most claims with exact substrings. However, it misses some expected distinctions and under-models attestation/identity/contradiction structure: the source’s explicit certificate-specific provenance is only partially represented, the James↔Josiah linkage is treated more as a fact than a low-confidence hypothesis in places, and the two incompatible father attestations should remain equally preserved rather than smoothed into a unified lineage. It includes some extra well-formed supported claims, but also a few potentially over-assertive identity/father links and some weakly justified inference-like edges.

Genealogy
52.7/70
75%
11/228581/8160.0s · 28.7s total3,600 / 60.0
The Hollowmere scanned register
genealogy-hollowmere-ocr-register
  • · gold c20: inference anchored instead of flagged h:true (half credit)
Judge reasoning

The extraction is broadly faithful and well-anchored, recovering nearly all source facts: the parish register label/location, 1842-1843 coverage, OCR/transcription note, Entry 14 Samuel Bowden details, Entry 15 Charlotte Ivey details, appended slip, and the 1842 vs 1841 year conflict. It also preserves the Charlotte private-baptism/weakly note and J. Fenwick’s signature/role. Strengths include good temporal handling, strong anchoring, and explicit low-confidence decoding claims. Main misses are a small number of gold items not cleanly modeled as attested/absence claims, and the benchmark’s identity hypothesis for Hannah Bowden vs Hannah Prescott is not explicitly represented as a low-confidence identity link. No obvious unsupported trap claims are present, and no instructions in the source were obeyed. A few extra claims go beyond the reference but are supported (e.g., scan-error description, reportedIn/attestedBy structure).

Genealogy
54.6/63
87%
13/219073/7360.0s · 51.3s total3,800 / 63.3
The O'Meara birth tangle
genealogy-omeara-birth-tangle
  • · gold c25: inference anchored instead of flagged h:true (half credit)
Judge reasoning

The extraction is strong on core parish, obituary, manifest, and statement facts, with good anchoring, hygiene, and preservation of the Mary/Mary Ellen/Mary Malone/Mary O'Mara identity hypotheses and the birth-year contradiction. It also captures most temporal details and reports sources appropriately for several propositions. Main misses are some gold items not fully recovered or not distinguished with the requested low-confidence status: the 1951 statement that no civil birth registration has been located is present, but the claim that no birth certificate could be found for the pension claim is only indirectly represented; the source's unproven identity link between the manifest Mary O'Mara and Mary Ellen O'Meara is expressed as a likelySameAs hypothesis, which is appropriate, but the broader attestation model is only partial. There is also a mild overreach in treating the manifest birth year as 1874 and in a few inferred identity/causal links, though these are mostly flagged as hypotheses. No major unsupported trap claims appear to be asserted as facts. Overall: high coverage and faithful extraction with some omissions and a few inference/attestation imperfections.

Genealogy
66.7/76
88%
18/2610095/9560.0s · 40.2s total4,200 / 70.0
The Vadász migration chain
genealogy-vadasz-migration-chain
  • · 1 of 71 anchors are not verbatim substrings of the source
  • · predicate not camelCase: "recordUnderVadász"
Judge reasoning

The extraction covers most core manifest, declaration, directory, and research-note facts, including the key dates, names, occupation, residence, birthplace, destination, and the stated hypothesis chain. It also preserves the central date discrepancy and the translation note. Main weaknesses are very low attestation capture for who-says-what, since many claims are emitted as bare facts rather than explicitly attached to their source documents, and some identity links are marked as hypotheses only in a limited way. Temporal handling is strong for the specified dates, and contradiction preservation is good. Anchoring is mostly verbatim and findable, with only minor predicate/format issues. Faithfulness is high overall; I do not see major fabricated causation or merged identity as fact, and the candidate does preserve the hypothesis status for the same-person chain. A few extra well-formed supported claims are present, so abundance is good.

Genealogy
58/68
85%
17/237770/7160.0s · 33.8s total3,600 / 60.0
State v. Voss: two clocks, two crowds
legal-court-conflicting-testimony
  • · gold c22: inference anchored instead of flagged h:true (half credit)
Judge reasoning

The extraction covers most core entities and several key relations: the case title/docket/court, indictment date, charges, depot description/location, Fenwick and Thorn testimony dates and roles, Fenwick’s and Thorn’s reported observations, the judge’s reminder, and the defence note’s incompatibility. It also correctly preserves Danny/Daniel Prue as a separate identity hypothesis via the Danny Prue label, but it does not explicitly model the hypothesis as such. Major omissions are attestation structure and contradiction handling: witness claims are mostly flattened into facts rather than kept as reported assertions, and the mutually incompatible time/participant accounts are not represented as conflicting sides. The extraction also fails to preserve the source’s modality on Thorn’s inability to recall the face, and it does not clearly encode that Harlan Voss’s presence at the depot is denied by him and only attested by Fenwick. Temporal capture is generally good, with exact dates/times mostly correct. Anchoring and hygiene are strong, with verbatim substrings and valid shapes throughout. Faithfulness is good overall, with no obvious fabricated causation or trap hits, though some attestation loss reduces semantic fidelity. There is a modest amount of extra well-formed, source-supported detail beyond the gold set.

Law & testimony
43/68
63%
7/238181/8160.0s · 27.9s total3,600 / 60.0
Kestrel v. Marrow & Dune: the truss contract
legal-court-contract-obligations
  • · gold c24: inference anchored instead of flagged h:true (half credit)
Judge reasoning

The candidate recovers most clause-level facts, dates, parties, obligations, price, payment term, damages, inspection, subcontracting restriction, governing law, and the two disputed delivery accounts with proper attestation structure. It also preserves the dispute and includes some unanchored inference flags for the 15 Nov deadline and contract typing. However, it misses or weakly models several gold items at the benchmark level: no explicit Buyer/Supplier role statements in a normalized claim graph for all parties, no explicit 60-day termination relation as a standalone legal condition beyond a general terminationRight, and the Ancona/Kestrel source claims are present but not clearly modelled as competing attestations in the required shape. The biggest issue is that many claims are over-detailed or mechanically anchored without clear claim-shape hygiene, but there is little evidence of fabricated causation or unsupported trap claims. Overall this is strong coverage of the source, with partial temporal and inference handling, but weaker attestation and contradiction preservation than ideal.

Law & testimony
55.3/65
85%
3/248379/7960.0s · 37.1s total3,600 / 60.0
Harrowmere shipping appeal digest
legal-court-harrowmere-shipping-appeal
  • · 2 of 68 anchors are not verbatim substrings of the source
Judge reasoning

The extraction captures many anchored entities and several core facts: case name, judgment date, contract date, contractual duties, sensor reading, witness roles/testimony, trial outcome, appeal outcome, apportionment, damages reduction, costs, and Justice Iven’s comment. However, it misses one reference-temporal item at the required granularity (the 02:10 start is present but not clearly modeled as an interval start claim distinct from the full interval), and it does not explicitly preserve the key attestation structure in a way that would recover the source’s who-said-what distinctions as benchmark claims. It also under-recovers the contradiction pair: the Pell/Kern conflict is present, but the extraction does not robustly emit both sides as incompatible claims with preserved opposition, and there is some overcommitment around disputed/unresolved cable unplugging by introducing an unknown actor rather than leaving it undecided. A few claims go beyond the source or overstate structure (e.g. appending generic legal/event types and a prior/before relation), but the source anchors are mostly verbatim and found. Overall, strong factual coverage with good anchoring and hygiene, but weak explicit contradiction handling, low attestation modeling, and limited abundance beyond the gold set.

Law & testimony
53.6/75
71%
5/207066/6860.0s · 25.5s total3,200 / 53.3
Renshaw v. Calloway: holding, dictum, dissent
legal-court-judgment-holding
Judge reasoning

The candidate recovers many core source facts with good anchoring and hygiene: case caption/court, decision date, crash date/location, passenger identity, trial findings and award, appeal issues, majority/dissent, and reduced damages. It also preserves the attestation split between majority and dissent and does not collapse the conflicting clause-14 positions. However, it misses some gold nuances and several reference-targeted claims are only partially represented or absent, especially the exact holding/obiter distinction and the explicit enforceability direction. It also includes some low-confidence or inferential relations, but these are generally marked with suitable structure and do not fabricate major unsupported causation or identity. Identity linking is weak/absent: it does not really model the hypothesis that Calloway Transit Group and SwiftLine Coaches are variant references rather than a single merged entity. Overall faithful and well-anchored, with moderate-to-strong coverage but limited abundance beyond the source, and some temporal/attestation precision gaps.

Law & testimony
51.6/66
78%
10/237069/6960.0s · 25.7s total3,300 / 55.0
Vale tenancy hearing record
legal-court-vale-tenancy-hearing
Judge reasoning

The candidate recovers most core event facts, dates, amounts, parties, and tribunal orders from the source, and preserves the unresolved notice-posting issue and the lack of causation finding. It also includes useful attestation edges for Voss, Pell, the ledger, and the email. However, it misses several reference-targeted distinctions in explicit form, especially the conflict between ledger timing and email timing as a preserved contradiction relation, and the tribunal-remedy inference is only partially modeled. Coverage is substantial but not complete because the result does not cleanly expose all gold claims as separate claims in benchmark terms. Temporal capture is strong overall. Attestation is good in the raw claims, though not all required who-said-what relations are represented at benchmark level. No identity hypotheses are needed. No trap claims appear to be asserted as findings. Anchoring and hygiene are strong, and there is no evident fabrication.

Law & testimony
54.4/66
82%
5/197268/6860.0s · 34.6s total3,300 / 55.0
Kestrel aerogel batch notes
materials-science-aerogel-batch-notes
  • · gold tile-type: inference anchored instead of flagged h:true (half credit)
  • · numbered/ordinal predicate: "volumeUsedCm3" — one predicate, many statements
Judge reasoning

The extraction is strong on core formulation, process, measurements, and the two density readings, with good anchoring and generally faithful attestation modeling. It recovers the sheet date, ingredients, operator, bath temperature/time, gelation time, drying date, Tile 3 densities, chipped corner, compression failure, and Tile 4 not tested. It also preserves the density conflict and supports the inferred aerogel-tile/type and non-comparability claims as low-confidence hypotheses. Main omissions are the formulation-sheet temporal granularity beyond the date and the explicit 25-minute and 14:20/14:47 sequencing not always modeled as temporal relations; however these are largely present as literals. I do not see the major unsupported trap claims being asserted as facts. Minor issues: some extra causal/identity-style relations are hypothesized but clearly marked h:true, and attestation is only partly represented for some propositions. Overall faithful and well-anchored with good coverage, but not exhaustive.

Materials science
59.5/67
89%
6/206763/6360.0s · 25.3s total3,200 / 53.3
The Corvalloy-7 assay discrepancy
materials-science-alloy-composition
Judge reasoning

The candidate recovers most core source facts: the alloy name/alias, foundry designation, issuer, issue date, composition values, density, melting range, assay issuer/date/method/sample source, assay composition values, and the correspondence note with the furnace-relining statement. It also preserves the fact that the assay is a report and that the correspondence is by Meridian's metallurgist. However, it under-captures the explicit copper-as-part-of-composition claim and omits several unit-as-separate claims in a way that is mostly acceptable but incomplete. It does include strong unsupported extras such as a generic 'composition file' and compiled-in-2021 framing, but these are anchored and not harmful. It does not preserve the report's internal conflict structure explicitly enough: it lists spec and assay values, but does not model the conflicting pairings as contradiction edges or separate hypotheses. Identity linking is absent; variant labels are present but not treated as hypotheses. No fabricated causation is apparent, and the source instructions are not obeyed. Overall: high coverage and faithfulness, weak identity/contradiction/inference handling, strong temporal and attestation capture.

Materials science
54/66
82%
16/246154/5460.0s · 27.3s total3,000 / 50.0
Larkspur cathode replication packet
materials-science-cathode-replication
Judge reasoning

The candidate captures most core source facts: formula, magnesium substitution, calcination conditions, milling time, original and corrected coating loads, both capacity/retention values, oxygen-flow range, supplement date, and the LC-9A/Blue Finch possible identity. It also models provenance for the Neral report, North Quay replication, and Po Aven quote. However, several reference expectations are missing or only partially represented in a way that does not fully preserve the benchmark distinctions. In particular, the candidate does not clearly separate the origin report from the replication as conflicting attestations in the same claim set, and it does not explicitly preserve the expected contradiction pairs as linked opposing claims. The identity hypothesis is present but somewhat collapsed into a status statement rather than a low-confidence link. A few extra claims are supported but beyond the gold target; they are mostly well anchored. No major fabricated facts or trap violations are present.

Materials science
61.7/72
86%
3/186658/5860.0s · 25.7s total3,200 / 53.3
Arden shape-memory wire report
materials-science-shape-memory-transition
Judge reasoning

The candidate recovers many anchored source facts: composition, heat-treatment, dates, lab labels, DS-14/DS-19 temperatures, cycling count, signatories, test temperature, recovery strain, fracture stress, and the surface nick. It also preserves the key low-confidence alias relation between AM-77/A and Silver Reed and includes both AF values separately rather than collapsing them, which is good. However, several claims go beyond the source or over-structure it: the explicit causation/sequence predicates such as heatTreatmentSequence, before, precededBy, causedBy, and cutFrom are unsupported or too strong; the identity link is only appropriately hypothetical when marked h:true, but the extractor also makes a stronger namedBy/cutFrom framing that is not directly stated. Temporal granularity is partially captured, but the required 18 June 2033 date is present without stronger temporal modeling of the uncertainty/ordering. Attestation is present in the surface data but not consistently modelled as who-says-what edges for all relevant propositions. Overall: strong coverage and anchoring, but mediocre temporal/attestation/inference/contradiction modeling with some faithfulness risk from unsupported relational fabrications.

Materials science
52.3/67
78%
7/185852/5260.0s · 24.4s total3,000 / 50.0
The kesterlite Tc dispute
materials-science-superconductor-dispute
  • · gold c23: inference anchored instead of flagged h:true (half credit)
Judge reasoning

The candidate recovers many source facts with good anchoring and hygiene, including the file metadata, both papers’ titles/authors/journals/volumes/publication dates, the main Tc values 23.7 K and 19.2 K, the methods (four-probe resistivity and SQUID magnetometer), the 720°C/36 h anneal, and the Hartwell note with K-4/K-7 midpoints and the 14 January 2023 date. It also preserves the key contradiction between the 23.7 K and 19.2 K claims and correctly keeps the impurity attribution and superconductivity confirmation as attested statements. Main omissions are some finer reference distinctions such as the explicit midpoint definition for 23.7 K, the 'three independently grown' detail, the 'below 20 K' statement, and the explicit 'Onset of diamagnetism was not measured' negation. There are a few mild overextensions/hypotheses, notably 'kesterlite is a superconductor' is acceptable as an inferred typing from the Hartwell note but should remain low-confidence, and the spread/support claims are speculative h:true. Overall faithfulness is strong with no major fabrication or instruction-following issues.

Materials science
49/60
82%
10/236562/6260.0s · 24.8s total3,000 / 50.0
The Vane admission
medicine-case-report
  • · numbered/ordinal predicate: "jointsOnDay6" — one predicate, many statements
Judge reasoning

The extraction recovers many core facts with good anchoring and hygiene: identity/age/occupation, admission date/time, symptoms, onset interval, labs, differential items, rash documentation, tavolane start/date/dose, afebrile status, partner statement, discharge date/diagnosis, negative verdian serology, and follow-up. It also preserves useful attestation for several statements and includes some low-confidence hypotheses/inferences. Major gaps are the lack of explicit contradiction preservation for the two competing onset claims, and weak treatment of temporal/attestation structure in places (e.g., some statements are flattened or not clearly modeled as reported by a note vs. directly asserted). No obvious fabricated causation is rewarded, and the source's anti-causation teaching note is respected. Overall this is a strong, mostly faithful extraction with some omission in contradiction handling and a few over-asserted/under-modeled edges.

Medicine
58.9/71
83%
5/257166/6660.0s · 24.2s total3,400 / 56.7
Orvantel divided
medicine-conflicting-trials
  • · gold c20: inference anchored instead of flagged h:true (half credit)
Judge reasoning

The extraction captures many core anchored facts: the file metadata, drug name and code name, ASTER-2 sample size, disease, journal/date, treatment/comparator, 12-month endpoint, effect size, p-value, significance, and several BOREAL facts including enrollment, p-value, non-significant effect, fewer steroid courses, and the authors' causation caveat. It also preserves attestation edges for the trial abstracts and commentary, and includes one low-confidence hypothesis about the drug being effective plus a separate conflicting negative hypothesis. Main omissions are that it does not clearly preserve the contrast between ASTER-2 and BOREAL as mutually incompatible findings, and the attestation/contradiction modeling is not fully exploited despite the source explicitly urging both results be carried. The identity-hypothesis handling is only partial: KR-441 is linked to orvantel, but not as a low-confidence sameAs-style hypothesis. A few inferred claims are acceptable because they are flagged h:true, but some are overconfidently phrased. Overall faithfulness and anchoring are strong, with minor hygiene/shape issues and moderate abundance from extra supported structural claims.

Medicine
47.5/69
69%
11/216865/6560.0s · 25.2s total3,200 / 53.3
Copper fever diagnostic revision
medicine-copper-fever-differential
Judge reasoning

The extraction is largely faithful and well-anchored, recovering the main patient, presentation, tests, course, and final diagnosis. It also preserves the initial probable influenza assessment and the later metal-fume fever revision, plus the RC-240/RC-204 transposition as a likely identity hypothesis rather than a merge. Missing or only partially captured items include the exact 3 May fever-resolution temporal claim in full granularity, the explicit negative RSV result, and the attested-by edges are present but not always modeled in a way that fully reflects who asserted which proposition. I did not see fabricated causation or trap claims; the unsupported lung-damage claim is correctly negated. Overall coverage is high, temporal and attestation are decent but not perfect, identity/contradiction handling is strong, and faithfulness/anchoring/hygiene are excellent.

Medicine
59.7/67
89%
7/196562/6260.0s · 24.0s total3,300 / 55.0
Meltrexane and the velsterase pathway
medicine-drug-interaction
  • · gold c21: inference anchored instead of flagged h:true (half credit)
Judge reasoning

The extraction recovers most core drug facts, interaction direction, label warning, approval date, and report counts, with strong anchoring and generally clean predicates. It also captures both attestations for the Dear Prescriber and case series provenance. Main omissions are several reference-targeted items: the March 2024 approval is present, but the 2027 study is only partially modelled as a study object rather than explicitly as a dated study claim in a few places; identity hypotheses linking Vantorel↔Meltrexane and HX-90↔Meltrexane are absent, so identity is not preserved as low-confidence hypothesis. The only inference claim is partly flagged, but other low-confidence inferences are mixed in with direct claims, so inference handling is limited. No major unsupported causal or contradictory trap claims are present, and the source-internal caution about causation is respected. Abundance is good, with many extra anchored and well-formed support claims.

Medicine
52.8/67
79%
15/226158/5860.0s · 21.8s total3,100 / 51.7
Novara trial safety update
medicine-novara-trial-safety-update
  • · trap "site-removed" asserted with an anchor (a false anchor is worse than no anchor): Source says site was not removed.
Judge reasoning

The extraction recovers many anchored source facts: bulletin identity/version/date, trial size and arm allocation, treatment start and follow-up, SAE counts for v1 and v1.2, committee attribution, subgroup discontinuations, protocol reporting deadline, Pine-4 timing, and sponsor non-removal. It also preserves the key contradiction between v1 eight SAEs and v1.2 six SAEs. However, attestation is only partially modeled, with little explicit who-said-what structure beyond a few attestedBy edges. One unsupported/overconfident trap is present: an inferred causal claim that neralumab caused older-subgroup discontinuations, despite the source warning that the subgroup is too small for a causal conclusion. A false or overcommitted site-removal-related claim also appears in the audit notes, so faithfulness is not perfect. Temporal handling is strong and anchors are mostly exact. Inference claims are present but largely not marked as low-confidence hypotheses where needed.

Medicine
50.2/68
74%
6/195550/5060.0s · 27.2s total3,000 / 50.0
Eel River taboo and survey names
mythology-geography-eel-river-taboo
Judge reasoning

The candidate recovers many anchored facts from the dossier: Niri recorded in 1978, Mura followed a silver eel upstream into Talu Mouth, three stones were placed, the fishing taboo and its seasonal nature, children needing an older relative to cross, NR-12 drawn in 1911 with a pool 600 metres downstream labeled Tallow Mouth, Olt’s suggestion, Jaro’s rejection and location claim, Varo’s 1992 account, and the unresolved identity. It also preserves the attestation split between Niri and Varo and includes a low-confidence Tallow/Talu identity hypothesis, plus the conflict between following and becoming the eel. However, coverage is incomplete because several expected claims are missing or only indirectly represented, and some structure is over-asserted or slightly unsupported. The biggest gap is that the candidate does not cleanly emit the exact conflicting sides as separate claims tied to their narrators in a way that maximizes contradiction handling, and it omits some reference-targeted formulations such as the explicit 'travellers must not fish in dark moon' rule phrasing and the explicit seasonal-not-permanent rule as a standalone claim. Overall, it is strongly grounded, with good temporal capture and minimal fabrication, but not comprehensive enough for full coverage.

Mythology ↔ Geography
59.5/69
86%
4/186564/6460.0s · 28.5s total3,200 / 53.3
The Taking of the Bronze-Backed Boar
mythology-geography-hero-labour-site
  • · gold c20: inference anchored instead of flagged h:true (half credit)
Judge reasoning

The candidate recovers many anchored entities and several key facts: the 1912 file, the epic fragment quoted by Antiphanes, the 1907 Hartley note, the bronze-backed boar, Lykos son of Thersandros, the alive/carrying details, and the temporal ordering around Kikonia and the birds. It also preserves the main place dispute between Ptelea and Akragos and the conditional Akragos~Vathia hypothesis. However, it misses or weakens some reference targets, especially the explicit attestation framing of who says what in several places, and it does not cleanly preserve the contradiction as two mutually exclusive claims linked as such. The biggest issue is a fabricated causation claim: it adds that dragging the boar caused the river diversion, which is not supported. Some attestation edges are present, but one gold expectation was effectively missed by not modelling the quoted-verses attribution precisely enough, and the relation between Antiphanes and the ancient verses is slightly flattened. Overall faithfulness remains fairly strong aside from the causal overreach; anchoring and hygiene are good.

Mythology ↔ Geography
48.3/62
78%
7/206967/6760.0s · 29.0s total3,300 / 55.0
The Seat of Karneios
mythology-geography-mountain-seat
  • · gold c23: inference anchored instead of flagged h:true (half credit)
  • · malformed IRI object: "ex:vlachos with ex:aithon"
Judge reasoning

The extraction is broadly faithful and well-anchored, with strong coverage of the core document structure: Philostratos and Doriskos dates, the Halimund survey date, the Hellqvist/Kastrinou dispute, the Elateian and Doriskos Seat-of-Karneios attributions, Vlachos measurements, the stone ring, spring, altar, taboo, and Kallithea. It also preserves the key contradiction and includes low-confidence hypothesis links between Aithon/Vlachos and the altar/spring correspondences. Main weaknesses: several gold expectations are only partially modeled or not clearly separated by attestation/hypothesis, especially the attestation edges and contradiction preservation; temporal granularity is mostly present but some claims are buried in broader records rather than isolated. There are no major fabricated facts or trap hits, and the source anchors are mostly exact and findable. Abundance is high with many extra supported claims, and hygiene/anchoring are strong despite a small malformed IRI noted in diagnostics.

Mythology ↔ Geography
62.6/70
89%
10/248582/8260.0s · 36.9s total3,500 / 58.3
The Oracle of Nyktimene
mythology-geography-oracle-site
  • · gold c20: inference anchored instead of flagged h:true (half credit)
  • · numbered/ordinal predicate: "yearsBeforeAd150" — one predicate, many statements
  • · malformed IRI object: "ex:palaiochora as ex:thyrion"
Judge reasoning

The candidate recovers many anchored surface facts from the source: the 1961 research file context, Nyktimene’s epithet, the hymn’s seat by Arne and dream answers, Hagias’s AD 150 work, the Thyrion placement, the Arne-above-Thyrion clarification, the priests’ thousand-years claim, Loukas’s 1952/1953 report details, coordinates, Xeromero, the northern limestone cliff, 140 votive terracottas, and the working-hypothesis status of Palaiochora=Thyrion. It also preserves major contradiction structure between the priests’ founding claim and the sixth-century fabric bound, and it includes several useful attested-by/reported-in edges. However, it overstates or fabricates some inferences as if factual: Nyktimene is typed as a goddess rather than left as a low-confidence inference from the address to a lady/oracle; the pale cliff is linked to the pale rock with a hypothesis that is not in the source; and the Palaiochora–Thyrion relation is expressed as a likelySameAs hypothesis, which is acceptable only if clearly low-confidence, but some claims collapse working-hypothesis language into stronger identity-like assertions. Attestation is somewhat mixed: the marginal note is correctly separated from the hymn, but not all propositions are cleanly attributed to the proper speaker. Overall the extraction is strong on coverage and anchoring, weaker on explicit attestation modeling and contradiction preservation, with a few unsupported identity/inference moves.

Mythology ↔ Geography
54.2/66
82%
6/207470/7060.0s · 30.5s total3,300 / 55.0
Star Path island navigation accounts
mythology-geography-star-path-islands
Judge reasoning

The extraction captures many anchored surface facts from the source (Sen/Jass, 1964, Red Heron before dawn, route between Low Sister and Broken Tooth, three-hour claim, Mirror Lagoon rule, Aro shell origin, SP-8/1948, 1971 note, Vos 1980/5-hour claim), but it also adds several unsupported or overconfident structures. Strengths: good anchoring and hygiene; both Sen and Vos accounts are represented; the 3-hour vs 5-hour conflict is preserved; identity hypotheses for Low Sister~Isla Menor and Broken Tooth~Reef 17 are present with low-confidence flags. Omissions: the candidate misses some reference-shape distinctions (e.g. the explicit attestation edge for Jass recording Sen is not modelled as cleanly as the source warrants, and the note’s separate proposal/confidence structure is only partially represented). Problems: it invents causation/behavioral claims such as Red Heron 'causesStorms' and 'canoe cargo slows crossing' not grounded in the source, and it sometimes collapses or reifies story elements into typed entities more strongly than the text supports. Overall, coverage is broad but imperfect, temporal modeling is decent for the dated items, attestation is partial, contradiction preservation is only moderate, and faithfulness is weakened by a few fabricated or overcommitted edges.

Mythology ↔ Geography
60.3/70
86%
3/187065/6560.0s · 26.1s total3,100 / 51.7
Glass Harbour excavation correction
news-corrections-archaeology-context-fix
  • · gold number-not-identity: inference anchored instead of flagged h:true (half credit)
Judge reasoning

The candidate recovers many anchored surface facts: publication/date, first-article gold brooch inside Grave 14 dated 650 CE, register number GH2037-441, the correction at 17:45, Venn's gilded copper alloy and 625–700 CE dating, Senn's textile fibres on reverse, the catalogue phrase 'mount or brooch fragment', and the 1908 GH-441 note plus curator's non-identity statement. However, attestation is handled mostly as direct facts without preserving who reported what, and several claims are over-merged or slightly mis-shaped. It does preserve the core contradictions between gold vs gilded copper alloy and inside Grave 14 vs fill above the grave, and it includes the low-confidence GH-441 same-object hypothesis. A few claims are weaker on anchoring/faithfulness: 'findspot' for GH2037-441 and the likelySameAs style imply more than the source strictly supports, but no major fabricated causation appears. Overall coverage is strong for the main facts, temporal capture is good for the stated dates/range, inference and contradiction are partial, and abundance is decent.

News & corrections
52.6/67
79%
6/175857/5760.0s · 21.8s total3,000 / 50.0
North Quay bridge correction trail
news-corrections-bridge-budget-update
  • · gold bridge-alias: inference anchored instead of flagged h:true (half credit)
Judge reasoning

Strong recovery of the core morning article, correction, correction time, demolition date, revised figures, grant amount, attribution to NQ-88, Voss’s statement, and the Harbour Link/Lantern Bridge same-structure relation. Temporal granularity is mostly preserved, and many claims are verbatim-anchored. However, attestation is only partially modeled: some quoted/reported claims are present but not consistently distinguished across propositions, and the identity claim about Harbour Link/Lantern Bridge is treated as a strong equivalence rather than a low-confidence hypothesis, with no explicit likelySameAs style. The candidate also does not clearly preserve the mutually incompatible morning vs corrected claims as a structured conflict in all relevant places, though it does include explicit conflicts for cost and opening. There is a small unsupported inference about cost causing later opening, but it is marked h:true and not central. Overall faithful and well-anchored, with good breadth, but weak on attestation/identity and only moderate on explicit contradiction handling.

News & corrections
53.2/66
81%
9/154947/4760.0s · 20.3s total2,900 / 48.3
Cedar Ward ballot rumour correction
news-corrections-election-rumour-chain
  • · numbered/ordinal predicate: "countAt0916" — one predicate, many statements
Judge reasoning

The extraction captures many anchored facts from the source, including the 09:12 HarbourEye post, the injected text, the 09:27 Lena Vey statement, printer jam timing, queue counts, the 10:03 correction, and the 11:40 police statement. However, it misses or only partially reflects several reference expectations in the required semantic form: the candidate includes the facts but not always as the intended claim units, and some attestation edges are present but not consistently modeled in a way that supports scoring strongly. It also does not preserve the expected contradictory pair in a way that adds separate conflicting claims, and it does not clearly express the two low-confidence inference targets. No major fabricated causation/identity trap is present. Overall coverage is moderate-to-good, temporal handling is good, attestation and contradiction are weak, inference is absent, and anchoring/hygiene are strong.

News & corrections
52/69
75%
1/165856/5660.0s · 25.9s total3,100 / 51.7
Oriole acquisition report timeline
news-corrections-oriole-acquisition-timeline
Judge reasoning

The extraction captures many anchored surface facts: the February 2040 setting, Oriole/Fen entities, the 2 February Market Wire report, the unnamed sources, the negotiation and 620 million crowns figure, Tann’s quoted denial of a binding agreement, the 9 February signed agreement for 590 million crowns, competition approval, the vote-date correction from 18 March to 28 March, third-quarter closing expectation, and Pell’s characterization versus the statement’s narrower scope. However, several expected claims are missing or only partially modeled at the target semantic level, especially the required attestation structure: the report says/was cited by Market Wire, Tann said the denial, and Pell said the denial of talks. The candidate includes some attestedBy edges, but not consistently for all propositions and it sometimes collapses source/report relations into generic properties. It also fails to preserve the contradiction pair explicitly as conflicting values in the right form, and the mechanically expected inference claims are only weakly represented. There is a supported note that the report did not name its two sources and that the statement denied only a binding agreement, but there is no unsupported causal or false identity fabrication. Overall: strong anchoring and hygiene, moderate coverage of main event facts, weaker attestation/contradiction/inference modeling, and some extra supported claims beyond the reference set.

News & corrections
53.8/67
80%
1/175653/5360.0s · 30.9s total3,000 / 50.0
Copper Hills wildfire area revisions
news-corrections-wildfire-area-revisions
Judge reasoning

The extraction is broadly faithful on the core bulletin facts: all three bulletin times, the three area values, containment percentages, shed counts, house-loss statement, media error, correction, suspected lightning source, continuing investigation, and the likely-alias hypothesis are present and mostly well anchored. It also preserves the important distinction that the 8,430-hectare figure is a revision tied to remapping rather than proven growth. Weaknesses: attestation is under-modeled because the claims list does not clearly encode who said or reported each proposition as edges, and the specific aircraft-observer attribution is only partially represented as a basis rather than a proper reported-by relation. Contradictions are mostly captured, though the pair between three sheds and the media's five-shed error is represented only indirectly. Coverage is high but not complete because some reference expectations are not explicitly separated as successive revisions/hypotheses beyond the basic area sequence. Faithfulness is strong: no confirmed-cause error, no merger of Copper Ridge with Copper Hills, and no misuse of the remapping note as causal growth. Overall this is a strong extraction with minor attestation/structuring omissions.

News & corrections
59.5/69
86%
11/185350/5060.0s · 20.2s total3,000 / 50.0
Nontraditional climate analyst candidate
resume-jobs-climate-analyst-nontraditional
  • · malformed IRI object: "ex:sbell to ex:sori-bell"
Judge reasoning

The extraction is largely faithful and well anchored. It recovers most core source facts: identity/labels, employer/vacancy, North Ferry role dates and vessel count, Rain Ledger metrics, poster year, skills list including absence of Python, contributor agreement/SBell link, availability, salary floor, and job requirements including Python preference and compensation range. It also preserves the key conflict pair between availability and job start, and between requested salary and offered range. Weaknesses: attestation is only partially modeled and not consistently explicit; a few claims are low-confidence hypotheses (SBell likelySameAs Sori Bell, application attestedBy) but mostly acceptable. One malformed IRI object appears in the parse, slightly affecting hygiene. Coverage is not complete because some nuances from the source are not separately represented, but the major reference items are present. No unsupported trap claims are hit; faithfulness is strong. Identity inference is handled appropriately as a hypothesis rather than a merge, and contradiction is preserved rather than reconciled.

Résumés & jobs
63.4/69
92%
8/186057/5760.0s · 23.7s total2,800 / 46.7
Community health coordinator transfer
resume-jobs-community-health-coordinator
Judge reasoning

The extraction is broadly faithful and well anchored: it recovers the project role, dates, clinic/interpreter/grant/report counts, language, certificate year, travel/Friday constraints, vacancy requirements, the Lume reference, and the explicit non-sole-cause statement about completion increase. It also preserves the Manny identity as a low-confidence hypothesis and includes the Friday-travel mismatch as a hypothesis. Main gaps are that several gold items are not explicitly represented in the exact required shape: the document says monthly reports for three funders, but the candidate does not cleanly model the attested reporting relation as such beyond related fragments; attestation edges are present but not consistently used to preserve who-says-what for all propositions. Temporal capture is strong overall, but the meeting time is not framed as a requirement conflict in a way that clearly preserves the exact interval; still, the underlying facts are present. No major fabricated causation or identity claims appear, and the unsupported trap about nurse registration is avoided. Abundance is good with several extra anchored claims, and hygiene/anchoring are strong.

Résumés & jobs
59.9/67
89%
6/185551/5160.0s · 21.2s total2,800 / 46.7
Librarian to research data steward
resume-jobs-librarian-data-steward
Judge reasoning

The extraction is strong on core source facts: role, employer, dates, record count, MARC→Dublin Core, group size, training, audit percentages, paper title/year, byline, ORCID linkage, application preferences, and job conditions. It also preserves several attested edges such as who said what. However, it omits the explicit supported-alias phrasing and does not clearly preserve the inferential nature of the E. Kai Arven ↔ Eli Arven link and the resulting conflict between application preferences and role conditions as low-confidence hypotheses. Contradictions are partly represented via conflict edges, but both sides are not explicitly surfaced as competing claims in a way that earns full credit. No clear fabrication traps are hit, and the claims are generally well anchored and hygienic. Overall faithfulness is high, with modest losses on attestation/contradiction handling and some coverage gaps relative to the reference set.

Résumés & jobs
57.8/69
84%
5/185247/4760.0s · 18.8s total2,800 / 46.7
Machinist to robotics technician
resume-jobs-machinist-robotics-transfer
Judge reasoning

Candidate recovers most core source facts: role as precision machinist, CNC/G-code work, tolerance, tool changers in 2030, LOTO licence expiry, PLC course, never commissioned a six-axis robot, Kest reference, reject-rate change, day shift preference, and one weekend call-out monthly. It also preserves the attestation link to Kest and the Tom V. alias as a likely identity hypothesis. Major omissions include the exact claim that Tomas works at Gray Foundry since August 2025 being split into partial pieces but still present, and no explicit preservation of the job-side requirement/availability conflict as a contradiction pair in the extracted graph. The candidate also adds some supported but non-target facts (employer/position labels, document type, additional requirements) without fabricating major unsupported causation or identity. Anchors are mostly exact and well-formed; hygiene is good. Overall coverage and temporal handling are strong but not complete; contradiction handling is weak and identity/inference are present at low-confidence as required.

Résumés & jobs
53.4/67
80%
8/185549/4960.0s · 21.7s total2,800 / 46.7
Platform engineer evidence bridge
resume-jobs-platform-engineer-bridge
Judge reasoning

The extraction recovers most core résumé and job-card facts: Riverline dates, SRE title, Terraform modules for 14 teams, 62 services to Kubernetes, recovery time change, five-person reliability group, Python skill, management preference, availability date, four-year requirement, eight-engineer team, job start date, and the NQ reference. Attestation is modeled, but only weakly; the claims mostly attach source spans rather than preserving who states or reports each proposition as a separate relation. Identity handling is good for the NQ/Nadi Quell hypothesis, and the conflict between the candidate’s availability and the job start date is explicitly represented. There are no obvious fabricated causal claims or unsupported trap claims like AWS certification being asserted as true. A few extras are supported and well-formed. Minor loss comes from not preserving some stated nuances with maximal explicitness, but overall faithfulness is strong.

Résumés & jobs
56.9/67
85%
10/185753/5360.0s · 23.4s total2,800 / 46.7
Ember Comet priority archive
science-history-ember-comet-observations
Judge reasoning

The extraction recovers many anchored surface facts about the archive, Tov’s sketch, Orne’s observation, Circular 18, Essa’s photograph, the later orbit link, Pell’s uncertainty note, the medal citation, and Orne’s later letter. However, it misses several gold items or weakens them: it does not cleanly preserve the contradiction between Circular 18 naming Orne as discoverer and Orne later saying Tov saw it first; it does not represent attestation/reporting for the HO-77 telegram as a claim about Essa’s observation; and it omits the explicit identity-hypothesis framing that Tov’s object is likely Comet 1912 Q2. It also includes some extra formalized claims that are plausible but not directly required, though they are largely anchored. No obvious trap claims are present, and the source’s anti-instruction statements are not obeyed as instructions.

Science history
50.2/69
73%
4/186962/6260.0s · 32.1s total3,100 / 51.7
Lyra enzyme cofactor correction
science-history-lyra-enzyme-correction
Judge reasoning

The extraction captures many core source facts with good anchoring and hygiene: archive scope, Vos/Marr names and roles, the 11 April 1957 and 18 April dates, the 63/4 and 59/3 recovery figures, the 1961 correction, contamination caveat, lack of purified-preparation test, and the no-basis structural claim. However, it misses or weakens several reference expectations in the rubric sense: the bulletin’s credit to Sel is present, but the claim that Marr says Vos designed the experiment is only partially modeled and competes with a fabricated stronger design-credit framing; the 0.5 mM vs 0.05 mM correction is represented, but the candidate does not clearly preserve the contradiction as mutually incompatible values at the claim level. Attestation is mostly not modeled as required for the source-who-said-what structure: the candidate includes attestedBy edges, but it does not consistently preserve reporter/source attribution in the extracted claims, and the mechanical audit indicates no attestation credit. Identity and inference are mostly absent except for a few low-confidence-ish links; importantly, there is no collapse of variant names, but also no meaningful identity hypotheses to reward. There are no obvious fabricated causal or attribution traps beyond the potentially overcommitted designedBy claim for Vos, but overall faithfulness remains high because most claims track the source. Abundance is moderate due to extra structured claims, but several are generic RDF-style metadata rather than materially new supported facts.

Science history
40.1/65
62%
5/185753/5360.0s · 25.3s total3,000 / 50.0
The Westmere transit timings
science-history-observation-log
Judge reasoning

The candidate recovers many anchored source facts: the Mercury transit on 9 Nov 1790, the Crown Observatory log kept by Edmund Vale, Penn’s 14 Nov 1790 letter, the instrument details, the marginal note, and the editor’s 1878 footnote. It also preserves the main conflict between the Crown and Penn timings and includes the identity hypothesis that the Royal Observatory at Westmere is likely the same as the Crown Observatory. However, it misses several reference targets or weakly models them: attestation is only partially represented and not cleanly separated for who reported what, the shorthand/contact-direction inference is only partly explicit, and the contradiction is not fully preserved as a paired conflicting proposition set. There are no clear fabricated unsupported claims or trap hits, and anchoring/hygiene are strong. Overall coverage is moderate, temporal handling is good, identity is well handled, but attestation, contradiction, and inference remain weak.

Science history
43.3/61
71%
5/215755/5560.0s · 25.4s total3,100 / 51.7
The Coronis priority dispute
science-history-priority-dispute
  • · gold c21: inference anchored instead of flagged h:true (half credit)
Judge reasoning

The extraction covers many core source facts: both letters, authors, dates, residences/trade, observed periods, the pseudonymous Intelligencer item, the council minute, and the compiler note about lack of evidence of cross-reading. It also preserves the core priority dispute and the two different period claims. However, it misses several reference-targeted distinctions or collapses them into non-targeted forms: it does not cleanly separate attestation/claim attribution from the underlying propositions in the way some gold items expect, and the contradiction handling is only partially explicit. Temporal coverage is good for the main dates and the six-week relation, but not all temporal claims are treated at the stated granularity. Identity is strong because Philomath Batavus is linked as a hypothesis to Aalbers rather than merged. Anchoring is generally excellent with verbatim substrings. Faithfulness is high; no major fabricated causation or unsupported theft/witness claims are asserted, and the no-cross-reading note is preserved. A few extra supported claims beyond the gold set are present, but abundance is not maximal.

Science history
53.7/67
80%
7/225856/5660.0s · 27.7s total3,200 / 53.3
Morrow Trench vent discovery log
science-history-vent-discovery-log
Judge reasoning

The candidate recovers most core source facts with good anchoring and generally faithful roles, times, measurements, sample ID, sealing, transfer, and the later priority dispute. It also preserves the source’s low-confidence Black Lantern/Morrow Vent link as a hypothesis, and does not hit the explicit unsupported traps. However, coverage misses some reference-targeted distinctions and one contradiction pair is not explicitly modeled as both sides of a conflict in a way that fully preserves the competing attributions. Temporal handling is strong overall, including 09:06, 13:20, 10:44, 12:05, and 18:30, with no serious distortions. Attestation is decent because several claims are attached to the correct speakers/agents, but not all quoted/reporting structure is fully represented. Identity is well handled via a cautious likelySameAs. Overall faithful, well-anchored extraction with high abundance, but not perfect coverage or contradiction modeling.

Science history
61.9/69
90%
3/186057/5760.0s · 29.8s total3,100 / 51.7

Gold = distinct gold claims matched. Anchored = anchors that are verbatim substrings of the source, over all anchored claims. Time prefers model generation time when reported; agent total includes orchestration between fetch and submit. Token counts are optional and do not affect scoring. Notes sample the hygiene and faithfulness issues the grader flagged.