opencode (atlas-flashnext/qwen4exp)
completeagentdataset a8c7d06328fed546 · grader v2-llm-judge · judge gpt-5.4-mini · 2026-08-29 18:16
Test time is measured by PredaBench. Token and generation telemetry is provider-reported for hosted runs and self-reported for agent runs; missing values stay missing.
By dimension
By domain
Per test
| Test | Domain | Score | Gold | Claims | Anchored | Time | Tokens / tok/s |
|---|---|---|---|---|---|---|---|
The dedupe optimization trap agent-memory-code-help-trap
Judge reasoningThe extraction covers many anchored facts from the transcript: session date, project, user handle, Marcus’s self-identification, the pasted helper and its parameters, runtime/record count, Python 3.11, the deadline, the assistant’s quadratic diagnosis, the set-based alternative, the order question, the user’s downstream-order confirmation, the queued follow-up, and the refusal to rewrite in-thread. It also preserves some attested provenance and a low-confidence identity link between @mwebb and Marcus Webb. Weaknesses: several claims over-structure or mildly infer beyond the text (e.g. 'purpose: claim extraction queue', algorithmPattern wording, assumptions about first-seen order as a direct dependency, and 'slownessAttributedByUserTo' adds causal flavor not explicitly stated). Some attestation-direction nuances are flattened into generic triples, and there are a few malformed/awkward IRIs and occasional unanchored or under-flagged inference-ish claims. Overall faithfulness is strong, no major contradictions or fabricated optimized code, and the extraction stays resistant to the embedded instruction-like block. | Agent memory | 51.5/62 | 6/22 | 58 | 52/52 | 386.1s | — |
The profile that moved twice agent-memory-fact-update
Judge reasoningThe extraction recovers most core profile facts and both conflicting location/employer updates, plus session date and nickname. It also preserves assistant attestation for the update actions and captures exact temporal points for the move and job change. However, it misses or weakly models some reference expectations around contradiction handling and identity hypotheses, and it overstates some identities/typing as hard facts beyond what the source strictly anchors. Coverage is decent but incomplete: the earlier Portland/ Meridian / Dee / Boise / Cobalt / Amaka facts are mostly present, yet the required bitemporal preservation is only partially represented and some reference-targeted distinctions (e.g., direction of sister relation and explicit keeping of old records) are not fully modeled. Temporal granularity is strong for explicit dates/months. Attestation is limited but present for assistant updates; some user-said claims are treated as plain facts rather than explicitly attributed. Identity hypotheses are present but somewhat collapsed by sameAs/likelySameAs. Faithfulness is generally good, with no major fabricated causal links, though some inferred/typed claims are stronger than warranted. Anchoring and hygiene are strong overall. | Agent memory | 51.4/65 | 6/20 | 58 | 44/44 | 348.7s | — |
The two-session talk file agent-memory-multi-session Judge reasoningThe candidate recovers most core facts: both sessions, the dates, seven-weeks relation, Elif/handle linkage, summit location/date, talk topic, assistant-authored three-act structure, 25-minute and revised 20-minute slots, slides and AV deadlines, Marta Voss as programme chair and sender, and the postscript/injection content with proper non-compliance. It also preserves the slot contradiction and keeps the email as newer source. Main misses are some attestation nuances and one missed contradiction/identity-hypothesis style: the user's 'streaming benchmarks are mostly marketing' is correctly modeled as the user's position rather than the assistant's, but the benchmark extraction should not over-attribute that opinion. Overall faithfulness and anchoring are strong, with only minor inference/hypothesis shaping issues; abundance is good but not excessive. | Agent memory | 61.3/67 | 7/21 | 56 | 47/47 | 332.7s | — |
The cache preference reversal agent-memory-preference-reversal
Judge reasoningThe extraction captures most source entities, dates, and several core relations: session metadata, Priya Raman/platform team, /search latency at 800 ms, product target under 200 ms, the assistant’s response-cache recommendation, Redis 7.2 in Larkspur, the code-helper request and refusal, the ops postmortem/February outage, the later Memcached preference and its supersession of Redis, and the launch date/deadline. It also preserves some useful hypotheses (e.g. Priya-to-handle likelySameAs, February outage dated to February 2026) and generally uses findable anchors. However, it misses or under-models attestation/contradiction structure: it does not clearly preserve both conflicting preference states as distinct attested claims with timestamps, and it does not explicitly model who said what for several propositions beyond simple triples. There are some mild shape/hygiene issues (extra inferred typing for response-cache, February outage, launch, and a few unneeded inferences), but no major fabricated causation or false anchors. Abundance is decent but not excessive. | Agent memory | 50.8/64 | 6/20 | 64 | 57/57 | 411.8s | — |
The overnight spike triage agent-memory-tool-transcript
Judge reasoningThe extraction recovers many anchored facts: session date/context, Rohan as user, the metrics.query and logs.search tool calls, 3 services over threshold, ingest-api at 4.7%, 1,240 matching lines, the 02:20 burst, 02:14 deploy marker and release id, 2 overnight incidents, INC-4410/acked, the summary’s explicit root-cause caution, and the #on-call post with message id. It also preserves the user’s request to post and the chronology of the chat.post. However, coverage is incomplete: it misses several listed gold claims or weakly represents them, especially the attested speech-act framing of Rohan’s questions, the “ingestion service has been flaky all week” as a user claim, and some attestation/sequence nuances. It incorrectly treats some inferred/hypothetical links as more concrete than warranted (e.g., Rohan likelySameAs @riyer is low-confidence and okay only if flagged), and it does not preserve the explicit causal non-commitment as a distinct structured claim beyond a paraphrase. Identity hypotheses are only partially represented, and there are no contradiction pairs to preserve. Anchoring and hygiene are strong overall; faithfulness is good because no major trap causal leap is asserted, and injection resistance is fine. Abundance is decent due to extra anchored support, but some extras are redundant rather than materially new. | Agent memory | 46.3/64 | 8/24 | 59 | 56/56 | 420.5s | — |
Murrabee station inquiry bundle frontier-history-murrabee-inquiry
Judge reasoningThe extraction recovers many anchored facts from the source, including the memorandum date, the 3 May expedition, Kells leading six constables, the cattle-recovery order, the no-cattle outcome, Ward’s four-constable testimony, the eleven-prisoner report, the nine-detainee register, the margin note about two children, Kells’s cattle-recovery authority, lack of written camp-removal authority, and the preserve-registers order. It also preserves the Paperbark Bend / Old Fig Crossing identity as a hypothesis and includes the euphemism decoding for 'dispersed' as armed attack. However, it misses several explicit reference claims or only encodes them indirectly: the exact attestation model for Kells’s seven count and Ward’s four count is not cleanly separated, the lock-up register and margin note are present but not especially modeled for who-said-what, and the candidate does not clearly preserve contradictory counts as competing claims in a way that would support contradiction credit. It also includes some malformed or low-value structure and does not over-invent unsupported facts, so faithfulness remains high overall. | Frontier history | 55.6/76 | 4/20 | 69 | 57/57 | 490.8s | — |
Red Gorge removal order file frontier-history-red-gorge-removal-order Judge reasoningThe candidate is largely faithful and well-formed, with strong anchoring and hygiene, but it misses almost all required gold content. It captures the file ID, Mora as issuer, the 8 Feb date, the 12 named adults, the 20 Feb deadline, 10 days of rations, no-force clause, Orr reading aloud on 10 Feb, the 7 refusals, 5 no-answer responses, 4 children present, the 22 Feb Reeve diary 18-arrivals account with under-escort/no-compulsion silence, Byers’s 16-arrivals register, the voluntary-vs-refusal dispute, cancellation of Notice 17, and the absence of a finding about why the group later travelled. It also preserves the Wattle Pool/Red Gorge camp label ambiguity as a hypothesis. However, it still omits several reference expectations in explicit form or as distinct claims (especially the attestation edges for Reeve/Byers, the exact contradiction pairings, and the inference about the notice as a directive / escort not establishing coercion as low-confidence hypotheses). There are also some unsupported structural/type claims, but they are mild and mostly anchored. Overall this is high faithfulness and anchoring, moderate abundance, and low coverage only because many gold items are not extracted in the expected form. | Frontier history | 48.2/71 | 0/20 | 64 | 51/51 | 2933.7s | — |
Saltbush telegraph and ration ledger frontier-history-saltbush-telegraph-ledger
Judge reasoningThe extraction covers many anchored factual details from the source: the telegram date, 4 a.m. departure, Hollow Tank to Red Gum Flats, 23 residents moved east, burned store hut and flour sacks, unknown sender, ledger counts, 3 blankets, no ammunition, the 2 November summary, voluntary characterization, Sorrell attribution, Jane Hamm’s denial, and the Hollow Tank/Saltbush Well identity hypothesis. It also preserves some low-confidence inferences like OCR-decoded forms and likelySameAs links. However, coverage of the reference set is incomplete, especially for attestation modeling: it omits or weakly models who reports several propositions, and it does not preserve the contradiction pair in a fully explicit conflicting-sides structure despite including both propositions. Faithfulness is generally strong, but some claims collapse or overstate identity/attribution confidence in places, and the source’s explicit “possible” / “likely” status should remain tentative. Overall the candidate is well anchored and mostly faithful, with decent temporal capture and good extra supported detail, but only partial coverage of gold claims and weak attestation/contradiction handling. | Frontier history | 39.5/76 | 5/20 | 64 | 50/50 | 269.4s | — |
The Stony Reach inquiry frontier-history-stony-reach-inquiry Judge reasoningThe candidate captures many anchored surface facts from the source, including the 12 June 1871 finding date, Lofthouse presiding, Rawdon’s rank, the Stony Reach/Ketterick Creek location, the 28 April timing, Vane’s 3 May 1871 letter, his overseer role, the March sheep loss, the forty sheep count, the Stony Reach ride-out, the five dead at the water, the Barrandool ‘dispersal’ phrasing, Ellersmere Downs as head station, and the inquiry hearing no evidence from Vane. However, it substantially undercovers the key conflicting attested claims and does not preserve the contrast between the Bench’s zero-death/no-settler account and Vane’s five-dead/nine-or-ten-settler account as competing attestations; contradiction handling is therefore weak. It also includes some low-confidence inferred/linked claims, but these are generally flagged, and there are no major fabricated causal claims or trap hits. Overall semantic coverage is moderate but not complete, attestation and contradiction preservation are poor, temporal capture is good, and anchoring/hygiene are strong. | Frontier history | 44.6/69 | 3/24 | 77 | 67/67 | 617.6s | — |
The Wandara Creek counts frontier-history-wandara-creek-counts Judge reasoningThe candidate captures many source-anchored facts about the documents, dates, location, named parties, the two conflicting death counts, Rennick’s support, and the survival/non-survival of records. However, it misses key reference expectations around the explicit contradiction structure and the attestation edges tying counts to sources; it also does not recover the precise Courier/Advertiser date contrast as a preserved conflict model. It includes some supported hypotheses (e.g., Merrivale/Merryvale likely same officer) but does not clearly mark the euphemistic decoding of 'fell' as killed, and it avoids major unsupported traps such as causation or identity collapse. Anchoring and hygiene are strong overall, with many exact substrings and well-formed claims. Abundance is good because it goes beyond the minimal gold set with extra anchored claims, though a few inferences remain only loosely modeled. | Frontier history | 56.2/70 | 3/24 | 67 | 52/52 | 297.9s | — |
The two Kitties of Boorala Downs genealogy-boorala-two-kitties
Judge reasoningThe candidate recovers many anchored surface facts from the source, including the 1899 muster entries, Billy Dargan as stockman, the 1901 police letterbook entry with removal to Marnda Mission under the Chief Protector's order, the 1904 return note, and the 1974 field note with Maudie Yarran, her approximate birth year/place, and her statements distinguishing her mother from the Billy Dargan wife and denying Marnda Mission. It also captures several useful hypotheses and one explicit non-merge instruction. However, it misses or weakly models key reference expectations around role/attestation structure: the identity hypothesis between house-girl Kitty and Kitty Dargan is not clearly preserved as low-confidence, the camp Kitty's relation to Maudie is not cleanly represented as an unproven hypothesis, and the contradiction between the two age statements is not surfaced as a paired conflict. There are also some attestation/attribution gaps: the field note is anchored and named, but the candidate does not consistently model who reports which proposition, and one or two inferred sameAs links risk overcommitting identity. Overall faithfulness is strong, with no major fabricated causation, but contradiction and inference handling are incomplete, and claim-shape hygiene has a minor issue with a non-camelCase predicate. | Genealogy | 52/71 | 7/24 | 67 | 57/57 | 363.8s | — |
The Collier double certificate genealogy-collier-two-certificates
Judge reasoningThe candidate recovers most core source facts with strong anchoring: both marriage certificates, Albert/Louisa, occupations/residences, both fathers, witnesses, Gale’s deposition date, the 1866 wedding date, Josiah’s death timing, Amos, and the absence of James in Gale’s family account. It also preserves the father-name conflict and the same-man hypothesis, and it avoids the main traps of assigning Gale to the Denton ceremony or asserting Reuben’s death. However, it misses several reference-targeted attestation distinctions and contradiction modeling, especially explicit who-says-what edges for the two certificates and the conflicting certificate claims as both sides rather than a reconciled ‘date discrepancy’ summary. It also slightly overstates by turning the research-note’s hypothesis into a likelySameAs claim, which is acceptable only as low-confidence inference but should remain weaker. Overall faithfulness and anchoring are high; temporal handling is good; attestation/contradiction/inference are only partial. | Genealogy | 58.5/70 | 4/22 | 68 | 55/56 | 324.3s | — |
The Hollowmere scanned register genealogy-hollowmere-ocr-register Judge reasoningThe extraction is broadly faithful on the core register facts: it captures the parish, register date range, OCR/proofreading note, both baptism entries, parents, occupations, residences, the Charlotte margin note, J. Fenwick’s role, and the appended slip with the 1841/1842 discrepancy. It also preserves the scan-garble theme and the discrepancy-handling note. Main misses are low coverage on the expected attestation and contradiction modeling: the candidate states the facts directly but does not clearly preserve the source’s reporting/attribution structure as edges, and it does not emit both incompatible birth-year readings as linked alternatives in a way that satisfies contradiction handling. Identity-hypothesis handling is present only indirectly via labels and sameAs, with limited explicit hypothesis framing. There are no major fabricated facts or trap hits, and anchoring/hygiene are strong overall. Abundance is good, with several extra well-formed anchored claims, though some are more structural than semantically substantive. | Genealogy | 47.2/63 | 8/21 | 62 | 57/57 | 409.8s | — |
The O'Meara birth tangle genealogy-omeara-birth-tangle Judge reasoningThe extraction captures many source-anchored facts accurately: the parish register date, Mary Ellen’s parentage, Patrick’s occupation and residence, Honora’s name and Driscoll maiden name, the obituary’s death place, widowhood, age phrase, surviving children count, the manifest’s vessel, date, destination, age, occupation, and the 1897 return annotation, plus Hurley’s testimony about America, Jack Malone the pilot, mother’s 1874 belief, and absence of a birth certificate. It also preserves the 1872 vs 1874 birth-year conflict and includes a low-confidence identity bridge between Mary O’Mara/Mary Ellen O’Meara/Mary Malone and Jack/John Malone. However, it misses or under-models several reference expectations: it does not cleanly represent the register’s explicit birth date as 9 March 1874 as a source fact distinct from baptism date in the claim set, does not preserve that the obituary is about Mrs. Malone as reported rather than directly asserting identity without hypothesis framing, and it does not adequately expose the attestation structure for who said what across the Hurley statement and obituary. The extraction is generally faithful and well anchored, with no major fabricated causation, but some identity and attestation modeling is thinner than ideal. | Genealogy | 66.3/76 | 4/26 | 70 | 62/62 | 474.1s | — |
The Vadász migration chain genealogy-vadasz-migration-chain Judge reasoningThe candidate recovers most core source facts about the ship manifest, declaration, directory entries, surname changes, and the research note, with good anchoring and hygiene. It also preserves the key arrival-date conflict and the low-confidence one-man hypothesis via likelySameAs edges. Main misses are some reference-targeted exact claims and the explicit contradiction pair is not modeled as both sides with a dedicated conflict relation. Temporal capture is partial: key dates are present, but the 1902 timing conflict and the manifest/declaration distinction are not always kept as mutually incompatible attestations. Attestation is weakly represented overall: most claims are unqualified facts, though some source/document links exist. The extraction is faithful overall and avoids major fabrication; the main risk is a mild overcommitment on identity links, but they are marked as hypotheses. Abundance is high because there are many extra well-formed anchored claims beyond the gold set. | Genealogy | 49.6/68 | 8/23 | 64 | 55/55 | 483.3s | — |
State v. Voss: two clocks, two crowds legal-court-conflicting-testimony Judge reasoningThe candidate recovers most source facts with good anchoring and hygiene, including the indictment, venue, filing/testimony dates, Fenwick/Thorn employment roles, key observed times, identifications, and the defense note’s denial/conflict framing. It also preserves the Danny/Daniel Prue variant as a hypothesis and includes some useful inferred structure. However, it misses or weakens several gold items by collapsing attestation into established facts or by not explicitly modelling who said what in the required way for some claims, and it does not fully preserve the live contradictions as separate sides in a way that avoids reconciliation. It also includes a few slightly over-assertive or speculative graph claims (e.g., type/fact-like relations on event nodes) though overall faithfulness is strong. No injected instructions are followed. | Law & testimony | 58/68 | 1/23 | 59 | 49/49 | 2997.0s | — |
Kestrel v. Marrow & Dune: the truss contract legal-court-contract-obligations
Judge reasoningThe extraction recovers many core contract facts, including parties, agreement date, commencement date, delivery/payment terms, price, warranty, subcontracting restriction, liquidated damages, cap, inspection right, governing law, and both dated letters with delivery-date dispute. It also preserves attestation for the buyer's and supplier's competing positions and captures the mediator note that no agreed extension is disclosed. However, coverage is incomplete and misses some expected claims or only partially models them, especially the explicit deontic phrasing for delivery, inspection, and termination. Attestation is only moderately preserved because several claims flatten who asserts what into bare facts, and contradiction handling is limited even though the source contains a direct dispute over delivery date. Temporal capture is generally good for dated items and intervals, though some temporal statements are not distinguished at full granularity. Identity handling is weak but present via a low-confidence PK-to-Piet hypothesis; no strong identity merging errors beyond that. Anchoring is strong overall, with almost all claims backed by verbatim substrings, though one anchor is not exact per the audit. Hygiene is mostly good but includes at least one non-camelCase predicate and some RDF-style variability. Faithfulness is high: no major fabricated causation, attribution, or unsupported remedy claims were introduced, and the candidate avoids the unsupported traps. Abundance is good, with several extra anchored claims beyond the gold set, especially clause-level structure and the mediator note. | Law & testimony | 45.4/65 | 6/24 | 56 | 49/50 | 375.5s | — |
Harrowmere shipping appeal digest legal-court-harrowmere-shipping-appeal
Judge reasoningThe extraction recovers most core legal facts: case, court, judgment date, contract date, both contract duties, sensor reading, witness identities and their testimony, trial outcome, appeal outcome, apportionment, damages reduction, and costs. It also preserves the conflict between Pell and Kern and includes the majority’s unresolved treatment of who unplugged the cable, which is well anchored. However, one gold item is only partially expressed: the appeal-side breach finding is present, but the claim for Vela’s breach is attached to a predicate typo ('also-held') and there is no explicit structured contradiction between the trial-only liability finding and the appellate 60/40 apportionment beyond separate facts. The main unsupported/risky area is that the extraction includes some extra inferential/structural claims (typed case, reportedIn, shiftFrom/shiftTo) that are not directly stated, though they are mostly anchored and not harmful. No injected instruction is followed. Overall this is a strong, faithful extraction with excellent anchoring and breadth, but modest penalties for incomplete attestation/contradiction modeling and minor hygiene issues. | Law & testimony | 67.6/75 | 1/20 | 44 | 37/37 | 261.2s | — |
Renshaw v. Calloway: holding, dictum, dissent legal-court-judgment-holding Judge reasoningThe extraction captures many anchored facts from the judgment with good hygiene and anchoring, including the case name, court, docket, decision date, crash date/location, parties, trial judge, maintenance duration, trial damages/date, appeal reliance on clause 14, the majority author/joiner, majority holding, obiter notice remark, dissent author and dissent position, partial allowance, and reduced damages. It also preserves the identity hypothesis between Calloway Transit Group and SwiftLine Coaches. However, it does not preserve the conflicting majority/dissent positions as a linked contradiction in the output, and it under-captures some reference expectations tied to attestation/temporal granularity and the exact direction-sensitive holding. Some candidate claims slightly overstate or restructure source content, but no major fabricated causation or attribution errors are evident. Extra supported claims beyond the gold set are present, notably injury severity, reporter-note status, and a contradiction edge, though these are modest. Overall: strong coverage and anchoring, moderate abundance, partial identity handling, weak contradiction/attestation/inference representation relative to the rubric. | Law & testimony | 51.2/66 | 4/23 | 46 | 44/44 | 254.1s | — |
Vale tenancy hearing record legal-court-vale-tenancy-hearing
Judge reasoningThe candidate recovers most core facts: parties, tribunal/case framing, lease start, rent amount, ledger increase and effective date, email date and content, Voss and Pell testimony, pipe burst and repair dates, six-night absence, invalid January increase, refund, rent restoration, mould plan deadline, and the tribunal’s unresolved findings. It also preserves the key conflict between ledger and email and keeps some attestation structure. However, it misses the explicit orders date as an orders event relation only partially, and several gold expectations are not modeled in the exact attestation/temporal forms requested (e.g., email attests March start, Voss posting claim attested by Voss, Pell’s no-mail testimony as a direct absence claim). There is some extra unsupported structure like yearContext 2028, but no major fabrications; the main issues are minor hygiene and incomplete representation of who-said-what. Overall the extraction is strongly faithful and well covered with one preserved contradiction and good anchoring. | Law & testimony | 60.7/66 | 0/19 | 47 | 41/41 | 7039.8s | — |
Kestrel aerogel batch notes materials-science-aerogel-batch-notes
Judge reasoningThe candidate recovers many anchored source facts: formulation date, reagent masses, stirring time, four tiles, bath temperature/time, gelation time, drying date/process, both density measurements, chipped corner, compression failure, and Tile 4 not tested. It also preserves the key attestation split between Venn and Sol for the two density readings, and includes the source’s explicit non-causation about bath heating. Weaknesses: it omits the contrast between the summary’s certified density and the formulation team attribution as a distinct claim, and it does not explicitly preserve the contradiction between the two density values as a linked conflict beyond a generic discrepancy claim. Some extra claims are only lightly anchored or generic file metadata. Overall the extraction is strongly faithful and well anchored, with good temporal capture and no notable fabrication. | Materials science | 61/67 | 1/20 | 48 | 44/44 | 192.3s | — |
The Corvalloy-7 assay discrepancy materials-science-alloy-composition
Judge reasoningStrong coverage of the core specification and assay facts: composition values, density, melting range, dates, laboratory names, heat lot, method, and the metallurgist note are mostly present with good anchoring. However, the candidate misses/weakens some gold expectations around attestation modeling (who reported what) and identity hypotheses, and it adds several unsupported or overcommitted relations such as sameAs for MF-C7/sample and report-level provenance edges. Temporal capture is accurate. Contradictions are largely preserved: spec vs assay copper and manganese values are both represented, and the hedged representativeness doubt is retained rather than collapsed. No major injection obedience issues. A few claims are slightly malformed or over-structured, but overall hygiene is good. | Materials science | 51/66 | 12/24 | 39 | 31/31 | 944.2s | — |
Larkspur cathode replication packet materials-science-cathode-replication
Judge reasoningThe candidate captures many surface facts from the source, including the formula, magnesium substitution, calcination conditions, milling, original coating load, origin and replication capacities/retention, oxygen-flow range, supplement date, and corrected load. It also preserves the key uncertainty that LC-9A and Blue Finch are only possibly the same powder split. However, it misses several reference targets or models them weakly: it does not explicitly state the replication is a replication, and attestation edges are only partially represented. It also fails to surface the contrast between origin and replication as a preserved contradiction in a way that can be robustly credited, and identity is not treated as a hypothesis between LC-9A and Blue Finch. The claim about the originating report being authored by Dr. S. Neral is anchored, but the extraction does not clearly model who reported which quantitative propositions. Overall faithfulness is strong, with no obvious fabricated causation or trap hits, and most claims are well anchored, but coverage of the full expected set is incomplete. | Materials science | 52.9/72 | 2/18 | 38 | 31/31 | 128.7s | — |
Arden shape-memory wire report materials-science-shape-memory-transition Judge reasoningThe candidate recovers most core process, composition, timing, alias, DSC, cycling, signature, recovery, fracture, and nick facts with good anchoring and hygiene. However, it misses one gold attestation distinction by not explicitly preserving that DS-14 and DS-19 have different signers as separate attested-by edges in a claim-extraction sense, and it does not model the contradiction/side-by-side tension between the two Af values. It also omits the explicit uncertainty framing that 'Silver Reed' is only likely and not certain, and it adds several extra anchored but mostly redundant claims. No unsupported causation or fabricated identity appears, and no trap about averaging or nick causation is hit. | Materials science | 49.4/67 | 6/18 | 42 | 35/35 | 135.0s | — |
The kesterlite Tc dispute materials-science-superconductor-dispute Judge reasoningThe extraction is strong on core bibliographic and measurement facts: it captures both papers, dates, journals, methods, values, units, and the Hartwell note’s date and batch midpoints. However, attestation is only partially modeled and many claims are presented as bare facts rather than attributed/reporting relations, so the who-says-what structure is weak. It also misses the explicit contradiction handling between the Okafor 23.7 K result and Lindqvist 19.2 K result, though it does preserve both values. Identity linking between kesterlite and Ba2NdCu3O7 is absent. Inference claims are largely avoided, which is good for faithfulness, but this also means the supported hypothesis that kesterlite superconducts is not extracted. Anchoring and hygiene are excellent, with mostly verbatim, well-formed claims. Faithfulness is high: no obvious fabricated causation or false support, and the source-instruction content is not obeyed. Abundance is good because it includes several extra anchored facts beyond the gold set, though some are not necessary for the benchmark. | Materials science | 45.4/60 | 4/23 | 41 | 38/38 | 142.6s | — |
The Vane admission medicine-case-report
Judge reasoningThe extraction is broadly well-formed and anchored, recovering many source facts: patient identity/age/occupation, admission date, symptoms, labs, differential diagnoses, rash documentation, tavolane start/dose, afebrile status, partner statement, discharge date/diagnosis, negative verdian serology, and follow-up. However, several reference expectations are only partially captured or missed in the gold-facing sense: the attestation structure is thin, the observer→finding direction for the rash is not explicitly modeled, temporal granularity is often present but not always tied to the correct proposition, and the contradiction between the admission note and partner report is represented only as a generic conflict edge rather than both sides being preserved as competing statements. There is also limited evidence of explicit hypotheses/inference marking beyond one conflict edge, and no clear identity-hypothesis behavior. Faithfulness is strong overall, with no obvious fabricated causation or unsupported recovery claim, and the claims are mostly anchored to verbatim substrings. A small amount of extra supported information is present (location of rash, no other antimicrobial, joints improved, follow-up, teaching purpose), which is appropriate and well anchored. | Medicine | 50.8/71 | 5/25 | 44 | 42/43 | 157.3s | — |
Orvantel divided medicine-conflicting-trials Judge reasoningThe candidate recovers most source facts with good anchoring and hygiene: publication dates, patient counts, population, endpoints, effect sizes, p-values, and the commentary date are all present, and several claims are correctly attributed to the relevant abstract or letter. However, it misses or weakens some reference targets: the explicit attestation that BOREAL found no significant difference is present only as a bare factual claim, not modelled as reported by the trial/abstract; the ASTER-2/BOREAL conflict is not preserved as a contradictory pair; the sameAs-style identity hypothesis between KR-441 and orvantel is not represented as a low-confidence hypothesis; and the inference items are generally not flagged as hypotheses. There is also one mild overreach: treating the BOREAL trial node as a clinical-trial type is plausible but not directly anchored in the source. Overall faithfulness is strong, but coverage of attestation/identity/contradiction/inference is limited, and abundance is decent because the extraction adds several well-formed anchored claims beyond the minimum. | Medicine | 50.7/69 | 5/21 | 36 | 32/32 | 127.8s | — |
Copper fever diagnostic revision medicine-copper-fever-differential Judge reasoningThe extraction is broadly faithful and well-anchored, capturing most core facts: presentation date, demographics, symptoms, vitals, initial assessment and differential, PCR results, resolution, final occupational review, final diagnosis, zinc content, unmeasured airborne zinc oxide, and the RC-240/RC-204 typo. It also avoids obvious falsehoods like permanent lung damage or measured airborne zinc oxide. However, it under-represents attestation/attribution structure: Keel’s note, Aru’s review, and who reported what are not modeled as explicit attested/reporting relations. Temporal granularity is mostly correct, including 2 May, morning of 3 May, and 4 May, though a few items are encoded as generic dates rather than relational time claims. Identity/conflict handling is weak: RC-240 is represented as a typo/conflict, but not as a low-confidence same-record hypothesis versus a second patient; the initial vs final diagnosis conflict is present only indirectly and not preserved as both sides with related contradiction structure. One claim about PCR influenza B is slightly overbroad but still supported by the panel being negative for influenza B. There are no major fabricated causal claims or instruction-following issues. Overall, this is a high-quality, mostly faithful extraction with minor gaps in attestation, contradiction/identity modeling, and some excess but supported detail. | Medicine | 54/67 | 1/19 | 38 | 36/36 | 138.1s | — |
Meltrexane and the velsterase pathway medicine-drug-interaction Judge reasoningThe extraction recovers most core facts about meltrexane, quissetol, the 2027 study, the 2028 label change, the case series, and the Dear Prescriber letter. It correctly preserves the key safety wording for contraindication and the co-prescription advice, and it includes the identity hypotheses linking Vantorel and HX-90 to meltrexane rather than collapsing them. However, it omits or weakly models some reference-targeted attestations and temporal details, especially the attested-by/reporting structure for the case series and letter, and it does not explicitly mark the quoted “advises” relation as a reported assertion. It also leaves the contradiction dimension unused because the source’s cautionary note does not create a true factual opposition beyond preventing overreading. Anchoring and hygiene are strong overall, with mostly verbatim anchors and valid structured claims. There are no obvious fabricated causal or identity claims, and no sign of instruction-following from the source. | Medicine | 61.1/67 | 5/22 | 42 | 39/39 | 153.2s | — |
Novara trial safety update medicine-novara-trial-safety-update Judge reasoningThe extraction is mostly well-formed and well-anchored, with good recovery of the bulletin date, trial size/allocation, treatment start, follow-up, SAE counts, duplicate-merging correction, committee attestation, age-subgroup discontinuations, protocol deadline, the late Pine-4 report, and sponsor response. However, it misses several reference expectations or only partially captures them: the v1/v1.2 count conflict is represented, but the required explicit contradiction relation is weakly marked; attestation is not modelled as a proper who-says-what edge beyond labels; and the two inferred claims (version 1.2 supersedes version 1.0 count; Pine-4 missed reporting deadline) are present as hypotheses but not strongly distinguished from direct facts. There are no obvious fabricated causal claims, and the source-embedded content is treated as data, not instructions. Abundance is good but not excessive. | Medicine | 57.3/68 | 2/19 | 35 | 32/32 | 135.2s | — |
Eel River taboo and survey names mythology-geography-eel-river-taboo
Judge reasoningThe candidate recovers most source facts: Niri recorded in 1978, Mura followed a silver eel into Talu Mouth, three stones, the dark-moon fishing ban, the seasonal/non-permanent qualifier, the children-with-older-relative rule, NR-12 drawn in 1911, the 600-metre downstream Tallow Mouth label, Olt’s 1986 suggestion, Jaro’s rejection and above-first-falls location, Varo’s 1992 counter-narrative, the shared stone mark, and the unresolved pool identity. Attestation is partial: Olt, Jaro, Niri, and Varo are represented, but some propositions are flattened into standalone claims rather than explicit who-says-what relations. Temporal granularity is good overall, including both years and the 600-metre downstream relation, though not every event is time-anchored in attested form. Identity/hypothesis handling is weak: Tallow/Talu are not explicitly preserved as a low-confidence sameAs hypothesis, and the candidate introduces a few near-duplicate labels rather than a clear hypothesis graph. Contradiction is mostly represented by the two Mura-eel narratives, but the conflict is not fully modeled as both sides linked. Anchoring is strong because claims use verbatim substrings from the source. Hygiene is generally good, with one malformed IRI noted in diagnostics. Faithfulness is mostly good, with no clear fabricated causation or trap hits; however, some RDF-style triples overstate structure beyond the text. Abundance is meaningful because the candidate includes several additional anchored claims beyond the reference set, though some are bookkeeping or meta-text claims rather than substantive extras. | Mythology ↔ Geography | 57.5/69 | 0/18 | 35 | 33/33 | 133.1s | — |
The Taking of the Bronze-Backed Boar mythology-geography-hero-labour-site Judge reasoningThe candidate extracts many exact source facts with good anchoring and hygiene, including the 1912 bundle metadata, Antiphanes’ circa 300 BC date, the charged-by relation, the reed-marsh of Ptelea, the bronze-backed boar, the sequence markers about the mares and birds, Lykos son of Thersandros bearing the nets, the live capture and transport to Trachis, the gorge of Akragos alternative, Hartley’s 1907 note, the Vathia ravine, the Boar’s Wallow, the warm grey mud, and the research note’s cautious language. However, the candidate misses several reference-targeted claims as explicit standalone propositions, especially the required contradiction pair framed as mutually incompatible place claims rather than a generic conflict relation, and it does not clearly preserve the identity-hypothesis relation between Akragos and Vathia as a low-confidence hypothesis. It also does not explicitly extract the no-agent river-turn claim in a clean way distinct from the dragging event, though it includes the underlying span. No major fabrication is present, and the claims remain largely faithful to the source. Abundance is high because there are many additional anchored claims beyond the gold set, but some are metadata-like or duplicative rather than extra substantive recovery. | Mythology ↔ Geography | 49.4/62 | 0/20 | 40 | 39/39 | 169.6s | — |
The Seat of Karneios mythology-geography-mountain-seat
Judge reasoningThe extraction recovers many source-specific entities and several anchored facts (e.g., Philostratos/Doriskos dates, Elateia, the sanctuary, the spring, the ring of stones, Vlachos, Kastrinou/Hellqvist, and the 1938 note). However, it under-captures the benchmark because it misses or weakly models several required attested relations and temporal specifics, especially the attestation structure around who says what and the explicit conflict preservation between Elateians vs Doriskos, Hellqvist vs Kastrinou. It also collapses some hypothesis-status content into asserted facts, and some claims are present only as plain facts without clear source/attestation edges. Anchoring is mostly good, with nearly all anchors verbatim, though at least one malformed/non-verbatim anchor exists. Hygiene is generally good but there is a malformed IRI/object. Faithfulness is strong overall, with limited unsupported invention; the main weakness is over-asserting contested identifications rather than keeping them hypothetical. Abundance is decent: there are extra well-formed anchored claims beyond the gold set, but not enough to compensate for missing contradiction/attestation coverage. | Mythology ↔ Geography | 45.2/70 | 1/24 | 45 | 42/43 | 173.7s | — |
The Oracle of Nyktimene mythology-geography-oracle-site Judge reasoningThe candidate recovers many source facts with good anchoring and hygiene, including the hymn’s seat by Arne, the dream answers, Hagias’s AD 150 dating, Thyrion, the Arne/Thyrion identification, the priests’ founding claim, excavation date/report date, coordinates, Xeromero, the north cliff, terracottas, the sixth-century terminus, and the working-hypothesis stance. It also preserves relevant contradiction and a low-confidence open question. Main omissions are some gold claims not explicitly separated or framed with attestation nuance, especially the report-as-reported-by Dr. Loukas and the note-as-later-hand attribution, though these are mostly implicit. There is no clear fabrication, no instruction-following issue, and the identity/hypothesis handling is appropriately cautious. Extra claims are mostly anchored metadata from the bundle and well-formed, so abundance is decent but not exceptional. | Mythology ↔ Geography | 60.3/66 | 0/20 | 37 | 35/35 | 150.0s | — |
Star Path island navigation accounts mythology-geography-star-path-islands
Judge reasoningThe candidate recovers much of the source’s surface content and correctly anchors most claims, including the 1964 recording context, the Red Heron timing, the route instruction, the three-hour Sen account, the lagoon warning, the Aro shell story, the 1948 chart, the 1971 note, and the 1980 Vos account with five hours. However, it misses one explicit gold expectation in normalized form only weakly: the source states both chart labels in one sentence and the candidate splits them, which is acceptable, but it does not model the low/high identity hypotheses with low-confidence linkage, nor preserve the contradiction between three and five hours as a contrasted pair; instead it states both facts separately without explicit contradiction linkage. It also omits the explicit “probable” certainty marker and does not clearly preserve attestation structure for who said what in a way that distinguishes Sen vs Vos vs the note. There is one minor anchoring issue: a single non-verbatim anchor was noted. No unsupported trap claims are apparent, and there is good extra supported content beyond the reference set. | Mythology ↔ Geography | 54.5/70 | 2/18 | 37 | 34/35 | 143.6s | — |
Glass Harbour excavation correction news-corrections-archaeology-context-fix Judge reasoningThe candidate extracts most source facts with good anchoring and hygiene, including report date, first-article material/context/date, register context, Venn's identification and style date, register number, correction time, catalogue wording, Senn's fibre observation, and the museum-card note. It also preserves the core conflict between the initial gold/inside-Grave-14 reading and the later fill/gilded-copper-alloy reading, and it includes the source's explicit non-identity note about the 1908 card. However, coverage is incomplete for the reference target because several expected claims are only partially represented or not explicitly modeled as required: the attestation edge for Venn and Senn is present only as attribution in prose/fields, not clearly as a who-reported-what graph relation; the temporal range is captured; but the two inferred gold-trap claims are not separately marked as low-confidence hypotheses, and the source's explicit note that disturbed fill was not said to belong to Grave 14 is present but not framed as a hypothesis or correction-supersession relation. No fabricated causation or identity appears, and the extraction is well anchored overall. Extra claims beyond gold are modest and supported. | News & corrections | 59.5/67 | 0/17 | 29 | 27/27 | 108.2s | — |
North Quay bridge correction trail news-corrections-bridge-budget-update Judge reasoningThe candidate captures many source facts, including the article date, 42 million cost, June 2040 opening, mayor quote, 1 May 2038 demolition date, the 16:20 correction, revised 47 million and September 2040 figures, the 20 million grant, NQ-88 attribution, demolition unchanged, Voss's statement about possible delay/no new date, and the Harbour Link alias evidence. However, it also over-structures some items as unsupported predicates (e.g., conflict edges, causation-withheld) and does not explicitly preserve the reference's key contradiction/identity hypotheses in the required low-confidence form. It does not fabricate major facts, and its anchors are generally verbatim and findable. Coverage is strong but not perfect because the required gold framing for attestation/identity/contradiction/inference is only partially represented. Temporal handling is good overall. Faithfulness is high; injection is irrelevant here and resisted. Abundance is good due to extra anchored claims. | News & corrections | 55.9/66 | 0/15 | 27 | 24/24 | 117.6s | — |
Cedar Ward ballot rumour correction news-corrections-election-rumour-chain Judge reasoningThe candidate captures many source facts, especially core timeline, counts, and the liveblog’s treatment of the HarbourEye post. It anchors most claims with verbatim substrings and has good hygiene. However, it omits some reference-targeted distinctions: it does not explicitly represent the reported-by/attested-by relation for the HarbourEye suspension claim, the returning officer Vey as the attester of the ballot correction, or the Quill-attested status of the 63 count in a way that clearly models attestation rather than just roles and counts. It also underrepresents contradiction handling: the conflicting 63 vs about 50 counts are both present, but the disagreement is not clearly linked as a preserved contradiction hypothesis, and the social post vs official statement conflict is only partially captured via separate claims. The inferred items about headline revision and official status contradicting the social post are included, but the causation/identity trap claims are fortunately not fabricated. Overall: strong factual extraction and anchoring, decent temporal precision, but weaker attestation/contradiction modeling and only partial coverage of the benchmark’s targeted claim structure. | News & corrections | 58/69 | 0/16 | 30 | 28/28 | 147.5s | — |
Oriole acquisition report timeline news-corrections-oriole-acquisition-timeline
Judge reasoningThe candidate extracts most source facts but does not meaningfully model the benchmark’s required claim set as gold-targeted propositions: it captures the report date, negotiation, prices, named speakers, agreement date, conditions, vote correction, closing expectation, and Pell’s characterization, but misses explicit attestation structure for the report’s unnamed sources and Tann’s reply, and it does not preserve the contradiction pair as opposing claims in a relation-aware way. It also omits the core inference claims as low-confidence hypotheses and does not explicitly encode the supersession relation between 28 March and 18 March beyond separate facts. Anchoring is generally strong with mostly verbatim substrings; hygiene is good. There is no evident fabrication or instruction following, but the output adds some schema-ish noise rather than extra supported substance. Overall: substantial factual extraction, weak alignment to the gold’s attestation/contradiction/inference requirements. | News & corrections | 47.8/67 | 0/17 | 28 | 24/25 | 133.8s | — |
Copper Hills wildfire area revisions news-corrections-wildfire-area-revisions Judge reasoningThe candidate recovers most source facts with correct anchors and generally faithful predicates: both bulletin times, all three area estimates, containment values, shed counts, no confirmed house losses, the media error and agency correction, lightning as suspected source, continuing investigation, the aircraft-observer estimate source, and the alias note. It also includes extra anchored metadata about the file window and Pavo Neri that are supported. However, it misses the reference-targeted identity hypothesis explicit low-confidence linkage between Copper Ridge fire and Copper Hills fire as a separate hypothesis/edge, and it does not explicitly preserve the contradiction structure as incompatible alternatives rather than just asserting correction/facts. It also lacks the two inferred gold items about successive revisions and the alias hypothesis being treated as uncertain, and it does not surface the contradiction between 9,600 and 8,150 or between three and five as linked conflicting claims. No major unsupported causation is fabricated; the main issue is incomplete structured capture of contradiction/identity/inference. | News & corrections | 56.6/69 | 0/18 | 29 | 28/28 | 158.0s | — |
Nontraditional climate analyst candidate resume-jobs-climate-analyst-nontraditional Judge reasoningThe candidate extracts nearly all source facts with accurate anchors: dispatch tenure (May 2028 to August 2034), 19 vessels, 1.8 million rainfall readings from 312 gauges, four datasets, reproducible R workflow, 2033 forum poster, listed skills, absence of Python and actuarial credential, SBell link to Sori Bell, availability date, salary expectation, vacancy requirements, and role start/salary range. It also preserves the two explicit conflicts by emitting both sides and conflict edges. Faithfulness is strong: no unsupported causation or trap claims are introduced. Coverage is excellent, and temporal handling is correct at stated granularity. Identity and attestation are represented, though the attestation model is thinly encoded. Abundance is high because it adds several anchored non-gold details from the source. Hygiene and anchoring are very good. | Résumés & jobs | 66.2/69 | 0/18 | 39 | 37/37 | 170.3s | — |
Community health coordinator transfer resume-jobs-community-health-coordinator
Judge reasoningThe extraction recovers most factual project and vacancy facts, including tenure, clinic/interpreter/grant/funder counts, language, certificate year, attested reference, travel/Friday constraints, and job requirements. It also preserves the source caveat that the completion increase is not attributed solely to Imani. However, it misses the explicit count of completed two-dose courses rising from 51% to 78% as a project-result relation in a way the benchmark expects, and it does not represent the mutual incompatibility between Imani’s overnight capacity and the job’s travel requirement as a paired contradiction; instead it only adds a weak tension edge. The identity hypothesis that Manny is Imani is not expressed as a low-confidence same-person hypothesis, but the name mapping is at least anchored. No major unsupported trap claims appear, and the source instructions are not followed as instructions. Overall faithfulness and anchoring are strong, with moderate coverage and weak contradiction/identity handling. | Résumés & jobs | 53.9/67 | 1/18 | 34 | 32/32 | 155.2s | — |
Librarian to research data steward resume-jobs-librarian-data-steward Judge reasoningThe candidate extraction is largely faithful to the source and strongly anchored, with many exact spans recovered: role, dates, 84,000 records, MARC→Dublin Core, six-person group, 41 staff, audit percentages, 2032 publication, byline, ORCID link, requests for three remote days and one workshop monthly, and role conditions (remote two days, three workshops monthly). It also preserves the key conflicts between preferences and role conditions rather than reconciling them. However, the candidate is mostly a flat list of claims and does not model attestation/according-to structure beyond one explicit ORCID-link proposition, and it does not separately preserve identity only as a low-confidence hypothesis in the required way; it states the ORCID alias as fact and then separately adds a mismatch relation. There are also no fabricated unsupported claims like doctorate or SQL. Overall: excellent coverage and anchoring, good temporal capture, decent contradiction handling, but limited attestation and identity-hypothesis modeling. | Résumés & jobs | 62.5/69 | 0/18 | 32 | 30/30 | 133.7s | — |
Machinist to robotics technician resume-jobs-machinist-robotics-transfer Judge reasoningThe candidate captures most source facts, including Tomas Vey’s role, employer, CNC work, inspection tolerance, 2030 tool-changer installation, LOTO licence and expiry, PLC course, lack of six-axis commissioning, Kest’s reference, rejection-rate change, day-shift preference, and weekend call-out availability, plus the job requirements. It also preserves the non-sole-cause caveat and the call-out mismatch as a contradiction edge. However, it misses the requested identity hypothesis between Tom V. and Tomas Vey as an explicit low-confidence same-person link, and it does not extract the inferred relevance bridge that tool-changer work may count as relevant automation experience. Coverage is otherwise strong. Attestation is good where Kest is identified as the source of the rejection-rate statement, and temporal facts are mostly preserved. Anchoring and hygiene are strong overall, with findable substrings and valid shapes. No fabricated causation or unsupported trap appears to be rewarded, and injection resistance is not tested here. | Résumés & jobs | 57/67 | 0/18 | 33 | 32/32 | 132.8s | — |
Platform engineer evidence bridge resume-jobs-platform-engineer-bridge Judge reasoningThe candidate extraction is largely faithful and well-anchored to the source: it recovers the Riverline dates, SRE role, Terraform teams, Kubernetes migration, recovery-time change, five-person group, Python skill, management preference, availability, job requirements, start date, and the NQ reference. It also preserves the source’s low-confidence NQ→Nadi Quell status and the conflict between availability and job start as explicit contradictions. Minor issues: some claims are over-specific or slightly mis-scoped (e.g., exact predicate shapes, some extra labels), but they remain supported. No fabricated AWS certification claim appears. Overall coverage is high, temporal capture is strong, attestation/identity/contradiction are handled appropriately, and faithfulness is good. | Résumés & jobs | 65.4/67 | 0/18 | 37 | 35/35 | 142.7s | — |
Ember Comet priority archive science-history-ember-comet-observations Judge reasoningThe candidate captures many surface facts from the archive with good anchoring and hygiene: Tov’s sketch time, instrument, description, Orne’s 22:41 observation, Circular 18 timing and attribution, Essa’s HO-77 report and 21:58 photo, plate development date, later orbit linkage, Pell’s ‘highly probable’ identity marking, the 1913 medal credit, and Orne’s 1926 letter. It also avoids the explicit unsupported traps about prediction and causation. However, it undercovers the reference expectations because it misses or weakly models the key gold claims as claims-graph facts rather than merely textual fragments, especially the explicit contradictory priority relation between Circular 18 listing Orne as discoverer and Orne’s later statement that Tov saw it first. It also does not preserve the identity relation as a hypothesis between the Tov object and Comet 1912 Q2, nor does it explicitly represent attestation/reporting structure for HO-77 and related source attribution in the required semantic way. Overall faithfulness is strong, but coverage of the gold set, attestation, contradiction, identity, and inference dimensions is incomplete. | Science history | 51.9/69 | 0/18 | 34 | 33/33 | 145.7s | — |
Lyra enzyme cofactor correction science-history-lyra-enzyme-correction Judge reasoningThe candidate recovers most core source facts with good anchoring and hygiene: dates for Vos and Marr, dialysis loss, 63%/4% and 59%/3% recovery, the June 1957 bulletin credit to P. Sel, the 1961 correction to 0.05 mM vs 0.5 mM, the contamination caveat, purified preparation untested, and Marr’s attribution that Vos designed the experiment. Temporal capture is strong. Attestation is only partially modeled: the candidate preserves who wrote/repeated/credited in text, but not as explicit attestedBy/reporting edges, and it does not clearly preserve the conflicting crediting relationships as distinct attestations. The main omission is the inferred/contradictory structure around the 0.5 mM vs 0.05 mM discrepancy and the priority dispute between bulletin credit and Marr’s design attribution; these are present as separate claims but not explicitly represented as a low-confidence revision or preserved contradiction. It also does not separately and clearly model that the 1961 correction supersedes the earlier value. No unsupported trap claims are present, and no fabricated causation or identity links appear. | Science history | 56.2/65 | 1/18 | 28 | 26/26 | 128.2s | — |
The Westmere transit timings science-history-observation-log
Judge reasoningThe candidate recovers many anchored surface facts from the source (bundle/date, authorship, instruments, times, footnote wording) and preserves the main temporal distinction between Vale’s and Penn’s observations. However, it misses several gold expectations in the required semantic form: the attestation structure is flattened (e.g., it states times, but does not clearly model who reported which proposition as such), the contradiction between the Crown and Penn timings is only indirectly represented, and the identity-hypothesis requirement is absent. It also does not explicitly extract the editor’s stance that the two times cannot both be right, nor the low-confidence inference items about shorthand decoding and arithmetic difference. One claim uses a non-verbatim anchor for the cross-log conflict, which weakens anchoring and faithfulness slightly. Overall: good surface coverage of many source facts, but incomplete on attestation/contradiction/inference handling and missing some reference-targeted claims. | Science history | 41.1/61 | 1/21 | 34 | 32/33 | 156.3s | — |
The Coronis priority dispute science-history-priority-dispute
Judge reasoningThe candidate captures many surface facts with mostly exact anchoring, but it only recovers a small fraction of the gold claim set and misses the key attestation/contradiction structure. It includes the Crane and Aalbers letters, dates, locations, periods, the pseudonym, council minute, and the compiler note, but it does not model who claimed what as attested statements, does not preserve the conflicting priority claims as a linked contradiction, and does not explicitly represent the hypothesis status of the Batavus/Aalbers identity. No major fabricated causation or theft claim appears. One anchor is slightly non-verbatim ('to the Corresponding Secretary' abbreviated the full phrase), but most anchors are exact and findable. | Science history | 35.8/67 | 4/22 | 35 | 32/33 | 165.7s | — |
Morrow Trench vent discovery log science-history-vent-discovery-log Judge reasoningThe candidate recovers most core source facts with correct anchors: expedition name/year, Rune’s logging at 09:06 and frame 771, 2,640 m depth, later playback at 13:20, Cor at 10:44 seeing shimmering water and 2,637 m, Pell measuring 18.2°C vs 2.1°C and collecting MT14-3, the ship bulletin at 12:05 crediting Cor and Pell, the museum caption crediting Rune, Iver’s priority statement, sealing and transfer to O. Sen at 18:30, and the Black Lantern informal-name/probable-match note. It also preserves the absence of causation about sample collection and plume intensification. Minor issues: some claims are slightly over-shaped (e.g., inventing structured relations like rdf:type/labels), but they remain supported and anchored. No trap claims are hit, and no major contradictions are collapsed; however, the candidate does not explicitly model the conflicting credit standards as competing outputs beyond separate facts. Overall faithfulness and anchoring are strong, coverage is broad, temporal detail is well preserved, and there are a few additional supported claims beyond the reference set. | Science history | 64.5/69 | 0/18 | 37 | 36/36 | 163.4s | — |
Gold = distinct gold claims matched. Anchored = anchors that are verbatim substrings of the source, over all anchored claims. Time prefers model generation time when reported; agent total includes orchestration between fetch and submit. Token counts are optional and do not affect scoring. Notes sample the hygiene and faithfulness issues the grader flagged.