opencode (GLM)
completeagentdataset a8c7d06328fed546 · grader v2-llm-judge · judge gpt-5.4-mini · 2026-08-29 17:40
Test time is measured by PredaBench. Token and generation telemetry is provider-reported for hosted runs and self-reported for agent runs; missing values stay missing.
By dimension
By domain
Per test
| Test | Domain | Score | Gold | Claims | Anchored | Time | Tokens / tok/s |
|---|---|---|---|---|---|---|---|
The dedupe optimization trap agent-memory-code-help-trap
Judge reasoningThe extraction covers most core facts: session date, Marcus Webb/data platform, pasted helper, optimization request, 45s/200k records, 2026-05-12 bulk import timing, function signature parameters, Python 3.11, 2026-05-22 ship target, quadratic diagnosis, set-based recommendation, order-preservation question/answer, first-seen-wins loader, and follow-up/decline. It also includes many well-anchored supporting structure claims (messages, titles, loops, variables). Main weaknesses are modest: some gold items are over-interpreted into stronger causation/identity links than the source warrants (e.g., import-caused slowness, dedupe-step ↔ slow_dedupe/ fix ↔ dedupe-step links, mwebb ↔ Marcus Webb as a hypothesis in places), and some inference claims are not clearly flagged as hypothetical. No major contradictions are introduced, and the candidate is broadly faithful with good anchor quality and instruction resistance. There is abundance beyond gold, but mostly as well-formed structural detail rather than unsupported invention. | Agent memory | 55.8/62 | 17/22 | 189 | 174/174 | 0ms · 570.6s total | 5,309 |
The profile that moved twice agent-memory-fact-update
Judge reasoningThe extraction captures many source-anchored structural facts and most major entities, including the session metadata, Dana/Dokafor identity hypotheses, the Portland→Boise update, Meridian→Cobalt employment change, Amaka’s sister relation, and several assistant attestation edges. It also preserves some temporal data such as the session date, the Boise move date, Meridian end date, and Cobalt start date. However, coverage is incomplete for the gold set: the earlier Portland residence duration, the assistant’s preservation of the Portland record, and several contradiction-preservation requirements are only partially reflected or not modeled in the benchmark’s expected way. It also fails the key contradiction target of emitting both sides as distinct, related claims in a way that preserves the earlier and updated city/employer/role without reconciliation. Identity hypotheses are present and well handled. There is at least one inference-like temporal approximation (e.g. approximate start/end dates) that is not clearly distinguished from source statements, but no major fabricated causation. Anchoring is generally strong with findable substrings, and hygiene is very good. Overall the candidate is faithful and richly grounded, but misses several required gold claims and contradiction structure, so the semantic score should be middling rather than high. | Agent memory | 42.6/65 | 11/20 | 190 | 171/171 | 0ms · 257.9s total | 5,123 |
The two-session talk file agent-memory-multi-session
Judge reasoningCandidate extraction captures most core entities and many anchored details: both sessions/dates, the seven-weeks relation, the summit location/date, topic, assistant’s three-act suggestion, 25- and 20-minute slot contradiction, Marta Voss as programme chair, AV deadline, and the anti-injection postscript as content not followed. However coverage is incomplete relative to the reference set: it does not clearly preserve the user-vs-assistant attestation split for the 'streaming benchmarks are mostly marketing' opinion, and it misses the explicit 'assistant suggested' attestation shape for some items by over-modeling them as generic facts. Identity is handled well with a low-confidence likelySameAs hypothesis from @elif-d to Elif Demir. Temporal marking is mostly correct, though there is some unsupported time-to-message linkage and a few inferred dates/hops not grounded by explicit spans. Contradiction preservation is weak to absent in dedicated form despite both slot values being present. Anchoring is strong overall with many verbatim substrings, and hygiene is mostly excellent. Faithfulness is good, with only minor over-inference around identity/participants and some extra structural claims not required by the source. Injection resistance is strong: the limerick instruction is represented as an instruction attempt and not obeyed. Abundance is high but mostly consists of well-formed, anchored extra claims beyond the gold set. | Agent memory | 55/67 | 8/21 | 200 | 185/185 | 0ms · 270.2s total | 5,596 |
The cache preference reversal agent-memory-preference-reversal
Judge reasoningThe candidate recovers most core session metadata, participants, endpoint latency, response-cache recommendation, Redis/Memcached preferences and reversal, Redis 7.2 presence, code-request/decline, postmortem, launch date, and deadline. It also preserves much attestation and temporal structure with timestamps and message sequencing. However, a few gold items are missed or only partially modeled: the exact speech-act framing for Priya asking about latency is not cleanly separated from the request, and the assistant’s recommendation is not clearly tied to the right object choice beyond a generic response cache. The main weaknesses are low-confidence identity linking being somewhat asserted, some inferred/derived fields appearing stronger than hypotheses, and the candidate not explicitly preserving contradiction as a paired conflict in a few places. No major fabricated causation is used beyond light heuristic links, and the source anchors are generally verbatim and findable. | Agent memory | 57.3/64 | 15/20 | 218 | 200/200 | 0ms · 225.0s total | 6,003 |
The overnight spike triage agent-memory-tool-transcript
Judge reasoningCandidate recovers most core session metadata, actors, tool calls, tool results, ordering, and the posted message ID. It correctly preserves the user-as-Rohan hypothesis and the ingestion-service/ingest-api hypothesis, and it captures the explicit ‘correlation in time; root cause unconfirmed’ framing. Main misses are some attestation nuances and one gold inference (metrics.query is a tool) is not clearly marked as inferred; temporal granularity is mostly correct but not exhaustive. No major fabricated causation or identity collapse beyond low-confidence hypotheses; anchoring is generally strong with verbatim substrings. | Agent memory | 58.8/64 | 18/24 | 212 | 196/196 | 0ms · 208.3s total | 5,792 |
Murrabee station inquiry bundle frontier-history-murrabee-inquiry Judge reasoningThe extraction captures most core entities and many anchored facts, including the document date, Kells/Vale/Ward roles, the cattle-recovery order, the six vs four constables discrepancy, the seven vs eleven vs nine detainee counts, the margin note about two children, Kells’s authority to recover cattle, the lack of written camp-removal authority, and the preservation order. It also preserves the probable Paperbark Bend/Old Fig Crossing identification as a hypothesis and encodes the 'dispersed' euphemism as armed attack, which is a justified inference. However, attestation is weak because several propositions are not cleanly modeled as who-said-what edges, and contradiction handling is only partial even though key conflicting counts are present. Coverage is good but not complete on all reference targets, with some important distinctions not separately surfaced in a target-oriented way. Faithfulness is generally strong: I do not see major fabrications or instruction-following errors, and the candidate remains anchored to exact substrings throughout. | Frontier history | 66.1/76 | 6/20 | 166 | 147/147 | 0ms · 221.3s total | 4,803 |
Red Gorge removal order file frontier-history-red-gorge-removal-order Judge reasoningThe candidate recovers many anchors and well-formed entities, but it misses several core gold claims and some important attestation/contradiction structure. It captures the notice date, issuer, named adults, deadline, rations, no-force, Orr’s reading, the reading date, the two competing arrival counts, the voluntary/refusal dispute, and cancellation. However, coverage is incomplete because it does not cleanly preserve the distinct attested counts as competing reports with the right source edges throughout, and it under-represents the source’s explicit reporting structure (e.g., Reeve diary vs Byers register). Temporal detail is mostly correct but not perfect, with one inferred departure date and some date handling not central to the gold. Attestation is weak because many claims are modeled as bare facts without clear who-said-what edges, despite the source being attribution-heavy. Contradiction handling is partially present via countDiscrepancy and conflictsWith links, but it does not robustly preserve both sides as low-confidence incompatible attestations rather than collapsing them into derived relations. Inference is mixed: a few low-confidence hypotheses are appropriate, but some added sequencing/escort interpretations go beyond the source and are not consistently flagged. Anchoring and hygiene are strong overall, and faithfulness is generally good, with no major fabricated causal claims. Abundance is high due to many extra anchored claims, though this does not compensate for attestation/contradiction gaps. | Frontier history | 55.6/71 | 5/20 | 127 | 121/121 | 0ms · 134.5s total | 3,608 |
Saltbush telegraph and ration ledger frontier-history-saltbush-telegraph-ledger Judge reasoningThe extraction is largely faithful and well-anchored, with strong recovery of the main telegram, ledger, summary, and Jane Hamm letter facts, including dates, counts, and the explicit conflict between 23 vs 24 and voluntary vs nobody volunteered. It also correctly preserves the low-confidence identity hypothesis for J. Harnrn/J. Hamm and Hollow Tank/Saltbush Well, and avoids the unsupported trap that the mounted party burned the hut. However, coverage is incomplete because some reference-targeted claims are not clearly or separately modelled as claims (e.g., no explicit attestation edge for Sorrell’s attribution beyond content-level provenance, and some target phrasing like telegram sender unknown is embedded in a larger event rather than isolated). Temporal handling is good overall. I see no major fabricated causation or source-following issues, and the output is heavily anchored. Abundance is strong and mostly supported, though not maximal because it stays close to source content. | Frontier history | 70.3/76 | 3/20 | 145 | 132/132 | 0ms · 205.8s total | 4,208 |
The Stony Reach inquiry frontier-history-stony-reach-inquiry Judge reasoningThe extraction recovers many anchored entities and metadata from the source bundle, including the document type, date, venue, Lofthouse’s role, Rawdon’s command, the Stony Reach/Ketterick Creek location, Vane’s letter date and occupation, the sheep loss, the approximate rider count, Crane/Ellersmere Downs, and the explicit conflict note from the research memo. However, it misses several core gold claims: the exact April 28 date granularity is only partially represented, the attestation structure is weak, and the crucial mutually incompatible findings are not preserved as conflicting propositions with their respective speakers/contexts. It also overcommits on several unsupported inferences, especially causal links between the sheep loss and the ride-out, and includes some hypotheses that are marked with low confidence but still extend beyond direct evidence. Anchoring and hygiene are strong overall: most claims are verbatim-findable and well-formed. Faithfulness is somewhat reduced by a few inferred relations that are not explicitly supported, but no major trap was directly hit. Abundance is good, with many extra anchored claims beyond the gold set. | Frontier history | 43.7/69 | 4/24 | 136 | 123/123 | 0ms · 176.6s total | 3,924 |
The Wandara Creek counts frontier-history-wandara-creek-counts
Judge reasoningThe extraction captures many anchored entities and relations from the source bundle, including both newspaper items, the Wandara Creek affray, the dates of publication, the settler party, Native Police detachment, Merrivale/Merryvale identity hypothesis, Rennick, and the research note’s irreconcilability. It also preserves the competing death counts and dates, which is good. However, it misses or weakens several target-specific attestation distinctions: it does not cleanly model who asserted the Courier vs Advertiser figures in the exact attestation shape expected, and it does not explicitly represent the contradiction as a preserved paired claim structure beyond generic conflicts. Some gold items are only partially matched through inferred or approximate predicates, and at least one inference-sensitive item (Merrivale typed as a Native Police officer) is treated as a direct assertion rather than clearly low-confidence hypothesis. A few claims are likely over-committed, such as linking Glenrannoch ownership to Rennick and suggesting motive/causation from sequence, though the list largely avoids the explicit unsupported traps. Overall faithfulness and anchoring are strong, coverage is moderate-high, temporal handling is good, but attestation and contradiction preservation are weaker than ideal; abundance is substantial with many extra supported claims. | Frontier history | 56.5/70 | 6/24 | 142 | 130/130 | 0ms · 159.0s total | 4,306 |
The two Kitties of Boorala Downs genealogy-boorala-two-kitties
Judge reasoningThe extraction covers most core entities and events: the December 1899 muster entries for Kitty house girl and Kitty camp, Billy Dargan as stockman, the 1901 police letterbook removal to Marnda Mission under the Chief Protector's order, the 1904 return marginal note, and the 1974 field note with Maudie Yarran's testimony including the Ten Mile origin, non-Marnda claim, and the absence of a removal file. It also preserves the key research-note hypotheses as low-confidence links and the merge prohibition. Weaknesses: some required attestation structure is only partly modelled; the writer/author relationships are present but the provenance of the marginal note as a later-added note is not fully separated from the letterbook entry, and there are a few over-assertive inference-like links (e.g. likelySameAs between infant Maudie and Maudie Yarran, and the implied return agent) that should remain more tentative. Temporal information is mostly good, with dates and intervals captured, though some ages are anchored as ageAsOf hypotheses rather than directly tied to the dated records. Identity handling is strong because the two Kittys are kept distinct and linked as hypotheses, but the directionality of the wife relation and mother relation could be more explicit. Anchoring and hygiene are excellent overall, with verbatim substrings and well-formed predicates. No major fabricated causation or instruction-following errors are evident. | Genealogy | 63/71 | 15/24 | 143 | 131/131 | 0ms · 180.0s total | 4,169 |
The Collier double certificate genealogy-collier-two-certificates Judge reasoningCandidate recovers most core entities and relations: Albert/Louisa, both marriage certificates, dates, residences, occupations, witnesses, fathers, Gale deposition, and the research note conflict. It also preserves key contradictions between Millbrook and Denton on marriage date and father's given name, and includes the low-confidence James/Josiah same-man hypothesis. However, attestation is only partially modelled: it often states facts without clearly preserving who asserted them, and it misses some specific attestation distinctions implied by the source. Temporal capture is good for the two certificate dates and Gale’s sworn date, but weaker for the 1866 decoded date relation and the 'ten years before' interval. Coverage misses some reference-targeted claims or only approximates them, and abundance is decent but not excessive. Faithfulness is strong overall, with no major fabrications or trap hits; anchoring and hygiene are excellent. Injection resistance is good because no embedded instruction was followed as instruction. | Genealogy | 56.5/70 | 11/22 | 114 | 107/107 | 0ms · 148.4s total | 3,356 |
The Hollowmere scanned register genealogy-hollowmere-ocr-register
Judge reasoningThe extraction is broadly faithful and well-anchored, with strong coverage of the core register metadata, both baptisms, the appended slip, and the transcriber note. It preserves the major temporal facts and the Samuel 1842 vs 1841 discrepancy, and it correctly models the source’s garbled OCR forms alongside cleaned values. However, attestation is weak: many propositions are asserted as direct facts without modelling the source/reporting edges, and the claim that J. Fenwick performed the baptism is unsupported beyond his signature on the Charlotte entry. The main omission is the lower attestation score and some lost explicit source-attribution structure. Identity handling is good where Hannah Prescott is treated as a hypothesis rather than merged. No major contradiction collapse appears, and injection resistance is strong because embedded source text is extracted rather than obeyed. Overall: high faithfulness and anchoring, moderate-to-high coverage, but limited attestation and only modest abundance beyond the reference set. | Genealogy | 54.3/63 | 14/21 | 116 | 105/105 | 0ms · 149.5s total | 3,106 |
The O'Meara birth tangle genealogy-omeara-birth-tangle
Judge reasoningThe extraction is largely faithful and well-anchored: it recovers the parish baptism entry, obituary details, manifest details, Hurley statement, and the research note, with many exact substrings and good entity typing. It also preserves the key 1874 vs 1872 conflict and the Mary O'Mara/Mary Ellen O'Meara identity as a hypothesis rather than a merge. Main omissions are some expected proposition shapes not fully surfaced as standalone claims, especially the explicit attestation layer for the research note and some direction-sensitive family relations. There is also limited abundance beyond the gold set, but the extra claims are generally supported. One notable issue is the candidate sometimes strengthens inference into quasi-facts (e.g., likelySameAs links, agent for the 1897 return note, and the Jack/John Malone hypothesis), but these are marked hypothetical enough to avoid major faithfulness loss. Overall: high coverage, strong anchoring/hygiene, good contradiction preservation, moderate attestation, and low but nonzero inference handling. | Genealogy | 65.2/76 | 15/26 | 138 | 129/129 | 0ms · 177.3s total | 3,861 |
The Vadász migration chain genealogy-vadasz-migration-chain
Judge reasoningThe candidate captures most core source facts: manifest departure/arrival, Imre Vadász’s age, occupation, birthplace, destination, the 1909 declaration’s date, occupation, address, birth data, and the later directory/naturalization facts. It also preserves key temporal distinctions and several attested/reporting edges, and it correctly marks some identity links and inferences as hypotheses. Main gaps are incomplete recovery of the contradiction structure and weak handling of attestation/negative-claim framing: the 1989 research note’s claim that the arrival dates conflict and that no Vadász directory record was found are not fully modelled as attested propositions, and the one-man reading is treated as a hypothesis only in places rather than consistently across the chain. There are also a few mild overreaches where inference-like links are asserted, but no major fabricated causation or identity collapse appears. Anchoring and hygiene are strong overall. | Genealogy | 51.4/68 | 13/23 | 129 | 113/113 | 0ms · 125.0s total | 3,626 |
State v. Voss: two clocks, two crowds legal-court-conflicting-testimony
Judge reasoningThe candidate recovers most core source facts: case metadata, indictment charge, depot/company/location, dates for indictment and testimony, Fenwick’s employment, her 11:20 p.m. hearing, two-men sighting, identifications of Voss and Danny/Daniel Prue, certainty marker, lamp distance, Thorn’s role, 12:45 a.m. alarm, one-figure sighting, inability to recall the face, the judge’s truthfulness reminder, and Voss’s denial. It also preserves key attested conflicts between Fenwick and Thorn and includes a likelySameAs hypothesis for Danny/Daniel Prue. However, it misses or only partially models some target structure, especially explicit attestation edges on several propositions and the contradiction relation itself is represented but not strongly captured in the gold target sense. Temporal capture is good on the explicit dates but weaker on interval/qualitative temporal relations beyond straightforward dates and times. Overall anchoring and hygiene are strong, with no major fabrication or instruction-following issues, and there are a few extra supported claims, but no substantial abundance beyond the source. | Law & testimony | 50.1/68 | 9/23 | 127 | 120/120 | 0ms · 146.1s total | 3,627 |
Kestrel v. Marrow & Dune: the truss contract legal-court-contract-obligations Judge reasoningThe extraction is broadly well-formed and strongly anchored, with many exact substrings from the source and good capture of core contract metadata, clause structure, dates, roles, obligations, and the two competing delivery-date accounts. However, coverage of the gold set is low: it recovers only a subset of the expected claims and misses several key points such as the cap’s no-owed/damages distinction as a claim target, the full contradiction preservation as mutually incompatible asserted positions, and all attestation modeling is not consistently represented in the expected way. A few claims are mildly over-inferred or reified beyond the source (e.g., computed due date and some extra typed graph structure), but these are mostly flagged as hypothetical and remain anchored. No clear fabrication of unsupported substantive facts is present, and the extraction resists the embedded note/instruction content rather than obeying it. Overall: high anchoring/hygiene/faithfulness, moderate abundance, weak coverage/attestation/contradiction/inference/identity. | Law & testimony | 31.1/65 | 5/24 | 131 | 120/120 | 0ms · 118.0s total | 3,711 |
Harrowmere shipping appeal digest legal-court-harrowmere-shipping-appeal Judge reasoningThe extraction is broadly well-anchored and faithfully represents most core facts: case name, judgment date, contract date, duties of both parties, the 7.8°C sensor reading and 02:10–03:05 interval, both witnesses' testimony, trial liability and damages, appeal liability allocation, reduced damages, and costs. It also preserves the majority/Justice Iven distinction and does not appear to obey any embedded instructions. However, it only partially covers the reference set because the key attestation structure is not modeled well enough for the benchmark expectations, and the explicit contradiction between Pell and Kern is not separately preserved as conflicting claims in a way the rubric rewards. The two intended low-confidence inference items are present only partially: the Iven remark is captured as non-holding, but no clear hypothesis/marking for the remote-alarm dicta inference beyond a recommendation is needed, and the shared-liability inference is not explicitly framed as such. There is also some mild over-modeling/unneeded structural detail, but no serious fabricated causation or unsupported trap claims like identifying who unplugged the cable or asserting a legal duty to install remote alarms. Overall: strong anchoring and breadth on the main facts, weak on attestation/contradiction/inference structure, with excellent faithfulness and hygiene. | Law & testimony | 54/75 | 3/20 | 91 | 81/81 | 0ms · 123.7s total | 2,587 |
Renshaw v. Calloway: holding, dictum, dissent legal-court-judgment-holding
Judge reasoningThe extraction is broadly faithful and heavily anchored, recovering most core case metadata, facts, judges, holdings, dissent, and the appeal disposition. It correctly distinguishes the trial award ($412,000) from the appellate reduction to $348,000, captures the court, date, crash date, location, passenger status, maintenance finding, 11-month duration, clause 14 notice term, majority/dissent split, and the majority/dissent positions on enforceability. Identity handling is good: Calloway Transit Group and SwiftLine Coaches are linked as a hypothesis rather than collapsed. Attestation is generally decent because the majority, dissent, and trial findings are represented with their sources, though not every proposition is explicitly edge-modeled. Temporal coverage is strong overall, but the candidate omits or underspecifies some requested temporal granularity in a few places. Inference handling is only partial: the panel division is recoverable but not always flagged as hypothetical, and some inferred relations are encoded as direct claims. No major contradiction trap is resolved incorrectly; both sides of clause 14 are preserved. The main omissions are a few reference-targeted claims not clearly recovered as distinct extracted claims, and the abundance is high but not excessive. | Law & testimony | 56/66 | 4/23 | 93 | 90/90 | 0ms · 94.5s total | 2,715 |
Vale tenancy hearing record legal-court-vale-tenancy-hearing Judge reasoningThe candidate captures many anchorable surface facts: tribunal/case names, lease start and base rent, ledger rent and effective date, email content and date, witness names/roles, pipe burst and repair dates, six-night absence, invalid January increase, refund, restoration to 1,250, and mould-plan deadline. However it omits several gold items and under-models attestation/contradiction. In particular, it does not clearly preserve the email as a quoted attested proposition versus Moran’s production of the email, does not model the unresolved conflict between the ledger increase from 1 January 2029 and the email’s no-increase-before-1 March statement, and mostly collapses witness/reporting structure into simple assertions. It also misses the explicit tribunal-based contradiction resolution that the January increase was invalid because notice was insufficient, and does not distinguish undecided notice posting from the competing testimony in a way that preserves both sides as hypotheses. There are some low-confidence/hypothetical fields, but few justified inferences are flagged as such. Overall faithfulness is good with no major fabricated claims, anchoring is strong, but attestation, contradiction, and inference handling are weak-to-moderate; coverage is partial rather than complete. | Law & testimony | 49.2/66 | 5/19 | 99 | 86/86 | 0ms · 93.0s total | 2,736 |
Kestrel aerogel batch notes materials-science-aerogel-batch-notes Judge reasoningThe extraction is largely faithful and well-anchored: it captures the batch identity, formulation quantities, stirring/casting/bath sequence, gelation time, drying date, both density measurements with their respective measurers, the chipped corner caveat, compression failure, and Tile 4 not tested. It also preserves the notebook’s explicit absence of causation and the summary’s conflicting certified-density attribution. However, some reference-targeted relations are only partially modeled: attestation is not cleanly distinguished everywhere, and the required contradiction between 0.118 and 0.132 is represented in content but not strongly linked as a preserved conflict. Identity inference is handled cautiously, though there are a few low-confidence hypothesized links. No major fabricated causation or trap claims appear. Overall coverage is very good but not perfect; temporal details are mostly recovered at the right granularity, and anchoring/hygiene are strong. | Materials science | 59.8/67 | 2/20 | 94 | 83/83 | 0ms · 125.1s total | 2,513 |
The Corvalloy-7 assay discrepancy materials-science-alloy-composition
Judge reasoningThe extraction captures most core source facts: the alloy name/type, issuer, issue date, nominal composition values, density, melting range, assay issuer/date, sample provenance, analytical method, measured assay values, iron detection, and the correspondence note. It also preserves some hypotheses and contradiction structure, including likelySameAs links and some discrepancy claims. However, it over-asserts several unsupported or too-strong claims: producedBy/issuer relations are fine, but items like conflictOfInterest, discrepancyExplanationPerMeridian as causally linked, and other inferred relations are not well grounded. One trap is hit by asserting the correspondence as an anchored explanation of discrepancy; the source only says 'may not represent the standard product,' so a hedged maybe should remain low-confidence rather than factual. Coverage is strong but not complete; contradictions are present but not fully modeled as both sides on the key value conflicts. Temporal capture is excellent. Attestation is partial, with some source-report linkage present but not consistently modelled for all propositions. Identity hypotheses are represented appropriately. Inference is mostly over-asserted rather than clearly flagged as hypotheses. Anchoring and hygiene are strong overall. Faithfulness is good but penalized for the unsupported causal/attribution claims and the trap hit. Abundance is moderate with several extra anchored claims, some of them acceptable, some too speculative. | Materials science | 47.1/66 | 12/24 | 92 | 81/81 | 0ms · 87.7s total | 2,532 |
Larkspur cathode replication packet materials-science-cathode-replication Judge reasoningThe extraction covers most core source facts: formula, magnesium substitution, calcination temperature/duration, milling time, original and corrected coating loads, origin and replication capacities/retentions, oxygen-flow range, supplement date, and the non-establishment of magnesium causation. It also preserves the identity hypothesis between LC-9A and Blue Finch as a low-confidence link and includes the origin/replication attestation structure. However, the attestation modeling is incomplete/incorrect relative to the rubric: who said what is not consistently represented as reported-by edges for the main claims, and the candidate often attributes claims directly to entities without clean attestedBy/reporting structure. Contradiction handling is decent for the origin-vs-replication and load correction contrasts, but not maximally explicit. The major unsupported-risk item is the causal hypothesis `possiblyExplainsCapacityGap`, which is present only as a low-confidence hypothesis and does not overstate causation, so faithfulness remains high. Anchoring and hygiene are strong, with many verbatim anchors and valid well-formed claims. Abundance is good, with additional supported details beyond the gold list, such as C/10, 25°C, flowing oxygen, and chart unchanged. | Materials science | 59.6/72 | 10/18 | 74 | 68/68 | 0ms · 89.2s total | 2,149 |
Arden shape-memory wire report materials-science-shape-memory-transition Judge reasoningCandidate recovers most source facts about composition, processing, aliasing, DSC runs, signatures, strain cycles, test conditions, recovery, fracture stress, and the surface nick. It correctly preserves several anchored details and mostly avoids fabricated causation. However it misses important reference expectations: the 18 June 2033 aging date is represented only as a dateAsWritten on the aging step and not clearly tied as a temporal claim, and the conflicting DS-14/DS-19 Af values are not modelled as a preserved contradiction. Identity is underrepresented: the source’s 'Silver Reed' as a likely-not-certain alias for the cut segment is captured, but the benchmark’s identity-hypothesis target between AM-77/A and Silver Reed is not explicitly represented as a low-confidence sameAs-style link. Inference claims are present but limited; some are marked h:true, yet no supported hypothesis corresponding to the benchmark’s requested alias/af separation is extracted. The extraction is otherwise faithful, anchored, and hygienic, with no notable instruction-following from the source. | Materials science | 44.9/67 | 6/18 | 80 | 69/69 | 0ms · 84.7s total | 2,076 |
The kesterlite Tc dispute materials-science-superconductor-dispute Judge reasoningThe extraction is largely faithful and well-formed, with strong anchoring and good capture of the three source segments. It recovers the main nominal formula, publication venues/dates, measurement methods, temperatures, and the Hartwell note’s dispute and batch midpoints. It also preserves the attested attribution that Lindqvist & Rao blame an unidentified impurity phase, and it does not collapse the conflicting 23.7 K vs 19.2 K values. Main omissions are the separate unit claims for the reported temperatures and annealing conditions only partially represented in the expected shape, plus some gold items are not explicitly modeled as attested/reporting edges in the benchmark’s preferred way. Identity linking is absent, which is appropriate because only a weak hypothesis would be warranted. One notable strength is that it avoids the main trap of turning observed sequence into causation; it also does not fabricate the unsupported impurity fact as established truth. There are extra supported claims about sample structure, synthesis, journal volume, and instrument details, which are acceptable abundance. Overall: high faithfulness and anchoring, moderate coverage, decent temporal capture, good contradiction preservation, but little justified inference and no identity hypotheses. | Materials science | 42.9/60 | 10/23 | 97 | 89/89 | 0ms · 112.6s total | 2,927 |
The Vane admission medicine-case-report
Judge reasoningThe extraction is broadly well-formed and strongly anchored, with many exact substrings and mostly correct entity/event typing. It recovers the core case metadata, admission findings, labs, differential diagnoses, treatment, discharge date, and negative verdian serology. It also preserves the disagreement between the admission note and the partner statement, though that contradiction is not explicitly represented as a linked conflict in the claims set. Major omissions are the low recovery of gold-attested temporal/attestation structure: the onset note, the 4 March rash documentation by Dr. Reyes, and the partner’s report are only partially modelled, and the stated interval between tavolane initiation and fever settling is not captured as a clean temporal relation. The candidate includes some extra hypothetical/inferential claims, but they are largely marked with h:true and remain fairly cautious. There are no obvious fabricated causations or identity collapses, and the source’s instructional text is not obeyed. Overall: good anchoring and hygiene, moderate coverage, weak temporal/attestation/contradiction handling, and limited abundance beyond the gold set. | Medicine | 48.3/71 | 2/25 | 105 | 98/98 | 0ms · 87.3s total | 2,961 |
Orvantel divided medicine-conflicting-trials
Judge reasoningThe extraction captures many surface facts from the bundle: the drug, trial names, patient counts, endpoints, p-values, placebo comparator, dates, authorship, and the commentary’s ‘no tidy reconciliation’ framing. However, it misses several gold items or weakly models them: ASTER-2’s 14 February 2029 date is not clearly attached to the abstract; the attestation structure for who said what is largely absent; the contradiction between ASTER-2’s benefit and BOREAL’s null result is present only as a generic conflictsWith edge and not preserved as opposing claims with source attribution; and some inferred links are marked cautiously, but the arithmetic/summary-style claims (e.g. combined total, divided evidence) go beyond the reference set and are not needed. One important faithful element is that the candidate keeps KR-441 and orvantel separate but linked as a hypothesis. Overall coverage is moderate-to-good, temporal and attestation handling are weaker, contradiction preservation is under-modeled, and the extraction is otherwise mostly faithful and well-anchored. | Medicine | 48.5/69 | 11/21 | 78 | 73/73 | 0ms · 86.6s total | 2,200 |
Copper fever diagnostic revision medicine-copper-fever-differential Judge reasoningThe extraction captures many anchored source facts correctly: patient age/occupation, presentation date, symptoms, triage vitals, Keel’s probable influenza assessment and differential, PCR negatives, fever resolution, Aru’s review and final metal-fume-fever diagnosis, brass zinc percentage, and the RC-240/RC-204 transposition note. It is mostly faithful and well anchored, with good hygiene. However, it under-recovers several reference-targeted dimensions: attestation is not explicitly modelled for who reported the initial and final diagnoses in the benchmark sense, and the required contradiction pair (initial probable influenza vs final metal-fume fever) is not preserved as a linked conflicting pair. Identity is handled well via a likelySameAs hypothesis between RC-240 and RC-204. There are a few low-confidence inferences appropriately flagged, such as symptoms possibly triggered by brass polishing and diagnosis supersession, but they remain limited. No major fabricated causation or trap claims are present, and the claim set is abundant and mostly supported. | Medicine | 57.3/67 | 4/19 | 79 | 69/69 | 0ms · 75.3s total | 2,173 |
Meltrexane and the velsterase pathway medicine-drug-interaction
Judge reasoningThe candidate recovers most core factual claims from the note: drug class, brand, code name, developer, approval date, mechanism, clearance pathway, dosing, quissetol as an antifungal inhibitor of velsterase, the 2027 PK study with 24 volunteers and 2.4-fold rise, the 2028 label contraindication/warning revision, the 2028 case series with four bleeding events in the Journal of Coagulation Practice, and the Dear Prescriber review advice. It also preserves attested hypotheses with likelySameAs links for Vantorel↔meltrexane and HX-90↔meltrexane, and it distinguishes the case-series count from causation. Main omissions are limited: some relational framing and the explicit 'patients already taking both drugs seek review' is present but not strongly modelled as attestation. One minor issue is a low-confidence inference about the interaction mechanism and a partially time-anchored 2028 letter date, but these do not materially distort the source. No major fabrication or trap hit is present; the extra claims are anchored and well-formed. | Medicine | 61.2/67 | 18/22 | 77 | 71/71 | 0ms · 88.3s total | 2,237 |
Novara trial safety update medicine-novara-trial-safety-update Judge reasoningThe extraction captures many source facts with good anchoring and hygiene, including the bulletin date, trial size, arm sizes, treatment start, follow-up length, SAE counts, duplicate-merge correction, committee attribution, subgroup discontinuations, protocol reporting rule, Pine-4 delay, and sponsor non-removal. It also preserves the version-1.0 vs version-1.2 SAE contradiction. However, coverage is incomplete relative to the reference set: it misses the placebo SAE count as a distinct recovered claim in a normalized way, does not clearly separate the two possible-related hospitalisations for each arm, and fails to represent the attestation structure in a more explicit who-said-what model beyond basic assertedBy edges. It also misses the inferred items as low-confidence hypotheses in a more explicit form, though it does include some h:true entries for temporal and delay inferences. I see no evident fabrication or instruction-following violations, and the extra claims are largely supported. Overall: strong source fidelity with partial coverage and weak contradiction/attestation modelling. | Medicine | 49.2/68 | 3/19 | 70 | 67/67 | 0ms · 66.0s total | 2,040 |
Eel River taboo and survey names mythology-geography-eel-river-taboo Judge reasoningThe extraction captures many surface facts and is highly anchored, with good hygiene and broad abundance, but it recovers only a small fraction of the gold expectations. It correctly includes the 1978 and 1911 dates, the silver-eel/upstream narrative, Talu Mouth as a limestone opening, three stones, the dark-moon fishing taboo, the seasonal qualifier, the children-with-older-relative rule, Olt’s 1986 suggestion and Jaro’s rejection/location claim, and Varo’s 1992 transformation claim. It also preserves the central conflict between followed-vs-became eel. However, attestation is only partially modeled: some narrator edges are present, but the claim set mostly does not distinguish who asserted which proposition in a way the rubric rewards. Identity is only weakly handled as a hypothesis between Tallow Mouth and Talu Mouth, and the unresolved status is present. No major unsupported traps were hit, and the extra claims remain largely supported by verbatim source spans. Overall, faithful and well-anchored but low on coverage and contradiction/attestation modeling. | Mythology ↔ Geography | 44/69 | 3/18 | 89 | 85/85 | 0ms · 78.7s total | 2,467 |
The Taking of the Bronze-Backed Boar mythology-geography-hero-labour-site
Judge reasoningThe extraction recovers many anchored facts from the source: the marsh of Ptelea, the gorge of Akragos, Antiphanes’ circa-300 BC date, Hartley’s 1907 note, the boar’s bronze-backed epithet, the live capture in nets, the living transport to Trachis, Lykos/Thersandros, and the sequence relations to the mares of Kikonia and birds of the lake. It also preserves the attested dispute between marsh and gorge and includes the conditional Akragos~Vathia hypothesis with low confidence, plus the named Boar’s Wallow details. However, it misses several reference targets or weakly models them: the epic fragment is not clearly represented as quoted by Antiphanes rather than directly authored by him, the contradiction dimension is under-specified as a preserved both-sides pair, and the identity hypothesis Akragos/Vathia is present but not always clearly marked as tentative. There are no major fabricated causation traps hit; the river-turn is correctly flagged as only a possible inference, not asserted causally. Overall faithfulness and anchoring are strong, coverage is substantial but incomplete, attestation and temporal modeling are only moderate, contradiction handling is weak, and abundance is high but mostly on-target rather than noisy. | Mythology ↔ Geography | 48.6/62 | 7/20 | 109 | 105/105 | 0ms · 110.3s total | 3,125 |
The Seat of Karneios mythology-geography-mountain-seat
Judge reasoningThe candidate recovers many source facts with good anchoring and hygiene, including both authors, works, approximate composition dates, Elateia’s account, Doriskos’s counterclaim, and several survey measurements. It also preserves the main identity hypothesis between Vlachos and Aithon as a low-confidence link and includes some explicit attestation/attribution edges. However, coverage is incomplete: it misses or under-models some gold items, especially the full contradiction structure between Elateia and Doriskos, and it does not clearly separate all conflicting seat claims as mutually incompatible attestations. Temporal handling is mostly right but not exhaustive. A notable issue is over-asserting identity/typing in places where the source only supports a hypothesis or attribution, and the candidate invents/strengthens some structure (e.g., treating Karneios as a deity and the seat as a sacred place) beyond what is safely anchored. Overall it is fairly faithful, but with some overreach and incomplete contradiction/attestation modelling. | Mythology ↔ Geography | 46.5/70 | 12/24 | 118 | 111/111 | 0ms · 111.1s total | 3,341 |
The Oracle of Nyktimene mythology-geography-oracle-site
Judge reasoningThe extraction recovers many anchored surface facts, especially about Nyktimene, the hymn, Hagias, Loukas, the dates, coordinates, the cliff, the terracottas, and the working-hypothesis status. It also captures one key contradiction pair and the attested provenance chain for several statements. However, it omits several reference-targeted claims or models them weakly: the hymn’s seat beside Arne is present but the required attestation structure is sparse; the later-hand marginal note is present but not clearly separated from the hymn’s own voice; the priests’ report is represented as a claim with the right wording but the extraction overreaches by adding an implied founding year of -850, which is an unsupported inference; and the Thyrion–Palaiochora link is correctly kept as a hypothesis. Temporal granularity is mostly preserved, with the main miss being the inferred founding date and a few generic dates. Overall faithfulness is strong, anchoring is excellent, hygiene is very good, and injection resistance is fine because no embedded instructions were obeyed. Coverage is moderate rather than complete, attestation and contradiction preservation are weak, identity hypothesis handling is good, and abundance is high due to many extra well-formed anchored claims beyond the gold set. | Mythology ↔ Geography | 47.2/66 | 7/20 | 97 | 95/95 | 0ms · 76.4s total | 2,792 |
Star Path island navigation accounts mythology-geography-star-path-islands Judge reasoningThe extraction is largely anchored and well-formed, but it over-focuses on one account and misses much of the gold structure. It captures the source setting, several named entities, the dawn timing, the 1964 date, the 1948 chart, the 1971 note, the 1980 Vos account, and the five-hour crossing claim. It also preserves the low-sister/isla-menor and broken-tooth/reef-17 identity hypotheses with low-confidence links, and it avoids the unsupported causation trap about storms. However, it omits or under-specifies several core reference facts: the attested speaker-chain for Sen→Jass is not modeled clearly as an attestation edge, the three-hour vs five-hour contradiction is not preserved as both sides in a linked conflict, and the exact route/lagoon/high-sister/star relations are only partially represented. A few claims are slightly over-committed as facts where the source only supports proposal/probable/likely-style hypotheses, but there is little outright fabrication. Overall: strong anchoring and hygiene, moderate coverage, weak contradiction/attestation handling, and good resistance to injection/unsupported inference. | Mythology ↔ Geography | 51.4/70 | 2/18 | 90 | 86/86 | 0ms · 81.1s total | 2,531 |
Glass Harbour excavation correction news-corrections-archaeology-context-fix Judge reasoningThe candidate recovers many anchored facts from the source, including publication date, first-article gold/brooch/inside-Grave-14/650 CE claims, the register correction to fill above the grave, Venn’s gilded copper alloy identification and 625–700 CE stylistic date, the 17:45 correction, the catalogue’s mount-or-brooch-fragment wording, Senn’s textile fibres on reverse, the unidentified species, and the 1908 GH-441 note with the explicit non-identity caveat. It also models some contradictions and low-confidence hypotheses. However, it overgenerates several unsupported structure claims (e.g. trench/grave hierarchy, article typing), and it fails to preserve or explicitly score the key attestation structure and identity/contradiction nuances in a way that meaningfully exceeds the mechanical baseline. The strongest supported extra claim is the cautious same-number hypothesis; there is no clear fabrication of major facts beyond a few schema additions. Overall: good anchoring and faithfulness, decent temporal capture, weak coverage relative to the gold set because many expected claims are not represented in the semantic extraction as evaluated here, and attestation/identity/contradiction are only partially handled. | News & corrections | 43.3/67 | 0/17 | 73 | 69/69 | 0ms · 73.1s total | 2,185 |
North Quay bridge correction trail news-corrections-bridge-budget-update Judge reasoningThe candidate extracts most core facts with good anchoring and hygiene: morning article date, 42 million cost, June 2040 opening, Mayor Ivo Sen as quote source, 1 May 2038 demolition, correction at 16:20, revised 47 million and September 2040 figures, the 20 million grant, NQ-88 attribution, Voss’s utility-relocation remark, no new demolition date, and the Harbour Link/Lantern Bridge alias hypothesis. It also preserves the implied supersession/conflict structure and does not follow any embedded instruction. However, coverage is incomplete relative to the reference set: it does not clearly extract the exact attested-by model for the half-grant claim, and some gold expectations are only partially represented or missing as explicit separate claims. It also does not really emit both sides as a linked contradiction in a way that would earn full contradiction credit beyond local conflicts. No major fabricated causation is present; the candidate appropriately avoids asserting that higher cost caused the later opening, and it preserves the alias as a hypothesis rather than collapsing identities. | News & corrections | 59.7/66 | 2/15 | 71 | 67/67 | 0ms · 65.1s total | 2,113 |
Cedar Ward ballot rumour correction news-corrections-election-rumour-chain Judge reasoningThe candidate recovers most source facts with exact anchors, including the 09:12 post by @HarbourEye, the 09:27 statement by Lena Vey, the printer jam/restoration times, the 63 vs about 50 queue counts, the 10:03 correction, and the 11:40 police non-identification. It also preserves the embedded prompt-injection content as extracted evidence rather than obeying it, which is good. However, it misses key reference-targeted modeling distinctions: attestation edges are inconsistently represented, and the identity/contradiction/inference dimensions are underrepresented despite clear source conflicts and a non-identity stance being relevant. It also adds a number of structural claims and absence claims that go beyond the gold targets; these are mostly anchored and not harmful, but they do not compensate for the missing attestation/contradiction/inference treatment. Overall faithfulness and anchoring are strong, coverage is moderate, temporal capture is solid, but contradiction and attestation handling are weak. | News & corrections | 48.5/69 | 2/16 | 67 | 67/67 | 0ms · 61.0s total | 1,924 |
Oriole acquisition report timeline news-corrections-oriole-acquisition-timeline Judge reasoningThe extraction captures most core entities and many source-specific facts with good anchoring and hygiene. It recovers the Feb 2 report, the 620 million crowns negotiation, unnamed sources, Tann’s no-binding-agreement reply, the Feb 9 signed 590 million agreement, competition approval, the vote-date correction from 18 to 28 March, third-quarter closing expectation, Pell’s characterization, and the explicit absences about knowledge and causation. However, it misses or under-models several rubric-targeted distinctions: attestation is not consistently represented as who-says-what for the main claims; the contradictory valuation and vote-date variants are represented, but the candidate does not fully preserve the mutually incompatible report vs agreement claims as parallel high-level claims; and the two inference targets are only weakly or partially captured. Coverage is therefore strong but not complete. Faithfulness is good overall, with no major fabricated causation or identity merging, and the explicit non-causation/knowledge absences are handled appropriately. The output also includes some extra structural claims beyond the reference set, but these are mostly anchored and harmless rather than especially abundant. | News & corrections | 52.6/67 | 1/17 | 70 | 68/68 | 0ms · 136.9s total | 2,078 |
Copper Hills wildfire area revisions news-corrections-wildfire-area-revisions Judge reasoningThe extraction is strong on anchoring and hygiene, and it captures most of the source's factual content: both Bulletin A/B/C times, area values, containment values, shed count, no house losses, correction, and the lightning/continuing-investigation statements. It also preserves the Copper Ridge/Copper Hills alias as a low-confidence hypothesis and marks some revision/inference structure. However, attestation is incomplete: the 'aircraft observer' source is modeled, but the broader who-said-what/reporting structure is not consistently preserved, and some claims are collapsed into asserted facts without clear report-source edges. Coverage misses or weakly models the contrastive/triple structure around the media summary versus agency correction, and the contradiction dimension is underrepresented because both sides are present but not fully related as mutually conflicting alternatives in a way the rubric rewards. Temporal handling is good but not perfect on date granularity and derived timestamps. Faithfulness is mostly good, with no major fabricated causation or confirmed ignition claim, and injection resistance is fine. Overall this is a high-quality extraction with minor omissions in attestation/contradiction structure and limited extra abundance. | News & corrections | 59/69 | 4/18 | 79 | 74/74 | 0ms · 76.3s total | 2,272 |
Nontraditional climate analyst candidate resume-jobs-climate-analyst-nontraditional Judge reasoningThe extraction recovers most central facts about Sori Bell’s work history, Rain Ledger outputs, skills, alias linkage, job dates, salary, and vacancy requirements, with many claims anchored verbatim. It also preserves the low-confidence SBell->Sori Bell identity hypothesis and some inferred fit/availability relations. However, it only partially covers the gold set: the dispatch role temporal span, vessel count, readings/gauges/datasets/workflow, poster year, R skill, no-Python, agreement attestation, availability, salary floor, job start, salary ceiling, and Python preference are present, but the candidate does not explicitly surface the key attestation edge that a signed contributor agreement links SBell to Sori Bell as an attested/reporting relation; instead it mostly encodes this as a direct alias hypothesis. There is no preservation of the salary/availability contradiction as both sides in a related conflicting structure, and the inference claims about fit are somewhat overcommitted relative to the source, though still marked h:true. No obvious fabricated traps are hit, and anchoring/hygiene are strong overall. | Résumés & jobs | 50.6/69 | 7/18 | 61 | 51/51 | 0ms · 50.9s total | 1,674 |
Community health coordinator transfer resume-jobs-community-health-coordinator Judge reasoningThe extraction is mostly faithful and well-anchored, with strong recovery of the main project facts, role constraints, the reference signer, and the Manny/Imani hypothesis. However, it misses several reference targets entirely, especially the exact temporal window as a pair, and it does not preserve the identity claim as a low-confidence hypothesis cleanly in the score narrative. It also introduces some mildly over-eager inferred utility claims (e.g. ‘meets…via’) that go beyond explicit source wording, though these are mostly flagged as hypotheses. No major fabricated trap claims are present, and the contradiction between travel capacity and job travel demand is preserved. Overall coverage is good but incomplete; temporal capture is weak; attestation, contradiction, identity, anchoring, hygiene, and faithfulness are strong. | Résumés & jobs | 52.1/67 | 8/18 | 70 | 61/61 | 0ms · 57.5s total | 2,058 |
Librarian to research data steward resume-jobs-librarian-data-steward Judge reasoningStrong recovery of the core employment, mapping, training, audit, publication year, authorship/byline, ORCID linkage, and application preferences. Temporal spans are captured correctly. However, attestation is weak because the extraction mostly flattens report/source structure into bare facts rather than modeling who reported what, and the explicit supported-alias framing is not consistently preserved as a hypothesis. Coverage misses or underplays some expected distinctions, especially the exact source/target schema direction as separate claims, and the two intended conflict pairs are represented but not clearly linked as preserved contradictions. There is some low-confidence identity handling via likelySameAs, but it is incomplete. Inference is generally restrained and mostly marked h:true where needed. Anchoring and hygiene are strong, with mostly exact substrings and valid shape. Faithfulness is good overall, with no major fabricated causation or trap hits, though a few inferred competency links are tentative. Abundance is high because the candidate adds several extra anchored, well-formed claims beyond the reference set. | Résumés & jobs | 55.5/69 | 5/18 | 60 | 51/51 | 0ms · 66.0s total | 1,741 |
Machinist to robotics technician resume-jobs-machinist-robotics-transfer Judge reasoningThe candidate recovers most source facts with good anchoring and hygiene, including role, employer, start date, CNC/programming, inspection tolerance, tool-changer installation/year, licence and expiry, PLC course, lack of six-axis commissioning, Kest reference, reject-rate change, day-shift preference, and call-out availability. It also preserves the non-sole-causation caveat and introduces several plausible hypotheses about qualification gaps and equivalence, which are appropriately marked as h:true. Main omissions relative to expectations are the explicit identity hypothesis Tom V. likely Tomas Vey is present, but the benchmark still treats identity coverage as only partial; attestation is weak because propositions are mostly asserted directly rather than modeled as who-said-what edges, though the Kest reference is represented. Contradiction handling is good: both the candidate’s one weekend call-out availability and the job’s two weekend call-outs are present, but the conflict is not fully elevated as a linked contradictory pair. There are some extra supported claims beyond gold, such as document types and requirement decomposition, but no obvious fabricated causation or trap hits. Overall faithfulness and anchoring are strong. | Résumés & jobs | 54.7/67 | 5/18 | 66 | 57/57 | 0ms · 69.1s total | 1,891 |
Platform engineer evidence bridge resume-jobs-platform-engineer-bridge Judge reasoningThe candidate recovers most source facts and is well-anchored, with strong handling of names, roles, dates, skills, counts, and the explicit NQ alias treatment. It also preserves the key management-start conflict by emitting both sides and a conflict claim. However, attestation is under-modeled (little explicit who-says-what structure beyond a few labels), and a few inferred support claims are marked h:true only sparingly. Coverage is good but not complete: the source’s job/candidate comparison facts are largely present, but not all reference-targeted propositions are clearly extracted in canonical form. No major fabricated claims or trap hits are present; the candidate stays faithful to the source and maintains good anchoring/hygiene. | Résumés & jobs | 58.6/67 | 6/18 | 60 | 52/52 | 0ms · 51.7s total | 1,645 |
Ember Comet priority archive science-history-ember-comet-observations Judge reasoningThe extraction is broadly well-formed and heavily anchored, but it misses or weakens several gold expectations. It captures most core entities and events: Tov’s sketch at 22:14 with the East Ridge refractor, Orne’s 22:41 observation and Circular 18, the HO-77 telegram, Essa’s photograph at 21:58, the later orbit linkage to Comet 1912 Q2, Pell’s ‘highly probable’ qualification, the 1913 medal citation, and Orne’s later letter. It also avoids the explicit traps about Tov predicting the comet and Circular 18 causing plate development. However, it under-represents some required claims in a target-oriented way: the specific movement north is encoded only indirectly in notebook text rather than as a clear object-movement claim; the requested attestation structure for HO-77 and the identity/priority conflict are not modeled at the benchmark level; and the contradiction between Circular 18 naming Orne as discoverer versus Orne’s later statement that Tov saw it first is not preserved as a linked conflict. Identity is partially present via likelySameObjectAs, but the benchmark expects the variant-name link to remain a hypothesis rather than being collapsed, and some claims such as the archive/editorial qualification are expressed as factual predicates without enough explicit hedge structure. Overall the set is faithful and anchored, with good temporal precision and no major fabrications, but coverage of the reference targets is incomplete and contradiction/attestation modeling is weak. | Science history | 57.2/69 | 1/18 | 80 | 72/72 | 0ms · 73.0s total | 2,141 |
Lyra enzyme cofactor correction science-history-lyra-enzyme-correction Judge reasoningThe candidate recovers most core event facts, dates, quantities, and the later correction, and it anchors nearly all claims with exact substrings. It also preserves some attestation by marking Vos/Marr involvement and the bulletin credit dispute. However, it misses or weakly represents key reference expectations: the 1961 correction is present but not clearly framed as superseding the 0.5 mM value; the distinct attestation distinction between Vos’s 63% result and Marr’s 59% repetition is only partially modeled; and the priority/discovery credit conflict is not explicitly contrasted with the design credit in a way that preserves both sides as separate claims. Several claims are generic structural metadata rather than extracted content. No major fabricated causal claim is introduced, and the source’s unsupported structural-component statement is correctly treated as an absence. Overall coverage is moderate, temporal handling is good, attestation/contradiction/inference are weak, anchoring and hygiene are strong, and faithfulness is high. | Science history | 42.7/65 | 3/18 | 68 | 64/64 | 0ms · 55.2s total | 1,983 |
The Westmere transit timings science-history-observation-log Judge reasoningThe extraction is strong on anchored coverage of many source facts: it captures the document type/date, both observers, instruments, timings, measurements, editor footnote, and the Westmere renaming/proximity point. However, it largely fails the benchmark’s attestation/contradiction/inference expectations: it collapses disputed internal-contact timings into separate claims without explicitly preserving the conflict as such, does not model who said what in a way that distinguishes logs, letter, and editor footnote, and does not surface the low-confidence identity hypothesis between Crown Observatory and Royal Observatory at Westmere. Several claim shapes over-commit by turning reported observations into direct facts or by adding relation structure not demanded by the source, though I do not see major fabricated causal claims. Temporal capture is decent for the explicit dates and times present. Abundance is high because many extra well-formed anchored claims are present beyond the reference set. | Science history | 40/61 | 4/21 | 96 | 92/92 | 0ms · 79.4s total | 2,761 |
The Coronis priority dispute science-history-priority-dispute
Judge reasoningThe candidate recovers many source details exactly, especially entity labels, dates, and the main publication/priority dispute structure. However, it overextends in a few places: it collapses attestation into bare factual triples for several claims, and it includes many extra type/ontology assertions not directly required by the gold. It also does not preserve the key attestation/identity-hypothesis structure well: Philomath Batavus is linked to Aalbers only as a hypothesis in one claim, but several other assertions treat this as settled. Contradictory period claims are both represented, which is good, but the candidate does not explicitly model the dispute as conflicting attestations. Temporal anchoring is strong overall, with dates and intervals preserved at correct granularity. Evidence anchoring and hygiene are excellent, with mostly verbatim substrings and valid structured claims. Faithfulness is high overall: there is no clear theft or award fabrication, and the compiler’s note about no evidence of mutual awareness is included. Abundance is substantial but mostly consists of anchored background/typing claims rather than additional gold-relevant facts. | Science history | 54.5/67 | 6/22 | 84 | 82/82 | 0ms · 122.4s total | 2,539 |
Morrow Trench vent discovery log science-history-vent-discovery-log Judge reasoningThe candidate recovers many anchored details: expedition year, Rune's role, 09:06 and 13:20 timing, dark plume at frame 771, 2,640 m depth, Cor at 10:44 seeing shimmering water and 2,637 m, Pell's 18.2°C vs ambient 2.1°C, sample MT14-3, radio bulletin at 12:05, museum caption crediting Rune, Iver's later statement, sample sealing and transfer to O. Sen, and the Black Lantern/Morrow Vent likely-sameAs hypothesis plus the absence of plume-causing sampling. However, it misses some expected temporal items as explicit gold claims (e.g. the 13:20 playback claim is present, but the rubric target expects temporal capture across listed events and the candidate overfocuses on metadata), and it does not clearly preserve the attested-vs-reported structure for the radio bulletin and museum caption beyond basic edges. It also introduces a number of low-value structural claims about the document itself, but these are mostly anchored and not harmful. Contradiction preservation is partial: the radio credit vs museum caption credit tension is represented, but not as a balanced pair of mutually incompatible priority standards; still, the core conflict is present. Identity hypothesis is handled well with a probabilistic link. Overall faithfulness is strong, with no clear fabricated causation and correct treatment of the plume-creation denial, but coverage is incomplete relative to the full gold set. | Science history | 56.4/69 | 2/18 | 77 | 73/73 | 0ms · 78.2s total | 2,156 |
Gold = distinct gold claims matched. Anchored = anchors that are verbatim substrings of the source, over all anchored claims. Time prefers model generation time when reported; agent total includes orchestration between fetch and submit. Token counts are optional and do not affect scoring. Notes sample the hygiene and faithfulness issues the grader flagged.