Every AI vendor claims accuracy. Almost none can prove it without asking you to trust their own grading. We wanted a harder test: an answer key written by someone with no stake in our product looking good.
Published court judgments turn out to be exactly that. In a construction delay case like Walter Lilly v Mackay [2012] EWHC 1773 (TCC), the judge spends hundreds of paragraphs establishing what happened and when: possession on 12 July 2004, a 23-week lead time flagged in November 2005, the instruction that came too late. That chronology is adjudicated ground truth - and nobody at datum wrote it.
The method has four steps. First, we reconstruct the project record strictly from the judgment's findings of fact: the letters, instructions, notices and minutes the court describes, re-created with the court's dates, parties and quoted words, every page marked as a reconstruction. Second, the engine ingests that record blind - OCR, classification, indexing, event extraction - exactly as it would a client's record. It never sees the judgment. Third, we score strictly: a court-established event counts as found only if the engine surfaced it on the right date with verbatim evidence carrying the operative phrase. Fourth, we publish - including the low scores, with their reasons.
The current board: eleven judgments, 153 documents, 100% document dating, 86% of court-established events found on the strictest matching. The hardest document shapes - NCR chains and claim letters recounting months of history - are stated on the page rather than hidden.
Two honest limits. Reconstruction corpora are real-shaped, not real: they rehearse the engine, they do not replace pilot validation on a genuinely real record with a practitioner's review. And a 12-event answer key means a single miss moves a score by eight points - which is why we publish the per-case detail, not just the average.
The full board, every number reproducible, is at datumclaims.com/benchmarks.