Headline Result
Across 12 leading LLM configurations from Anthropic, Google, and OpenAI — tested in each provider’s consumer chat interface against an LLM-generated earnings summary with 51 known numerical errors — the best F1 score was 0.41; the median was 0.26. Rosey, evaluated on the same task under the same protocol, scored an F1 of 0.76.
About This Study
Generative AI has become the default first-draft tool for numerically dense professional documents — investment committee memos, earnings summaries, consulting deliverables, scientific manuscripts. But when the same models are asked to verify what they (or their peers) have drafted, how well do they do?
To find out, we took an LLM-generated first-draft Q2 FY2026 earnings summary for Cracker Barrel Old Country Store (CBRL) and had a human reviewer check every numerical claim against the company’s earnings release, investor presentation, and earnings call transcript. The result is a Ground Truth set of 51 verified numerical errors — 21 contradictions and 30 unsupported claims — out of 330 total numerical claims (a 15.5% error rate).
We then asked each of 12 leading LLM configurations — spanning three providers, three Model Tiers, and two Reasoning Tiers — to find those errors. Each ran three times, in each provider’s consumer chat interface, with no access to the Ground Truth. A human auditor then matched every finding to the Ground Truth to compute Precision, Recall, and F1.
We also ran Rosey, InSummary’s production document-verification system, on the same task under the same protocol, and adjudicated by the same process.
The findings are clear: today’s LLMs cannot reliably verify the documents they draft. A purpose-built verification system that is structurally separate from the generator can.
Explore the Study
Three companion artifacts: the full whitepaper (methodology, results, discussion), an interactive data viewer (per-model leaderboards, charts, and per-error coverage), and a claim viewer (every numerical claim in the test document with its Ground Truth verdict and supporting source quotes).
Replication Kit
In coordination with this report, we release everything required to reproduce the study end-to-end. Researchers, practitioners, and skeptics are invited to run, audit, and extend our findings. We intend to regenerate this benchmark periodically on new test documents to maintain its relevance as LLMs continue to progress.
Replication Kit (v1.0)
The downloadable archive contains the complete bundle of test materials, evaluation scaffolding, and Ground Truth needed to reproduce the study.