A Rosey Research Report · April 2026

Benchmarking Large Language Models
on Document Accuracy Verification

A Precision–Recall–F1 Evaluation Using an LLM-Generated
Cracker Barrel Old Country Store Quarterly Earnings Summary

01

Headline Result

The Verification Gap
No frontier LLM tested could reliably verify a document another LLM had drafted. Rosey, a purpose-built verification system, closes the gap.

Across 12 leading LLM configurations from Anthropic, Google, and OpenAI — tested in each provider’s consumer chat interface against an LLM-generated earnings summary with 51 known numerical errors — the best F1 score was 0.41; the median was 0.26. Rosey, evaluated on the same task under the same protocol, scored an F1 of 0.76.

🥇 Rank 1 · Best Overall (F1, P, R)
Rosey by InSummary
0.76
P 87% · R 67% · ~5 min/run
🥈 Rank 2 · Best LLM (F1, P)
Claude Sonnet · Extended Anthropic, Balanced
0.41
P 87% · R 28% · ~9 min/run
🥉 Rank 3 · Best LLM (R)
GPT Thinking · Baseline OpenAI, Max
0.40
P 54% · R 33% · ~5 min/run
02

About This Study

Generative AI has become the default first-draft tool for numerically dense professional documents — investment committee memos, earnings summaries, consulting deliverables, scientific manuscripts. But when the same models are asked to verify what they (or their peers) have drafted, how well do they do?

To find out, we took an LLM-generated first-draft Q2 FY2026 earnings summary for Cracker Barrel Old Country Store (CBRL) and had a human reviewer check every numerical claim against the company’s earnings release, investor presentation, and earnings call transcript. The result is a Ground Truth set of 51 verified numerical errors — 21 contradictions and 30 unsupported claims — out of 330 total numerical claims (a 15.5% error rate).

We then asked each of 12 leading LLM configurations — spanning three providers, three Model Tiers, and two Reasoning Tiers — to find those errors. Each ran three times, in each provider’s consumer chat interface, with no access to the Ground Truth. A human auditor then matched every finding to the Ground Truth to compute Precision, Recall, and F1.

We also ran Rosey, InSummary’s production document-verification system, on the same task under the same protocol, and adjudicated by the same process.

The findings are clear: today’s LLMs cannot reliably verify the documents they draft. A purpose-built verification system that is structurally separate from the generator can.

0.76
Rosey F1
P 87% · R 67%
0.41
Best LLM F1
Claude Sonnet, Extended
0.26
Median LLM F1
12 included configurations
5 / 12
LLMs below F1 0.20
Most errors missed; majority of findings false
51
Ground Truth Errors
21 contradictions · 30 unsupported
330
Total Numerical Claims
15.5% baseline error rate
03

Explore the Study

Three companion artifacts: the full whitepaper (methodology, results, discussion), an interactive data viewer (per-model leaderboards, charts, and per-error coverage), and a claim viewer (every numerical claim in the test document with its Ground Truth verdict and supporting source quotes).

04

Replication Kit

In coordination with this report, we release everything required to reproduce the study end-to-end. Researchers, practitioners, and skeptics are invited to run, audit, and extend our findings. We intend to regenerate this benchmark periodically on new test documents to maintain its relevance as LLMs continue to progress.

Replication Kit (v1.0)

The downloadable archive contains the complete bundle of test materials, evaluation scaffolding, and Ground Truth needed to reproduce the study.

CBRL_Q2_FY2026_EA.pdf CBRL_Q2_FY2026_IP.pdf CBRL_Q2_FY2026_Transcript.pdf evaluation_prompt_ui.md generation_prompt.md test_document.docx VMH_Q4_2024_Earnings_Summary.docx
Download .zip 3.9 MB
05

Cite This Work

Suggested Citation
InSummary, Inc. (2026). Benchmarking Large Language Models on Document Accuracy Verification: A Precision–Recall–F1 Evaluation. Rosey Research Report v1.0. Retrieved from insummary.com/prf-cbrl.