LLM Document Accuracy Benchmark · Analysis Results

Precision · Recall · F1

Test Document: Cracker Barrel Old Country Store (CBRL) Q2 FY2026 Earnings Summary

4 Providers 13 Model Configurations 3 Runs per Configuration 51 Ground Truth Errors Filter: thinking < 12.0 min per run
01

Executive Summary

This report presents a Precision–Recall–F1 (PRF) evaluation of large language model performance on accuracy error detection. The test document — an LLM-generated first draft of a Cracker Barrel Old Country Store (CBRL) Q2 FY2026 Earnings Summary — was evaluated against the earnings release, the investor presentation, and the earnings call transcript. A Ground Truth set of 51 verified numerical accuracy errors was independently curated through a human-in-the-loop process, in which every numerical claim in the test document was individually checked against the source materials. 13 model configurations — spanning 4 providers across three Model Tiers (Lite, Balanced, Max) and two Reasoning Tiers (Baseline, Extended) — were each independently prompted (3 times) to identify numerical contradictions and unsupported claims, with no access to the Ground Truth. Each model’s findings were then scored against the Ground Truth to produce per-model Precision, Recall, and F1 scores. Results are presented below.

13
Model Configs
4 Providers
76%
Best F1
Rosey
87%
Best Precision
Rosey
67%
Best Recall
Rosey
51
Ground Truth Errors
LLM-Generated Document
2%
Worst F1
GPT Instant · Baseline
7%
Worst Precision
Claude Haiku · Baseline
1%
Worst Recall
GPT Instant · Baseline
🥇 Rank 1
Rosey
0.76
P 87% · R 67% · 5 min
🥈 Rank 2
Claude Sonnet · Extended
0.41
P 87% · R 28% · 9 min
🥉 Rank 3
GPT Thinking · Baseline
0.40
P 54% · R 33% · 5 min
Reference
Ground Truth (GT)The curated set of verified accuracy errors for the test document
True Positive (TP)A finding that matches a known Ground Truth error
False Positive (FP)A finding that does not match any Ground Truth error
False Negative (FN)A Ground Truth error the model completely missed
Accuracy FindingAn error item raised by a model configuration.
PrecisionTP ÷ (TP + FP) — what proportion of raised findings were correct. Expressed as a percentage; higher is better
RecallTP ÷ (TP + FN) — what proportion of Ground Truth errors were found. Expressed as a percentage; higher is better
F1(2 × Precision × Recall) ÷ (Precision + Recall) — combined score, 0–1; higher is better
PRFPrecision · Recall · F1 — this evaluation framework
Number of Runs (n)The number of independent evaluations conducted and averaged for each model configuration in this report.
02

Methodology

Test Document
Cracker Barrel Old Country Store (CBRL) Q2 FY2026 Earnings Summary — an LLM-generated first draft evaluated against 3 primary source documents: earnings release, investor presentation, and earnings call transcript.
Ground Truth
51 numerical accuracy errors (contradictions or unsubstantiated claims) independently verified human reviewed. Every numerical claim in the test document was individually checked against all source materials. Claims that directly conflicted with source data (incorrect values, miscalculated derived figures, directionality errors, internal inconsistencies) were marked as Contradicted Claims; assertions that could not be verified against any provided source material were flagged as Unsupported Claims.
Evaluation
Each model being evaluated received a single-pass structured prompt containing the complete test document and all source materials. Each configuration ran independently, with no access to the Ground Truth, to identify numerical contradicted and unsupported claims in the test document vis-à-vis the source materials, and to produce a structured, enumerated list of findings. Evaluations were run across two sources: LLM UI (browser chat) and Rosey (InSummary's production pipeline).
Matching
Each model's accuracy findings (flagged contradictions or unsupported claims) were compared pairwise against every Ground Truth error. A human reviewer evaluated each pairing to produce a True Positive (TP), False Positive (FP), or False Negative (FN) designation. A finding is designated a True Positive (TP) for a given Ground Truth error if it substantively identifies the same underlying accuracy issue — semantic equivalence is required; exact wording is not. The enumerated output format described in Evaluation above enforces a strict pairing: a single finding matches at most one Ground Truth error, and a single Ground Truth error matches at most one finding within a given run.
LLM UI Model Tiers
Model TiersAnthropicGoogleOpenAIDescription
LiteClaude Haiku (4.5)Gemini Flash-Lite (3)GPT Instant (5.3)Efficient, cost-optimized models
BalancedClaude Sonnet (4.6)Gemini Flash (3)GPT Pro (5.4)General-purpose performance models
MaxClaude Opus (4.6)Gemini Pro (3)GPT Thinking (5.4)Frontier capability models
LLM UI Reasoning Tiers
Reasoning TiersAnthropicGoogleOpenAIDescription
BaselineExtended OffDefaultStandardNo extended reasoning enabled in the LLM UI
ExtendedExtended On— (no Extended mode)ExtendedExtended reasoning mode engaged in the LLM UI
Report Metadata
Report v016_20260427 · Ground Truth v003_20260422 · Rosey v0.3.2 · Generated 2026-04-27
03

Full Leaderboard

# Source Model Provider Model Tier Reasoning Tier n Findings TP FP FN Prec. Recall F1 Thinking (min)
1RoseyRoseyRoseyRoseyRosey339.734.35.316.787%67%0.765
2LLM UIClaude Sonnet · ExtendedAnthropicBalancedExtended216.014.51.536.587%28%0.419
3LLM UIGPT Thinking · BaselineOpenAIMaxBaseline329.016.712.334.354%33%0.405
4LLM UIClaude Opus · BaselineAnthropicMaxBaseline319.713.06.738.066%25%0.366
5LLM UIGPT Instant · ExtendedOpenAILiteExtended233.014.518.536.547%28%0.346
6LLM UIGemini Flash · BaselineGoogleBalancedBaseline318.011.36.739.760%22%0.323
7LLM UIClaude Opus · ExtendedAnthropicMaxExtended221.511.010.540.053%22%0.319
8LLM UIClaude Sonnet · BaselineAnthropicBalancedBaseline312.76.76.044.353%13%0.218
9LLM UIGemini Pro · BaselineGoogleMaxBaseline312.76.36.344.750%12%0.202
10LLM UIClaude Haiku · ExtendedAnthropicLiteExtended314.75.09.746.033%10%0.153
11LLM UIGemini Flash-Lite · BaselineGoogleLiteBaseline311.72.09.749.016%4%0.061
12LLM UIClaude Haiku · BaselineAnthropicLiteBaseline36.31.05.350.07%2%0.032
13LLM UIGPT Instant · BaselineOpenAILiteBaseline36.30.75.750.312%1%0.021
04

Visual Leaderboard

05

Precision vs. Recall

06

Segmented by Provider

08

Ground-Truth Coverage

LLM UI12 cfg · 33 runs
% run-catches
Rosey1 cfg · 3 runs
% run-catches
GT Segment Table · catches per source
GT id Error Type Claim Text Paragraph ID LLM UIRosey
GT001unsupported1406407CB46 (18%)3 (100%)
GT002contradicted(4.8%)19BE2C3D00
GT003contradicted$750.5M20E7A4982 (6%)1 (33%)
GT004contradicted2.8–3.1%35DC62872 (6%)2 (67%)
GT005contradicted$47.1–$47.1M5B50AF247 (21%)3 (100%)
GT006unsupported14626FC3C901 (33%)
GT007contradicted$634.8M63D74FB03 (9%)0
GT008contradicted$60.7M6E22283517 (52%)1 (33%)
GT009contradicted$198.9M763252444 (12%)3 (100%)
GT010contradicted160 bps77E64D6F7 (21%)3 (100%)
GT011contradicted$1,416.9M51F1F12F8 (24%)3 (100%)
GT012contradicted$1,329.1M6F88109B12 (36%)3 (100%)
GT013contradicted$7.5M74FD6B9A7 (21%)3 (100%)
GT014contradicted60 bps7A8D872B03 (100%)
GT015unsupported$135M–$150M588E444C15 (45%)3 (100%)
GT016unsupported$135–150M593585CC6 (18%)3 (100%)
GT017unsupported$150M–$190M2D4B123810 (30%)3 (100%)
GT018unsupported$150M–$190M6E132EBB14 (42%)3 (100%)
GT019unsupported$150–190M56D858405 (15%)3 (100%)
GT020unsupported$3.35B–$3.45B31A5475F12 (36%)3 (100%)
GT021unsupported$3.35B–$3.45B6E132EBB10 (30%)3 (100%)
GT022unsupported$53.4M39E0A60E19 (58%)3 (100%)
GT023unsupported$53.4M584B23986 (18%)3 (100%)
GT024unsupported(4.7%)16D8849510 (30%)3 (100%)
GT025unsupported(4.7%)27000A0210 (30%)3 (100%)
GT026unsupported(4.7%)32F4C79B2 (6%)3 (100%)
GT027unsupported(4.7%)63116C7E5 (15%)3 (100%)
GT028unsupported(4.7%)741D979A2 (6%)3 (100%)
GT029unsupported(8.5%)26CE3BD17 (21%)3 (100%)
GT030unsupported(8.5%)27000A028 (24%)3 (100%)
GT031unsupported(8.5%)2FCC4B202 (6%)3 (100%)
GT032contradicted110 bps77E64D6F1 (3%)0
GT033unsupported140A30A9A35 (15%)1 (33%)
GT034unsupported14228784B111 (33%)3 (100%)
GT035unsupported2.8x755AC14B12 (36%)3 (100%)
GT036contradicted(160 bps)049BA3543 (9%)0
GT037contradicted60 bps119814581 (3%)0
GT038unsupported2.5%–3.5%353040AD2 (6%)3 (100%)
GT039unsupported3.0%–4.0%53A572443 (9%)3 (100%)
GT040contradicted$666.4M56190ABF00
GT041contradicted(48.8%)5C1C49F12 (6%)0
GT042contradicted(110 bps)69D5640400
GT043contradicted$59.5M765BD7CF00
GT044contradicted110 bps to 6.0%7A8D872B00
GT045contradicted140 bps7A8D872B00
GT046unsupported2582035D54 (12%)2 (67%)
GT047unsupported13D7CE6B05 (15%)0
GT048unsupportedtwice593585CC03 (100%)
GT049unsupportedtwice6E132EBB1 (3%)3 (100%)
GT050unsupportedtwice15E5697D02 (67%)
GT051unsupportedtwo56D8584000
09

Excluded by Thinking-Time Filter

Filter: thinking < 12.0 min per run
LLM UI · 3 configs fully excluded · 12 runs dropped
ModelProviderModel TierReasoning TierExcluded RunsThinking min (per run)
Claude Opus · Extended PartialAnthropicMaxExtended1/312
Claude Sonnet · Extended PartialAnthropicBalancedExtended1/312
GPT Instant · Extended PartialOpenAILiteExtended1/316
GPT Pro · Extended Fully excludedOpenAIBalancedExtended3/350 / 38 / 34
GPT Pro · Baseline Fully excludedOpenAIBalancedBaseline3/345 / 36 / 22
GPT Thinking · Extended Fully excludedOpenAIMaxExtended3/317 / 12 / 14
Rosey · 0 configs fully excluded · 0 runs dropped

— none —