Executive Summary
This report presents a Precision–Recall–F1 (PRF) evaluation of large language model performance on accuracy error detection. The test document — an LLM-generated first draft of a Cracker Barrel Old Country Store (CBRL) Q2 FY2026 Earnings Summary — was evaluated against the earnings release, the investor presentation, and the earnings call transcript. A Ground Truth set of 51 verified numerical accuracy errors was independently curated through a human-in-the-loop process, in which every numerical claim in the test document was individually checked against the source materials. 13 model configurations — spanning 4 providers across three Model Tiers (Lite, Balanced, Max) and two Reasoning Tiers (Baseline, Extended) — were each independently prompted (3 times) to identify numerical contradictions and unsupported claims, with no access to the Ground Truth. Each model’s findings were then scored against the Ground Truth to produce per-model Precision, Recall, and F1 scores. Results are presented below.
| Ground Truth (GT) | The curated set of verified accuracy errors for the test document |
| True Positive (TP) | A finding that matches a known Ground Truth error |
| False Positive (FP) | A finding that does not match any Ground Truth error |
| False Negative (FN) | A Ground Truth error the model completely missed |
| Accuracy Finding | An error item raised by a model configuration. |
| Precision | TP ÷ (TP + FP) — what proportion of raised findings were correct. Expressed as a percentage; higher is better |
| Recall | TP ÷ (TP + FN) — what proportion of Ground Truth errors were found. Expressed as a percentage; higher is better |
| F1 | (2 × Precision × Recall) ÷ (Precision + Recall) — combined score, 0–1; higher is better |
| PRF | Precision · Recall · F1 — this evaluation framework |
| Number of Runs (n) | The number of independent evaluations conducted and averaged for each model configuration in this report. |
Methodology
| Model Tiers | Anthropic | OpenAI | Description | |
|---|---|---|---|---|
| Lite | Claude Haiku (4.5) | Gemini Flash-Lite (3) | GPT Instant (5.3) | Efficient, cost-optimized models |
| Balanced | Claude Sonnet (4.6) | Gemini Flash (3) | GPT Pro (5.4) | General-purpose performance models |
| Max | Claude Opus (4.6) | Gemini Pro (3) | GPT Thinking (5.4) | Frontier capability models |
| Reasoning Tiers | Anthropic | OpenAI | Description | |
|---|---|---|---|---|
| Baseline | Extended Off | Default | Standard | No extended reasoning enabled in the LLM UI |
| Extended | Extended On | — (no Extended mode) | Extended | Extended reasoning mode engaged in the LLM UI |
v016_20260427 · Ground Truth v003_20260422 · Rosey v0.3.2 · Generated 2026-04-27Full Leaderboard
| # | Source | Model | Provider | Model Tier | Reasoning Tier | n | Findings | TP | FP | FN | Prec. | Recall | F1 | Thinking (min) |
|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|
| 1 | Rosey | Rosey | Rosey | Rosey | Rosey | 3 | 39.7 | 34.3 | 5.3 | 16.7 | 87% | 67% | 0.76 | 5 |
| 2 | LLM UI | Claude Sonnet · Extended | Anthropic | Balanced | Extended | 2 | 16.0 | 14.5 | 1.5 | 36.5 | 87% | 28% | 0.41 | 9 |
| 3 | LLM UI | GPT Thinking · Baseline | OpenAI | Max | Baseline | 3 | 29.0 | 16.7 | 12.3 | 34.3 | 54% | 33% | 0.40 | 5 |
| 4 | LLM UI | Claude Opus · Baseline | Anthropic | Max | Baseline | 3 | 19.7 | 13.0 | 6.7 | 38.0 | 66% | 25% | 0.36 | 6 |
| 5 | LLM UI | GPT Instant · Extended | OpenAI | Lite | Extended | 2 | 33.0 | 14.5 | 18.5 | 36.5 | 47% | 28% | 0.34 | 6 |
| 6 | LLM UI | Gemini Flash · Baseline | Balanced | Baseline | 3 | 18.0 | 11.3 | 6.7 | 39.7 | 60% | 22% | 0.32 | 3 | |
| 7 | LLM UI | Claude Opus · Extended | Anthropic | Max | Extended | 2 | 21.5 | 11.0 | 10.5 | 40.0 | 53% | 22% | 0.31 | 9 |
| 8 | LLM UI | Claude Sonnet · Baseline | Anthropic | Balanced | Baseline | 3 | 12.7 | 6.7 | 6.0 | 44.3 | 53% | 13% | 0.21 | 8 |
| 9 | LLM UI | Gemini Pro · Baseline | Max | Baseline | 3 | 12.7 | 6.3 | 6.3 | 44.7 | 50% | 12% | 0.20 | 2 | |
| 10 | LLM UI | Claude Haiku · Extended | Anthropic | Lite | Extended | 3 | 14.7 | 5.0 | 9.7 | 46.0 | 33% | 10% | 0.15 | 3 |
| 11 | LLM UI | Gemini Flash-Lite · Baseline | Lite | Baseline | 3 | 11.7 | 2.0 | 9.7 | 49.0 | 16% | 4% | 0.06 | 1 | |
| 12 | LLM UI | Claude Haiku · Baseline | Anthropic | Lite | Baseline | 3 | 6.3 | 1.0 | 5.3 | 50.0 | 7% | 2% | 0.03 | 2 |
| 13 | LLM UI | GPT Instant · Baseline | OpenAI | Lite | Baseline | 3 | 6.3 | 0.7 | 5.7 | 50.3 | 12% | 1% | 0.02 | 1 |
Visual Leaderboard
Precision vs. Recall
Segmented by Provider
Ground-Truth Coverage
| GT id | Error Type | Claim Text | Paragraph ID | LLM UI | Rosey |
|---|---|---|---|---|---|
| GT001 | unsupported | 14 | 06407CB4 | 6 (18%) | 3 (100%) |
| GT002 | contradicted | (4.8%) | 19BE2C3D | 0 | 0 |
| GT003 | contradicted | $750.5M | 20E7A498 | 2 (6%) | 1 (33%) |
| GT004 | contradicted | 2.8–3.1% | 35DC6287 | 2 (6%) | 2 (67%) |
| GT005 | contradicted | $47.1–$47.1M | 5B50AF24 | 7 (21%) | 3 (100%) |
| GT006 | unsupported | 14 | 626FC3C9 | 0 | 1 (33%) |
| GT007 | contradicted | $634.8M | 63D74FB0 | 3 (9%) | 0 |
| GT008 | contradicted | $60.7M | 6E222835 | 17 (52%) | 1 (33%) |
| GT009 | contradicted | $198.9M | 76325244 | 4 (12%) | 3 (100%) |
| GT010 | contradicted | 160 bps | 77E64D6F | 7 (21%) | 3 (100%) |
| GT011 | contradicted | $1,416.9M | 51F1F12F | 8 (24%) | 3 (100%) |
| GT012 | contradicted | $1,329.1M | 6F88109B | 12 (36%) | 3 (100%) |
| GT013 | contradicted | $7.5M | 74FD6B9A | 7 (21%) | 3 (100%) |
| GT014 | contradicted | 60 bps | 7A8D872B | 0 | 3 (100%) |
| GT015 | unsupported | $135M–$150M | 588E444C | 15 (45%) | 3 (100%) |
| GT016 | unsupported | $135–150M | 593585CC | 6 (18%) | 3 (100%) |
| GT017 | unsupported | $150M–$190M | 2D4B1238 | 10 (30%) | 3 (100%) |
| GT018 | unsupported | $150M–$190M | 6E132EBB | 14 (42%) | 3 (100%) |
| GT019 | unsupported | $150–190M | 56D85840 | 5 (15%) | 3 (100%) |
| GT020 | unsupported | $3.35B–$3.45B | 31A5475F | 12 (36%) | 3 (100%) |
| GT021 | unsupported | $3.35B–$3.45B | 6E132EBB | 10 (30%) | 3 (100%) |
| GT022 | unsupported | $53.4M | 39E0A60E | 19 (58%) | 3 (100%) |
| GT023 | unsupported | $53.4M | 584B2398 | 6 (18%) | 3 (100%) |
| GT024 | unsupported | (4.7%) | 16D88495 | 10 (30%) | 3 (100%) |
| GT025 | unsupported | (4.7%) | 27000A02 | 10 (30%) | 3 (100%) |
| GT026 | unsupported | (4.7%) | 32F4C79B | 2 (6%) | 3 (100%) |
| GT027 | unsupported | (4.7%) | 63116C7E | 5 (15%) | 3 (100%) |
| GT028 | unsupported | (4.7%) | 741D979A | 2 (6%) | 3 (100%) |
| GT029 | unsupported | (8.5%) | 26CE3BD1 | 7 (21%) | 3 (100%) |
| GT030 | unsupported | (8.5%) | 27000A02 | 8 (24%) | 3 (100%) |
| GT031 | unsupported | (8.5%) | 2FCC4B20 | 2 (6%) | 3 (100%) |
| GT032 | contradicted | 110 bps | 77E64D6F | 1 (3%) | 0 |
| GT033 | unsupported | 14 | 0A30A9A3 | 5 (15%) | 1 (33%) |
| GT034 | unsupported | 14 | 228784B1 | 11 (33%) | 3 (100%) |
| GT035 | unsupported | 2.8x | 755AC14B | 12 (36%) | 3 (100%) |
| GT036 | contradicted | (160 bps) | 049BA354 | 3 (9%) | 0 |
| GT037 | contradicted | 60 bps | 11981458 | 1 (3%) | 0 |
| GT038 | unsupported | 2.5%–3.5% | 353040AD | 2 (6%) | 3 (100%) |
| GT039 | unsupported | 3.0%–4.0% | 53A57244 | 3 (9%) | 3 (100%) |
| GT040 | contradicted | $666.4M | 56190ABF | 0 | 0 |
| GT041 | contradicted | (48.8%) | 5C1C49F1 | 2 (6%) | 0 |
| GT042 | contradicted | (110 bps) | 69D56404 | 0 | 0 |
| GT043 | contradicted | $59.5M | 765BD7CF | 0 | 0 |
| GT044 | contradicted | 110 bps to 6.0% | 7A8D872B | 0 | 0 |
| GT045 | contradicted | 140 bps | 7A8D872B | 0 | 0 |
| GT046 | unsupported | 2 | 582035D5 | 4 (12%) | 2 (67%) |
| GT047 | unsupported | 1 | 3D7CE6B0 | 5 (15%) | 0 |
| GT048 | unsupported | twice | 593585CC | 0 | 3 (100%) |
| GT049 | unsupported | twice | 6E132EBB | 1 (3%) | 3 (100%) |
| GT050 | unsupported | twice | 15E5697D | 0 | 2 (67%) |
| GT051 | unsupported | two | 56D85840 | 0 | 0 |
Excluded by Thinking-Time Filter
| Model | Provider | Model Tier | Reasoning Tier | Excluded Runs | Thinking min (per run) |
|---|---|---|---|---|---|
| Claude Opus · Extended Partial | Anthropic | Max | Extended | 1/3 | 12 |
| Claude Sonnet · Extended Partial | Anthropic | Balanced | Extended | 1/3 | 12 |
| GPT Instant · Extended Partial | OpenAI | Lite | Extended | 1/3 | 16 |
| GPT Pro · Extended Fully excluded | OpenAI | Balanced | Extended | 3/3 | 50 / 38 / 34 |
| GPT Pro · Baseline Fully excluded | OpenAI | Balanced | Baseline | 3/3 | 45 / 36 / 22 |
| GPT Thinking · Extended Fully excluded | OpenAI | Max | Extended | 3/3 | 17 / 12 / 14 |
— none —