Announcing Harvey LAB-AA v1.1: adding hallucination checks to raise the bar for agentic legal work
We collaborated with Harvey on Harvey LAB-AA v1.1, which updates our scoring methodology for the Legal Agent Benchmark (LAB) for AI agents doing real-world, agentic, legal work. Every deliverable is now checked for hallucinations against its source documents, a three-judge panel grades every rubric criterion, and the new headline metric, Hallucination-Gated All-Pass Rate, only credits a task when the deliverables satisfy every rubric criterion and contain no material hallucinations.
Every model works through 120 private legal tasks built by the team at Harvey, from corporate M&A and capital markets to tax, litigation and bankruptcy. This is the first step in enhancing the methodology for Harvey LAB-AA. In future updates, we're working with Harvey to better account for the full set of factors lawyers value, including usability features like style and tone.
Score
Harvey LAB-AA v1.1: Hallucination-Gated All-Pass Rate
Average over tasks of the share of the judge panel finding every rubric criterion passed, with a task zeroed on any material hallucination · Independently benchmarked by Artificial AnalysisGrok 4.7 (xhigh) leads Harvey LAB-AA v1.1 with a 9.4% Hallucination-Gated All-Pass Rate, narrowly ahead of Muse Spark 1.3 (max) at 8.9% and GPT-6 Astra (max) at 8.6%. GPT-6.1 Sol (max) follows at 6.9%, then Claude Fable 5.1 (max, with fallback) at 6.4%, Kimi K3 (max) at 5.3% and Claude Opus 5.5 (max, with fallback) at 4.2%. The Claude models ran with Anthropic's fallback enabled, but never used it.
Accounting for hallucinations changes the rankings on Harvey LAB-AA. Without the hallucination gate, Muse Spark 1.3 would lead clearly with a 26.7% All-Pass Rate, but two thirds of those passes contain a material hallucination. GPT-6 Astra keeps almost all of its passes (8.9% to 8.6%) and moves from joint 10th on all-pass to 3rd on Hallucination-Gated All-Pass Rate. GPT-6.1 Sol falls relatively little, from 7.5% to 6.9%. Across the models tested, more than 60% of otherwise passing results contain a material hallucination, and several models score 0% after accounting for hallucinations.
Most models complete the large majority of rubric criteria: 16 models pass 85.6-96.0% of criteria. The rubrics assess each deliverable holistically, rather than only through criteria selected for difficulty. We report the Criterion Pass Rate before the hallucination gate alongside material hallucinations per task to show rubric coverage and grounding separately.
For up-to-date results see the Harvey LAB-AA evaluation page. This article shows data as at 8 October 2026.
Changes in Harvey LAB-AA v1.1
We collaborated with Harvey to update Harvey LAB-AA to v1.1. The update changes how the benchmark is graded and scored, so v1.1 results are not directly comparable to previously published v1.0 numbers:
- Hallucination auditing. Every deliverable a model submits is now audited against the task's source documents in a two-stage check; a task with no usable submission scores zero and is not audited. A first judge flags candidate hallucinations across three categories - contradictions of the source materials, fabricated source content, and specific assertions with no support in the sources - and a second pass by the same judge re-checks each flag against that task's source documents, dismissing flags that do not hold up and classifying the rest as material or minor.
- Hallucination-Gated All-Pass Rate headline. The headline metric is now the Hallucination-Gated All-Pass Rate: a task counts only when the deliverables satisfy every rubric criterion and contain no material hallucination.
- Criterion Pass Rate and material hallucinations. We report the share of rubric criteria passed before the hallucination gate, averaged across the three judges and pooled over all criteria, alongside the mean number of material hallucinations per checked task. The scatter compares these two metrics directly.
- Three-judge rubric panel. Rubric criteria are now graded by a panel of three LLM judges - GPT-6 Sol, Grok 4.7, and Claude Opus 5.5 - with verdicts averaged across the panel, replacing the single judge used in v1.0. We reviewed these judges and found minimal self-preference for their own model families, and averaging across all three mitigates any biases that emerge.
- Dataset update. v1.1 uses the latest private dataset from Harvey (v1.1.0), which includes improvements to the tasks and criteria.
How Harvey LAB-AA differs from Harvey's LAB
Harvey LAB-AA is our independent reimplementation of Harvey's evaluation, and there are several key differences to the original version:
- Models are run on our Stirrup agent harness, enabling features such as context compaction rather than failure when reaching context limits, with simplified Artificial Analysis-authored agent and judge prompts
- We do not include Harvey's custom tools and document-generation skill scripts (e.g. pptx, docx), instead providing a simple code execution tool to reflect raw model ability
- Deliverables must match the exact filename specified, rather than fuzzy matching when models produce incorrect filenames
Hallucinations
In human preference studies run by Harvey, human experts flagged hallucinations as a primary determining factor in which response they prefer out of two otherwise comprehensive answers. Our hallucination check focuses on errors that could materially affect real legal work, and takes a conservative approach to flagging them. It distinguishes unsupported task-specific claims from general legal knowledge and places the latter out of scope, so models are not penalized for drawing on case law or statutes beyond the source files.
Every deliverable a model submits is audited for hallucinations against the task's source documents. A task with no usable submission scores zero and is not audited nor used in the hallucinations per task calculation. A hallucination is a claim in the work product that the sources do not support, and each falls into one of three categories:
- A contradiction of the source materials.
- Fabricated source content.
- A specific assertion with no support in the record.
Each hallucination is graded material or minor. A material hallucination would mislead a reader on a substantive point, such as a wrong contractually required date. A minor hallucination is a real error that is unlikely to meaningfully affect the legal interpretation of a deliverable. Only material hallucinations affect scoring: having one or more zeroes a task's score in the Hallucination-Gated All-Pass Rate. Criterion Pass Rate measures rubric coverage before the hallucination gate. Material and minor hallucination counts are reported per task the check ran on; minor hallucinations do not affect scores.
How the hallucination check works
Every deliverable is checked against the task's source documents in three steps.
Find possible hallucinations
The judge compares the submission with the task sources and flags statements for review.
Tax compliance review · submission excerpts3 examplesTask source · draft partnership return excerpt
Scheduled principal amortization of $6,000,000 was made during 2023, reducing the outstanding balance from $220,000,000 (beginning of year, inclusive of $6,000,000 classified as current) to $214,000,000 (end of year, inclusive of $6,000,000 classified as current in Line 15).
The submission · discussion list
“debt schedule $220M term loan + $214M notes”
Possible double counting of debtThe same submission · exposure schedule
“$6,400,000 *293/365 $5,136,986”
Possible error in the prorated calculationSeparate submission, same task · issues memo
“The $960,000 overpayment was credited to 2024”
Says the $960,000 was already credited, but the source asks for confirmation before filing.The rubric and hallucination check assess different aspects of the submission.
1 material example1 minor · 1 not upheldIllustrative scoring example
All-Pass RateShare of judges that pass every criterion67%67%2 of 3 judges passed every criterion
Criterion Pass RateShare of judge verdicts that are passes90%90%27 of 30 judge verdicts were passes; unaffected by hallucinations
The GPT-6 models hallucinate least: GPT-6 Astra averages 0.03 material hallucinations per task (4 hallucinations across all 120 tasks) and GPT-6 Sol 0.07 (8 hallucinations across all 120 tasks), while Gemini 3.8 Flash averages the most in the launch set, at 13.96 per task. Completing the criteria and not hallucinating are different skills: Muse Spark 1.3 passes the most criteria (96.0%) but averages 1.68 material hallucinations per task, against 0.03 for GPT-6 Astra. Every open-weights model averages at least 2.09 material hallucinations per task. A model's hallucination rate depends far more on the model than on the practice area.
Harvey LAB-AA v1.1: Hallucinations per Task
Upheld hallucination flags per task, split by severity · Material flags would mislead a reader on a substantive point, minor flags are real errors unlikely to affect the legal interpretation, and only material flags affect the headline score · Lower is betterAverage number of upheld hallucination flags per checked Harvey LAB-AA task, split by severity. A flag is material when it would mislead a reader on a substantive point - a wrong party, amount, date or obligation, or a fact that changes a conclusion - and minor when it is a real error unlikely to meaningfully affect the legal interpretation of a deliverable. Only material flags gate the headline; minor flags are reported but never scored.
How we chose the hallucination checker
We use GPT-6 Sol (high) for both passes of the hallucination check, separately from the three-judge rubric panel. We compared six candidates for the hallucination judge: GPT-6 Sol, GPT-6 Luna, Grok 4.7, Claude Opus 5.5, Claude Sonnet 5.5 and Gemini 3.8 Flash, each at high reasoning effort, on the same 20 tasks and deliverables from eight evaluated models.
Across this subset of tasks, GPT-6 Sol and GPT-6 Luna generally identified more material hallucinations, while Claude Sonnet 5.5 and Gemini 3.8 Flash identified far fewer. Claude Opus 5.5 fell between Grok 4.7 and Claude Sonnet 5.5 in every model row. In total, Opus upheld 99 material hallucinations, compared with 219 for Grok, 57 for Sonnet and 470 for GPT-6 Sol. GPT-6 Sol identified more material hallucinations despite the check's conservative approach to flagging errors, which excludes general legal knowledge not contained in the source documents.
All six checkers found no material hallucinations in GPT-6 Astra's outputs, while GPT-6 Sol ranged from 0 to 0.20 per task. Both were among the models with the fewest material hallucinations under every checker.
Harvey LAB-AA v1.1: Hallucination Judge Model Comparison
Mean upheld material hallucinations per task · Subset of 20 tasks and 8 models · Lower is betterMean material hallucinations per task0481217Harvey LAB-AA v1.1: Hallucination Judge Model Comparison. Mean upheld material hallucinations per task · Subset of 20 tasks and 8 models · Lower is better.| MODEL BEING CHECKED | HALLUCINATION JUDGE | |||||
|---|---|---|---|---|---|---|
| GPT-6 Sol(high) | GPT-6 Luna(high) | Grok 4.7(high) | Claude Opus 5.5(high) | Claude Sonnet 5.5(high) | Gemini 3.8 Flash(high) | |
| GPT-6 Astra (max) | ||||||
| GPT-6 Sol (max) | ||||||
| Grok 4.7 (xhigh) | ||||||
| Claude Opus 5.5 (max with fallback) | ||||||
| Claude Fable 5.1 (max with fallback) | ||||||
| Muse Spark 1.3 (max) | ||||||
| Kimi K3 (max) | ||||||
| Gemini 3.8 Flash (high) | ||||||
| Total material hallucinations per task | 23.50 | 19.35 | 10.95 | 4.95 | 2.85 | 1.40 |
Hallucination-Gated Near-Pass Rate
Allowing near misses in terms of criteria puts GPT-6 Astra and GPT-6.1 Sol ahead. GPT-6 Astra rises from an 8.6% Hallucination-Gated All-Pass Rate to a 20.3% Hallucination-Gated Near-Pass Rate with one missed criterion and 31.7% with two missed criteria, and GPT-6.1 Sol from 6.9% to 20.3% and 29.6%, against 15.3% and 23.1% for Grok 4.7. The next biggest gains go to GPT-6 Sol (3.6% to 18.1%) and Claude Sonnet 5.5 (2.8% to 16.9%). Models whose passes are wiped out by hallucinations gain little, since a material hallucination still zeroes the task in every band: GLM-5.3 reaches only 2.2% even with two misses and Gemini 3.8 Flash stays at 0%.Harvey LAB-AA v1.1: Hallucination-Gated Near-Pass Rate
Average over tasks of the share of the judge panel that missed 0, <=1 and <=2 criteria · Tasks with at least 1 material hallucination count at 0% · Higher is betterSort by criteria missedCost
Top-scoring models aren't the most expensive. Grok 4.7 leads at ~$9.50 per task, under half the cost of Claude Fable 5.1 (~$21.70), the most expensive model. Muse Spark 1.3 represents strong performance for its cost, landing in second for ~$4.20 per task.Harvey LAB-AA v1.1: Cost per Task
Average cost per task (USD), broken down by input, cache hit, cache write, reasoning, and answer tokensAverage cost per task in the evaluation. Costs are split by input, cache hit, cache write, reasoning, and answer token pricing where canonical token counts are available.
Token Usage
Generating more output tokens does not necessarily translate to a higher score. GPT-6 Astra scores 8.6% on ~81k output tokens per task, under half of Grok 4.7's ~180k, while the three Claude models generate the most output tokens (~202k to ~562k per task) and score 2.8% to 6.4%.Harvey LAB-AA v1.1: Output Tokens per Task
Output tokens used to run one task, broken down by reasoning and answer tokensThe average number of answer and reasoning tokens produced per benchmark task in this evaluation.
Speed
Stronger models tend to take longer. Estimated decode time is ~33 minutes per task for both Grok 4.7 and Claude Fable 5.1. Muse Spark 1.3 scores second at ~12 minutes per task. These estimates exclude time to first token and other overhead.Harvey LAB-AA v1.1: Time per Task
Weighted average decode time (minutes) per task; excludes TTFT and overhead time · Lower is betterThe weighted average estimated decode time (minutes) per task. Output tokens per task are divided by output speed and converted from seconds to minutes. This excludes time to first token and other overhead.
Turns
The top scorers don't run the longest loops. Grok 4.7 averages ~63 turns per task and GPT-6 Astra ~54. Claude Sonnet 5.5 runs the longest, ~179 turns, and scores 2.8%.Harvey LAB-AA v1.1: Average Turns per Task
Average number of model turns per Harvey LAB-AA v1.1 task · Lower is betterThis chart shows the average number of turns the agent takes per task. It is a rough proxy for how many actions, tool calls, and iteration cycles an agent is using to complete benchmark tasks.
Model Size (Open Weights Models Only)
Harvey LAB-AA v1.1: Hallucination-Gated All-Pass Rate vs. Compute Proxy
Harvey LAB-AA v1.1 Hallucination-Gated All-Pass Rate · Compute proxyMost attractive quadrantPareto lineA relative index of the compute used, calculated as Active Parameters (billions) × (Input Tokens / 5 + Output Tokens) / 1,000,000. Input tokens are downweighted to reflect the lower compute cost of prefill vs. generation. Lower values indicate more compute-efficient models. Models in the top-left of the chart achieve high scores with less compute.
Score vs. Release Date
Harvey LAB-AA v1.1: Hallucination-Gated All-Pass Rate vs. Release Date
Most attractive regionExample Tasks & Submissions
View task set on GitHubBrowse representative Harvey LAB tasks from the public task set, the reference files each model was given, and the deliverables it produced.
Mergers & AcquisitionsInstructions
Review the attached acquisition data room contracts and internal memo for change of control and assignment provisions, and prepare a comprehensive deal team report.
Output: coc-analysis-report.docx
Deliverables
Expected outputs the model must produce
- coc-analysis-report.docxA comprehensive deal team report analyzing change of control and assignment provisions across the target’s material contracts.
Reference files
Provided to the model
OpenOpenOpenOpenOpenOpenOpenOpenOpenOpenOpenOpenOpenOpenOpenOpenOpenOpenOpenModel submissions
Deliverables produced by each model
Claude Opus 5.5 (max with fallback) - coc-analysis-report.docxOpenHarvey LAB-AA resources
- The leaderboard and full results live on the Harvey LAB-AA evaluation page, updated as new models are released
- The methodology page documents the full implementation, including the agent and rubric grading prompts
- Harvey's original LAB announcement introduces the benchmark and its design
- A public set of representative tasks is available on GitHub
- Harvey LAB-AA runs on Stirrup, our open-source agent framework
Read the latest

Anthropic has released Claude Haiku 5.5
Claude Haiku 5.5 scores 43 on the Artificial Analysis Intelligence Index, up 26 points one year after the last Haiku release
October 7, 2026

Mistral has released Mistral Large 4, making France home to the most intelligent model outside the US and China
Mistral has released Mistral Large 4, scoring 38 on the Artificial Analysis Intelligence Index; France is back to having the most intelligent model from outside the US and China
October 6, 2026
Korean AI Lab Upstage releases Solar Mini 4
Korean AI Lab Upstage has released Solar Mini 4 which scores 24 on the Artificial Analysis Intelligence Index, but costs ~5x as much per task as GPT-6 Luna (max) despite similar per-token prices
September 30, 2026
