Artificial Analysis · XOriginal · English

In human preference studies run by Harvey, human experts flagged hallucinations as a primary determining factor in which response they…

In human preference studies run by Harvey, human experts flagged hallucinations as a primary determining factor in which response they prefer out of two otherwise comprehensive answers. Our new hallucination grading…

In human preference studies run by Harvey, human experts flagged hallucinations as a primary determining factor in which response they prefer out of two otherwise comprehensive answers. Our new hallucination grading pipeline focuses on errors that could materially affect real legal work, while taking a conservative approach to flagging them. It distinguishes unsupported task-specific claims from general legal knowledge and places the latter out of scope, so models are not penalized simply for drawing on case law or statutes beyond the source files.

The check has two stages, both using GPT-6 Sol (high). First, it compares every deliverable with the task’s source documents. A task with no usable submission scores zero and is not audited or included in the hallucinations-per-task calculation. A hallucination is a claim in the work product that the sources do not support, and each falls into one of three categories:

- A contradiction of the source materials.
- Fabricated source content.
- A specific assertion with no support in the record.

It then re-checks each flag against that task’s source documents, dismisses flags that do not hold up, and classifies confirmed hallucinations as material or minor. A material hallucination is one that would mislead a reader on a substantive point. One material hallucination sets the task’s contribution to Hallucination-Gated All-Pass Rate to zero; minor hallucinations are reported separately and do not affect the score.

Rubric criteria are now graded by a three-judge panel of GPT-6 Sol, Grok 4.7 and Claude Opus 5.5, replacing v1.0's single judge. We are also running v1.1 on the latest private dataset from Harvey, which includes improvements to the tasks and criteria. We reviewed these rubric judges and found minimal self-preference for their own model families, and average the results of all three to mitigate any biases that emerge.

Original source

Artificial Analysis · X

Content notes

Original publication and rights belong to the source.