Artificial Analysis · XUpdated Original · English

We compared six hallucination checkers on the same deliverables from a fixed subset of 20 tasks and eight models. We ran our two-stage…

We compared six hallucination checkers on the same deliverables from a fixed subset of 20 tasks and eight models. We ran our two-stage check using GPT-6 Sol (high), GPT-6 Luna (high), Grok 4.7 (high), Claude Opus 5.5…

Image source · Artificial Analysis · X

We compared six hallucination checkers on the same deliverables from a fixed subset of 20 tasks and eight models. We ran our two-stage check using GPT-6 Sol (high), GPT-6 Luna (high), Grok 4.7 (high), Claude Opus 5.5 (high), Claude Sonnet 5.5 (high) and Gemini 3.8 Flash (high), and selected GPT-6 Sol (high) for production usage.

Across this 20-task subset, GPT-6 Sol and GPT-6 Luna generally identified more material hallucinations, while Claude Sonnet 5.5 and Gemini 3.8 Flash identified far fewer. Claude Opus 5.5 fell between Grok 4.7 and Claude Sonnet 5.5 in every model row. In total, Opus upheld 99 material hallucinations, compared with 219 for Grok, 57 for Sonnet and 470 for GPT-6 Sol. GPT-6 Sol identified more material hallucinations despite the check’s conservative approach to flagging errors, which excludes general legal knowledge not contained in the source documents.

All six checkers found no material hallucinations in GPT-6 Astra’s outputs, while GPT-6 Sol ranged from 0 to 0.20 per task. Both were among the models with the fewest material hallucinations under every checker. This comparison shows differences in checker behavior; the counts alone do not establish accuracy or rule out self-preference.

Original source

Artificial Analysis · X

Content notes

Original publication and rights belong to the source.