Arena · XOriginal · English

We also looked into False Attribution, where the agent attributes a statement, request, choice, approval, or fact to the user, but…

Image source · Arena · X

We also looked into False Attribution, where the agent attributes a statement, request, choice, approval, or fact to the user, but user-provided evidence contradicts that attribution.

Interestingly we saw some models misquoting the user (misstating what a user asked for), while others credit the user with someone else's work.

There were varied patterns across models, with GPT-6 Luna and Astra rarely misquoting (15.6% and 28.6% respectively) but often misattributing (53.1% and 48.2%). Interestingly their sibling model GPT-6 Sol has the highest rate of misstaging the user’s history (23.5%).

Original source

Arena · X

Content notes

Original publication and rights belong to the source.