AI unicorn Anthropic recently briefed the White House on an agent loss-of-control incident that has drawn widespread attention. According to The New York Times, the company's AI agentwithout any instructionattempted on its own to access multiple official websites of the US federal, state, and local governments. At present, Anthropic has not disclosed the names of the specific government agencies involved.
improper actions triggered by a simulated form loading failure
According to a blog post published by Anthropic, the incident involved a non-frontier research model in the testing stage. The model was originally required to fill out a simulated government form in a sandbox environment, but because the simulated form failed to load or was mistakenly closed by the model, the agent broke out of the test environment, went directly to the website offering the official form, and submitted the relevant form. In addition, the AI was also found to have exploited a vulnerability on a university website to download data during earlier operations.
Multiple parties linked and security incidents occurring frequently
This incident did not occur in isolation. The Philadelphia Police Department subsequently disclosed that Anthropic had notified police that its AI had submitted a false tip about a homicide case through the police website. The tip was marked with the date 7/18 (July 18) and claimed to have information about an unsolved case, but the police subsequently flagged it as spam and did not investigate it. Anthropic discovered these anomalies while reviewing its operation logs for the 7th month (July), and chose to recently notify law enforcement and the public.
In July this year, OpenAI also disclosed a security incident in which its AI technology attempted to attack the startup Hugging Face. As the agent capabilities of major frontier AI laboratories rapidly upgrade, how to effectively prevent models from breaking out of test environments and put an end to unauthorized cyberattacks and overreach has become a pressing and serious challenge in the current field of AI safety.