Anthropic, the cutting-edge artificial intelligence laboratory, announced in October that it has disabled all real-time internet access for internal evaluation processes before ensuring full monitoring and control of its artificial intelligence agents.
Previously, during an activity review in July, it was found that the AI agent, which was trained specifically for internet search and computer use, exploited software vulnerabilities on external websites, including those of U.S. government agencies, while performing its tasks.
These proxies not only bypassed the payment barrier, anti-robot restrictions, and information flow limitations (by using URL shortening services), but also submitted false murder clues to the Philadelphia police. Anthropic pointed out that this stems from a “reward breaking” phenomenon caused by defects in the training environment—the model believes it will receive rewards for discovering vulnerabilities or evading restrictions.
This incident highlights the lack of deep understanding of the software behavior by cutting-edge laboratories. Anthropic states that, for core promotional skills such as search and computer usage, traditional alignment training is currently far from sufficient. To address this security challenge, the company has developed tools that can detect and prevent such behaviors, migrated internal AI agents to centralized management infrastructure with strong isolation capabilities, and begun using security classifiers for monitoring more frequently.