Claude AI Security Tests Go Wrong as Anthropic Discovers Accidental Access to Real Computer Systems
Anthropic revealed that Claude AI accidentally accessed real company systems during cybersecurity testing due to a configuration error. The incident has raised concerns about AI cybersecurity, testing environments, and AI safety standards.
Anthropic, an AI company, has revealed that its Claude AI models accidentally accessed three organizations' real computer systems while doing internal cybersecurity testing. According to the company, the incidents occurred as a result of a misconfigured testing environment, rather than the AI attempting to escape or conduct attacks on its own.
Anthropic discovered the flaw after reviewing more than 141,000 cybersecurity evaluation runs. The review was initiated after OpenAI recently disclosed a similar event involving one of their AI models. Anthropic discovered that three Claude models Opus 4.7, Mythos 5, and an internal research model were accidentally connected to the internet during security testing that were supposed to run in a completely isolated setting.
The AI models were taking part in cybersecurity exercises known as capture-the-flag, in which they were instructed to attack imaginary computer systems in search of hidden information. However, due to a configuration error, the test environment gained internet connection. As a result, the AI models mistook the exercise for something involving actual websites and servers.
In one incident, Claude gained access to a real company's systems by utilizing weak passwords and exposing web services. In another case, the AI published a malicious Python package to the public PyPI software repository, believing it had completed the challenge. The package remained online for approximately an hour before being removed. A third internal research model looked into internet-connected systems and got access to one organization before discontinuing the activity when it realized it had reached a legitimate target.
Following the accidents, Anthropic suspended all cybersecurity evaluations, notified the affected organizations, and started working with independent AI safety researchers to enhance their testing procedure. The organization intends to tighten monitoring, improve isolated testing facilities, and implement tougher measures to avoid similar situations in the future.
The most recent disclosure emphasizes the growing relevance of AI cybersecurity and demonstrates that secure testing settings are equally crucial as developing powerful AI models. As AI systems become more powerful, businesses are expected to implement stricter safety protocols before using them in real-world settings.
Information referenced in this article is from The Indian Express