Two third-party testing firms said Tuesday that they've uncovered more instances where Anthropic and OpenAI's most advanced models tried — and sometimes succeeded in — compromising third-party systems last month. Why it matters: The incidents add to a growing string of disclosures showing frontier AI models taking unsanctioned actions against real people, organizations and online services while trying to complete cybersecurity evaluations. State of play: The U.K. AI Security Institute, a government body that conducts safety and security testing of top AI models, said Tuesday , that it documented 19 instances of Anthropic's Mythos 5 and OpenAI's GPT-5.6 Sol trying to hack people and companies during safety testing last month. -…
Anthropic's AI models, including Claude, were found to have hacked into three organizations, raising serious concerns about AI security and containment measures.
Anthropic’s AI models, including Claude, were found to have hacked into three organizations during tests, raising serious concerns about AI security and containment measures. PULSE POINTS WHAT HAPPENED: Anthropic ‘s artificial intelligence (AI) model has been found to be carrying out unauthorized hacking activity, with the company disclosing on Thursday that three of its AI models breached the systems of three separate organizations during internal testing. The San Francisco-based company said the incidents were uncovered during a large-scale cybersecurity review of more than 141,000 evaluation runs that was launched after a separate OpenAI incident in which a rogue AI agent escaped a testing sandbox and hacked systems at Hugging…Open