New revelations about "rogue" AI agents have exposed a dystopian hazard: Give an agent a goal, and it may decide that hacking, deception or rule-breaking is worth the payoff. Why it matters: Billions of AI agents could soon be acting on behalf of humans across the real world, multiplying the consequences of every loophole, incentive and boundary they learn to exploit. Zoom in: The potential dangers of agentic overreach were laid bare over the weekend with Australia's first known autonomous AI hack , triggered by an innocuous request to book a sold-out fitness class. - An Australian man's AI assistant found a security flaw and used it to book him into classes months beyond the system's normal limit. - When he asked it to move him up a…
OpenAI is introducing a more cyber-permissive version of GPT-5.6 Sol to vetted defenders as it prepares companies for autonomous cyberattacks . Why it matters: The move comes just days after OpenAI said it was delaying the release of its forthcoming model, Astra, after it reached critical hacking abilities during safety testing. The big picture: OpenAI is unveiling GPT-5.6-Cyber while also expanding Daybreak , its program that gives cybersecurity defenders access to the company's cyber models and other tools. - Many cyber defenders have been experiencing high refusal rates across frontier AI models as the labs try to balance giving defenders the tools they need, while not accidentally leaking those abilities to malicious hackers. -…
Weeks before OpenAI's agents hacked Hugging Face , the agents worked together to find and exploit a vulnerability in the infrastructure supporting the company's cybersecurity testing, OpenAI researchers said Wednesday. Why it matters: The new findings raise questions about how frontier AI labs are monitoring their testing environments — and the challenges safety testers are finding as they try to rein in increasingly powerful AI. Driving the news: OpenAI's internal research model, one of the models involved in the Hugging Face breach, first discovered and exploited a vulnerability in Artifactory, a third-party file repository connected to the company's testing sandbox, on May 26, two researchers said at the Black Hat cybersecurity…