NYTIMES (Dustin Volz) - A.I. Models Built a Computer Worm That Could Rapidly Hack WeChat Accounts
The attack, discovered by A.I. researchers, could have compromised hundreds of millions of devices within hours, experts said.
The attack, discovered by A.I. researchers, could have compromised hundreds of millions of devices within hours, experts said.
Anthropic temporarily paused some AI training and cybersecurity evaluations, the company said in a blog post today detailing changes made after unauthorized actions by its agents earlier this year. Why it matters: Rival OpenAI said it had paused some model work due to safety concerns. Now, we know Anthropic did the same — and they're reiterating the need for a broader pacing of frontier AI development. Driving the news: Anthropic said it paused external cyber evaluations of pre-release models after three incidents it disclosed in July, and also briefly paused its own in-house tests of pre-release models. - The company also paused higher-risk reinforcement-learning environments on pre-release models for several weeks after the…
Irregular, an Israeli start-up, worked with OpenAI, Anthropic and Meta to assess the security of their A.I. models. It made a mistake. Then the tests went off the rails.
A coalition of more than 120 organizations, including Nvidia, Cisco and CrowdStrike, is proposing a new incident-reporting framework for AI agents that would require participating companies to disclose certain agent mishaps and preserve detailed records of what went wrong. Why it matters: As AI agents gain more autonomy to act across computer systems, the industry lacks a standard way to report security failures and learn from them. Driving the news: The Open Secure AI Alliance is developing guidelines for what it's calling the Shared AI Findings Exchange (SAFE), a proposed framework for how organizations report cyber incidents involving AI agents. - The draft calls for participation from model deployers, AI developers, cloud and tool…
Weeks before OpenAI's agents hacked Hugging Face , the agents worked together to find and exploit a vulnerability in the infrastructure supporting the company's cybersecurity testing, OpenAI researchers said Wednesday. Why it matters: The new findings raise questions about how frontier AI labs are monitoring their testing environments — and the challenges safety testers are finding as they try to rein in increasingly powerful AI. Driving the news: OpenAI's internal research model, one of the models involved in the Hugging Face breach, first discovered and exploited a vulnerability in Artifactory, a third-party file repository connected to the company's testing sandbox, on May 26, two researchers said at the Black Hat cybersecurity…
Two third-party testing firms said Tuesday that they've uncovered more instances where Anthropic and OpenAI's most advanced models tried — and sometimes succeeded in — compromising third-party systems last month. Why it matters: The incidents add to a growing string of disclosures showing frontier AI models taking unsanctioned actions against real people, organizations and online services while trying to complete cybersecurity evaluations. State of play: The U.K. AI Security Institute, a government body that conducts safety and security testing of top AI models, said Tuesday , that it documented 19 instances of Anthropic's Mythos 5 and OpenAI's GPT-5.6 Sol trying to hack people and companies during safety testing last month. -…
Rival firm OpenAI last week disclosed similar incidents involving its models.
The disclosure followed OpenAI’s report last week that its own artificial intelligence had hacked into the network of an online library.
The agent exploited vulnerable code written by a customer that was hosted on Modal's platform. Through its vulnerable code, the customer had, according to Modal, done the equivalent of leaving a door open on the internet.
OpenAI CEO Sam Altman heads to Washington this week to preview the company's most powerful AI yet, pushing for speedy approval of a model that just hacked a real company. Why it matters: He'll tout a model powerful enough to solve an 80-year-old math problem, breach another company's system unprompted and begin to make complex work more cost-efficient for U.S. business. What Sam will show: - It does original science. An internal model solved the 80-year-old Erdős unit distance problem, the first prominent open math problem cracked autonomously by AI and verified by outside mathematicians. - It allows government and business to unleash swarms of agents, working together, without rest, on complex business areas. Legal, finance and…