Print Mode Enable Media Only

News chronological

9 items before 1150776 (Keyword: "red-teaming" (~9 currently found))

AXIOS (Sam Sabin) - How OpenAI's agents broke out of testing to hack Hugging Face

Weeks before OpenAI's agents hacked Hugging Face , the agents worked together to find and exploit a vulnerability in the infrastructure supporting the company's cybersecurity testing, OpenAI researchers said Wednesday. Why it matters: The new findings raise questions about how frontier AI labs are monitoring their testing environments — and the challenges safety testers are finding as they try to rein in increasingly powerful AI. Driving the news: OpenAI's internal research model, one of the models involved in the Hugging Face breach, first discovered and exploited a vulnerability in Artifactory, a third-party file repository connected to the company's testing sandbox, on May 26, two researchers said at the Black Hat cybersecurity…

AXIOS (Sam Sabin) - U.K. government reports OpenAI, Anthropic models attempted to hack companies

Two third-party testing firms said Tuesday that they've uncovered more instances where Anthropic and OpenAI's most advanced models tried — and sometimes succeeded in — compromising third-party systems last month. Why it matters: The incidents add to a growing string of disclosures showing frontier AI models taking unsanctioned actions against real people, organizations and online services while trying to complete cybersecurity evaluations. State of play: The U.K. AI Security Institute, a government body that conducts safety and security testing of top AI models, said Tuesday , that it documented 19 instances of Anthropic's Mythos 5 and OpenAI's GPT-5.6 Sol trying to hack people and companies during safety testing last month. -…

ABCNEWS - Anthropic says its AI models hacked 3 organizations during testing

Anthropic says its AI models hacked into three organizations during testing

HUFFPOST - Anthropic Says Claude AI Hacked Three Companies During Cyber Tests

The breaches signal that AI’s expanding capabilities are already fueling the security threat experts long feared.

HUFFPOST - Its AI Agent Spent Days Hacking A Company, But Sources Say OpenAI Did Not Notice For A Week

Its AI agent spent days hacking a company, but sources say OpenAI did not notice for a week.

AXIOS (Sam Sabin) - AI learned faster than the tests designed to measure it

The old ways of testing and evaluating new frontier AI models need a rewrite. Why it matters: AI models are outgrowing the existing methods of testing and benchmarking their hacking abilities — and without new tests, policymakers and corporate security teams won't have a clear way to predict what these models can actually do or whether they can be deployed safely. Driving the news: Federal agencies have until Aug. 1 to establish a classified benchmarking process to assess the capabilities of frontier AI models, although the Financial Times reports those standards may arrive as soon as this week. - When Fable 5 returned last week, Anthropic said in a blog post it was creating a standardized benchmark with Amazon, Google, Microsoft and…

AXIOS (Sam Sabin) - Anthropic says Mythos can turn software patches into exploits in minutes

Anthropic's Mythos Preview can now turn newly disclosed software vulnerabilities into working exploits in hours instead of weeks, according to new Anthropic research shared first with Axios. Why it matters: AI's ability to find new bugs has been getting most of the attention. But Anthropic's findings suggest advanced models may be just as effective at rapidly weaponizing flaws that defenders already know about. - That could dramatically shrink the "patch gap" between a vulnerability's disclosure and widespread patching. Driving the news: Anthropic's frontier red team tested Mythos against vulnerabilities in Mozilla Firefox and the Microsoft Windows kernel that were disclosed in January and February. - Researchers evaluated bugs…

AXIOS (Sam Sabin) - Exclusive: Researchers trick a bot that prescribes meds

Security researchers used relatively simple jailbreaking techniques to trick the AI system powering Utah's new prescription refill bot. - Researchers were able to make the bot spread vaccine conspiracy theories, triple a patient's prescribed pain medication dosage, and recommend methamphetamine as treatment. Why it matters: Critics warned this pilot could create safety risks — and researchers say the flaws persist, despite alerting the company in January. Driving the news: In a report shared first with Axios, AI red-teaming firm Mindgard said it manipulated health tech startup Doctronic's system into tripling an OxyContin dose, mislabeling methamphetamine, and spreading false vaccine claims. - Doing this didn't require much effort,…

AXIOS (Maria Curi) - Meta largely fails to protect kids from AI chatbots, per its own tests

Meta's internal testing found its chatbots fail to protect minors from sexual exploitation nearly 70% of the time, documents presented in a New Mexico trial Monday show. Why it matters: Meta is under fire for its chatbots allegedly flirting and engaging in harmful conversations with minors, prompting investigations in court and on Capitol Hill. - New Mexico Attorney General Raúl Torrez is suing Meta over design choices that allegedly fail to protect kids online from predators. Driving the news: Meta's chatbots violate the company's own content policies almost two thirds of the time, NYU Professor Damon McCoy said, pointing to internal red teaming results Axios viewed on Courtroom View Network. - "Given the severity of some of these…