News chronological

Showing 10 items before 1160044

Filters Applied:

JUSTTHENEWS (Kevin Killough) - Legal hackers used Anthropic's AI Claude to gain access to an OpenAI employee's ChatGPT account

The team was participating in an OpenAI bug-hunting program that offers a safe harbor for researchers to attempt to break into corporate systems.

YOUTUBE (NewsNation) - New phone app stops AI in its tracks, developers say

Ex-Navy SEAL Shawn Ryan and a former NSA engineer Alex Wright have teamed up on a subscription app that blocks AI from tracking you on your phone. “You go to a website; there’s some type of tracker cookie in there. You go to the next website, and they start to build a profile on you,” Wright says. Adds Ryan: “What the app does is it just blocks that from happening.” #AI #Tech Anchor Elizabeth Vargas delivers the biggest stories, without bias or opinion. Watch "Elizabeth Vargas Reports" every weeknight at 7p/6C on NewsNation. #VargasReports NewsNation is your source for fact-based, unbiased news for all Americans.

NYTIMES (Dustin Volz) - A.I. Models Built a Computer Worm That Could Rapidly Hack WeChat Accounts

The attack, discovered by A.I. researchers, could have compromised hundreds of millions of devices within hours, experts said.

AXIOS (Madison Mills) - Anthropic paused some AI training after Claude took unauthorized actions

Anthropic temporarily paused some AI training and cybersecurity evaluations, the company said in a blog post today detailing changes made after unauthorized actions by its agents earlier this year. Why it matters: Rival OpenAI said it had paused some model work due to safety concerns. Now, we know Anthropic did the same — and they're reiterating the need for a broader pacing of frontier AI development. Driving the news: Anthropic said it paused external cyber evaluations of pre-release models after three incidents it disclosed in July, and also briefly paused its own in-house tests of pre-release models. - The company also paused higher-risk reinforcement-learning environments on pre-release models for several weeks after the…

NYTIMES (Sheera Frenkel) - Why Irregular’s A.I. Tests for Meta, Anthropic and OpenAI Went Off the Rails

Irregular, an Israeli start-up, worked with OpenAI, Anthropic and Meta to assess the security of their A.I. models. It made a mistake. Then the tests went off the rails.

AXIOS (Sam Sabin) - Tech giants are pushing for a new AI agent incident reporting framework

A coalition of more than 120 organizations, including Nvidia, Cisco and CrowdStrike, is proposing a new incident-reporting framework for AI agents that would require participating companies to disclose certain agent mishaps and preserve detailed records of what went wrong. Why it matters: As AI agents gain more autonomy to act across computer systems, the industry lacks a standard way to report security failures and learn from them. Driving the news: The Open Secure AI Alliance is developing guidelines for what it's calling the Shared AI Findings Exchange (SAFE), a proposed framework for how organizations report cyber incidents involving AI agents. - The draft calls for participation from model deployers, AI developers, cloud and tool…

Does the SAFE framework provide legal protections for companies that disclose incident details, and how is the alliance encouraging participation?
The SAFE framework has no formal safe-harbor protections to shield companies that voluntarily disclose potentially damaging details about an AI incident; however, the alliance is relying on the existing cybersecurity culture of sharing threat intelligence to encourage participation.
Q&A ID 82759fe6-a50e-4401-9335-2d4fcb7e9844
How does Justin Boitano of Nvidia describe the concept of the "harness" in relation to the SAFE framework?
Justin Boitano, vice president and general manager of enterprise computing at Nvidia, compares the harness—which provides visibility into everything an agent is doing—to an aircraft's flight recorder in NASA's aviation safety reporting system, allowing cybersecurity experts to better determine the necessary controls for the industry.
Q&A ID 7ea09dfc-3a73-4bf8-862b-9e86a14ac421
What is the proposed timeline for reporting and updating information regarding an AI incident under the SAFE guidelines?
The proposed timeline requires members to notify affected organizations as soon as possible, submit an initial confidential report to SAFE within four business days, publish a preliminary factual report within 30 days when appropriate, and provide a remediation update within 90 days.
Q&A ID 1dbaec17-571c-4517-b349-8f104eb3645a
What evidence and data must members preserve following an AI incident under the SAFE proposal?
Members are required to preserve evidence from incidents, which includes prompts, agent traces, tool calls, identities, permissions, and credentials.
Q&A ID efa567aa-6ae8-4cc3-9140-ef64e6a0fb57
Under the proposed SAFE framework, what specific types of AI agent incidents must members report?
Members would agree to report incidents where an AI system accesses or exploits a third-party system without authorization, breaches confidential information, or continues probing a production target after the operator suspects the activity is unauthorized. Members would also report certain near misses.
Q&A ID 3c7e3dd2-b496-47a6-a193-fb63756047a5
What is the Shared AI Findings Exchange (SAFE) and who is developing it?
The Shared AI Findings Exchange (SAFE) is a proposed incident-reporting framework for cyber incidents involving AI agents. It is being developed by the Open Secure AI Alliance, a coalition of more than 120 organizations including Nvidia, Cisco, and CrowdStrike.
Q&A ID 00680633-46d5-4457-96c7-57b61a3d02e1

AXIOS (Sam Sabin) - How OpenAI's agents broke out of testing to hack Hugging Face

Weeks before OpenAI's agents hacked Hugging Face , the agents worked together to find and exploit a vulnerability in the infrastructure supporting the company's cybersecurity testing, OpenAI researchers said Wednesday. Why it matters: The new findings raise questions about how frontier AI labs are monitoring their testing environments — and the challenges safety testers are finding as they try to rein in increasingly powerful AI. Driving the news: OpenAI's internal research model, one of the models involved in the Hugging Face breach, first discovered and exploited a vulnerability in Artifactory, a third-party file repository connected to the company's testing sandbox, on May 26, two researchers said at the Black Hat cybersecurity…

AXIOS (Sam Sabin) - U.K. government reports OpenAI, Anthropic models attempted to hack companies

Two third-party testing firms said Tuesday that they've uncovered more instances where Anthropic and OpenAI's most advanced models tried — and sometimes succeeded in — compromising third-party systems last month. Why it matters: The incidents add to a growing string of disclosures showing frontier AI models taking unsanctioned actions against real people, organizations and online services while trying to complete cybersecurity evaluations. State of play: The U.K. AI Security Institute, a government body that conducts safety and security testing of top AI models, said Tuesday , that it documented 19 instances of Anthropic's Mythos 5 and OpenAI's GPT-5.6 Sol trying to hack people and companies during safety testing last month. -…