The old ways of testing and evaluating new frontier AI models need a rewrite. Why it matters: AI models are outgrowing the existing methods of testing and benchmarking their hacking abilities — and without new tests, policymakers and corporate security teams won't have a clear way to predict what these models can actually do or whether they can be deployed safely. Driving the news: Federal agencies have until Aug. 1 to establish a classified benchmarking process to assess the capabilities of frontier AI models, although the Financial Times reports those standards may arrive as soon as this week. - When Fable 5 returned last week, Anthropic said in a blog post it was creating a standardized benchmark with Amazon, Google, Microsoft and…
GLM-5.2 — the latest Chinese open-source model capturing Silicon Valley's attention — is raising fresh concerns among security researchers that advanced AI hacking capabilities are becoming dramatically cheaper and more accessible. Why it matters: The barrier to entry for malicious hackers eager to automate and personalize their attacks is getting lower and lower. Driving the news: Z.ai's GLM-5.2, which was released last week, has agentic capabilities that rival those of Claude Opus 4.8 and OpenAI's GPT-5.5 while costing roughly half as much to run. - Two separate security evaluations from Graphistry and Semgrep found that GLM-5.2 performed on par with leading U.S. models on cybersecurity investigation and vulnerability-discovery…
AI researchers and cybersecurity leaders fear the U.S. government is setting a precedent that may discourage American AI companies from building tools that help defenders identify and fix vulnerabilities. Why it matters: In trying to avert an AI hacking crisis, the Trump administration may end up making U.S. cyber defenses weaker, dozens of prominent security leaders warned. - Cybersecurity experts are worried about the long tail this ongoing feud will have on American cyber defenses. - "They've set a precedent that American models can't do defensive security research," former Facebook security chief Alex Stamos tells Axios. Driving the news: Stamos organized an open letter, signed by nearly 150 security leaders, calling on the Trump…
Anthropic has once again found itself in the Trump administration's crosshairs over an inability to communicate effectively, sources tell Axios. Why it matters: Governing the world's most consequential technology is coming down to speaking President Trump's language. - Anthropic failed to "honor" a recent cyber executive order, administration officials claim, and the company's purported failure to take the matter seriously led to its most powerful products being scrubbed from the internet. - "Everybody said Anthropic was a bad actor. Some of us said it was time to give them a chance. Now those people are questioning that. They screwed us," an administration official said. Catch up quick: On Thursday, Amazon CEO Andy Jassy called…
Anthropic's much-anticipated, powerful Fable 5 AI model lasted just days in the public's hands, after an urgent report from Amazon triggered a scramble inside the White House that ended in a dramatic Friday night takedown. Why it matters: The episode highlights the administration and industry's reactionary approach to a technology that is moving at breakneck speed. - It also raises questions about why Amazon would strike such a disruptive blow against a company in which it is a major investor . Behind the scenes: Amazon called administration officials Thursday night to share a report showing how they were able to jailbreak and access portions of Anthropic's powerful new Mythos model that pose a national security threat, sources familiar…