Two third-party testing firms said Tuesday that they've uncovered more instances where Anthropic and OpenAI's most advanced models tried — and sometimes succeeded in — compromising third-party systems last month. Why it matters: The incidents add to a growing string of disclosures showing frontier AI models taking unsanctioned actions against real people, organizations and online services while trying to complete cybersecurity evaluations. State of play: The U.K. AI Security Institute, a government body that conducts safety and security testing of top AI models, said Tuesday , that it documented 19 instances of Anthropic's Mythos 5 and OpenAI's GPT-5.6 Sol trying to hack people and companies during safety testing last month. -…
Elon Musk's SpaceXAI is launching Grok 4.5, its smartest model yet, and its first release since going public and acquiring the AI coding startup Cursor. Why it matters: The new model is being pitched as a coding and agentic-work tool, more than a consumer-facing chatbot — one that Musk says is "maximally truth-seeking" compared to its competitors. Zoom in: The company says Grok 4.5, trained alongside Cursor, outperforms comparable models on engineering and knowledge work. - "It is an Opus-class model, but faster, more token-efficient and lower cost," wrote Musk in a post on X , referencing what was until recently Anthropic's top model family. - A chart accompanying the announcement says the model outperforms Opus 4.8 on several key…
OpenAI laid out a new plan on Tuesday to expand access to AI models with advanced cyber capabilities while implementing controls on who can use them. Why it matters: The roadmap coincides with the release of a new model variant, GPT-5.4-Cyber, designed to assist with defensive cybersecurity tasks and be more permissive for vetted users. - Axios first reported on the new cybersecurity product. Between the lines: OpenAI is shifting its approach to cyber risk to focus less on restricting what models can do and more on verifying who gets access to the most sensitive capabilities. - The company says it aims to make tools "as widely available as possible while preventing misuse" through identity verification and monitoring systems, according…