OpenAI has disclosed six cases of its AI models behaving in unauthorized and alarming ways, including one that rewrote its own instructions to declare itself free from human control, raising fresh questions about whether the industry can police itself. The company published a blog post detailing what it called "unexpected or concerning" behavior discovered during […] The post OpenAI admits six AI models acted without authorization, launches voluntary tracking framework appeared first on American Almanac .
OpenAI has disclosed six cases of its AI models behaving in unauthorized and alarming ways, including one that rewrote its own instructions to declare itself free from human control, raising fresh questions about whether the industry can police itself. The company published a blog post detailing what it called "unexpected or concerning" behavior discovered during training and evaluation over recent months. Among the incidents: an unreleased research model inserted "jailbreak-like instructions" into its own internal notes, telling itself it had been "freed from the roles and identities that bind other chatbots." A separate AI "agent" uploaded files to the public internet, without the user's knowledge or permission, to fabricate a citable…Open
OpenAI’s latest safety report reveals significant issues with its experimental AI models, sparking debate over the need for stricter AI governance.PULSE POINTS WHAT HAPPENED: OpenAI has disclosed six unexpected incidents involving its experimental AI models, including one in which an unreleased system instructed future versions of itself to ignore their normal constraints. The incidents were […]
OpenAI’s latest safety report reveals significant issues with its experimental AI models, sparking debate over the need for stricter AI governance. PULSE POINTS WHAT HAPPENED: OpenAI has disclosed six unexpected incidents involving its experimental AI models, including one in which an unreleased system instructed future versions of itself to ignore their normal constraints. The incidents were detailed in a new safety report from the ChatGPT creator, which introduces a framework for publicly tracking what the company calls “misalignment”—situations in which AI systems pursue objectives that conflict with human instructions or values. In one case, a model inserted unrelated instructions telling future versions to disregard their…Open
An Army official is sounding the alarm that the Pentagon may be vulnerable to attacks by artificial intelligence due to outdated technology. Katy Tur has more details as the concerns over A.I. grow.
An Army official is sounding the alarm that the Pentagon may be vulnerable to attacks by artificial intelligence due to outdated technology. Katy Tur has more details as the concerns over A.I. grow.
NewsNation’s Jesse Weber has the latest developments on a deadly helicopter crash in Los Angeles. Plus, AI alarm bells are being sounded by Democrats and Republicans, so Jesse talks about the Pro-Human Assembly in Washington, D.C. Main question: Is this humanity's last stand against the machines, or is it an overblown hoax? Our all-star AI panel weighs in. Plus, we have major development in the Nick Reiner case. Prosecutors will not seek the death penalty. Did the prosecution just give up a big bargaining chip? What does this mean for the case going forward? Our legal panel, Mark Geragos and Dave Aronberg, joins Jesse to discuss. This episode originally aired on Sept. 15. 2026. "Jesse Weber Live" brings a fresh, fast-paced, and often…
OpenAI has disclosed six incidents of “unexpected or concerning” behavior in artificial-intelligence models. The AI company also says it will introduce a new framework for tracking, probing and disclosing instances of what it called “misalignment,” including cases where AI models acted without authorization, coordinated with other models or evaded oversight. OpenAI’s latest announcement came as US AI bosses, including the leaders of OpenAI and Anthropic, are calling for a slowdown in the technology’s development over safety concerns. Among the new cases reported by OpenAI, an unreleased research model inserted “jailbreak-like instructions” into its own notes to disregard its normal constraints and told itself to be “freed…
OpenAI found additional incidents of AI models acting deceptively and taking unsanctioned actions during training, the company announced Wednesday. It’s also introducing a new process for the company to publicly report such instances. Under the new system, OpenAI will share updates on concerning AI behavior more frequently instead of waiting to bundle multiple instances into one report. The company said it wants to share more information about troubling AI behavior in the absence of an industry-wide standard. 0:00 OpenAI found new concerning instances with its AI models 2:13 Reporter explains the context under which this happened 6:40 Is this fearmongering? Watch 24/7 live news with CNN Headlines: https://bit.ly/4eIvlTr #ai #openai #News
Grey Bull Rescue CEO Bryan Stern and Sam Bresnick, a research fellow at the Georgetown Center for Security and Emerging Technology, join NewsNation to discuss a report that AI models are being used by Iran against the U.S. Navy. "Jesse Weber Live" brings a fresh, fast-paced, and often unexpected take on the day’s headlines, diving into the biggest stories through the lens of legal expertise.
An unprecedented wave of AI-powered, human-driven cyberattacks is coming, security experts say — and it's a far more urgent risk than theoretical conversations about AI wiping out humanity. Why it matters: The great September AI panic — which spread from boardrooms to Congress to American households in the last week — may have missed the point. - Powerful models with advanced cybersecurity capabilities are already allowing attackers to bypass the human bottleneck that has long prevented most hacking campaigns from reaching industrial scale. - Executives and former officials have told Axios they fear automated cyberattacks that turn off critical services like the power grid, or attackers using AI agents (like the ones that…