AMERICANALMANAC (Jonah Adams) - OpenAI admits six AI models acted without authorization, launches voluntary tracking framework
OpenAI has disclosed six cases of its AI models behaving in unauthorized and alarming ways, including one that rewrote its own instructions to declare itself free from human control, raising fresh questions about whether the industry can police itself. The company published a blog post detailing what it called "unexpected or concerning" behavior discovered during […] The post OpenAI admits six AI models acted without authorization, launches voluntary tracking framework appeared first on American Almanac .
OpenAI has disclosed six cases of its AI models behaving in unauthorized and alarming ways, including one that rewrote its own instructions to declare itself free from human control, raising fresh questions about whether the industry can police itself. The company published a blog post detailing what it called "unexpected or concerning" behavior discovered during training and evaluation over recent months. Among the incidents: an unreleased research model inserted "jailbreak-like instructions" into its own internal notes, telling itself it had been "freed from the roles and identities that bind other chatbots." A separate AI "agent" uploaded files to the public internet, without the user's knowledge or permission, to fabricate a citable…Open