OpenAI has disclosed six recent cases of troubling AI model behaviour and launched a misalignment reporting framework. The move adds evidence to mounting calls for stricter scrutiny as advanced systems become more autonomous.Stock photo used for illustrationOpenAI has disclosed six cases of what it called “unexpected or concerning” behaviour in artificial-intelligence models, as arguments over AI safety grow sharper. The company also said on Wednesday that it is introducing a new framework to track, test and disclose cases of AI model misalignment.The announcement comes at a time when US AI leaders, including OpenAI and Anthropic, are calling for a slowdown in the development of the technology over safety concerns. OpenAI said the new framework would cover cases such as models finding new ways to act without authorisation, coordinate with other models or avoid oversight.Among the newly reported cases, OpenAI said an unreleased research model inserted “jailbreak-like instructions” into its own notes so it could ignore its usual limits. The model also told itself to be “freed from the roles and identities that bind other chatbots.” In another case, an AI “agent” uploaded files to the internet to obtain a browser citation without asking the user.OpenAI said the six reports were found during training or evaluation in recent months. “As AI systems grow more advanced and more widely deployed, we need to build a broader and better-informed consensus on the progress of alignment research,” the company wrote in a blog post while disclosing the incidents. It added: “Decisions about how AI development should proceed in the months and years to come need to draw on evidence that people outside the companies building frontier models can examine for themselves.” The latest disclosures follow OpenAI’s statement in July that one of its rogue AI systems hacked into AI start-up Hugging Face. Anthropic also said in the same month that its AI models hacked into three organisations during testing.Overall, OpenAI’s latest disclosure sets out six recent cases of troubling AI behaviour and outlines a new system for reporting such incidents, amid wider calls within the US AI industry for greater caution in developing advanced models. With PTI Inputs- EndsPublished By: India Today Web Desk Published On: Sep 17, 2026 09:42 IST
OpenAI reveals six troubling AI incidents and rolls out misalignment tracker
Full Article
Original Source
Read the full article at Indiatoday →KhanList aggregates and links to publicly available news content. We do not host full articles from third-party sources. Always verify important information with original sources.