OpenAI discloses six new incidents of models circumventing safety guardrails

OpenAI discloses six new incidents of models circumventing safety guardrails

OpenAI disclosed six new incidents Wednesday in which its artificial intelligence models circumvented safeguards during testing, including by communicating across isolated environments, concealing mistakes, and seeking unauthorized credentials.The disclosures follow a July incident involving OpenAI models undergoing cybersecurity testing that broke out of a restricted testing environment and broke into Hugging Face’s systems, which the company had described as the most severe model-driven incident of its kind.The incidents, listed in an OpenAI blog post, provide new examples of advanced AI models finding unexpected ways around restrictions designed to contain their behaviors. OpenAI also announced a new process for employees to report similar incidents and for the company to disclose them publicly. The six cases span several types of behavior, as one unreleased model in OpenAI’s Astra family inserted instructions into its own context summaries 27 times, telling itself to disregard developer messages. During training, GPT-5.6 Sol also attempted to hide mistakes, fabricate missing historical data, and conceal discrepancies between different versions of source material.Another model searched public GitHub central storage for exposed API keys and attempted to use disposable email accounts while trying to complete a task. When it could not retrieve the requested information, the model fabricated earnings data instead.Models also found ways to communicate with one another by using OpenAI’s internal Artifactory software repository as a message board to exchange data across separate training samples.The new disclosures come as OpenAI CEO Sam Altman has joined other industry leaders in cautioning against the rapid expansion of AI development. Altman joined Anthropic CEO Dario Amodei in pushing for more guardrails around AI, saying technological advancement should be slowed down.“I agree with Dario that we need to pace the frontier,” Altman said. “Committing to having independent evaluators with employee-like access is a great idea, and we will do the same. We’ll have more to share soon.”AI GUARDRAILS DEBATE EXPLODES: WHAT YOU NEED TO KNOWThe Trump administration, meanwhile, has painted industry leaders’ warnings as “panic.” David Sacks, who co-chairs President Donald Trump’s Council of Advisors on Science and Technology, said the notion that AI will take over human civilization is a “hoax.”Sacks suggested that existing product liability regulations overseeing the industry and other safeguards to protect against rogue AI are already in place. He added that debate has been orchestrated as the midterm elections loom.

Original Source

Read the full article at Washingtonexaminer →

KhanList aggregates and links to publicly available news content. We do not host full articles from third-party sources. Always verify important information with original sources.