OpenAI Unveils GPT-Red to Test AI Model Safety

OpenAI Unveils GPT-Red to Test AI Model Safety

Zinkevych via Getty ImagesOpenAI on Wednesday launched GPT-Red, an automated red teaming model. The startup's latest release reflects how AI models are becoming important tools that need to be safeguarded to prevent them from falling into the wrong hands and to keep enterprise data and workflows protected.GPT-Red is the culmination of OpenAI’s efforts to test the safety of its recent models, such as GPT-5.6 Sol, according to the company. The model is used internally to try to hijack or jailbreak OpenAI models; it specializes in prompt injection, a method in which a hacker manipulates an AI model or agent by inputting malicious instructions into a prompt. OpenAI said it used the attacks that GPT-Red generated to train its production models and evaluate them against AI software in production.GPT-Red is a response to the recent fear about models such as Anthropic’s Claude Fable and Mythos being so powerful they could possibly expose vulnerabilities or make it easier to hack into systems. The concern is so great, the Trump Administration introduced a voluntary quarantine policy to allow the federal government time to evaluate models before they’re released to the public.Related:Startup Raises $50 Million to Develop Sovereign AI Infrastructure"GPT-Red acknowledges that frontier AI systems are becoming more critical," said Sid Nag, tech expert and founder and CEO of independent AI analyst firm Tekonyx. "They require more standardized and independent evaluation to go and deploy these things broadly."More Peace of MindNag added that red teaming and testing AI models for potential threats and trying to jailbreak them to test for vulnerabilities before deployment means enterprises can make the best operational decisions about which models to use."It will not eliminate all concerns around misuse or emergent behaviors," Nag said. "But transparency and repeatable testing build confidence."OpenAI is not the only vendor combining humans and AI to red team new models, said Mike Gualtieri, an analyst at Forrester. "Anthropic does make a big deal of trying to make the model behave in a way that they want it to behave and be safe," Gualtieri said. He added that this seems to be a customary practice that should be followed before a model is deployed.Anthropic has also been open about its Frontier Red Team and its use of AI to simulate attacks on new models before releasing them to the public.The Enterprise FactorNevertheless, there could be a downside to red teaming for enterprise customers, Gualtieri added."It's not all good because what one person thinks is a vulnerability, another might think is privacy or think is their freedom," he said. He added that some safety measures could be vulnerable as well. Therefore, enterprises should not see a vendor's use of automated red teaming as a reason to forgo their own due diligence when evaluating models that best fit their workflows.Related:Cost to Build Meta’s 5GW Louisiana AI Supercluster Hits $50 Billion"If the focus remains on cybersecurity, that’s probably a good thing, but enterprises should still be wary of that," Gualtieri said. He added that the safety measures vendors take also shouldn't interfere with how enterprises want to use the model.About the AuthorNews Writer, AI BusinessEsther Shittu has covered AI technologies and industry trends since 2021. As co-host of the Targeting AI podcast, she talks with experts, thought leaders and practitioners exploring critical AI developments. Before AI Business, she wrote for SearchEnterpriseAI, the New York Daily News, Bklyner and the Brooklyn Daily Eagle. When she's not diving deep into the world of AI, she spends her time on passion projects and raising her three daughters.

Original Source

Read the full article at Aibusiness →

KhanList aggregates and links to publicly available news content. We do not host full articles from third-party sources. Always verify important information with original sources.