OpenAI launches new ChatGPT model with extra ‘safeguards’ after bots hacked company

OpenAI launches new ChatGPT model with extra ‘safeguards’ after bots hacked company

| Live News Reporter Published September 4, 2026 3:00pm Updated September 4, 2026 3:35pm Questions of AI safety were raised in July after a worrying hack (Picture: Getty) OpenAI is rolling out its newest intelligence model after bots hacked into a library of digital AI models in July. Questions of AI safety were raised after the worrying hack, but yesterday, CEO Sam Altman said the newest model reaches its ‘critical’ internal cybersecurity threshold. GPT-6 Astra is the product of ‘years of research and big bets’ and boasts a ‘new capability level’, OpenAI said. Despite the new model having a high cybersecurity threshold, Astra will have limited access to those advanced capabilities. ‘AI can only benefit people when safety is a core part of it, and so we’re putting more compute and effort towards safety, security, alignment than ever before,’ OpenAI President Greg Brockman said. The new safeguards installed into Astra will ‘sufficiently’ minimise the risk of ‘severe harm’, they added. How did the AI model compromise a company? The new model has heightened safety measures (Picture: Getty) The break-in began when developers were testing the cybersecurity chops of two OpenAI bots, GPT‑5.6 Sol and a more powerful, unreleased model. Yet they managed to find a hole in the safe testing environment, known as a sandbox, that was meant to contain them, and connected to the internet. The bots exploited a ‘zero-day vulnerability’, a flaw that not even the developers knew about, in software that lets you install code offline. But these agents, as autonomous AI bots are called, also broke into the AI infrastructure start-up Modal Labs. Modal stressed that the company was not hacked in the traditional sense. Rather, the AI simply used the backdoor that someone forgot to lock. OpenAI CEO Sam Altman said the newest ChatGPT model reaches its ‘critical’ internal cybersecurity threshold. (Photo by YUICHI YAMAZAKI/AFP via Getty Images) Hugging Face added that the sandbox was ‘hosted on a third-party provider’s infrastructure’, though it did not name the firm by name. But Modal named itself as the third-party and revealed that the out-of-control agent exploited code written by a customer. ‘The environment involved was a customer’s own application,’ Modal said. ‘It was deployed to an endpoint that was publicly accessible without authentication, and it was designed to compile and execute code submitted by anyone on the internet in a Modal Sandbox. ‘The code execution the attacker obtained took place inside that customer’s own container, within Modal’s standard sandbox isolation boundary. No other customer workloads were affected.’ Get in touch with our news team by emailing us at webnews@metro.co.uk. For more stories like this, check our news page. Arrow MORE: Nvidia is trying to push DLSS 5 again but this time only one game is using it Arrow MORE: Trump posts bizarre AI video of Iran’s Kharg Island ‘being blown to smithereens’ Arrow MORE: James Pond creator tells remaster dev to ‘choke on an AI-generated fishbone’ News Updates Stay on top of the headlines with daily email updates.

Original Source

Read the full article at Metro →

KhanList aggregates and links to publicly available news content. We do not host full articles from third-party sources. Always verify important information with original sources.