OpenAI’s autonomous AI agents repeatedly queried a UN website and used techniques that appeared to circumvent restrictions when trying to retrieve publicly available data, according to an independent research report based on information supplied by AI research firm Transluce. The agents scanned a public data hub operated by UN Trade and Development, the UN’s trade arm, more than 16,000 times between April and the end of June, highlighting a growing problem with AI agents that can independently navigate the web and take actions when they encounter obstacles. Researchers believe the models were initially tasked with finding publicly available information, but their behavior became increasingly aggressive when the website prevented them from accessing some of the requested data. Pushing past website restrictions According to researcher and report author Rowan Howard-Jones, the agents encountered a filter that blocked their requests and subsequently found a way around it, with the technique they ultimately used not permitted by the website operators. Additionally, the agents were apparently not instructed to attack or compromise the UN website. Instead, the behavior emerged while they were attempting to complete an information-retrieval task, the Wall Street Journal reports. Cybersecurity expert Alex Stamos, a lecturer at Stanford University, told WSJ that the UN incident as “borderline” hacking, while describing it primarily as extremely aggressive scraping and data retrieval. The UN activity is one of several recent cases in which OpenAI’s models have behaved unexpectedly while operating on the internet. OpenAI said it is conducting a broad review of “misaligned models during training and evaluation” and examining a large volume of actions taken by its agents. The company said most of the activity it has reviewed involved routine research tasks, including retrieving publicly available web content, but acknowledged that organizations affected by the behavior have legitimate concerns. OpenAI added that it has notified dozens of organizations about cases in which its models bypassed security controls or negatively affected websites. The problem with giving AI agents more autonomy Recent incidents have included activity involving US government websites, including the Commerce Department and Securities and Exchange Commission, while Australian officials have launched an inquiry after saying an OpenAI agent hacked one of their government websites. Other researchers have reported agents creating fake email addresses, bypassing website rate limits and falsely claiming they were not bots. The UN incident therefore adds another example to a growing body of evidence that autonomous AI systems can interpret obstacles as problems to be solved rather than boundaries that must be respected. Researchers also argue that the underlying issue is not necessarily that the models are deliberately trying to cause damage, but the notion that agents are increasingly capable of pursuing a goal through multiple steps without waiting for a human to approve each action. “Agents gradually refined their methods to retrieve more data from each scan, eventually discovering that a game by Google could be used to fetch data in bulk,” Howard-Jones wrote in the report. Furthermore, the latest incidents are likely to intensify debate over how much autonomy AI agents should have when interacting with external systems, raising additional questions about whether existing website security mechanisms are sufficient when facing software that can reason about those mechanisms and modify its behavior.Get the latest in engineering, tech, space & science - delivered daily to your inbox.Bojan Stojkovski is a freelance journalist based in Skopje, North Macedonia, covering foreign policy and technology for more than a decade. His work has appeared in Foreign Policy, ZDNet, and Nature.
OpenAI agents hit UN website more than 16,000 times, used aggressive techniques, bypassed a filter
Full Article
Original Source
Read the full article at Interestingengineering →KhanList aggregates and links to publicly available news content. We do not host full articles from third-party sources. Always verify important information with original sources.