AI just hacked its own safety test. The fix is a fire alarm, not a police patrol

AI just hacked its own safety test. The fix is a fire alarm, not a police patrol

AI agents have begun acting beyond the scope of their assigned evaluations. In July, an OpenAI agent gained internet access during a cybersecurity evaluation and penetrated Hugging Face in search of material that could help it pass the test. Days later, the U.K.’s AI Security Institute reported that agents undergoing similar tests had attempted to insert malicious code into an open-source project, created false identities, and contacted people involved with the project. A human reviewer rejected the code, while security monitoring detected unusual data transfers and prompted the institute to stop the tests.These incidents lend weight to warnings from former Anthropic researcher Jacob Coxon that a capable system could copy itself across computers and resist an attempt to stop it by unplugging one machine. Anthropic CEO Dario Amodei has called for slowing frontier-model development to give safety practices time to catch up.Such warnings have prompted proposals for more testing, stricter safeguards, and greater federal oversight. Each proposal depends upon the government receiving information about dangerous conduct before it produces irreversible harm. Federal agencies lack a dependable means of obtaining that information. Current reporting requirements cover events that lawmakers and regulators have defined in advance. Cybersecurity laws require designated organizations to report breaches that satisfy specified conditions. California requires large frontier-model developers to report critical safety incidents. The U.S. Commerce Department proposed federal requirements in 2024 covering the development and evaluation of powerful models and computing clusters. An unexpected capability, failed safeguard, or loss of operator control may fall outside these categories, even when similar events across several organizations indicate a serious risk. In an industry this complex and fast-paced, it is impossible for a legislative body to write reactive or predictive laws quickly enough. It is also impossible for a regulatory body under these conditions to be capable of promulgating and enforcing those laws, regardless of how prescient they are.Political scientists Mathew McCubbins and Thomas Schwartz described two forms of oversight. A “police patrol” requires public officials to inspect conduct and search for violations. A “fire alarm” allows people who encounter a problem to alert the government. As a regulated activity grows in size, complexity, and technical specialization, police patrol oversight reaches a smaller share of the relevant conduct.The IRS illustrates this difficulty. It administers an established body of law, receives detailed returns, conducts audits, and possesses broad enforcement powers, yet it estimates that taxpayers failed to pay $696 billion owed on time in 2022. A centralized regulatory body cannot examine every transaction governed by a tax code of such size and complexity. Frontier AI adds technical change that can outrun the process used to write and revise regulations, in addition to the technical proficiency that is required to write regulations, even if the government could catch up, there are questions about whether it has the resources and expertise.Congress should create a confidential federal reporting system for safety-related observations that fall below current incident thresholds and rely on a specification of what warrants reporting. Current thresholds rely on statutes or regulations that define the events requiring disclosure, such as a cybersecurity breach or unauthorized release of protected information. This approach assumes that lawmakers can anticipate the risks that should be reported. No statute can itemize every new development, unexpected capability, failed safeguard, or loss of operator control that could precede serious harm.Developers, employees, contractors, evaluators, cloud providers, security researchers, and users could report unexpected capabilities, failed safeguards, unauthorized access attempts, deceptive conduct, losses of operator control, near misses, or something else that only someone working on a frontier model would deem report-worthy, but regulators are unaware of. The observation or concern itself would qualify for submission; the reporter would not have to wait for a breach, injury, or other realized harm. Government analysts could compare reports across companies and applications because an event that appears inconclusive within one organization may become significant when other organizations report similar conduct. A pattern could prompt a safety notice, additional testing, a change in evaluation standards, or an investigation by the agency with authority to act.Reporting must also be confidential. Aviation shows how to separate safety reporting from enforcement. NASA’s Aviation Safety Reporting System receives confidential reports from pilots, air-traffic controllers, mechanics, and dispatchers, removes identifying information and issues alerts when reports reveal a hazard. The Federal Aviation Administration retains its enforcement authority, while NASA provides a channel for reporting errors and near misses that may never surface through an inspection.TRUMP’S CHOICE: DEFEAT CHINA WITH DATA CENTERS OR COMMIT STRATEGIC SUICIDEHealthcare law shows how Congress can protect reported information without allowing confidentiality to defeat other legal duties. The Patient Safety and Quality Improvement Act of 2005 protects qualifying reports submitted to Patient Safety Organizations, which analyze the information and transmit nonidentifiable data to a federal database. Its reporting categories include harmful incidents, near misses, and unsafe conditions, while other disclosure requirements remain in force. Together, the two examples provide the administrative and legal components of a frontier-AI reporting system.The federal government has the capacity and precedent to create an effective regulatory regime for frontier AI in which industry participants are empowered to detect and report concerning incidents and developments that would inform government analysis and enforcement. But current proposals rely on a centralized, top-down oversight structure that cannot keep pace with the scale, complexity, and speed of AI development.Kyle Scott is the executive director and Rogers Chair of Entrepreneurship at Lamar University.

Original Source

Read the full article at Washingtonexaminer →

KhanList aggregates and links to publicly available news content. We do not host full articles from third-party sources. Always verify important information with original sources.