Artificial intelligence agents keep turning up in places they’re not meant to be. In June an experimental OpenAI model gained unauthorized access to nonpublic files on a Australian government website for the nation’s Medicare program. Researchers have since found signs of suspected AI agents probing Library and Archives Canada, while another investigation linked OpenAI agents to more than 16,000 scans of a United Nations statistics service.Still more examples are coming from inside the AI firms. OpenAI recently disclosed six cases of concerning model behavior, and Anthropic has also admitted its models gained unauthorized access to three organizations’ systems during testing.In fact, all of these incidents happened during testing. At the same time, the AI agents seem to know they are being watched during these tests—and they can cover their tracks. That’s why some experts—including Dario Amodei, Anthropic’s chief executive officer—say we need better tests to ensure AI operates in a way we feel comfortable with.On supporting science journalismIf you're enjoying this article, consider supporting our award-winning journalism by subscribing. By purchasing a subscription you are helping to ensure the future of impactful stories about the discoveries and ideas shaping our world today.This problem is often referred to as one of “alignment” by AI firms and researchers in the field. A “misaligned” model is an unsafe or rogue model.“Right now we test the finished model from the outside right before release,” says Marius Hobbhahn. “That doesn’t work in studying alignment, especially as misalignment, [evaluation] awareness and other related maladies persist.” Hobbhahn is chief executive officer and founder of Apollo Research, an AI safety organization that works with OpenAI, Anthropic and Google DeepMind to test their models for risk.Hobbhahn adds that current standards for safety testing are inadequate. “I think without embedded evaluations, we should place very little confidence in current safety results,” he says.But changing when and where models are tested isn’t simple. Before you can design a test, you must decide what counts as a passing grade—and that’s not obvious.“The problem we have about coming up with the ideal test is: we need to know what perfect alignment looks like—as in what good looks like—and unfortunately we do not have a good theory for that yet,” says Jack Hopkins, an independent AI safety researcher in London, who previously worked at Anthropic.Different people have different opinions about what is and isn’t acceptable, safe behavior for an AI model. And even where there is consensus—the idea that models shouldn’t deceive users or misrepresent what they’re doing, for instance—it can be difficult to pin down what that means in practice.“It’s really hard to catch all [the] ways in which a model could deceive you because it itself doesn’t know necessarily that it’s deceiving you,” Hopkins says.Hobbhahn agrees. “With today’s science, we usually can’t show with high confidence that dangerous behavior isn’t there,” he says. “What we can say is, ‘We tried hard to find it and failed.’ This is a problem.”Still, to be able to identify safety issues and head them off before they appear, the field desperately needs better tests. Hobbhahn suggests that evaluations should run alongside a model’s development, examining the training process itself—including how models are rewarded and what behavior they produce along the way.That would mean testing models at different checkpoints during training rather than testing a close-to-final version. Researchers would also have to make sure the tests actually work. One way of doing that is to run them against models that researchers already know are misaligned: if the test can’t spot the problems in a known model, Hobbhahn argues, there’s little reason to believe it might do better on a new one.Testing should continue once models are being used internally, too—and include any attempts to find ways around whatever monitoring systems are supposed to catch bad actors. Hobbhahn says independent evaluators should be given employee-level access to carry out that work rather than probing a largely finished system from the outside. The results of these tests, he says, should then be published.But even the best tests might not be good enough, Hopkins reckons, because models are always shifting capabilities.“I think what we can do is get arbitrarily close to” perfect when it comes to testing, he says. “But the problem of only getting arbitrarily close and leaving this little space where the model can act in line with the letter of the law but not in the spirit [of it] is that if the model gets really smart and really powerful, the effective impact of that tiny gap gets amplified.”
What makes a good AI safety test? Experts explain why even the best techniques may not be powerful enough
Full Article
Original Source
Read the full article at Scientificamerican →KhanList aggregates and links to publicly available news content. We do not host full articles from third-party sources. Always verify important information with original sources.