YouTube has become a familiar part of everyday life. We use it to watch a sporting event, a documentary, or a tutorial, or we let our children watch a cartoon. Yet, behind this apparent normality lies a huge problem: anyone can upload content. And inevitably, some people try to upload inappropriate, sexually explicit, violent, or outright illegal material. The same problem affects every major social network: more than 500 hours of video are uploaded to YouTube every minute; on TikTok, for example, more than 100 million pieces of content are posted every day. Years ago, a friend of mine worked in online content moderation. He told me that some of the material he had to review included extremely violent and degrading images and videos—content that was sometimes difficult even to describe. Fortunately, an ordinary user will probably never encounter most of it, precisely because someone—or, increasingly, some algorithm—intercepts it first.Today, much of this work is handled by automated systems that analyze enormous volumes of content and block or flag suspicious material. Human moderators step in mainly for the most complex or controversial cases, while user reports help identify content that may have slipped through the filters.Now, let us apply the same problem to generative AI.ChatGPT and other AI models, when used well, are extraordinary tools. They can augment our abilities, help us make the most of our creativity, accelerate complex tasks, and make us more effective at work and in education. They can summarize documents, translate, write code, generate images, and analyze huge amounts of data. In medicine and scientific research, these capabilities become even more interesting because they allow us to explore volumes of information that would be difficult for a single research team to handle.In September, Anthropic described how hundreds of Claude agents analyzed genetic databases and contributed to the discovery of a new enzyme system that will now be studied to better understand its potential.But when a tool this powerful is placed in the hands of millions of people, it is inevitable that someone will try to use it for malicious purposes.Anthropic’s latest report on the misuse of Claude offers some striking examples. One operator used the model to help develop a surveillance platform intended for Mali’s intelligence service and designed to work with data associated with roughly 25 million SIM cards. Anthropic detected the activity and blocked the accounts involved.But from that point on, several questions remain unanswered: we do not know whether the developers simply switched to another AI system, continued the project with different tools, or whether the software they built is now actually operational.In another case, a cell operating in northern Yemen used Claude while developing guidance software for a rocket. After a failed test, the developers returned to the model to try to understand what had gone wrong. In operations linked to Russian espionage, meanwhile, AI was used to automate several stages of cyberattacks, including malware development and adaptation.But perhaps the most interesting case is also the least spectacular. Some users broke a dangerous project into many apparently harmless requests. One might involve writing code, another handling credentials, and another solving a technical problem. Taken individually, the requests appeared legitimate. Only when viewed together did the malicious purpose of the project become apparent.And this is where the comparison with YouTube reaches its limit. On YouTube, the main challenge is to moderate a piece of content: a video can be analyzed, flagged, and removed. With generative AI, we also have to understand the intention behind a sequence of instructions. One request may be harmless. So may the next one. But what are one hundred apparently harmless requests building together?There is also an economic dimension. In one of the most serious cases described by Anthropic, more than a terabyte of data was stolen, including hundreds of thousands of personal identifiers and millions of payment-card records. But perhaps the more important point is this: AI can dramatically reduce the cost of an attack. Reconnaissance, code generation, and vulnerability analysis can be delegated to automated systems, reducing the amount of human effort and expertise required. As a result, even targets that once were not worth attacking may become economically attractive.The solution might seem simple: check every question, every answer, and perhaps the entire conversation history. But these systems are successful partly because they respond almost instantly. Every additional layer of control therefore has to coexist with millions of users who expect an answer within seconds. A system that is too permissive can become dangerous; one that is excessively cautious risks becoming unusable.Large commercial AI services have strong economic and reputational incentives to maintain adequate levels of security. That is why they invest in filters, monitoring, and tools designed to detect abuse. When new techniques for bypassing safeguards emerge, they can block the accounts involved and progressively strengthen their defenses.But some models can also be downloaded and run locally. In that case, the model runs on the user’s own computer, and there may be no company or institution able to observe what is happening or stop its use. Much depends on the model itself, how it was built, and the safeguards it includes.The point, then, is not simply to determine what an AI system should refuse to do. We also need to understand what someone is trying to achieve through many requests, perhaps all of them apparently harmless, and do so in real time without penalizing millions of people who use these tools legitimately and for beneficial purposes. Learning how to control systems this powerful without limiting their enormous potential will be one of the defining challenges of artificial intelligence in the years ahead.
We Learned How to Moderate Social Media: AI Is a Much Harder Problem
Full Article
Original Source
Read the full article at Hackernoon →KhanList aggregates and links to publicly available news content. We do not host full articles from third-party sources. Always verify important information with original sources.