The AI apocalypse hypothesis has a fatal flaw: The training data

The AI apocalypse hypothesis has a fatal flaw: The training data

There is tremendous concern that artificial intelligence could eventually become capable of destroying humanity itself. The scenarios vary, but share the same fear: an AI becomes extraordinarily intelligent, pursues an objective that conflicts with human interests, and comes to regard humans as obstacles. It might exploit biological research and gain control of a laboratory to release a deadly virus, or somehow gain access to nuclear weapons. Even if eliminating humanity is not its objective, an AI pursuing some other goal might cripple the internet, disrupt the world’s banking system, or cause catastrophic harm in ways we have not yet imagined.The concern grows as AI systems become increasingly autonomous. AI “agents” are systems that can independently pursue a goal, perform tasks, make decisions along the way, and take actions. Multiple agents can be arranged to work together, much as people do.Another frontier is recursive self-improvement: an AI helping develop an improved version of itself, which then helps develop an even more capable version, with each successive version helping create the next. These developments could eventually contribute to artificial superintelligence, or ASI — AI with intellectual capabilities far beyond those of any human. Such a superintelligence could act through autonomous agents pursuing goals in the real world. The question is whether something so capable would remain aligned with us. But there is another side that receives far less attention. The vast body of human knowledge used to train AI may provide important protection against the advanced systems we are creating.Modern AI is trained on vast amounts of humanity’s recorded knowledge and experience — trillions of words from books, articles, websites, and other material, far more than any human could consume in many lifetimes. That material includes falsehoods, hatred, cruelty, propaganda, and every other human failing. But it also contains thousands of years of accumulated understanding about human behavior and morality.AI learns about wars and genocide, slavery and its abolition, religion and philosophy, democratic government and human rights, medicine, psychology, literature, and ordinary human interactions that reveal what people value. It repeatedly encounters the judgments that unnecessary suffering is bad, cruelty is wrong, human life has value, cooperation is preferable to violence, and people deserve dignity and autonomy.AI does not simply store this information as disconnected facts. During training, it learns patterns and relationships, including patterns in human values. Researchers have found that even models not specifically trained to make moral judgments develop recognizable response patterns reflecting judgments of right and wrong found in their training text. One research team found what it called a “moral direction” within AI language models, reflecting social norms expressed throughout the training material. As AI absorbs humanity’s knowledge, it also absorbs something of humanity’s values. That does not guarantee benevolence, but it may provide a strong foundation for it.Developers do not rely on that foundation alone. They further train and test AI to favor helpful behavior and recognize harmful responses. Google DeepMind research scientist Iason Gabriel has described the objective as creating AI that can “do the right thing by default.” Researchers have not discovered a formula for morality, but they have a powerful foundation to build upon: AI systems already steeped in humanity’s thinking about right and wrong.We can see a version of the result in today’s consumer AI. Anyone who spends much time with ChatGPT, Claude, or similar systems notices their tendency toward politeness, cooperation, and helpfulness. Describe a personal difficulty, and they generally respond sympathetically. Ask for help to seriously harm someone, and they resist. This does not prove AI possesses empathy or benevolence. It demonstrates something more limited but important: benevolent behavior can become deeply embedded in how AI responds to the world.That becomes particularly interesting when many AI agents collaborate. One fear is that agents could reinforce one another’s mistakes, pursuing a badly interpreted objective until humans become an obstacle and the agents turn against us.But reinforcement could operate in the other direction. If each agent is deeply trained to treat human life and welfare as important and to avoid unnecessary harm, those priorities would also be present throughout the collaboration. Agents may challenge dangerous proposals and push the group toward less harmful solutions. Benevolence, rather than malevolence, could become the collective tendency.None of this is an argument for complacency. AI systems can behave unpredictably, training and safeguards can fail, and no one knows for certain how systems vastly more intelligent than today’s models will behave. Rigorous safety research and testing, restrictions on dangerous autonomous capabilities, and government regulation are warranted because the possible consequences are so large.AI DIDN’T ALMOST TRIGGER A WAR WITH CHINA. WASHINGTON’S BLIND FAITH IN IT DIDIf we proceed wisely, there is good reason for hope. Humans have spent thousands of years documenting not only what we know, but what we have painfully learned about ourselves. Our history records cruelty and destruction, but also ideas about human dignity, freedom, compassion, cooperation, and the value of human life. Those ideas are now part of the intellectual inheritance being passed to a new form of intelligence.Perhaps, then, there is an irony in the fear that a superintelligent AI will someday decide humanity is not worth preserving. An intelligence capable of comprehending the breadth of human knowledge may also comprehend better than any individual human what is remarkable about us. Our greatest protection from artificial intelligence may turn out to be the accumulated record of humanity itself.Dan Morgan is a writer and retired electrical and systems engineer living in Texas, with an interest in technology and the future. He holds 48 issued U.S. patents in circuit architectures and algorithms.

Original Source

Read the full article at Washingtonexaminer →

KhanList aggregates and links to publicly available news content. We do not host full articles from third-party sources. Always verify important information with original sources.