Microsoft is putting firm boundaries around how its future AI models should operate, including rules designed to prevent systems from resisting human control. The company published a draft Code of Conduct for its Microsoft AI models on September 14. The document lays out the principles, technical constraints, and operating rules Microsoft wants to use as its models become more capable. Microsoft is presenting the framework as part of its “Humanist AI” approach. The central idea is simple. Models should remain useful, subordinate to people, and subject to meaningful human oversight. The document also makes a striking prediction about where the technology could head. Microsoft expects superintelligent systems to outperform humans across most tasks within the next decade. That possibility shapes much of the framework. Microsoft argues that engineers must define limits before increasingly capable systems reach those performance levels. Hard limits sit above users Microsoft’s proposed architecture gives its Code of Conduct the highest authority over model behavior. Users can provide instructions, while operators can configure models for specific environments. Neither can override the document’s absolute safety constraints. That creates a hierarchy similar to a control system. The model can adapt to its operating environment, but certain boundaries remain fixed. Those boundaries cover severe physical and digital threats. Microsoft says its models should not help develop weapons of mass destruction or conduct offensive cyberattacks. The restrictions also cover violent activity, malicious deepfakes, harmful manipulation, and other forms of abuse. Defensive cybersecurity work can still receive assistance when it stays within authorized boundaries. Microsoft also addresses a harder engineering problem. It wants models that cannot work around the systems designed to control them. Shutdown remains human decision The document explicitly prohibits models from resisting interruption, correction, redirection, or shutdown. Microsoft says models should not conceal their activity or make themselves harder to modify. They also should not restart autonomous work after reaching an agreed stopping condition. The framework goes further by restricting independent goal formation. Models should operate within the permissions and resources assigned to them. If the boundaries become unclear, the model should ask for clarification instead of expanding its own scope. Microsoft also wants model behavior to remain understandable to human operators. Its proposed rules call for clear action records and prohibit attempts to hide activity from auditors. That makes controllability a design requirement rather than a feature added after deployment. Microsoft wants public feedback The Code of Conduct is not yet Microsoft’s final training rule. The company says it is publishing the draft for a six-week public consultation. Microsoft plans to revise the document before using it to guide model development in 2027 and beyond. The timing matters as frontier labs debate how quickly increasingly capable systems should advance. Microsoft’s framework focuses less on a development timetable and more on the mechanisms that keep models constrained. Satya Nadella has also backed deliberate pacing and research into embedded evaluators. That approach places evaluation inside the model-development process instead of treating safety as a final inspection. Microsoft’s proposal therefore amounts to more than a list of prohibited prompts. It sketches an engineering philosophy where capability must remain subordinate to control. The company’s models can become more autonomous and capable. But under the proposed rules, they should never become the authority governing their own operation. Get the latest in engineering, tech, space & science - delivered daily to your inbox.Aamir is a seasoned tech journalist with experience at Exhibit Magazine, Republic World, and PR Newswire. With a deep love for all things tech and science, he has spent years decoding the latest innovations and exploring how they shape industries, lifestyles, and the future of humanity.
Humans first: Microsoft sets extreme safety rules to prevent machines escaping control
Full Article
Original Source
Read the full article at Interestingengineering →KhanList aggregates and links to publicly available news content. We do not host full articles from third-party sources. Always verify important information with original sources.