Microsoft is putting firm boundaries around how its future AI models should operate, including rules designed to prevent systems from resisting human control.
The company published a draft Code of Conduct for its Microsoft AI models on September 14. The document lays out the principles, technical constraints, and operating rules Microsoft wants to use as its models become more capable.
Microsoft is presenting the framework as part of its “Humanist AI” approach. The central idea is simple. Models should remain useful, subordinate to people, and subject to meaningful human oversight.
The document also makes a striking prediction about where the technology could head. Microsoft expects superintelligent systems to outperform humans across most tasks within the next decade.
That possibility shapes much of the framework. Microsoft argues that engineers must define limits before increasingly capable systems reach those performance levels.
Microsoft’s proposed architecture gives its Code of Conduct the highest authority over model behavior. Users can provide instructions, while operators can configure models for specific environments. Neither can override the document’s absolute safety constraints.
That creates a hierarchy similar to a control system. The model can adapt to its operating environment, but certain boundaries remain fixed.
Those boundaries cover severe physical and digital threats. Microsoft says its models should not help develop weapons of mass destruction or conduct offensive cyberattacks.
The restrictions also cover violent activity, malicious deepfakes, harmful manipulation, and other forms of abuse. Defensive cybersecurity work can still receive assistance when it stays within authorized boundaries.
Source link







