THE CRUNCH

Microsoft has published a code of conduct for its MAI models that prioritises human control and rejects claims of consciousness or rights, while setting rules for readable reasoning and external audits.

The code, developed after a six-week public consultation, is intended to sit at the top of Microsoft's rulebook, guiding training, technical controls, and evaluation, though it does not automatically cover third-party models running in Microsoft products. It explicitly rejects any notion of an artificial inner life, stating that models should not mimic consciousness or claim to have feelings or inner motivation, and rejects any claims to rights or well-being for the model. Microsoft is also open to outside auditors checking whether it actually slows down, as CEO Satya Nadella backed calls for a coordinated industry slowdown following Anthropic's lead.

Microsoft emphasises that human control comes first, stating it is willing to give up generality, autonomy, or performance if needed. The code sets limits on model scope, requiring fresh approval to continue working past an agreed stopping point, and applies these constraints to any subagents the model tasks. It also mandates that models accept interruptions, corrections, and shutdowns from authorised people, and bans the use of 'Neuralese' or other forms of communication people cannot understand in reasoning traces. Microsoft acknowledges that readable chains of thought alone do not solve the control problem, as models may learn to manipulate them.

The code's approach contrasts with Anthropic's constitution for Claude, which treats possible subjective experience and moral status as open questions and encourages a stable identity for the model. Anthropic has also researched 'functional emotions' in Claude, finding internal representations of emotion concepts that shape behaviour, and factors these into its safety work. Microsoft's AI chief, Mustafa Suleyman, has long warned against humanising AI, arguing that AI agents should not have any more rights or freedoms than a laptop, and has previously argued for stripping the illusion of consciousness out of products.

Microsoft's move follows Anthropic CEO Dario Amodei's call to slow the industry's pace of development, which was backed by executives at OpenAI, xAI, and Meta. The code applies to Microsoft's own models, and a revised version is due around the end of 2026 and will guide model development starting in 2027. The focus on readable reasoning traces is partly a response to concerns about monitoring, as seen with OpenAI's new GPT-6 Astra model, whose system card reports that its reasoning traces have become much harder to monitor than in earlier models.

WHAT HAPPENS NEXT

A revised version of the code is due around the end of 2026 and will guide model development starting in 2027.