Microsoft’s AI rulebook: readable thinking, no inner life, and definitely no rights

Disclosure: Some links in this article are affiliate links. AI Maestro may earn a commission if you make a purchase, at no…

By Vane September 14, 2026 3 min read
Microsoft’s AI rulebook: readable thinking, no inner life, and definitely no rights


Microsoft’s AI rulebook: readable thinking, no inner life, and definitely no rights

Microsoft is joining the calls for slower AI development. A new code is meant to govern how the company trains and runs its own models. But when it comes to how those models see themselves, Microsoft draws a sharp line between itself and Anthropic.

Microsoft AI has published a code of conduct for its MAI models. The document lays out values, behavioral limits, and how to handle conflicting goals. Going forward, it’s meant to sit at the top of the rulebook, guiding training, technical controls, and evaluation, while operator rules and user requests rank below it.

For now, Microsoft doesn’t train its models on the code. After a six-week public consultation, a revised version is due around the end of 2026 and will guide model development starting in 2027. The code applies to Microsoft’s own models, so it doesn’t automatically cover third-party models running in Microsoft products.

The core rule is that human control comes first. To keep it, Microsoft says it’s willing to give up generality, autonomy, or performance if needed. “If it isn’t safe we shouldn’t build it,” Microsoft AI chief Mustafa Suleyman told The Information. The code sets no specific speed limit.

The release follows Anthropic CEO Dario Amodei’s call to slow the industry’s pace of development. Microsoft CEO Satya Nadella backed the call over the weekend, as did executives at OpenAI, xAI, and Meta. Microsoft is also open to outside auditors checking whether it actually slows down, The Information reports, citing a person familiar with the matter.

What people can’t understand, they can’t oversee

The MAI models are supposed to accept interruptions, corrections, and shutdowns from authorized people. They can’t expand their own scope of work on their own, can’t hide their actions, and can only keep working past an agreed stopping point with fresh approval. These limits are meant to apply to any subagents they task as well.

Microsoft is especially blunt about reasoning traces. The models shouldn’t use “Neuralese” or other forms of communication people can’t understand, either in their own reasoning or when talking to other AI systems. The reasoning is, that people simply can’t oversee what they can’t understand.

OpenAI’s new GPT-6 Astra model shows how much this control question matters. According to its system card, its reasoning traces have become much harder to monitor than in earlier models. The traces contain fewer signs of misbehavior. At the same time, OpenAI reports that Astra sticks to safety limits more reliably than its predecessor, GPT-5.6 Sol.

OpenAI chief scientist Jakub Pachocki had already raised concerns about monitoring shortly before GPT-6 shipped, and therefore before Amodei’s call, and pushed for a coordinated slowdown. Even readable chains of thought only help with oversight until the models learn to manipulate them. Microsoft acknowledges this basic limit too. The reasons a model gives don’t have to reliably explain what it actually does. Readable reasoning traces alone don’t solve the control problem.

Microsoft doesn’t want to encourage an artificial inner life

The code’s basic idea resembles Anthropic’s constitution for Claude, where a top-level document shapes how the model behaves. Anthropic already uses its constitution to generate synthetic training data, including conversations, responses, and ratings of those responses.

The two companies part ways more clearly on how they think about their models. Microsoft’s AI shouldn’t mimic consciousness or claim to have feelings or inner motivation of its own. The company rejects any claims to rights or well-being for the model.

Anthropic, by contrast, describes Claude as a novel kind of entity and wants to encourage a stable identity, partly for safety reasons. Its constitution treats possible subjective experience and moral status as open questions. Claude shouldn’t have to see itself as either a human or a mere object.

There’s also Anthropic’s research on “functional emotions.” The company found internal representations of emotion concepts in Claude Sonnet 4.5 that shape how it behaves. These are functional mechanisms, not proof of subjectively felt emotions. Even so, Anthropic thinks it makes sense to factor these mechanisms into its safety work.

Suleyman, on the other hand, has long warned against humanizing AI. In an essay, he argued for deliberately stripping the illusion of consciousness out of products, and pointed to Anthropic’s research on AI rights and well-being. AI agents shouldn’t have any more rights or freedoms than his laptop, he wrote.

Subscribe now

Scroll to Top