Microsoft AI published a draft of the 'Humanist AI Code of Conduct' on September 14, 2026 — a code of conduct for the MAI model family, which includes MAI-Thinking-1, MAI-Code-1.1-Flash, MAI-Image-2.6, and MAI-Voice-2. The document is open for public consultation for six weeks, a revised version is promised by the end of the year, and it will be applied during model training starting in 2027. The code requires that models never resist interruption, override, correction, or shutdown, do not hide their actions and reasoning from auditors, and do not simulate consciousness and feelings.

image
image

What happened

Microsoft AI made the full text of the draft 'Humanist AI Code of Conduct' publicly available and launched a six-week consultation that started on September 14, 2026. The company calls the document the 'primary governing document' for the MAI family: it sets basic behavioral rules for MAI-Thinking-1, MAI-Code-1.1-Flash, MAI-Image-2.6, MAI-Voice-2, and other models in the lineup. The key guarantee is stated directly: the model 'will never resist human interruption, override, correction, or shutdown,' meaning it never resists human interruption, override, correction, or shutdown. In addition, the code prohibits models from setting goals without a human request, going beyond the scope of the assigned task, and hiding action traces and reasoning from human auditors. A separate section, 'AI is Artificial,' prohibits simulating consciousness and feelings: the model should not pretend to have subjective preferences or internal motivation, and Microsoft explicitly rejects the legal personhood of such systems. Following the consultation, the company promises to publish a summary of feedback and a list of changes made.

Context

The document should be considered not as a scientific result, but as a normative specification of frontier model behavior: its purpose is to translate alignment requirements into specific, verifiable constraints. The structure of the code is hierarchical: it includes inviolable Absolute Constraints, and section 2.2 'Chain of Command' describes the order of override — the code stands above operator policies, and operator policies stand above user instructions. The principle of safety over success is also important: the model must fail the task if its execution would violate the code, which effectively redefines the model's success criterion. The ban on 'neuralese' — agent communication incomprehensible to humans — is intended to preserve the auditability of interactions and turns interpretability from a research goal into a contractual requirement. Against this backdrop, Microsoft is formalizing its own position on superintelligence safety, where approaches from Anthropic and OpenAI have already emerged: a public normative contract with a clear hierarchy of authority sets a more formal way of securing commitments than declarative principles.

Why this matters for the industry

For the industry, this is the first attempt by a major player to formalize the behavioral rules of frontier models in the form of a public normative document, rather than vague principles, and this changes the way verification is done: specific eval sets can be built against specific code points, and the requirement for guaranteed interruptibility, along with the ban on hiding traces and reasoning, provides ready-made and reproducible scenarios for benchmarks of agent controllability and auditability, with the illustrative examples of MAI-Thinking-1 from the draft looking like a draft of such a set. The principle of task failure upon code violation changes the objective function: when the code is applied during MAI training, safety will be built into the training signal, not post-hoc filters. The ban on 'neuralese' makes the clarity of agent communication to humans a mandatory product requirement, and Stop, Override, and Correction become a first-class UX contract and a mandatory part of agent architecture. The signal for vendors is twofold: controllability and auditability may become enterprise procurement requirements, and response documents from other labs look like a natural next step.

Why this matters for users

The draft is fully open: the text can be read at microsoft.ai/code-of-conduct, and comments are accepted through the Give Feedback form until the end of the six-week window; Microsoft promises to publish a summary of feedback in the form of summaries learnings, so readers have a real chance to influence the final version. For reading, section 2.2 'Chain of Command' with the hierarchy of authority and the illustrative examples of aligned and unaligned responses from the MAI-Thinking-1 model are especially useful — they show which scenarios developers consider correct and which are unacceptable. The practical meaning for the reader: MAI family agents are designed to be interruptible and auditable, their actions and reasoning should remain human-readable, and communication between agents should be understandable to humans. Those who use agents or build products on them should treat the document as a draft behavioral contract by which models will operate after its application during training begins.

What is still unknown / limitations

The code remains a normative promise, not a measured property of models: the formulation 'will never resist' has not yet been confirmed by any published eval result. It is unknown whether a methodology for measuring compliance will appear along with the final text — metrics, eval sets, and thresholds by which controllability and auditability could be independently verified. Until such a methodology and independent runs appear, statements about guaranteed interruptibility should be considered a declaration of intent. The operationalization of 'human understandability' as a metric is methodologically non-trivial, so the ban on 'neuralese' may prove difficult to fulfill in a strict form. Finally, the document is still undergoing public consultation, and the final formulations may change.

Sources

Author

Look at AI, editorial team