IT Home On October 11, Microsoft CEO Satya Nadella said that in an era where AI is developing rapidly, simple strategies such as using superintelligent self-monitoring are not enough.
Nedera stated in a long article on the evening of October 10, Beijing time, that traditional software systems have been widely deployed over the past few decades, and humans now have the tools and capabilities to track the behavior of specific code paths. However, AI models are more powerful than traditional software, but they are a black box. We deploy these complex intelligent system and model, access our most sensitive data, and give them the ability to perform critical tasks on our behalf.

Nadella believes that we should take a step back and re-evaluate the trust architecture of this new era. Model providers simply cannot outsource their responsibilities, and we cannot treat super intelligence as a nested black box—simply accepting or rejecting its suggestions, answers, and actions.We must establish closed systems that can observe its behavior, test its limits, and always accommodate its actions.
Nadella said that treating the closed and open weight models as internal risks is a way to build such systems. This is not because AI models must be malicious, but because anyone with sufficient capabilities that allows them to access important systems may make mistakes or be exploited. The frameworks for containment and control must take this into account.
Nedera mentioned that developers should design these systems around the observability principle. IT Home lists the specific principles as follows:
Model diversity: No model should be the only dependence for important results, nor should it be responsible for verifying its own work.
Observe everything: every meaningful model action must leave tamper-proof, human-readable evidence. Without observation, there is no trust. We need to be able to reproduce how the results were achieved without relying on the model for proof.
Verifiability: We need to continuously test the entire system, including failures, attacks, edge cases, and system changes, not just successful tasks.
Independent control: Organizations should be able to independently decide what the model can access and what actions it can take.
Independent auditability: The verification must be independent of the intelligent system being verified. No single model should control both the behavior of the system and the evidence required to determine whether that behavior is consistent with the original intent.
Containment: We must assume that the model has been damaged and control it from the beginning. It can be imagined as an emergency brake. Authorizers should always be able to pause or shut down the model during a task. More advanced models will require more advanced containment techniques, and we need to standardize this.
Event disclosure: When these systems fail or are breached, we need to promptly disclose information to the affected parties, and establish mechanisms to share the reasons for the errors, which controls failed, and how to prevent recurrence, as well as share experiences with the entire industry. This should include the implementation details of changing the behavior of the agents during operation.
