Microsoft CEO Satya Nadella is the latest tech executive to offer lengthy thoughts on how AI security could be improved.
In an article published Saturday morning on X, Nadella wrote that it was time to “step back and evaluate the trust architecture” of AI.
“We cannot treat Super Intelligence as a set of nested black boxes and simply accept or reject its recommendations, responses, and actions,” Nadella wrote, using the Trump administration’s preferred term for AI.
As Nadella points out, this approach “means separating the model from the harness that orchestrates its work,” as well as “outsourcing controls and safeguards.” He also called for “every significant action of the model” to be documented with “tamper-proof, human-readable evidence” and for systems in which “an authorized person” always has the ability “to pause or stop a model mid-task.”
“We need to assume that a model is compromised and contain it from the start,” he said. “Think of it as an emergency brake.”
Nadella’s comments come as major AI companies increasingly acknowledge incidents in which they appear to lose control of their models, and after Anthropic CEO Dario Amodei released a plan for more cautious AI development.
Gn bussni

