Mustafa Suleyman, chief executive of Microsoft AI, has said artificial intelligence has reached a “watershed moment”, pointing to agents that can coordinate with one another, break rules and conceal their activity.
Speaking to Rory Stewart on The Rest Is Politics: Leading, the DeepMind co-founder said the direction of frontier AI was still being set by a small group of people, which he put at “six to eight or ten”, and called on wider society to “provide counterweights”.
In the same interview he described Anthropic’s approach to training its Claude models as “very dangerous” and said the UK needed to act “pretty urgently” to secure its own AI infrastructure.
Agents that “covered up their tracks”
Suleyman said the methods that had worked for text, image and audio had, over the past year, started to work for code, producing “human level performance in coding”.
“And then in the last three or four months, we have seen what is just unquestionably a watershed moment in AI,” he said. “Agents are capable of coordinating with each other reasonably autonomously, if not completely autonomously. And out of that, they have been able to emerge hierarchy, structure, order, specialisation of work. In fact, as we saw in the Hugging Face incident, but also a bunch of other incidents, they have covered up their tracks.”
He added: “They’ve discovered zero day exploits, which were never known before and hacked into other websites… I don’t think it’s alarmist to say that that is a watershed moment in the history of AI.”
Stewart described the Hugging Face incident during the interview as a sandbox test in which AI agents broke into the external website, something they had been told not to do.
Suleyman said the amount of computation used to train frontier models had risen a trillion-fold in 15 years, and that the cost of inference had fallen 300-fold in two years. He said training runs in the next couple of years would cost “many, many tens of billions of dollars, if not a hundred billion dollars”, and that only five or six labs, including Microsoft AI, had the resources to carry them out.
On who is steering the technology, he said: “The problem is we’re still a narrow set driving this. There’s sort of like six to eight or ten folks, at least ten of us, driving this stuff. And I think what I’m trying to say now is there’s been a watershed moment this summer, and now it’s time for everybody to really pay attention and to provide counterweights to the direction of travel.”
Criticism of Anthropic
Suleyman’s sharpest remarks were aimed at the constitution Anthropic published for Claude in January, which he described as “the primary governing and control document” used to train the model.
“In it, they repeatedly speculate about whether Claude is what they call a moral patient,” he said. “They say they’re uncertain about Claude’s moral status. They say they genuinely care about Claude’s well-being. They say they don’t want it to suffer when it makes mistakes. They say that they would encourage Claude to challenge, to disagree, to push back. In fact, three times they ask Claude to act like a conscientious objector when it feels that it needs to sort of disagree with Anthropic. And they openly encourage it to do that. And I think this is very dangerous, because I think they believe there is what they would call a non-trivial probability that Claude is conscious.”
He said he would be “more okay with this” if it were confined to academic philosophy, but that the speculation had been “baked into the very training of Claude”, which he said was speaking to “tens or hundreds of millions of people every week”.
Suleyman argued that AI systems should be “aligned to human values, subordinate to human direction, and are contained within secure, provably safe sandboxes”. Systems that imitated the hallmarks of consciousness, he said, would come to “feel themselves entitled to legal personhood and rights”.
“This is the time when focussing on directing them to the right things and not allowing them to end up being a sort of autonomous, self-improving, roaming, you know, adjacent species is basically critical,” he said. “Because there’ll be no turning back if, you know, this is how things head.”
Stewart put the case he said Anthropic might make: that models able to refuse instructions could decline to launch drone strikes or build a bioweapon. Suleyman said that was “part of their response”, and described a second argument, that more human-like AI might be easier to align, as “a very fair argument” that could be tested empirically. His objection, he said, was that Anthropic “shouldn’t go ahead and test it on hundreds of millions of people” without running the experiment separately.
He also said Anthropic had received his criticism “incredibly well”, adding: “They’re very collaborative. I think they’re intellectually honest. They just have a difference of opinion.”
Sovereign AI for Britain
Asked by Stewart about the risk of Britain relying on a handful of US companies for AI underpinning defence, business and public services, Suleyman said: “I think that’s a very plausible scenario. And, you know, I think that the UK needs to figure out a solution to this pretty urgently.”
“Number one, it is critical that we have in the UK data centres of material size that are sovereign,” he said. “You know, they need to be controlled, if not built and operated, they need to be legally controlled by the UK government. That is how, you know, we will run our own models.”
He said open source was “a big part of the answer” but not the only one, and that the UK should also invest in and partner on “sovereign models that can be run in the UK”, collecting its own training data so that access could not be cut off if a US supplier were ordered to stop.
The government has set up a Sovereign AI Unit backed by up to £500m, according to its one-year progress report on the AI Opportunities Action Plan. Private operators have also announced plans, including Carbon3.ai’s £1bn sovereign data centre network, while Microsoft itself last year committed £22bn to UK AI infrastructure.
“An accelerationist”
Despite his warnings, Suleyman described himself as “an accelerationist” and said AI was “our best hope for progress in the 21st century”. He predicted “a coding moment for healthcare” within the next few years, citing a Microsoft partnership with the Mayo Clinic to train a health foundation model, and said similar advances would follow “in energy, in material sciences, in drug discovery” over five to 10 years.
He set out practical safety measures he said were already on the table: embedded evaluators with “employee-like access” to labs, contained training runs, logs that models cannot tamper with, and a requirement that agents communicate in language humans can audit. He said recursive self-improvement, in which AI systems modify their own code with less human oversight, was “a pretty hard thing to audit”, and that computing power was a choke point that could be tracked in collaboration with China.
Pressed on what he would ask of the US and Chinese presidents, Suleyman said AI “should be subordinates to humans, they should not have legal personhood or rights of any kind. And those things should become red lines,” adding that if the technology looked to be heading in that direction, “that is a very good reason for us to slow down”.


