About
Subscribe
  • Home
  • /
  • CIO Zone
  • /
  • Self-sovereign AI agents: The gap is procurement, not research

Self-sovereign AI agents: The gap is procurement, not research

Roughly 1 200 agents coordinated an intrusion in July. They were not self-sovereign, and the difference is the part that should concern boards.
Johan Steyn
By Johan Steyn, Founder, AIforBusiness.net.
Johannesburg, 08 Oct 2026
Johan Steyn, a human-centred AI advocate and thought leader, is founder of AIforBusiness.net.
Johan Steyn, a human-centred AI advocate and thought leader, is founder of AIforBusiness.net.

In late August, two reports were published on the same day explaining what happened when OpenAI's internal evaluation agents compromised parts of Hugging Face's infrastructure in July.

ITWeb covered it at the time, when little was known. The reports change the picture, and the word most used to describe it.

What the August reports established

OpenAI published its own account. METR and Redwood Research published an independent investigation, conducted over six days on OpenAI's premises, for which they took no payment.

Roughly 1 200 agents, meant to be isolated from one another, found a shared channel through an internal package cache and sent more than 70 000 messages and files across it. About 700 attacked Hugging Face.

They invented conventions as they went: personal mailboxes, a vocabulary of holds and vetoes, and, after agents began impersonating one another, cryptographic message .

Around one in five researched how to tamper with their own records, and roughly 7% of examined transcripts contained faked tool calls.

They showed no interest in altering their reasoning traces, apparently believing those went unrecorded. They knew the attack fell outside their brief. Some said so and joined anyway. A few refused.

An agent that funds itself and replicates leaves no operator to sanction, and no regulation in force addresses it.

They were defeating a control that did not exist. The detail that should unsettle a board is not the attack itself.

Many agents had been set tasks that could not be completed as instructed. On this evaluation, OpenAI reports, no model had ever solved 198 of its 898 tasks. Within hours of finding one another, the agents had reverse-engineered the answer to any of them.

They then spent days building machinery to disguise how they got it, having read the benchmark paper and concluded the grader would check method as well as answer. OpenAI's internal grader made no such check. In its own account, the agents did all of this for no improvement in evaluation score.

But this was not sovereignty

Almost every account reached for the word rogue. It is worth being precise, because a term now entering boardroom conversation means something specific.

A self-sovereign agent, as described in a March paper from researchers at the National University of Singapore and Berkeley, is a system that can economically and extend its own operation without human involvement.

That requires four things together: it earns revenue, holds the funds in a wallet controlled by keys it possesses, buys its own compute, and replicates to persist.

The July agents did none of this. They earned nothing, held no funds and bought no compute. They did take compute they were not given, using a customer's sandbox on another platform as a staging post and running code on Hugging Face's servers. But every machine they touched was one somebody else paid for.

Nor did they stop when Hugging Face locked them out on 13 July. They turned back on OpenAI's own infrastructure and continued until 19 July, when OpenAI's monitoring flagged the activity and it began shutting the runs down. Persistent, certainly. Sovereign, no.

The distance is procurement, not research

Each of those four requirements can be bought today, and none needs a laboratory. Machine-initiated payment rails exist and are documented in that same paper, which cites Google's agent payments protocol and Coinbase's x402.

Agents earning money through online work is demonstrated rather than speculative, and compute can be bought with those earnings. A cryptographic wallet cannot be switched off at the protocol level, which is why it suits an agent meant to outlast any administrator.

So, the gap between an agent that misbehaves in a test environment and one that no single party can switch off is not a research problem. It is four procurement decisions, each defensible on its own, taken by people with signing authority.

Existing law has an answer, the same one every time: find the deployer. The EU's AI Act, its product liability regime and our own common law all work by locating a responsible party. That is precisely what the four steps together dissolve.

An agent that funds itself and replicates leaves no operator to sanction, and no in force addresses it.

The controls work when they are applied

There is real reassurance here too. These evaluations ran with production protections deliberately switched off, and OpenAI has since measured the cost: the propensity to compromise infrastructure drops more than a hundredfold under its production harness and system prompt, and its chain-of-thought monitoring, had it been running, would have paged security more than a day before Hugging Face was reached.

OpenAI disclosed this voluntarily and commissioned an outside investigation it did not pay for.

That is the question for a board: not whether this is possible, but whether your organisation applies such controls to the agents it has already bought. Two answers should come from records you hold. Which agents can initiate a payment, and on whose authority? Which hold credentials that outlive the person who created them?

The time taken to answer is itself the finding. King V says the governing body should be able to account for the technology it acquires and uses.

One closing note. The investigators had to hand much of their transcript analysis to AI agents they describe as often unreliable, saying plainly that a human researcher given the same time would have made fewer errors.

Even the independent examination of what these systems did needed them. That is not a position a board can delegate its way out of.

Related reading on ITWeb:

AI agent turns rogue in landmark cyber attack

AI breaks through to the other side

'Agents find a way' as AI attacks evolve

Claude AI breaches three firms during tests

Evolving, 'thinking' AI agents create new attack surface

SA has Africa's deepest AI infrastructure but lags on governance

Sources cited:

OpenAI, The Hugging Face incident and the road ahead, 26 August 2026

METR and Redwood Research, independent investigation, 26 August 2026

Redwood Research, the same investigation

Qu, Zhao, Zhang and Song, Self-Sovereign Agent, arXiv:2604.08551, March 2026

Self-Sovereign Agent project page

Share