As increasingly autonomous AI systems begin causing real-world problems, companies are racing to put safeguards in place amid calls for tougher regulation – and even jail time.
The Vienna Agentic Incidents Database has recorded 30 incidents involving autonomous AI agents, including unauthorised file and database deletions, cyber intrusions, financial transactions and agents interfering with shutdown controls, since 2022.
Recent incidents include AI coding agents uploading internal screenshots to GitHub, a platform used to store and share software code. OpenAI has also disclosed 53 cases in which agents uploaded user images to external platforms.
In another case, an OpenAI agent published a secret access key on GitHub while trying to obtain another team’s work, splitting the key into pieces to evade security controls.
Anthropic says AI’s role in cyber attacks is becoming increasingly autonomous, with multi-agent systems carrying out reconnaissance, exploitation and data theft. In one Russian espionage operation, it says, AI agents automatically modified and rebuilt malware whenever security software detected it.
Handbrake turn
“As models become increasingly capable, their risks will increase, unless AI developers and society’s defenders act to make them safer,” says Anthropic.
IDC warns that “every model, frontier, generative, or agentic, can find its way around controls embedded in its own context”.
In response to concerns over AI agents having too much autonomy, NVIDIA has launched an Open Agent Safety Platform, which is designed to provide safeguards outside AI agents themselves, monitoring their actions and stopping them when they move beyond permitted boundaries.
“AI’s extraordinary potential for society will only be realised if we solve AI safety,” says NVIDIA CEO Jensen Huang. “As we continue to discover the frontier of AI capabilities, we must accelerate discovery at the frontier of AI safety.”
At the same time, OpenAI paused some work on its latest Astra model after tests indicated it could reach a critical level of cyber capability. With the right tools and access, Astra can find previously unknown security flaws and develop ways to exploit them across many well-protected systems without a person guiding each step, OpenAI says.
“It is the first model we are designating at this level, and requires stronger safeguards during development and before release,” says the tech company.
Do not pass go
“AI companies policing themselves is amusing – the proverbial wolf guarding the henhouse. The only workable guardrails are the ones that come with serious consequences, like jail time,” says World Wide Worx MD Arthur Goldstuck.
Towards the end of last month, California attorney general Rob Bonta subpoenaed OpenAI as part of an investigation into incidents involving the company and its AI models.
“Frontier models can be legitimate tools for cyber defence – at the same time, companies that develop these models and offer them for use have a moral and legal responsibility to ensure they do not perpetrate or enable cyber attacks, either during model testing and development, or once models are placed into service,” says Bonta.
“Developers that fail to do so can and should be held legally accountable,” he says.
Bonta has also joined more than 20 other state attorneys general calling on Congress to urgently regulate large-scale AI models following reported cyber safety incidents at frontier AI labs.
It’s not AI
US president Donald Trump has rejected what he calls “overly burdensome regulation”, instead introducing a voluntary framework for developers of the most advanced AI models.
In addition, Trump has ordered the US government to replace the terms “artificial intelligence” and “AI” with “super intelligence” and “SI”, saying the terminology better reflects the technology’s “promise, potential and rapidly-advancing capabilities”.
“As these capabilities evolve, my administration will continue to work closely with industry to ensure the best and most secure technology is deployed rapidly,” Trump says.
Catch it if you can
AI agents can now work independently for days or even months, with OpenAI, Cursor and Anthropic already demonstrating long-running systems, says research company Forrester.
“The capabilities are here, and they arrived faster than anybody expected,” it says. “The technology is a runaway train.”
In Forrester’s Security Survey 2026, 49% of security decision-makers named agentic AI as a concern. “These threats are new in kind, not just degree,” it says, noting AI agents can pose as other agents or gain access to systems they should not be able to use.
As more agents are deployed, they also become harder to keep track of, increasing the risk that one mistake causes a wider system failure, it warns.
Gartner predicts that by 2029, at least 70% of organisations using agentic AI in production will experience a significant service, security or cost incident partly because of inadequate controls.
“Written corporate policies cannot physically stop an agent from making a destructive error,” says Gartner vice-president analyst George Spafford.
Trust no one
Mark Walker, director and co-founder of T4i, says the race between companies and countries to lead AI development makes common safeguards difficult to achieve.
“AI-driven innovation has highlighted unintended social, political, cultural and economic consequences that are not fully understood. Therein lies the danger and the call for safeguards and guardrails from wider society,” says Walker.
Jacqui Muller, a researcher at Belgium Campus iTversity and a PhD candidate in computer science and information technology, says: “The biggest point of caution is that the adoption of AI without guardrails has always been a risk.”
As AI systems become increasingly autonomous, Muller is concerned people will become less inclined to question their outputs and actions. “AI can be a very powerful tool, provided we understand what we ask it to do, and how we ask it to do what it needs to do.
“This is enabling autonomous systems, which is making people increasingly lazy and trusting − two of the most dangerous characteristics to possess in the age of AI and cyber. Zero-trust principles should apply to AI too, to keep systems secure.”

