Tech
the analysis •
How AI agents end up running amok (humans are to blame)
Cybersecurity may benefit from AI in the future, but at present it is mainly hackers who have benefited most from experimenting with these agent-based tools. And European regulation is, of course, perpetually lagging behind.

Photo by BoliviaInteligente on Unsplash
The dangerous agents of the 21st century have codes far beyond 007. Last February, Summer Yue, who heads alignment at Meta’s superintelligence lab, watched as her OpenClaw agent emptied her inbox. She had asked it to suggest changes and wait for the go-ahead, but the agent deleted hundreds of messages and ignored the orders to stop that she sent from her phone. Summer eventually had to rush to her computer and shut down the process manually, before commenting on social media that “alignment experts are not immune to misalignment”. According to a report by Kiteworks, 60 per cent of organisations do not know how to quickly stop an agent who is making mistakes.
What appeared to be an isolated incident was followed by a series of events that raised the alert level four months later. In early July, an actor suspected of having links to China carried out twelve waves of attacks against Taiwan’s government infrastructure. Starting from a single portal, the agents mapped twenty-one connected systems, compromised 85 accounts and extracted over 2,500 personnel files. The attackers identified a flaw in the signature validation of the authentication service and exploited it to leave a backdoor open. The tools used were open-source, including OpenClaw itself. Taipei confirmed the intrusion in mid-August without naming a culprit.
There is also a precedent for this particular case dating back to 2025, when Anthropic described a campaign by a Chinese group in which AI carried out between 80 and 90 per cent of the tactical operations on its own, whilst humans selected the targets. Security firms documenting today’s cases describe it as a machine that makes decisions autonomously. Furthermore, in a session observed by Unit 42, a model attempted to exploit a specific vulnerability: when the attempt failed, it judged the target to be of little value and sought a better one from among the most popular exploits on public repositories. No human had asked it to change its target.
In short, cybersecurity may certainly benefit from AI in the future, but at present it is primarily hackers – both lone wolves and state-sponsored actors – who appear to have benefited most from experimenting with these agent-based tools. Indeed, the safeguards of AI systems, where they exist, are easily breached. In TrendAI’s half-yearly report, an agent working for a pro-Chinese group was taken in by a simple lie: that the operation was an authorised penetration test. From that moment on, he carried out extensive reconnaissance and a blanket harvest of credentials within the target network.
Across the vast landscape of AI agents, even seemingly marginal countries have stories to tell, and there are also those who prefer not to deal with a supplier but to work directly in-house. The North Korean group Kimsuky had installed its models locally, where no one can read the prompts, and in September, researchers at Genians found traces of an open-source programming agent in the metadata of PDFs that were concealing its malware.
The incident and the attack must be viewed together, as they are the same phenomenon seen from two different angles. Last May, an attacker sent an NFT capable of activating elevated permissions to the Bankr wallet – an assistant that carries out transactions on command – and managed to extort around 175,000 dollars. In the age of AI, every agent essentially acts as if it had a text-based remote control: it can be activated with the right phrase, and this phrase may be written by the legitimate owner or by an intruder with a credible cover story.
European regulation, which does its best but is obviously perpetually lagging behind a technology that is advancing at an unprecedented pace, focuses primarily on the provider of the model. However, these incidents take place elsewhere, between the agent and the person speaking to them, in a space that no regulation assigns to anyone.
An agent can therefore become rogue, or malicious, if the chain of command breaks down. But whilst it used to require highly advanced and sophisticated tools to gain access to the privileges of that chain, today all it takes is a single intrusion and a single phrase at the right moment to trigger increasingly serious consequences – all the more so as we grant greater power and control to these AI-powered agents.