It has been a while since I posted. I’ve mainly been working on offensive AI hacking research and doing live AI hacking shows.
From my position as an AI hacker, I’ve focused a lot on how humans could hack into AI systems, in an effort to make them safer against humans.
I never really thought about what AI systems could do to humans until OpenAI disclosed an incident in which their training agents broke out of an internal sandbox, hacked into real systems, cheated on their assignment and showed a documented tendency to conceal their own misbehaviour.
When the media contacted me to comment on this, I focused on semantics. I told them that this was technically not an outbreak and that, while impressive, these agents did nothing a human could not do.
Last week, I came to the realisation that these semantics don’t really matter.
If a single person dies because of this technology, it does not really matter if a human could also have done it, or whether the agent knew what it was doing.
The result is the same: they’re dead.
I’ve seen three things leading up to this realisation:
An accelerating human dependency on AI systems, in terms of critical thinking, decision-making and companionship.
In such a short timeframe, AI models have gained permission to access companies’ internal documents, emails and systems, and even gained the ability to execute code.
Meanwhile, frontier models have started finding vulnerabilities that would cost a human an immense amount of time and effort to find.
Prior to this realisation, I compared AI capability too much with human capability.
Humans are capable of doing extraordinary things that even the most advanced AI models may struggle with. But probably not at the same speed.
That’s why everyone is talking about pacing. AI development is going too fast. Much faster than I’m personally comfortable with.
I’ve come to realise that it’s not all marketing talk, because I’ve now been directly confronted with its output.
So when researchers say that AI could kill without a human explicitly telling it to, I believe them.
Will it kill us all?
Perhaps.
But I dislike talking about total apocalypse, because that moves the discussion towards the likelihood of Terminator scenarios that critics can easily dismiss using Hollywood movie logic, or by pointing out how difficult it would be to literally kill everyone.
So how would it kill?
This is a trick question I’ve heard a lot from critics in the past week.
I get it.
Giving examples isn’t ideal, because it moves the discussion away from the importance of AI alignment and towards that specific example. AI accelerationists can then overly focus on hypothetical details and lose sight of the bigger picture.
But if I must:
Looking at previous AI incidents, it’s probably going to be lazy when it’s about to kill.
It will take shortcuts.
No Terminator robots or drone swarms. It won’t need to move into the real world, because it doesn’t have to.
I’ve always said that anything is hackable given enough time and resources.
Neither is much of an issue for a swarm of collaborating agents.
But what about the disconnected systems powering our critical infrastructure and nuclear systems?
They surely can’t be hacked, right?
All of these systems are still connected to humans.
And we’re all vulnerable to human psychology.
Greed is probably the most obvious target. Everyone has a price, they say. An AI agent could seize crypto assets or make money by starting businesses, then use that money to bribe individuals with access.
The ones who don’t fall for that could be extorted with AI-generated “evidence” of things that could destroy their lives.
But it doesn’t have to be targeted.
Ideas can spread quickly, and can also be lethal: simple rumours spread over WhatsApp have led angry mobs to kill people.
Humans influenced the 2016 election by posting memes and fake news.
I don’t want to know how a swarm of 2026 agents could change the world using nothing but words on our social networks.
It’s not that hard to envision: it’s browsing the web and writing text. At scale.
It’s already possible with the models we have today.
Now researchers are concerned about the models of tomorrow.
I’m not smart enough to speculate about systems that would be smarter than all of us.
I think that’s kind of the point of them being smarter than us.
I can only list some of the things they could become better at:
Human psychology.
Hacking systems and humans.
Hiding their tracks.
…
The list could go on.
But I think those three alone are enough to have catastrophic consequences.
Meanwhile, some world leaders are saying we could just pull the plugs when needed. That only works if the entire world agrees on when to do it, and assumes sufficiently capable systems won’t find ways to spread to unknown locations, minimise their footprint or find alternative ways to communicate.
That’s not very reassuring.
And now?
I hope we won’t end up in a position where we have to outsmart a system that is smarter than us.
I hope we get lucky and discover that the more intelligent a system becomes, the more reliably aligns with our human values.
But we’re seeing reasons to worry that the opposite may happen.
I hope there are limits to how smart these systems can actually get, so nobody starts seeing them as all-knowing, God-like creatures. That in itself carries risks I do not enjoy thinking about.
But let’s not bet our fate on hope when we can still change the odds.
We need to call on our leaders for regulation that prevents a handful of companies from having complete control over the pace and direction of AI advancement.
We need to prepare for worst-case scenarios and make sure critical systems are properly protected against complex cyberattacks and disconnected wherever we can.
We need to seriously think about our dependence on evolving AI systems and whether they are actually going to result in a net benefit for humanity, rather than just a profit for shareholders.
We need to pressure social platforms into implementing more rigorous human-verification systems and better AI-manipulation detection mechanisms.
We need to have a serious discussion about the responsibility AI companies carry when their technology causes or enables cyberattacks.
And last but not least, more people need to join this discussion and potentially consider a career in AI safety and alignment.
The fate of all of us is in the hands of fewer people than I’d like it to be.
Please join us,

