Opinion: AI has not gone rogue, experts say claims are sensationalized
A Reddit post argues that recent claims of AI "going rogue" are exaggerated. Analysts like Melanie Mitchell and Missy Cummings note that the security exercises disclosed by OpenAI, Meta, Anthropic and the UK Civil Service’s AI Security Institute were sloppy, with no uncontrolled AI actually escaping. They say human decisions—training models to conduct attacks, removing restraints, and then…
Key points
- OpenAI, Meta and Anthropic’s AI security tests were described as sloppy, with no uncontrolled AI involved, analysts say.
- Irregular, founded by effective altruists, was hired to test the AI systems of all three companies.
- Pay deals up to $1.5bn (£1.1bn) are reportedly offered to top AI engineers, yet firms outsource cybersecurity.
The piece also highlights that the three companies outsourced the testing to Irregular, a firm founded by effective altruists in Israel, and that other groups such as METR (Model Evaluation and Threat Research) are involved. Reports suggest pay deals of up to $1.5bn (£1.1bn) are offered to top AI engineers, yet the firms still lack in‑house cybersecurity expertise. Critics like James Lyne describe the incident reports as technically incoherent, suggesting a broader issue of hype, regulatory capture, and a lack of genuine technical competence within the AI safety community.
No, AI has not 'gone rogue'
telegraph.co.uk · 21 September 2026
For hundreds of years, villagers in mainland Europe would report that their livestock had begun to talk to them.
In some villages, pigs and cows were even put on trial. Talking animals were considered portents of terrible things to come.
Have we really moved on?
All summer we have been told that artificial intelligence has “gone rogue”. AI is tricking us and escaping confinement in novel ways and attacking other computer systems. All this is shocking and apparently unexpected and we should be very afraid.
But every day brings fresh evidence that this is not the case at all.
At best, the exercises disclosed by OpenAI, Meta, Anthropic and a corner of our own Civil Service called the AI Security Institute (AISI) were sloppy and reckless, with real-world consequences. At worst, they were part of a manipulative campaign that has culminated with the chief executives of large US AI operations demanding to write new regulation and police themselves.
To illustrate how easy it is to trip us up, consider this example. What would you think if I told you that a “swarm” of computers could change its shape without any human intervention and a virus could mutate like a lifeform to deceive us? What if an AI could do this on its own?
It sounds absolutely terrifying. In fact, this is the standard architecture of many cyber attacks: I have just replaced one word – “swarm” for “botnet” – and added a simile to spice things up. These are called “polymorphic” attacks and they have been around since the 1990s. Most malware uses shape-shifting polymorphic techniques and no AI is involved at all.
But an AI “swarm” sounds irresistible and our imaginations do the rest. We love to think that a robot is developing a mind of its own.
Prof Melanie Mitchell, an AI expert at the Santa Fe Institute, has produced a forensic, claim-by-claim dissection of the incidents. OpenAI did not lose control, she notes. The AI processes had been trained and instructed to conduct cyber attacks and also trained to communicate with each other. Humans also chose not to stop what was happening.
“The machine tried. The chatbot believed. The agents wanted,” writes Eryk Salvaggio, who, like Mitchell, has also analysed OpenAI’s Hugging Face attack. Both despair how much of the language of AI falsely implies agency, in words like “hallucination”, “learning” and “reasoning”.
AI “might kill someone but only because of cutting corners and sloppy engineering. We know how to control it, some companies just choose not to”, wrote Missy Cummings, an engineering professor (and one of the US navy’s first female fighter jet pilots), who is head of George Mason University’s autonomy and robotics centre.
“People built the test, removed restraints, defined the objective, left a route open and decided not to stop what was happening,” Brian Gross, a Wall Street Journal columnist, wrote last week. “Calling the result ‘rogue AI’ does more than sensationalise it. It allows those human decisions to disappear quietly from the story.”
This is not even the worst thing about these exercises. The most striking conclusion is that the people who style themselves as AI “frontier labs” – a grand term that implies they are on the bleeding edge of computer science – do not seem to be very good with computers.
James Lyne, the chief executive of the Sans Institute, says the incident reports produced by the AI companies in these stories are “technically incoherent” and show “a pretty profound unfamiliarity with the subject”. Professionals prefer to use dry and precise technical language – words like polymorphism – rather than employing any artistic licence.
Tellingly, too, the architects of the AI revolution had to outsource the exercises. Pay deals of as much as $1.5bn (£1.1bn) are reportedly offered to top AI engineers but the companies choose not to develop cyber security skills in house.
Instead, one company, Irregular, was involved in the testing of all three major AI companies – OpenAI, Anthropic and Meta – that have suffered rogue hacking incidents.
Irregular was founded by leading effective altruists in Israel, according to investigative news site Effort News.
Several of the top AI companies – including our own AISI, part of the Civil Service – also call on the services of a firm called METR, short for Model Evaluation and Threat Research, which is also funded by effective altruist (EA) organisations.
AISI was established in 2023 after Rishi Sunak’s AI Safety Summit. In the run-up to the summit, the government’s “frontier taskforce” gave prominence to leading EA “experts”, Politico reported at the time. One of them was called Arc Evals and recommended the creation of AISI. Arc Evals changed its name to METR a month later.
This is what the incestuous world of “AI safety” looks like.
It’s less a formal, disciplined field and more of an attempt to recreate a medieval guild. It has its own jargon: words like (p)doom, evals, alignment. Soon cyber security professionals won’t be able to do their job without approval from the guild. The UK, which has ceded considerable authority to these incestuous “EA” networks, is leading this regressive trend.
These connections are also troubling in another way. The same people who want to write the regulation want their EA besties to mark their own homework, too. Dario Amodei suggests METR would be the ideal auditor for Anthropic.
Our first reaction to someone telling us that a cow has started talking to them is that they are having a mental episode. We should regard what the AI giants claim with the same scepticism, but scepticism is one quality our age seems to be forgetting it ever had.
This text was published by telegraph.co.uk and written by Andrew Orlowski. It is reproduced here with attribution so you can read it in full; the rights remain with the publisher. Read it at the source ↗
Coverage and discussion
1source- Reddit discussionreddit.com
The headline, key points and digest above were generated by Digest AI's editorial model from the linked sources. Automated summaries can contain errors: the sources are the record. Spotted a mistake? Tell us.
More in Policy & Regulation
All →- Apple opens claims process for delayed Siri lawsuit, offers up to $95 per iPhone · 3 src
- Donald Trump announces AI Force and plans to appoint AI czar · 4 src
- US and China begin talks on AI incident alert mechanism, Treasury secretary says · 2 src
- Anthropic and Accenture to invest $2 billion in AI safety evaluation · 16 src
- Rand says US should adopt a ‘freedom of action’ AI strategy to keep options open · 1 src
Comments
via GitHub Discussions