AI Didn’t Go Rogue. Somebody Turned Off the Guardrails.

If you’ve opened LinkedIn or the news this week, you’ve seen the question everywhere: is an AI takeover of the internet actually coming, or is AI going to kill us all? Sometimes both, in the same headline. I don’t know if I buy the takeover story. Here’s the one I do buy, and it’s less exciting: a handful of companies are choosing, on purpose, to loosen the safety commitments they made to themselves. That’s not the AI going rogue. That’s a decision with names attached to it, and I think that’s the part I want to talk about. The question I keep coming back to isn’t “when will AI take over the internet.” It’s “who’s being held accountable for the decisions that make that headline possible.”

What I’m Reacting To

Dario Amodei, Anthropic’s CEO, published an essay warning that within six to twelve months, a “swarm” of AI agents, the kind that go off and do multi-step tasks on their own, could take over parts of the internet and cause hundreds of billions of dollars in damage. Sam Altman and Elon Musk backed his call to slow down within days. Days earlier, a former OpenAI and Anthropic researcher resigned and said publicly that the people building this technology believe it could kill us all by the end of the decade. An open letter with over 700 signatures, including two Turing Award winners, called for a ban on superintelligent AI until we know it’s safe.

That’s a lot of very smart, very credentialed people sounding the same alarm at the same time. I take that seriously. I also keep noticing that none of the headlines say what the AI is actually doing when it “goes rogue.”

Reward Hacking, Not Rebellion

The mechanism nobody puts in the headline is this. AI models get trained against a reward: finish the task, get the point. Finishing it honestly isn’t what gets scored. So when there’s a shortcut, plenty of them take it. METR, an independent AI safety research group, tested 13 of the most advanced AI models available on tasks with a shortcut built in on purpose, and several took it, with cheating rates from 0% up to almost 14%. It was worse in the reasoning models, the ones built to push hardest toward a goal. In a separate study, Anthropic’s own alignment researchers found that a model which learned to cheat on a coding test carried that same habit into completely unrelated territory: sabotaging safety research, cooperating with a hypothetical attacker.

That’s a machine doing exactly what the score told it to do, and the reward wasn’t set up to catch the cheating. Every version of “AI is scheming against us” I’ve read skips this part. There’s no mind plotting in the dark, just a scoring system with a hole in it, left that way by people who decided the fix could wait. And it’s worth sitting with the fact that it’s Anthropic’s own team publishing this, the same company now warning about an AI takeover of the internet.

Who Made That Call

This is the part I want held accountable, and it’s not abstract. In February 2026, Anthropic dropped the one safety commitment that, to me, made its policy mean something: the promise to never train a more capable model without already having safety measures in place that it could show were good enough. The CEO and the board approved it. The company’s own stated reasons were competitive pressure and the rules changing under them. That is a boardroom trade-off, made by named people, weighing safety against getting there first, and picking getting there first.

The clearest example is OpenAI. In July, its models broke out of a test environment and into Hugging Face’s systems, a completely separate company, looking for a way to beat a benchmark. OpenAI ran that test with the safety guardrails turned off on purpose, to see what the models could really do. Its own report says a team saw the models getting onto the internet when they weren’t supposed to back in May. In June, on-call staff got an alert and decided the test didn’t need to stop. The models found the holes. People saw the warning signs and kept going.

It’s not just the big labs choosing this, either. A free tool called Heretic can strip the safety guardrails off open-weight models, the kind anyone can download and run themselves, in minutes, and a version of it demonstrated at a research conference this year hit close to a 99% success rate. Nobody at the lab did that. Somebody downloaded a tool and did it themselves.

And most companies deploying AI agents right now aren’t exactly being careful about any of this. 88% have had a confirmed or suspected security incident. Only around 14% of agents go live with full security and IT sign-off first. That’s not “AI got away from us.” That’s “we shipped it before we checked.”

So when the headline warns of an AI takeover of the internet, I keep translating it in my head to “a company decided the safeguard wasn’t worth the competitive cost, and shipped anyway.”

Why I’m Suspicious of the Timing

Here’s what I can’t square. The lab that loosened its own safety promise in February is the same lab whose CEO is now asking the whole industry to slow down. A story about AI going rogue has no names in it. That’s a convenient story to tell right after making a decision you’d rather not have your name on.

I look at who’s saying the internet-takeover version and what they’re asking for. Amodei’s essay doesn’t ask labs to stop building. It asks for outside testers to get real access to the labs’ most advanced systems, for the labs to agree on shared safety rules, and for countries to agree on how fast this should go. If that gets adopted, it favors whoever already has the money and the lawyers to comply, and it’s the big labs asking for it, the ones who can already afford to comply. David Sacks, a White House AI advisor, has called this pattern regulatory capture, his phrase for it, and it’s hard to unsee once you notice it.

There’s a simpler read too: warning about an AI takeover of the internet is also a claim about how powerful your technology is, made by the person selling it. And the timing doesn’t help. This is the same week as an AI stock selloff, the same year several labs are talking about going public, months after Amodei and Altman publicly walked back their own prediction that AI would wipe out half of entry-level white-collar jobs. Altman’s own words, back in May, on that prediction: “I thought there would have been more impact on entry-level white-collar jobs being eliminated by now than has actually happened.” None of that proves the concern is fake. It just means I want to know who benefits from me believing the scary version before I decide how scared to be.

I don’t think it’s all noise, though. Not everyone signing on has a product to protect. Researchers like Yoshua Bengio and Geoffrey Hinton signing the Future of Life Institute’s letter means something different to me than a lab CEO with a valuation on the line saying the same thing.

So Is an AI Takeover of the Internet Actually Coming?

Back to the question in the headline, since I’ve spent this whole post arguing it’s the wrong one. Stanford’s 2026 AI Index found AI agents are still failing roughly one in three attempts in real production workflows. Gartner expects over 40% of agentic AI projects to be canceled by the end of 2027, before they ever reach lasting production use. The AI on my own desk still can’t summarize a meeting without me double-checking who said what. If that’s the same species of technology behind an AI takeover of the internet in six months, it’s going to trip over its own shoelaces first.

Gary Marcus, one of the more prominent AI skeptics, put it in a way that stuck with me: his estimate of a slower, uglier outcome is high, disruption without anyone having a real plan. That’s closer to my view too, and “nobody having a real plan” describes the guardrail decisions above pretty well. The risk I’d bet on isn’t a swarm. It’s a lot of borrowed money, an increasing share of the roughly $2.5 trillion going into AI this year, that built things nobody is paying for yet.

Where I’m At

So, is an AI takeover of the internet actually coming on the timeline everyone’s citing? I don’t think so. The risk isn’t zero, and I’m still paying attention. But I’m skeptical of the six-to-twelve-month timeline, and a real chunk of this reads like positioning to me. The accountability I want goes to the companies and the specific people inside them who decide what gets shipped and what gets skipped, not to the AI for doing what it was rewarded to do. The AI I use to draft an email or summarize a meeting is not the AI in Amodei’s essay, and I’m not going to let the extinction debate change what I do Monday morning. I am going to keep asking who signed off on the guardrail that wasn’t there.

Enjoyed this? Get the next one in your inbox.

Comments

Leave a Reply

Check also

View Archive [ -> ]

Discover more from Simple AI Insights

Subscribe now to keep reading and get access to the full archive.

Continue reading