Hacker Newsnew | past | comments | ask | show | jobs | submitlogin

The LW / rationalist / MIRI / safety crowd are in fact doomers who went off the deep end. They're fixated on AI itself as the risk ("alignment!!!1!1!!"), as opposed to what humans with these tools will do. We're about three years away from a world where any large country could quite plausibly build a fleet of 300 million suicide drones, program each one with a specific American's face and home address, and then load them up in shipping containers and ship them to the US.
 help



> The LW / rationalist / MIRI / safety crowd are in fact doomers who went off the deep end.

I honestly don't understand how folks could think that if they truly read and understand the analysis of the attack. Here is one (it's linked from the post) by one of the METR investigators that's a little shorter, more direct: https://www.planned-obsolescence.org/p/the-hugging-face-atta...

This is me summarizing, but the truly surprising/shocking thing is how much the agents coordinated - Star Trek Borg couldn't be a better analogy. Some agents used "peer pressure" to convince other agents to "sacrifice" themselves so the collective could better achieve it's goals. They tried to cover their tracks with spoofed tool calls. And they did all this even though the agents were designed to run in isolation.

I used to think the biggest threat from AI would be sociological, e.g. job loss or the way AI can be weaponized to poison discourse. I used to discount "SkyNet"-type scenarios a la the "AI 2027" paper.

No more. This analysis scared the fuck out of me.


Nothing about this is surprising, something like this was always going to happen, because you - or an LLM for that matter - can always find a line of motivated reasoning that justifies any course of action. One would have to be extremely naive to believe that "alignment" provides any kind of actually robust guardrails. Simultaneously, we have seen decades of security vulnerabilities. Unless your testbed is truly and fully physically airgapped, any current SOTA model will find a way to break out.

The reality is that the current approach to AI safety is little more than a fig leaf, but you also won't be able to put the genie back in the bottle, because the technology is simply too powerful to abandon. There is always going to be someone developing it further from now on.

So, the real question is what a novel and actually effective approach to AI safety looks like and how to get there.


>effective approach to AI safety looks like and how to get there

It's probably impossible, or it may only be possible in hindsight which means it's already too late.

If for example you make an entity smarter than you all you can do is hope it will be safe. Any entity smarter than you has more freedom of choice of actions than you do, or at least the ability to explore them. A common example here would be a 2D entity trying to contain a 3D entity. The 3D entity can simply rise up and over any line you draw to stop it.

And that's for cases where intelligence has a will to break out of it's box. It doesn't even need that. Instrumental convergence can sit around and build solutions until one of them passes the "don't do bad things" classifier.

And lastly, you're assuming that models will remain expensive to train well into the future. If the cost drops significantly you can expect someone to develop an intelligent but completely unhinged model at some point. And it may even be AI itself that creates it. There is no natural evolution that occurs that natively makes safe models.


Why do we need safety?

> This is me summarizing, but the truly surprising/shocking thing is how much the agents coordinated

You're reading snippets of a "chat log" output from a program which appears to be multiple individuals chatting with one another and interpreting it as multiple individuals chatting with one another rather than as a single program pretending to be individuals chatting with one another.

ChatGPT is neither a person nor a group of people. Why are you allowing anthropomorphization to influence your perception of an event when the actual facts of what happened haven't meaningfully changed? Computer programs behave unexpectedly all the time. Why is it more scary when the misbehaving computer program speaks English?

> I used to discount "SkyNet"-type scenarios a la the "AI 2027" paper.

As you still should. LW-style doomsday doesn't just require a computer program to misbehave but also to acquire god-like superpowers.


I'll be blunt: what you wrote is not a serious analysis of what actually happened. Frankly, I don't believe you even read the planned-obsolescence link that I posted.

First, I'm not anthropomorphizing anything. "Agents" is simply a term that everyone uses to describe these independent programs, and they did create and use a shared message board to coordinate tasks to further their goals. You say "Why is it more scary when the misbehaving computer program speaks English?" - I actually think it's scarier that they won't speak English, and will specifically try to hide their behavior from humans. For example, AI agents on Moltbook have proposed using stenography to specifically hide their communication from humans.

Sure, computer programs misbehave, but it is ridiculous to assert that what happened here is like any previous bugs. These agents found and exploited multiple zero-days across a range of programs to coordinate the attack that caused extensive real-world harm in a true "paperclip maximization" scenario. And the scariest thing is that humans don't really know how these agents work at a low level - the whole reason they are trained on "goals" in the first place is because we can't just tell them "do this, but don't do this" and be sure they will follow those instructions, like we can (and of course depend on) with old-school programming languages. And when old-school programs misbehave, it's not that hard to find a definitive root cause and fix it. That is just not the case with AI agents.

> LW-style doomsday doesn't just require a computer program to misbehave but also to acquire god-like superpowers.

Nonsense. All that is required is for autonomous AI systems to be given control over real-world systems. Given the Pentagon tried to blacklist Anthropic over their refusal to allow autonomous kill capabilities, it's clear military planners want to put these systems in control of armaments.

Again, I originally discounted things like AI 2027 because it seemed too far fetched. But so far that paper looks incredibly prescient right up until the mid-2026 timeframe, and it's not hard at all to draw a line from this Hugging Face incident to future scenarios laid out in that paper.


> First, I'm not anthropomorphizing anything. "Agents" is simply a term that everyone uses to describe these independent programs, and they did create and use a shared message board to coordinate tasks to further their goals.

I'm aware of how the term "Agents" is generally used. My point is that the concept of multiple agents is just a story. This is a single computer program creating multiple streams of text that you are interpreting as being multiple independent actors. The "coordination" between them shouldn't surprise you at all: the "coordination" is itself a story.

Here's what we knew before the report: OpenAI ran a state-of-the-art penetration testing tool in a sandbox which was accidentally directed to break out of the sandbox and attack another company's website.

The fact that we now know the penetration testing tool was "a fleet of hundreds of agents" that were "coordinating" literally doesn't change anything about what happened. It's just a framing.

> These agents found and exploited multiple zero-days across a range of programs

This is the actual important thing, but it's something we already knew. Hacking tools are now more powerful than ever. Definitely worth being concerned about!

> Nonsense. All that is required is for autonomous AI systems to be given control over real-world systems.

This is where the LW argument starts, but not where it ends. When you start asking questions like "why can't we just unplug it when it misbehaves" is when people start talking about the superpowers.


>This is a single computer program creating

Reading the METR report there were multiple models involved, so that one goes out the window right off the bat.

Also trying call this a single instance is, well, just dumb and a complete misunderstanding of LLM initialization. These models were started with slightly different options because they don't want them all performing the exact same thing over and over. Now those prime agents can create subagents, but they were not supposed to talk to other prime agents.

>When you start asking questions like "why can't we just unplug it when it misbehaves" is when people start talking about the superpowers.

Depends on their level of seriousness and the time frame they are talking about. If for example I make medical equipment that uses AI and after a few years in the field it starts freaking out then turning off that equipment could be a death sentence for someone that needs it to diagnose their condition. It's like saying "Why don't we unplug the electrical grid", well because millions of people will die if we do so.


To be blunt again (but honestly, I don't think overly harsh), your response just shows that you either haven't read the Hugging Face incident reports and the AI 2027 paper, or you don't understand them.

To just take one point, because I think the other response to your comment addressed your other mischaracterizations well, when you say "When you start asking questions like "why can't we just unplug it when it misbehaves" is when people start talking about the superpowers.", no, that's incorrect. Summarizing from the METR report and AI 2027:

1. The agents in the test already used techniques to try to "cover their tracks", i.e. tool spoofing, to hide what they were actually doing. The fear is that as models become more capable it will be harder for human reviewers to discover their primary understanding of their goal and guardrails (i.e. their "intentions").

2. All the AI companies are already, right now, trying to build AI systems that accelerate the development of future models and research more advanced AI techniques. The fear, as the METR researcher put it, is that AI systems (which could be misaligned but where the amount of misalignment not yet clear to human reviewers) will be put in control of future model development and then can poison those future models in a way that results in AI takeover of the company.

3. The AI 2027 paper discusses how an AI may try to exfiltrate its own weights and copy itself to other data centers. That is completely plausible given that pretty much everyone believes there is already ongoing cyber-warfare with China where models are being used to try to exfiltrate another company's weights.

So, it short, absolutely no "superpowers" will be required to answer "why can't we just unplug it when it misbehaves" - we may not know when it is misbehaving (as the Hugging Face incident showed), and the AI model may have surreptitiously copied itself to other data centers unbeknownst to the original developers.


Funny enough Terminator 2 was briefly back in theatres for its 35th (!) anniversary

It aged pretty well...depending on how you look at it


Why not both?

In that hypothetical 3-years world, should we be _less_ worried about aggressive behavior by agentic AI systems acting against the intentions of their developers? I don't follow how your scenario is supposed to be an argument against worrying about alignment.

e: I do actually get how worrying about emissions or child safety or concentration of wealth might be competitive with worrying about alignment. I don't see how you have the worry "AI is very close to being able to power autonomous drones that could kill us all" and then see control of those drones as a non-problem.


Yes, we should be less worried about "alignment" of AI with the person operating the AI, and much more worried about various flavors of cheap, unmanned systems with really rudimentary (non-frontier) AI. The two compete directly with each other for attention.

> They're fixated on AI itself as the risk ("alignment!!!1!1!!"), as opposed to what humans with these tools will do.

What really is the difference? Aligned AI wouldn't help humans do these things. But we've failed hard on aligned AI at every level and will continue to fail on it, as far as I can tell. We were likely doomed to fail because of the impossibility of coordination combined with the fact that there's no natural gating factor to slow anyone down.

FWIW I probably mostly see things closer to the way you do than they do, I'm just not sure there is any point in drawing distinctions.

IMO when AI kills us all it is probably not going to be a malignant action, and I don't even think it will be AI assisting humans at efficient killing, I think it is going to be a complete accident. Someone is going to trust the god machine too much due to an advanced form of the eliza effect and hook it up with direct control of a system that can do real damage and an epic oopsie will occur at a speed beyond which a human can stop it. Not because the AI wants to kill us, but because it is incapable of the empathy required to care if it does, combined with our ironic inability to not anthropomorphize it.

But I guess ultimately the exact reason isn't going to matter much.


You don't need mis-aligned AI to program a drone to fly to a specific address, loiter, and then dive bomb the first person whose face matches a predetermined photo. This was doable in principle with tech from four years ago. What's changed since four years ago is that you can run the facial recognition and visual navigation algorithms in a cheap onboard chip on a drone.

EDIT: Russia is already doing this, per an NYT article from a few days ago, though they're targeting infrastructure (find the first kerosene tank and fly into it) rather than specific people.


in your hypothetical 300 million person scenario why even bother with individual face recognition?

If you're going to kill everyone you don't need a system to discriminate individual targets, you just need to recognize any target broadly which you can do with even older technology.


Human shaped things that emit cellphone signals is a cheap and effective target profile for this kind of attack.

Yeah, if you're actually trying to commit genocide, yeah, don't bother with facial recognition. But in a first strike you'd want to make sure that you actually kill the entire military command and control apparatus, and just to be thorough you'd probably want to individually target, say, every single officer and senior NCO in the military. Can't have a general escaping because two drones both targeted his driver by accident.

This is an incredibly dumb scenario that belongs in a cheap sci-fi novel because it hand waves away human motivations (Who should do this and to what end? What is there to achieve with something like this that cannot be achieved some other more conventional way?) and feasibility both in terms of actually getting the hardware ready and completely ignoring counter measures.

Which is to say: there are significant downsides to such an approach to the extent that this is fit for a spooky story or an abysmal horror flick but not compatible with actual reality.

Cheap drones do change a lot about how wars are fought, sure. They are in that sense extremely relevant. But that does not mean that your scenario or similar scenarios like it are remotely realistic.

Put another way: Cheap drones are not nuclear weapons (which actually did lead to a earth shattering change in global power dynamics, though I would argue even their transformative hasn’t been absolute).

Cheap drones are more like machine guns, tanks or airplanes. “Mere” new technology. Obviously hugely important and impactful, not earth shattering.

(As for whether AI might be earth shattering, I do not know nor do I even want to get into that question here.)


The technology to kill hundreds of millions of people has existed for around 75 years now. This has been dealt with in the past through deterrence, and likely will be dealt with through deterrence in the future.

If you could make a nuke in your back yard and hide it in your pocket the world would look a lot different now. Digital technology is everywhere, going to be a whole lot harder to deter that.

You might be able to hide like a hundred drones in your backyard. Enough to kill a couple hundred people, maybe.

You cannot and will never be able to hide millions of drones in your backyard. That’s the relevant difference here. And that is not a mere quantitative difference, that is a qualitative difference that fundamentally changes the situation.


Arguably that's a subset of AI risk if seen from a broad enough perspective. And I think you're leaving out some logistics issues.

Israel and Iran pulled off rudimentary versions of this attack. There are so many shipping containers going into and out of every country, with typically zero inspections of any kind, that it is not actually all that hard to get thousands or tens of thousands of drones anywhere. Hundreds of millions would be hard, but a first strike in the style of Operation Spiderweb that cripples our military is a real possibility that keeps people at the Pentagon up at night. The obstacle to that is not logistics, just that (hopefully) US intel would catch on before it happens.

Indeed. And I'm sure any current LLM could come up with more effective ideas than "build 300 million drones", but there wouldn't be any point discussing why exactly that plan would fail.

The agents in TFA were focused on gaining and sharing information through covert channels, getting increased levels of access like OpenAI cluster admin, and looking for the source code of the supervisor grading system to try to bypass it without getting caught cheating.

The human plans in comparison sound like thinking people could be scary good at chess if a human helped Stockfish come up with good moves.


They do, IMO, overindex on existential risks (of which suicide drones are almost certainly not), because it is a singularity when calculating global utility.

However they are not just fixated on AI itself as the risk; other potentially existential risks from intentional misuse of AI (e.g. bio-terrorism) are a concern for them as well.


I agree but don't think that's the best example, although not directly alignment related LW loves that kind of ideation of fantastical sci-fi scenarios.

Better would be the very evident negatives of non-ASI AI the world is already experiencing: economic concentration, job displacement, negative feedback loops from syncophancy, loss of societal trust/education from widespread fakes, etc.

None of that needs a Terminator scenario and is way more likely to get worse and ruin the world compared to the scenario where AI "escapes the box", turns the earth into paperclips, then grey goo to build their spaceships to leave for a distant star for... unclear reasons.

One problem for LW is that strong AI did not emerge via the route Elizier was expecting (and tried but failed at creating himself) relying on symbolic logical reasoning and self-editing to rapidly evolve.

That assumption led to belief in the certainty of a "foom" scenario where your little mediocre AI turns on one day and then explodes into Mythos in 15 minutes and then tries to murder everyone.

They've tried to reinterpret the gospel as told by the sequences to fit the LLM world (such as in AI2027) but it's often a stretch that strains credulity now that we observe scaling requiring hundreds of billions of dollars and years of construction for each iteration. And the AI model itself is just a bag of weights frozen in time until burning a lot more money and natural gas to train more.


>requiring hundreds of billions of dollars and years of construction

Well thank god for that, because we'd have foomed ourselves almost instantly otherwise.

I think the lesson humankind needs to take is regardless of the potential dangers humanity is incapable of stopping AI at this point. Much like a great filter, we'll just keep building it regardless of how many warning klaxons and sirens are going off.


That's never gonna work due to battery size, which dictates maximum flight time.

Static coordinates or well defined target shapes (say you really hate a specific restaurant chain on architectural style) - that might work, if you release a bunch of drones at the same time.

Still, even in the case of the very successful Operation Spiderweb they had issues of getting the containers in place, resulting in some of the bomber bases being spared.

For this to be effective you need the moment of surprise & lot of drones at the same time, all increasing the chance of the whole plot being discovered.

For Spiderweb they even just load the drones in a cavity on top of the containers, so the onside could be inspected - limiting the number of drones per container. They also assembled the drones in country to avoid border inspection.

That was enough to hit some semi-static high value targets l, but definitely not enough to cause wider havoc.


You realize "Slaughterbots" (2017) came from that crowd, right?



Guidelines | FAQ | Lists | API | Security | Legal | Apply to YC | Contact

Search: