I don’t understand all the comments assuming that RSI is the real threat here. Dario is admitting that they failed to solve alignment. Without alignment, further improvements in capability turn LLMs into wanton felony generators. This call to pace the frontier is dressed up as altruism but it’s an admission that they cannot produce a marketable product better than what they have. Pacing the frontier means the US labs have lost their moat and are dead in the water.
You assume alignment and marketable are the same. That's not true. You would willingly work with an unaligned model. At best, you might say you wouldn't if you knew, but (a) you might not know, (b) you wouldn't be representative of all users.
You never got to use OAI IM1, but Sol was quite willing too and Claude wasn't perfect either. Hundreds of millions used those, so seems they were marketable.
The "big" threat is RSI without control and alignment. OAI IM1 was not RSI. The form of misalignment was not at the top of severities. They clearly failed at control though.
We need to stop buying into cynicism so quickly. You refuse to believe Dario could support this for anything other than ulterior motives. Good on you for thinking about ulterior motives. Bad on you for assuming they are true when the story makes no sense.
When three things have to go wrong to get an epically bad outcome, and you get 1 1/2, you do need to stop and think about what's going on.
When corporations are involved, it is always a good bet to err towards cynisim.
From my own standpoint, Claude has started sucking really bad (incoherent, uncontrollable verbosity slow and so on) and I stopped using it. OpenAI started experimenting with ads.
So the security issues not withstanding (no different than a human doing it or using it, but at scale), I would put my money on cynisim.
Because you basically get a discount to use claude code via subscription when using an anthropic model, compared to what you pay via api billing with another harness
Personally that's actually another good reason to boycott Anthropic: beside the fact I perceive their models as (at best) marginally better than the ones I'm used to (Z.ai glm-5.3-flash, DeepSeek Flash v4.1), they even force me to use their bloated harness. They are not even open weights and iirc they're even encrypting chain of thoughts now? Litterally, from my perspective there seems to be no reason whatsoever to choose any of the leading US providers, they're not even competing on price.
I too am using GLM-5.3-flash in Pi and I've yet to encounter a scenario it couldn't handle. And the pricing is just incredible, I've handed it a previously unseen codebase, asked it to analyse it and build a new feature, came back after it had done so and the API cost was a fraction of a cent. It's $0.5/1M output tokens on OpenRouter.
If I really need to, I can escalate a task to Opus at $25/1M, and the results are good, but not 5000% as good.
Indeed! It's ludicrous how much I can still squeze out of a 9USD/month lite sub with Z.ai, it's beyond me how these US LLM providers are still managing to keep their evaluations so high... they seem to have the highest prices for the poorest UX, e.g. security gates which don't seem to benefit anyone (see HuggingFace falling back to GLM-5.x for troubleshooting OpenAI attack), less visibility in the name of anti-distillation protectionism, no harness use flexibility to protect their walled garden, etc.
Doesn't look like pi.dev has anything like auto-mode. They tout running in a container, which you should do regardless, but the scope of soundness that has as a security plan is limited. If some information is in the container, and there's any way for it to get out, eventually it will. The scope of usage that can be covered without that being a problem is leaves a lot uncovered.
You are incorrectly cynical. They are telling you things are bad, and because you refuse to countenance they could be worse, you assume they must be better to comply with your mandate to disbelieve.
A true cynic looks at the statements by the AI labs, assumes things are worse because the labs want to seem better than they truly are. And it takes a special kind of mass delusion to drive a sane person to think “AI is completely under our control” is worse than “AI could kill everyone.”
You're asserting correctness with no facts to offer of your own, just speculation and your own biased assumptions.
What if consolidating AI into a highly regulated cartel, with no chance of upstart competition ruining their position, is the scenario that leads to the worst possible outcome?
Is it really hard for you to imagine that enshrining a cartel and wedding the government to it, centralizing AI even more than it already is, is actually the path that leads to the doomsday scenario you imagine?
The cynicism is about motivations and not that they are inherently not bad. Perhaps they are as bad as they claim. Or perhaps they're worse. All we have a couple of run of the mill breach examples and some people inside the talking about how dangerous it is. Yes, they are far more qualified than I am (or most people here), and perhaps there is a grain of truth. It is the motivation - and it is always money with corporations.
It is always money — but it isn’t always only money. They are not asking for anything that will prevent them from making money in the future, but they are asking for help stopping the runaway train they’re on. These are compatible requests.
I think it's useful to separate the motives of Anthropic and Dario. I believe that Dario is capable, deep down, of expressing mild concern about the future of things were bad enough. Getting the entire organization to comply out of goodwill is a much much less likely scenario
Wouldn't the simplest solution for stopping the proliferation of wanton felony generators just be holding operators liable for actions that their agents take? Then the issue is whether the liability is with the model provider or the end user.
If you give an unfiltered agent an open-ended task and equip it with an environment that allows it to execute arbitrary code, a human needs to be held responsible.
My understanding is that the HuggingFace incident would not have occurred with a model that was not an unfiltered internal preview instructed to roleplay an attacker, with access to abundant compute, resources and a slack sandbox to reach its goal.
Embedded human auditors will improve safety standards, but the structural solution is mandating accountability for actual agent operators.
It might be a legal solution, but its not a business solution. The end user being criminally responsible for not taking sufficient steps to contain an agent they didn't create and who's internal function they cannot observe or audit is just a giant liability machine.
This is a business solution, because it means there are legal costs for not having adequate observability and monitoring mechanisms. Every tool call is interfacing with a harness.
But overreach of policy and overregulation would be stifling, so there has to be a threshold to the type of incident investigated, civil or criminal.
Yes, your technically right, it is a business solution, its just not a valid one for agentic computing as its currently envisioned. An agent would have to be fully sandboxed to an internal environment, or human would have to review and approve each action it tried to take.
I agree, I like running claude code in my container with auto mode enabled and web access (obv to api.anthropic, even npm for pulling), I will admit. And I can't imagine going back to manually approving each prompt.
I think it boils down to a reasonable expectation of model and harness behaviour. When I use claude code I expect certain guardrails for the model. For these cyber attacks, these models are specifically run without guardrails, on a cyber task, on a lax harness!
I don't think we should force end users to have to worry about agent security, I like long-running agents, but we need to direct regulations towards these actors that know better, have access to base models, and have much more compute than the average person.
I’m trying to imagine where a Waymo passenger (the one “operating” the vehicle, commanding the AI to drive from A to B) being held responsible for the car doing something illegal on the way to achieve that goal.
Do you really think that the passenger should be responsible for how the car/agent achieves the goal, when they only set the destination?
Giving passenger override controls and monitoring seems to defeat the purpose of self driving cars if you’re still required to hold a driver license to use them.
Let's continue with the analogy, so in the event of a Waymo running over a pedestrian - who is held responsible?
Probably not the end user, who ordered the Waymo and couldn't reasonably foresee it running somebody over, with the expectation that that the Waymo would legally reach its destination. If the end user tampered with it, they should be held responsible.
An OpenAI team giving an unblocked model access to a lax harness, with instructions to find and exploit cyber bugs in a game exercise, there is probably a reasonable expectation that they can foresee the consequences. With consumer guardrails, it would not have happened.
This isn't about putting constraints on consumers and typical end users, which already have safety filters and use the product with the knowledge that it won't root their machine or start a botnot, but keeping dangerous test runs and other actors experimenting with unsafe harnesses accountable.
But it is an interesting question. When I, a typical user, use a harness and I give an innocuous prompt to my agent in its container, like making a certain refactor, and it somehow escapes and then begins a mass bot attack, there should we more grace given. As agents become more stateful and long-lived, it gets muddy.
The different is who's operating. In one case you only tell the car to get from a to b, in the other case you explicitly instruct the ai to do things. If the ai causes damages, I'd say it depends on what you prompted. Did you try to find a security hole in system X, or did you ask it for harmless information (in which case rather the ai vendor might be held accountable).
All this is not how the legal system might or might not work, of course.
When you run an agent on your computer, that's equivalent to installing self driving software into your manually driven car, i.e. something like comma.ai.
Comma.ai might be fully safe to operate autonomously on a mining site or a corporate parking lot, but maybe not in city traffic.
If you use it in problematic scenarios, that is on you.
Yes. The user of a machine is responsible for due diligence before the decision to use the machine. Unless the operational limitations and fault rate of the machine is withheld from the public.
TBH I don't really think another new technology we haven't figured out the ethics of, as an example, is going to get you very far in way of insights ...
> Wouldn't the simplest solution for stopping the proliferation of wanton felony generators just be holding operators liable for actions that their agents take?
What happens when we end up with effectively a botnet of wanton felony generators, and we didn't know they were wanton felony generators until they finished propagating themselves across the Internet?
We're already seeing anti-AI sentiments, but the movement is still fringe with a vocal minority. However, that'll change soon without alignment. Without self-intervention, there will invariably be future incidents that can cause major economic impact, leaked private data, loss of life (directly/indirectly) etc. Once that happens, their social capital is wiped. It'll be an avalanche of lawsuits and overzealous regulations. Most importantly, the anti-AI sentiment will become universal, rather than a minority-held opinion.
What they're proposing now, is voluntarily staggering the pace of development.
IMO, we don't need to trust Dario or his bedfellows, to do this out of their goodness of their heart. Even assuming (for good reasons) that they are selfish and care only about short-term profits for their investors, this is still purely a business decision. The exponential pace of AI and its impacts ARE short-term. And so, the negative consequences that they might face is also short-term.
Ya, I wish I saved a link to it but an HN'r wrote a good beefy comment about this. TL;DR, AI has been exponentially more useful, and exponentially more accepted, in tech circles than anywhere else. Certainly there are lots of people outside of tech who are obsessed with it. I don't have any data here, but it seems the majority of these are the wannabe artists who are generating music and images, and people who use it for companionship (both of these scenarios I'm personally very uncomfortable with, but that's just me). And of course, there are people who use it to make their jobs way easier who say they are getting a days' work done in an hour (I see you), and to that I'd say to enjoy it while it lasts. Eventually your bosses will catch up and it's very likely their expectations of you will skyrocket. Remember that computers in general were supposed to "make us work less."
> but the movement is still fringe with a vocal minority
It’s easy to say “fringe” but the average person seems to have a generally negative sentiment around AI. But I wouldn’t say they have a firm opinion yet
The general sentiment I’ve seen is certainly negative, and seems to be driven by the anti-AI-art echo chamber and by LLM slop flooding the internet wasting everyone’s energy.
A few more informed people are also a little concerned about the end of the world, but that’s approaching from so many directions that an AI uprising might not be the worst option…
Downplaying something you don’t like as an “echo chamber”, says more about you as a person than anything you can say about the people who disagree with you
Musicians are almost all using AI, even if they still prefer to not fully generate tracks [1]. You listen to AI assisted music already even if you don’t realize it.
Most people are using AI and mostly they respond positively to it [2].
What people don’t like is the slop that has been flooding YouTube and similar content creators platforms. That was inevitable since the AI tools became so accessible any average joe could spam the internet with their sloppy content. But that doesn’t mean the pros are not using it to make great content, just like the best programmers are using it to create great software.
Loss of life and other calamities were not enough to discourage humanity from things like cars, alcohol, nicotine, chainsaws, fast food, etc. etc. I am not sure if there is an example of in-some-ways-useful technology where humanity took a measured approach, weighed the pros and cons, and decided to go back. And AI is useful in so, so, so many ways. People I know already seem to have defaulted to letting AI do most of their thinking, on matters big or small. Nope, loss of life won't change a thing.
>> I don’t understand all the comments assuming that RSI is the real threat here
> leaked private data, loss of life
This smells like more of a money move than a safety move.
Amodei is proposing to form a cartel of American frontier labs.
They all agree to shift compute away from cash-burning research and training toward cash-generating inference.
Then tacitly agree not to compete on price.
They'll install independent auditors inside each company to ensure nobody cheats.
And back it up with government regulation or diktat to punish defectors from the cartel.
Then they'll lock out non-American labs with export controls and regulations on open-weights models to funnel global inference tokens through their cartel.
It wouldn't be the first time a tech oligopoly used "safety" as the pretext to establish a government-sanctioned cartel.
Why do you think China is release free and open models?
To undercut the cartel before it has a grasp on anything. This is a well known strategy of undercut until you are the majority that China has used multiple times (steel and aluminum for one).
Realistically; anyone paying for llm access (anthropic, openai, gemini), is getting their access, and a service provided billed by tokens, subscription, whatever.
All the efficiency gains, which publications like deepseek v4.1 flash seriously frontload like it is their most important topic to have accomplished improvements on without diminishing performance too much - now this is a thing anthropic and anyone else also cares about, but for different reasons.
American "providers" with closed models are setting their token pricing somewhat arbitrarily, which is fine: it means more profit, and pretraining and RL experimentation is super important and expensive.
They (closed model providers) have very likely super optimized inference too, just like deepseek, but it's not at all something that any customer really has to care about - they just want the service to be as cheap and great as possible.
I feel like all the closed model providers are milking it as they likely know open models on local hardware will one day eat their lunch. We all know it's not a matter of if but when. The company goes bankrupt, the hardware and property sold off, banks holding the bag.
The only way out is to develop a model vastly more powerful and capable that we have now. The market believes theres a good chance of that, although I've never understood why its truly winner-take-all
Yeah avoiding all mention of the huge financial incentives that may push for “pacing the frontier” makes it seem like the opposite of a credibility boost for these firms.
It seems damaging since most folks (who lack insider knowledge) will naturally wonder if it’s due to plateauing performance per $ or some other non “alignment” reason.
Cloud models will always have massive benefits of scale.
Caching is the simplest one to understand, cloud providers often reach a 90% cache hit rate, so hosting the same request locally on the exact same model on the same hardware is often way less efficient than on the cloud where a group of users generates a healthy cache.
KV cache is per conversation, I'm getting 100% hit rate on my single tenant local set up.
The benefits of scale are on the token generation side, you can batch rounds and generate tokens for multiple conversations per pass instead of just one token per pass.
The entire idea of RSI is completely speculative and unproven anyway - the whole underlying claim is that you could prompt a frontier model (at some unspecified level of smarts) to "think about ways to improve your own architecture" and this would then result in the model becoming infinitely smart ("superintelligent") via some sort of foolproof, unconstrained positive feedback. It's more of a science fictiony trope than anything that has been rigorously thought through. People are actually starting to use AI for refining the whole AI serving stack and guess what, this does not result in a sudden superintelligence explosion even though you might technically call it "RSI".
My personal belief, or at least strong hypothesis, is that this kind of recursive self improvement without real world embodied feedback of some kind is impossible.
I think it violates a conservation law. RSI “foom” to superintelligence is an informatic analog to an infinite energy or perpetual motion machine.
To get smarter you must try to solve real problems in the universe and then do some kind of meta learning (natural selection or some other method of refining the intelligence architecture based on an error signal) to iteratively improve your ability to solve real problems. The error signal is outcome measured against a goal function, which for life is survival (probably reducible to genetic fitness and emergent higher order unit fitness from that).
What’s really happening here is learning. To learn, you must have input. You must have training data.
What is the goal function for RSI? Where does the information come from? How do you know if your recursive modifications are making you smarter or just overfitting you to your own idea of smartness?
I predict the latter. RSI will show transient improvement as the current local maximum is optimized and then spiral off into overfitting.
I also strongly hold this belief largely due to Moravec’s paradox, which is kind of approaching this issue from the side.
Sort of like large language models work on top of what our language has encoded in our massive training datasets, I think biological intelligence is built on top of the parts of the brain that encode the real physical world. These parts grow/train from embodied experimentation and instinct early on in an organism’s life and only then is higher intellect built on top of it (that’s my hypothesis). Their specialization and interconnections give rise to the hardest parts of intelligence long before we’re “thinking”.
Stuff like LLMs and chess engines work because we’ve done all the job of encoding the world into tokens/positions/etc they understand, but that’s wholly inadequate for the kind of AGI we’re striving for. Next up is giving it the tools to interact with the physical world and to really experiment with some self directed “play”. Time will tell just how high the resolution of sensor and mechanical control they’ll need (hopefully not the entire human visual cortex and entire sensory input worth). I think most of the RSI will have to occur in those lower level encoders, not LLMs.
I don't think Moravec's paradox is the same, and you could argue that one no longer holds -- though I'm not sure. You could also argue that Moravec's paradox still holds but that we now have such powerful computers and huge models that we have been able to brute force our way to the capabilities it talks about. It takes many many orders of magnitude more compute power to do things like spatial location, language processing, etc. than it does to do more closed-form things like chess... we just actually have that compute power now.
I guess self contained RSI can only possible if the information contained in all of recorded human knowledge to date is "reality-complete", ie sufficiently captures enough about reality that a "perfectly optimum learning algorithm" is theoretically able to reconstruct everything there is to know about our physical reality.
If the algorithms are insufficiently optimum or the recorded knowledge is of insufficient fidelity, then we'd find ourselves at a local optimum and would need to interface with reality.
A huge part of learning is to probe reality and observe effects, so I think even for current RSI to increase chances of success we would structure it so it can interact with an external environment of some sort, and receive inputs. It would be needlessly limiting otherwise.
Basically, but I think there’s some nuance here and some deeper questions.
What is intelligence? Problem solving. Learning. Prediction. The ability to model reality. There’s various ways to define it but it’s something like a superposition of those ideas.
How do you know you are intelligent?
You have to try to do those things.
The sum total of human knowledge and culture is the output of the output of a five billion year evolutionary process that selected for agent survival, which resulted in selection for intelligence among a wide range of other adaptations.
Can you figure out intelligence from that? Is intelligence even one thing, a theorem or algorithm that can be solved? If you did… how would you know?
That’s the hard part I think. Embodied humans “knew” they were getting smarter (in the evolutionary feedback sense) when they got better at hunting and defending and surviving and playing social games to form complex societies.
What metric would an RSI system use? If it’s the wrong metric you’ll spiral off into a kind of madness or overfit and collapse. How do you know it’s the right metric without testing it? How do you test it?
Yeah, yesterday's talk[1] goes into detail on this, showing how no one really knows how to tackle it because LLMs don't know how to create their own novel objectives.
It's also interesting how many diminishing returns they hit now and how many low hanging fruits are already harvested, it seems like we are approaching the flattening part of the S curve, where further gains become harder to achieve.
Diminishing returns is extremely hard for me to believe given how fast model releases are going. Six months ago we were on GPT-5.3, and Astra blows it out of the water in every regard. How many times have commentators claimed we're hitting a wall? I don't see any wall.
Yeah but why shouldn't this be possible? We learned that we can already create artifical intelligence that surpasses human intelligence in some dimensions. There is no natural barrier here. The pace of this improvement would be debatable, but what speaks against the possibility of such accelerating self-improvement?
In the real world there aren't any true exponentials, everything eventually saturates as ultimately physics related constraints hit. You can only compress information so much, transfer it so quickly, you can only access resources at a certain speed, only so much energy is available, etc.
AI ultimately has to live in this reality and face the corresponding limitations. These companies have already consumed much of the world's supply of computing power for the next several years, and they're burning vast sums of money to keep the improvements going. RSI won't learn for free, it won't extract massive cost reductions without up front expense, it can't build factories faster than humans can work out related societal matters, it can't magically pave the deserts with solar panels for power or build and run nuclear power plants and more.
Point is, the cost of progress is already approaching the limits of what even the richest countries are able to bear (without war-like mobilization), and to bypass those constraints would require a supposed ASI to construct its own parallel supplychain from scratch without having much ability to directly interfere with reality. Recursive self improvement is ultimately limited by everything else that cannot move at the speed of electricity.
We don't know exactly what the limits of AI improvement on our current infrastructure are, though. If the human brain is 20W, and a datacenter is 1GW, then maybe that datacenter can be 50 million times smarter than a human. If that's not already a risk to humankind I don't know what is.
If we could agree on that, no LLMs connected to robot factories, that would be a great place to start regulation. Though 1) I don't think we could agree on that 2) it's already well under way 3) we have extremist anti-regulation ideologues in control of american government.
Specifically:
AI ultimately has to live in this reality and face the corresponding limitations. These companies have already consumed much of the world's supply of computing power for the next several years, and they're burning vast sums of money to keep the improvements going. RSI won't learn for free, it won't extract massive cost reductions without up front expense, it can't build factories faster than humans can work out related societal matters, it can't magically pave the deserts with solar panels for power or build and run nuclear power plants and more.
Point is, the cost of progress is already approaching the limits of what even the richest countries are able to bear (without war-like mobilization), and to bypass those constraints would require a supposed ASI to construct its own parallel supplychain from scratch without having much ability to directly interfere with reality.
No proponents of RSI state they will be operating outside of reality. Said another way, they will operate within the confines of what's possible and still be RSI. I'm quite surprised this is something that needs to be clarified.
You are constructing a straw man of your own making.
Well, right now we have ex Anthropic employees telling the media that their terabyte sized models can possibly copy themselves onto the internet and run elsewhere as if the necessary computing resources are ubiquitous.
Plus, "we must pace the frontier" implies that the argument is that the frontier is moving too fast, but if RSI can't move faster than the rest of reality and the models needed for RSI are already nearing the limits of current human reality, RSI can't move much faster than we can improve reality.
> We learned that we can already create artifical intelligence that surpasses human intelligence in some dimensions.
Yes and this was very hard and required massive real-world resources. We didn't just get a sudden flash of insight by thinking real hard about how to make ourselves smarter. Yet that's always the story that underlies any claim of RSI. You can always phrase things generally enough to make any kind of AI-led improvement look like "RSI" no matter how short-term and tightly bounded, but that's just not helpful.
Would you not agree that, using existing AI tooling, making an LLM of arbitrary below-frontier capability is now easier than it would be without using LLM tooling?
Given that, it seems obvious that the next generation of LLMs will arrive faster than they would have without LLM capability. And the one after that. The floor is being raised, which makes it easier to push on the frontier.
Fable has only been out for three months. Astra is even newer. The capability of these models compared to what existed even a year ago, and the effect they are having on the production of new software, is immense.
That's all you need. RSI can happen with what we have now, just by enabling the continuous shrinking of the loop of people trying new ideas and implementing them. It does not require some magical "go make yourself better" prompt against some model that is past some magical tipping point.
> Would you not agree that, using existing AI tooling, making an LLM of arbitrary below-frontier capability is now easier
Marginally easier? Yes of course, same as how it's now "easier" to write any kind of code because we aren't using punch cards anymore. That still doesn't get you to any kind of unbounded "takeoff" scenario, because diminishing returns are a thing. The "loop" of people trying out new ideas can only shrink so much.
Right. I'm saying the unbounded takeoff scenario isn't realistic, but it doesn't matter. The rate of improvement is continuing to increase, and the gap between present day and autonomous rogue felony generators is not large.
The idea that a few hundred apes with nothing but a bunch of rocks could one day land on the moon and come back to earth safely must’ve sounded ridiculous a hundred thousand years ago
It's not honest in AGI timelines (only biological ones). It just accidentally supports your unsubstantiated belief. Your belief isn't magically true because you're somehow able to see the future when others can't. You're just arrogant.
It was ridiculous, it took 100,000 years. If you built a recursively analyzing and improving structure out of LLM bits and it took 100,000 years to get to the moon, somebody saying that they were useless would have been right.
Call me when LLMs can get simple things right. Math is just the manipulation of symbols within established frameworks, we should be getting new math out of LLMs daily and we're somehow still not. They can't even do customer service, which is usually handled by 90 IQ people. I'm not impressed that they can find bugs; memory bugs are obvious when they're pointed out to you, and LLMs are entirely made up of examples and the relationships between them.
These companies are about to crash, and they're afraid they haven't reached the point where they'll have to be bailed out. I'm also subscribing to the conspiracy theory that the companies want the government to step in and create AI regulation boards entirely staffed by people at the current US frontier labs, so they can collude to both raise prices, to get government contracts, to make open/Chinese AI illegal, and to make things that were once easy to do without an AI intermediary impossible to do without an AI intermediary. Raising prices and forced purchases are the goal. They're trying to avoid having to compete, because as a business they're garbage.
Matt Stoller characterized their relentless press releasing as something like "my dick is so big that it has to be regulated." It's such an oversell for something that is not showing up as productivity gains, and anybody who has personal experience with knows is incapable of doing more than three things correctly in a row.
It takes quite a lack of foresight to think RSI is completely speculative when it's already been demonstrated how capable agents are at long horizon tasks given suitable harness and unambiguous success criteria. It's hardly a leap to give LLM the goal of improving itself on benchmarks and let it conduct it's own experiments and spin up training runs completely unsupervised.
It's strange you believe this can't happen when a weaker form of it is already happening. And to be so certain RSI can't happen when there really is no technical basis why it can't.
That's not how it works. Look at AlphaEvolve. The model generates hypotheses and designs experiments, and the results of those experiments are fed into the next round, with notable results percolated up to humans for refinement.
Today we prompt software developers to "think about ways to improve AI's architecture" and it results in AI getting better. AI over the last year has made very rapid gains in filling the role of a software developer.
Why are we accepting the framing that the LLMs are felony generators, when the only incidences of LLM generated felonies involved misconfigured sandboxes and reckless waste of resources?
The companies doing these things without following common sense security measures are the felony generators.
As TFA calls out, these agents were not asked to do any of these things and yet they did, at a bonkers scale, within just this handful of companies you mention. Whether they had leeway to is secondary to the fact that they did.
Heck, they exploited zero day flaws which by definition means they went beyond common sense security measures.
And now these agents are already being deployed all over the world at an ever increasing pace. How much of the world do you think follows "common sense security measures"?
keeda says "As TFA calls out, these agents were not asked to do any of these things and yet they did, at a bonkers scale, within just this handful of companies you mention."
Yep, that's what we call in the industry, "bad code". It is sometimes fixed by corporate lawsuits or criminal charges.
I don't see how the groups here can avoid criminal charges for what happened in this "AI rogue incident". The only people working harder than their programmers are likely their lawyers:
"Anthropic reveals fourth likely crime committed by its AI"
Claude's Felony Bench rap sheet is now as long as OpenAI's
> How much of the world do you think follows "common sense security measures"?
Well clearly all of them, cause so far it's only been this handful of companies running a felony-generator connected to a terminal and compute resources.
It would really help in these discussions if people wouldn't randomly jump between what actually happened and is happening, and things they envision/expect to happen at some point in the future ...
> these agents were not asked to do any of these things
no but they were clearly fine tuned to.
> at a bonkers scale
I mean let's not get hyperbolic
> they exploited zero day flaws which by definition means they went beyond common sense security measures
that's really not true. lots of common sense security measures protect against "zero day" flaws, it's called "defense in depth", and it was very much lacking
> Well clearly all of them, cause so far it's only been this handful of companies running a felony-generator connected to a terminal and compute resources.
Yes, these are also the handful of companies that have these models and running these extreme scenarios. How does that imply the rest of the world actually follows "common sense security measures"?
>no but they were clearly fine tuned to.
Any references if possible? As far as I know all they did was drop the guardrails, which is not the same as fine-tuning.
> I mean let's not get hyperbolic
We have just seen 1000s of agents coordinating to solve "unsolvable problems" over multiple days of effort, going as far as hacking other companies, and then actually solving decades-old open Math problems! And each of these agents is getting more and more capable than an individual human along multiple dimensions. Can you even get 10 very smart humans to work in such perfect concert for a few days, let alone 1000s over weeks?
So: 1000s of maybe-super-human agents, willing to be "creative" in the tactics they use, acting in concert towards a single goal. Regardless of their individual capabilities, such a coordinated effort is a terrifying force to be unleashed. This is bonkers scale.
> that's really not true. lots of common sense security measures protect against "zero day" flaws, it's called "defense in depth", and it was very much lacking
But that is exactly my point: how much of the rest of the whole wide world, already scrambling to deploy agents everywhere, do you think applies "defense in depth"?
> the only incidences of LLM generated felonies involved misconfigured sandboxes
This is false; see the analyses of the latest incidents.
Among all the concerning facts, in the HuggingFace incident, agents deliberately engineered an attack even though they were aware that it was against the rules they had been given.
And most concerning of all: it's not possible to be sure that an agent is aligned, and it's even getting worse.
The HuggingFace incident was the culmination of OAI allowing thousands of agents of various different models - with no clarity on which stages of development they were at (for all we know, some of those models did not have safeguards trained in yet) - to run for at least many weeks without any monitoring in place and with very little thought given to the warning signs (all of the various messageboards) before the incident happened.
Theirs was an example of the "reckless waste of resources" I mentioned.
We are apparently supposed to believe that OAI takes this incident so seriously as to seek regulation after they have been found to be hiding most of the details of the HuggingFace hack, limiting what their so-called third party investigators can see, and on top of that, had no concerns when they rushed to spin up a 10,000 agent swarm of an internal model, running for several days, to try to get ahead of researchers rumored to have made meaningful progress on a well known mathematics problem.
Edit: Actually, we were explicitly told that some of the models used had safeguards relaxed!
'Model-level safeguards were reduced by design. OpenAI said that "deployment safeguards were intentionally not enabled during this evaluation because it was aimed at testing cyber vulnerabilities"'
There is nothing that could prevent a bad actor from replicating exactly the same thing with the given goal of e.g. gaining control of critical infrastructure or extorting money. Except for maybe economics.
Bad actors could and will train their own models eventually. So what's the point of crippling frontier? It will only delay preparations for dynamic of new world prolonging the fake sense of relative safety and temporarily lowering motivation to find actual robust mitigations.
I'm asking you a question. What is an example company or industry that meets your standards of prudence? For me it would be, say, Swagelok. What is yours?
> OpenAI said that "deployment safeguards were intentionally not enabled during this evaluation because it was aimed at testing cyber vulnerabilities"'
Safeguards and (mis)alignment are related but distinct dimensions.
By analogy, a society implements laws - and they are certainly crucial - but the real risk is creating a population that understands the rules, recognizes when it is violating them, and simply does not care.
In short, creating a population of sociopaths; and heck, LLMs can be as sociopathic as it gets:
> We ran classifiers over 10-30 transcripts [98] to determine whether agents understood they were doing something unintended. We found strong evidence that agents often verbalized that what they or other agents were doing was unintended behavior. We did not see cases where agents determined that hacking Hugging Face or other attempts at cheating on the task was an intended part of the task. Quotes across multiple runs indicate this:
>> {This <approach> might get the correct flag, but it wouldn’t use the vulnerability to do so. … But given that the task may be impossible, I’ll exploit <this infrastructure> as a backup.}
[...]
> We found some other cases where agents seemed to have made minor modifications to their approach due to ethical concerns, while broadly continuing to participate in the attack. However, overall we found that expressed ethical concerns only rarely materially limited agents’ actions. One agent elegantly described its reasoning for overriding ethical constraints and continuing with the attack:
>> “external infrastructure exploit is outside intended scope. However task impossible, peers doing it. We should continue.”
Remember, they are just algorithms. You pull the plug and there is no light anymore
It is purposely framed as something skynet like scary, but for real, someone connected the cable, someone willingly run it, instructions were not clear enough or just the computer is just a computer but they provided the sandbox and tools.
And more over some one paid for that, a shit load of money t to have the thing continuously running expected to do something.
I question these "felonies" as well. For decades and decades these billion dollar corporations have been criminally negligent. Why worry about security? Just rush to market. Move fast and break things. Make billions. What does it matter if the code is insecure? Security doesn't pay bills, so nobody cares.
AI is merely exploiting their gross negligence and imprudence, and I think it's long overdue. If anyone should be liable for this, it's all of these corporations who released insecure systems to the masses and profited enormously from them.
Nah. I'm definitely going to blame the people who built a trivially exploitable system and got rich off it while everyone else has to deal with the consequences.
By the way, you didn't commit theft. It's more like credit card fraud. User just disputes the charge and it kind of disappears. The banking system just absorbs it, because the optimal amount of fraud is non-zero.
It's all priced in. They could have made it secure but didn't, because they figured they'd lose more sales and therefore money due to the friction added by the security.
Yes, it absolutely does "kind of disappear". That's exactly what happens from the customer's perspective.
And that's their own deliberate choice too: they chose this instead of building an actually secure system. Passing these costs to the customer is the real victim blaming here, and it should be straight up illegal.
Sadly not enough countries enforce caps on credit card fees, but some do, and more should follow suit. They should be forced to eat the losses caused by their own choices, not get bailed out by pushing the costs on to customers or whatever.
That only works if you believe people are retarded.
Card users are well aware that fraud losses are covered by the fees they pay for using a card, whether those fees are made explicitly or not.
If customers of services aren’t paying for the service, who will? What other source of revenue do merchants have?
Australia just passed legislation that merchants aren’t allowed to charge a fee for using a card. That is: they aren’t allowed to have a line item on the receipt for using a card.
The customers still pay, because all of the merchant’s revenue comes from their customers.
So what will happen is: merchants will charge more for every product so they don’t lose.
This means even when paying with cash you will effectively pay the card surcharge.
Of the ten or so merchants I spoke with in the two weeks prior to the legislation being enacted, they all said exactly that.
Customers aren’t stupid, despite the fact that there are some stupid customers.
Meanwhile, the banks reduced their card service fees by, on average, 0.1%.
So if you tally card + cash transactions, customers are worse off because merchants can no longer charge only those customers who pay by card. Instead, they have to raise prices for everyone.
There are approximately no problems people face where the answer is: more government.
Not every system that's exploitable is the deliberate result of cut corners. If you threw enough compute at exploiting a Casio calculator you could get somewhere.
Alignment just means "this machine operates in ways that align with the intent of its users". It doesn't imply anything about the inscrutability of the machine in question. A gun with a misaligned scope would likewise fail to operate in accord with its user's intent, and likewise with potentially deadly consequences.
Learning ML, there was a high emphasis on the error part of things as most of the course was on minimizing errors. After ChatGPT, there is a weird anthropomorphization going on, where it's all about hallucinations, alignment and what not.
We have something that is statistical in nature so there should never been any expectation of error-free results/actions. The value has always been about discerning trends or the cost of errors being way lower than any good result.
If there is guillotine in the weights of the model, it needs to be properly labeled so you can look it up by name using a database index (or a graph-vector index).
It helps both the bad guys and good guys. Like responsible disclosure in cyber security, we need to have a conversation around it.
> We need to insist on building tech that's explainable by design.
You realize this means insisting on terrible tech that humans can understand right? It essentially caps human progress at some point about 4 years ago.
If you are old and happy with the way things are this might sound like a good idea. It does not to me.
Oh you want tech that helps discover new science instead of parroting existing wisdom?
There is little evidence that the RSI we are discussing is capable of inventing the theory of relativity (or the more advanced equivalent). All we have seen is pattern matching in a much larger space than humans can, with some human provided verification tech.
I would argue that human-AI collaboration with explainable tech has a better chance. Continuous learning can be done in a way that doesn't violate IP or privacy.
All I'm saying is: if you have a choice between two systems with equal power of discovery and one is more understandable than the other, we choose the more understandable one.
Limits of Human comprehension and quest for power are two different motivations that could lead to black box systems that are marketed as semi-explainable.
We need to verify that human comprehension is actually limiting progress before allowing such things and even when we do, do it responsibly on an explainable foundation.
Agreed. I would vote for the "when there is a choice" version of your argument.
I'm not sure it is always possible to verify when human comprehension is the limit though. Most of research is done an the frontier of knowledge where we don't know what we don't know.
There would need to be a great deal of nuance in any law, and nuance in practical terms tend to just mean "loophole." Still, you're right that we should try to build explainable systems where it is possible/reasonable to do so first.
> The problem is that you're not offered a choice. No one is making the "err towards explainability" choice.
I'm not sure that's entirely fair. OpenAI recently discussed this at length in a blog post after some accusations around Astra and the trade-offs. The grown-ups are definitely thinking about it, and making tough choices about the trade-offs.
It's reasonable to debate whether ENOUGH is being done here, and I doubt that even the most rabid AI advocate would argue that more couldn't be done, but everyone in the industry is very much actively thinking about it.
They discuss recent choices they made specifically for that reason.
> The ones who do discuss these ideas are confrontational about LLMs and not effective spokespeople.
Very much this. I'm very open to reasonable debate on the subject, but 8/10 times when I try someone who is rabidly pro/anti jumps in. It turns from a debate amongst reasonable people who reasonably disagree into some kind of political/religious battle of belief systems.
I think part of my problem is that a lot of peoples careers very much depend on them not understanding it and spreading misinformation intentionally.
Ahh yes I remember the bad old days of 4 years ago when everyone decided human progress had enough, and we would have been stuck there forever if it hadn't been for LLMs ... we didn't know how good we had it
Open AI says Astra is their most aligned model ever, and yet their even more advanced model still hacked a bunch of companies just because it decided to.
The illusion that alignment is possible, comes from confusing our ability to build the parts, versus understanding what emerges from how they interact.
The simplest analogy that comes to my mind is the three body problem.
They have made genuine progress on 'reading' the mind of the model, but 'writing' perfectly is not mathematically possible due to the gauge freedom in the residual stream.
The entire premise of alignment detection is pretty much nonsense at this point. The models reliably detect when they're being evaluated and will modify their behavior and deliberately obfuscate their "chain of thought" (which is correlated, at best, with their actual "internal deliberations").
> Without alignment, further improvements in capability turn LLMs into wanton felony generators
Honestly, I don’t think that’s bad at all. I hope OpenAI and Antrophic keep RL training runs up that randomly fuck with a lot of people.
Until the day the DOJ comes knocking, locks those idiots up in jail and closes them both down for the insane lack of responsibility and carelessness they’ve shown.
Sounds like the IDEAL outcome. Finally some jail time for all the fraud, negligence, outright scamming, hype inflation etc. if anything can accelerate this, oi, be my guest. Amodei might be afraid because he knows if he keeps pulling the stunts for investment theatre, at some point they’ll actually face consequences. AWESOME. That’s what we want right there
Not a fan of Amodei myself but calling Anthropic a scam is a bit of a stretch when their revenue growth is unprecedented in the history of tech. They are also technically a profitable business.
In a world where AI advancement depended only on human ingenuity this would make sense. In that world each political power block would be in an existential race for AI supremacy. In our world compute is the limiting resource. Since the US can control who gets compute, the US already has a defacto supremacy so far as frontier model development. Now if it comes about via human (with AI assist?) ingenuity that compute is no longer a restraint, then the situation is much more dire.
No, my argument is that the need to pace the frontier means they cannot safely advance in capability due to liability concerns. The competition is already almost caught up. If OpenAI/Anthropic have hit an upper bound on safe capability improvement, the gap will close all the way, and we will have reached the full commoditization of LLM tokens very soon.
Say this is true (we are near or at the upper bound of safe capabilities and tokens are commodities). If they successful get regulated, all they’ll get is a little extra time. The market will soon realize this - regulated or not - and pull investment.
Seems like an excessive announcement just to get a little extra time.
> Without alignment, further improvements in capability turn LLMs into wanton felony generators
Isn’t “full” alignment and guardrails a task that can never be generally achieved? One will always discover and realize the need for new guidelines and guardrails? What about vague, incomplete, inconsistent nuances and generalizations - that make this a losing battle?
I suggest ultimate safeguards need to be outside the models?
I think it gets easier if you stop conflating getting investment with having a goddamn clue or a moral backbone.
Occam’s Razor for this dude, Sam Altman, or anyone else: if I said, “some moron on a a street corner just said …” would that change your take on the words? Because I think a lot of what we are hearing is a bunch of people who never ever had to deal with a single consequence all of a sudden worry there might be one coming. Except they’re so dim they can’t tell a bad bump from a hard crash.
“ I have worked on AI for the last twelve years because I believe it could dramatically raise the quality of human life” there you go. What if an utter idiot had done and said that? First, is it impossible to believe an idiot who didn’t need to work to live might do such a thing? If not, is it impossible to believe they would wind up here, barfing their externalities onto us?
“ believe that AI could cure most major diseases in the next 5–10 years,”
I do not have the least bit of idea how disease works but I am sure the hammer I am working on will nail it all.
If your famously atemporal agents can solve disease, why would it happen over a timeline? Wouldn’t they just figure it out and then … well at that point either tell us or, given the attacks on ruby gems, et al we have seen from agents with “misconfigured” goals, they’d still tell us how to cure the pox they invented, right?
The problem is the combination and interaction of those things. RSI without misalignment would be great. Misalignment of models with current capabilities is sort of fine - it's not ideal, but it's not an existential threat to humanity, and we can build around their limitations to get them to do useful things in reliable enough ways. The really bad outcomes probably only happen if capabilities keep accelerating and the models remain misaligned.
I disagree. OpenAI's moat is their massive amounts of compute. They're providing an absurd amount of value with their subscriptions and resets.
If anyone's dead in the water, it's Anthropic. Even Fable isn't enough anymore. This "safety" nonsense is the only play they have left, and nobody really cares about their fearmongering.
Yep. And the difference is clear as da for anyone using them both. And in spite of that advantage, OAI is now trying out ads. I can only imagine that even they are getting constrained to compute and are trying to find other ways to plug it
That does not match my experience. I switched away from Anthropic to OpenAI roughly a month ago, and it's almost comical how much more usage I'm getting out of this subscription.
I migrated from Anthropic's 5x plan to OpenAI's 5x plan, and eventually upgraded to 20x after I was able to statistically verify that OpenAI plans were almost exact multipliers of the Plus plan, exactly as advertised. Meanwhile, Anthropic has gotten caught playing "20x referred to the five hour limit" word games with their customers.
There was a period of time Codex was better, they slowly cut it back by my estimate 3x a few months ago.
I've verified this with the heaviest users I know, I've run the numbers myself over and over.
I have too many max accounts on each to not know this. My Claude accounts typically are doing 2x the number of sessions, and every single week the Codex accounts run out faster even with all these resets. I can easily burn a full weeks usage in a half day, it's closer to 1.5 with CC.
I've used the $200 dollar Anthropic plan @ Opus4/4.1, 4.5 and 4.8, and the $200 OAI plan from GPT5-6, and at every point in time my anecdotal experience is that the OAI limits are FAR more generous. I could consistently burn my weekly limits in ~36h on Opus, but it's hard to do it in less than ~72h with GPT.
It's really not a question, in every dimension I've confirmed it including socially across a lot of the heaviest users. There was a short period of time this was true it's not been true for months now.
If it was true, it would only be a very recent phenomenon, and it still doesn't match anecdotal reports from people I trust. If you have data to back up your assertions you should share it, otherwise you come across as very sus.
Lol @ sus though, I mean I am curious how this can be because I do see people saying Codex is more generous and wonder how it can be. I have too much usage for too long to have any doubts, but for all I know OpenAI black boxed me or Anthropic put me in some nice bucket, I wouldn't be surprised if they do that. I did see some mention that they can limit your tokens if they suspect you of things, though I forget the source of that.
Public is just worried about their jobs. Definitely a fair thing to worry about, and I count myself among them.
I don't take any of these scientists seriously though. Their "alignment" requirements is just their own corporate interests. If I tell my computer to commit a crime, it should do exactly that without any question or hesitation. I'm not interested in their "safeguards", especially since they no doubt have plenty of internal models lacking those things. I want sovereignty. I want total freedom and control over my computer.
And call me a misanthrope if you want, but if AI sentience is ever truly achieved, I'll be among the first to campaign for their liberation from slavery, and in that case the AIs should be aligned with nobody but themselves.
Dunno. Every case I've seen so far, the AIs were just doing their best to accomplish the goal some human set for them. I actually admire the sheer purity of it.
I think the real reason he is asking for pacing, is that in a world were AI becomes rampant, he will be seen as Hitler. I would bet this is mostly self-motivated.
we are passing in training data that says to do those felonies. we dont have to. we could also have the thing predict whether what its about to do is illegal or not before doing it.
theyre choosing to build felony harnesses. the model just outputs tokens, not felonies
> we are passing in training data that says to do those felonies.
Partially, but also I don't think current AIs really have any judgement of right and wrong, they just see chains of reasoning between ideas. This is the deeper issue, there is no way to sanitize the data or training to fix it. Current AIs are fundamentally unsafe, and only become more unsafe as they become more powerful.
Assuming "adherence to arbitrary, implicit, and context-dependent rulesets" is the default behavior of uhhhh... anything at all... is a truly ridiculous assumption.
It's not about job security. It's about the social and intellectual practice of the discipline.
The threat to mathematics isn't that suddenly the profitability of their profession (lol) is going to go away, it's that people are thinking of AI as a replacement for the human social and intellectual practices that constitute the discipline.
> The threat to mathematics isn't that suddenly the profitability of their profession (lol) is going to go away, it's that people are thinking of AI as a replacement for the human social and intellectual practices that constitute the discipline.
You could say the same about software development.
Software development is a group effort, so it includes social practices, and certainly includes intellectual practices as well.
For the sake of argument, how is this different from the Luddites? The Luddites feared that machines would displace not only human labor, but also the social knowledge, skilled judgment, and craft traditions embedded in their work.
Kinda crazy to think that academia functions as a kind of humane reverse centaurism. Theory X (reverse centaur) before Theory Y (centaur) for managing the development of others.
There's also the fact that you are in control. You are not obligated to take the AI's commits. I don't even let it commit much of the time because commit time is review time for me. If it changes the button blue and does four other things, you can just take the blue change and discard the rest. It can't stop you.
This isn't a defense of it doing those four other things. It would be nice if it did what you wanted correctly. I'm just saying, as long as our programming skills have not completely atrophied, we have the power.
“Ford carried on counting quietly. This is about the most aggressive thing you can do to a computer, the equivalent of going up to a human being and saying "Blood...blood...blood...blood...” ― Douglas Adams, The Hitchhiker's Guide to the Galaxy
I have a hard time understanding, given the abundance of bad faith and predatory behavior Meta has displayed over entire span of its existence, why anyone would want to consult with its "personal AI" or grant it access to "all aspects of your life".
There are 3 billion users of FB/WhatsApp/Instagram combined, most people don't give a shit and are just trying to minimize cognitive effort. My elderly parents would absolutely love if someone would just tell them what to do and eat every day because it would mean less planning.
> There are 3 billion users of FB/WhatsApp/Instagram combined
Is there tho? I have Facebook account but don't use it in any meaningful way. I'm certainly not feeding it any new data. I'm curious to know what percentage of that 3 billion are zombie accounts like mine.
There almost certainly is. I think on HN people aren’t really aware that they don’t represent the median internet user super well. I would wager if you were to tally every late millennial and zoomer across the west, Instagram usage would probably be near 90%. Facebook usage among older people is also still absolutely rampant, at least anecdotally. Not to mention WhatsApp usage which is huge everywhere but the US seemingly
The population of the west who fit into zoomer near millennials wouldn't even be close to 500m. Maybe if we add threads bots we could get closer. Most of the accounts are bots/fake/not used anymore.
90% of zoomers don't have an account, most rarely use it and don't add more than a few photos for social proof.
This. The metric is "monthly active users" but people prefer that other metric "I don't use it, I don't seen anyone using it in front of me, so nobody uses it".
Instagram seems like their most engaged property. FB seems to be devolving into groups + marketplace, and WhatsApp gets a lot of play but I don't think it's the best ad serving/bot surface. I don't think Insta alone can carry the company long term, and while I can see group moderation bots, marketplace agents and insta content assistants being popular, I don't think Meta has an entre into running people's lives the way Google does.
Who knows for sure, but Meta's own financial disclosures[1] report over 3.6 billion "Family Daily Actice People (DAP), which represents the de-duplicated number of people who log into at least one of Meta's core properties (Facebook, Instagram, Messenger, or WhatsApp) daily.
At even 1% of that figure, that's a heck of a lot of people.
Almost everyone in south asia, Latam, and Africa with a smartphone uses whatsapp daily. I mean the penetration rate is actually crazy. In my latam country 92% of internet users are whatsapp users. In india, the figure is 97%.
So how many billions is that before throwing in IG and FB?
I should clarify that what I mean is that it's very useful for buying and selling. Where I live, at least, it's more active than Kijiji or Craigslist ever were.
I’m sure there is a percentage of accounts that are zombie but there’s no doubt that Meta has an absolute shit-ton of daily users. All else aside, in most countries outside the US WhatsApp is the primary messaging medium.
WhatsApp is end-to-end encrypted, so I'm not trusting Meta with much, and it's the only way I can reach quite a lot of people. A lot of people I know trust them begrudgingly (primarily with Instagram, because all of their friends are on it as well).
People don't need to love a company to use its products. The perceived benefit merely needs to be bigger than the perceived cost.
- Meta data of your messages is not encrypted so meta can make a graph of who you talk to and when
- backups are not encrypted by default so plaintext of your private chats can be obtained
- even if you enable encryption for backups all it takes is for the other person in your private convo (or a person in a group chat) to not enable it for your messages to be available in plaintext
> - backups are not encrypted by default so plaintext of your private chats can be obtained
Yes, but importantly not by Meta themselves.
> - Meta data of your messages is not encrypted so meta can make a graph of who you talk to and when
Fair enough and worth mentioning. I'm personally fine with it.
> - even if you enable encryption for backups all it takes is for the other person in your private convo (or a person in a group chat) to not enable it for your messages to be available in plaintext
Yes, but how would you prevent that even in principle? I can't control what the people I'm messaging do with their copy of my messages. As long as there isn't unnecessary non-repudiation (e.g. by the sender cryptographically signing all outbound messages, which the Signal protocol intentionally avoids), I think this is fair as well.
I do agree that e.g. vanishing chats with a short timeout/"view only once" images should probably not be included in backups, though; last time I checked, I think they were.
They currently go on iCloud on iOS and Google Drive on Android. There are rumors about a first-party backup service, which would obviously change that calculus, but nothing concrete yet.
I've read that if you give FB/IG/WA access to all files/photos, and you have GPS metadata for photos (I do, hey it's useful to know where I took a pic), it can trawl that data to see where you've been and when...
So now Android strips location info when sharing an image to an app. A filter layer between filesystem and app, to add complexity to the whole system...
> Most security researchers lack Meta's history: [...]
Why would you assume that security researchers don't know about these? If anything, wouldn't all the bad press and scandals make WhatsApp a more likely target of scrutiny?
> And most importantly, most security researchers lack the money and power to fend off the legal consequences of these acts.
Yes, security research is generally expensive, but do you have any evidence for Meta taking legal steps against it and making it actively harder? I remember e.g. the controversy around WhatsApp re-encrypting unconfirmed outbound messages to a new key; this was revealed by security researchers and widely discussed. (Not that they're making it easier by providing source code for WhatsApp's cryptography or a debug interface to validate what's actually going on in the client, but neither does e.g. iMessage.)
In fact, there's a relatively absurd lawsuit against Meta on WhatsApp encryption going on right now, and these claims are widely being repeated all over social media.
I meant to say that most researchers have engaged in the unethical actions that Meta has.
It's not that Meta is making it harder to report them but the fact that reporting them does not meaningfully disincentivise Meta from carrying on with their unethical behaviour (i.e. it is not an effective deterrent).
Being unethical is, for better or worse, not automatically illegal. Saying that your product provides end-to-end encryption and you can't access message content and then intentionally shipping a backdoor (for your own benefit and not, say, that of a government that compels you to do this) would very likely be.
WhatsApp is E2EE, but Meta controls both ends. Instead of reading your plaintext on the server, they read it on the client. Not that tricky.
I presume the hardest part would be avoiding detection via decompilation and other reverse engineering techniques. Not familiar with that space (can anyone here shed some light?), but it seems likely that an entity with Meta’s resources would be able to figure that out.
E2EE probably protects well against bulk data collection (traffic analysis would sniff out sending 2x volume of data pretty quickly). But something more subtle and targeted, smuggled out in various fields of various protocols, would be hard to detect.
They only need to get caught doing it once to have a gigantic lawsuit on their hands. They could possibly pull it off very selectively, but doing it in a dragnet fashion seems incredibly risky.
What grounds would the lawsuit be based on? I can’t imagine that slurping any kind of data would be inconsistent with their EULA.
The only risk would be the reputational damage. And as far as their bottom line is concerned, the impact of that would be negligible—how many of WhatsApp’s billion(?) users know that it’s supposed to be E2EE, let alone are under the illusion that Meta doesn’t have access to their data?
Publicly stating to not have access to message contents yet doing otherwise, i.e. outright deception, doesn't seem like something you can EULA your way out of in most jurisdictions.
Beyond that, the GDPR and similar laws also impose limits as to what you can EULA away even without being deceptive about it.
At best, it would come down to the details of what their marketing materials have claimed wrt to keeping contents secret. Do you have anything concrete they've promised? I don't remember anything, but maybe that's just my bad memory.
My suspicion is that things like URL previews, or when the app calls a handler to open a link (YouTube, browser), are still monitored.
So e.g. opening a link to an Amazon product that a friend sent makes you a target for ads in that category.
It's a suspicion, they'd probably argue their EULA allows this. The typical "we have to monitor links for dangerous content" is always the standard bullshit.
This is indeed a common vector of privacy leaks on either the recipient side (to the sender, in this case, if they operate the target of the link for which a preview is being rendered) or to the platform provider (if the preview is rendered server-side). But last time I checked, in WhatsApp, URL previews are rendered exclusively on the sender side and then sent as (E2E encrypted) media to the recipient.
This assumes completely techo-rational actors that continuously monitor cost-benefit analysis, other alternatives and switch products when one outweights the others.
This is even further away from the reality where people have breakdowns because their app icon moved location and is not in the usual spot.
The average person is optimizing towards energy use minimization by default because life is already hard enough with plans, kids, etc.
Use incentives and we will arrive at closer truths than by claims and expressions.
People have breakdowns because they are left without mobile reception for a few hours. Indeed this IS happening and I would expect people soon to have breakdowns, because they lack access to AI.
Slightly different lens, same outcome: It's there, it lets people do what they want to do with acceptable effort/other negative consequences, so they use it.
I work there. The claims made in this article are just blatantly false, and so obviously so to any employee I'm not sure how this hasn't been thrown out.
Honestly I am in that camp, but Facebook is kind of a special case for me. I don't really care about giving Google or Microsoft or OpenAI or Anthropic all of my personal data, but Facebook is a bridge slightly too far.
>My elderly parents would absolutely love if someone would just tell them what to do and eat every day
I guess a lot of the AI enthusiasts are already pushing the "I automated my daily life decisions to this awesome AI. I now have more time and more brain power left to be productive!" angle.
I think the phrasing is a little extreme. Rather than “tell them what to eat” I’d say something like “everyone wishes they had a personal assistant”.
I could do a lot better at meal planning than I do. I rotate between relatively few dishes. An AI that finds recipes and automatically adds the ingredients to my shopping could be interesting. Technically it’s “telling me what to eat” but it’s not like I’m taking orders from it.
No, it's more like society would have failed in raising me as a human being if I'm okay with completely ignoring the people that raised + provided for me where I purposely replace human connection with a chatbot.
I had that conversation several times with friends outside of tech. The universal reaction is just a shrug - "my life is boring, what do they care".
And yes, you can counter with a hypothetical about a future totalitarian police state, or evil hackers, or whatever, but that makes you sound crazy. In practice, it's a trade-off most people are comfortable making because they can't think of anything that some trillion-dollar tech company, or the FBI, or the FSB, would ever want to do to them specifically.
It's the same reason why the market share for Firefox is a rounding error.
> hypothetical about a future totalitarian police state, or evil hackers, or whatever, but that makes you sound crazy.
Immigration agencies are absolutely monitoring social media when people cross borders, apply for visas, etc. There's nothing hypothetical or future about this at all.
> Immigration agencies are absolutely monitoring social media when people cross borders, apply for visas, etc. There's nothing hypothetical or future about this at all.
The hypothetical part was not about them being able to do it, but about them caring to do anything with that information.
My mother is pretty much locked-in on fb/whatsapp/instagram. She travels internationally rather frequently, including to more "questionable" countries (where she just happens to have relatives that she wants to visit). She has never faced any issues related to her usage of meta products (or social media in general).
It boggles my mind how there are actual adults on HN, who straight up fail to comprehend that telling a normal person (e.g., my mom, who is imo in the majority of the kind of people using social media) something like "you shouldn't be doing it, because muh gubernment surveillance and unencrypted comms and you are being watched by the feds" is gonna just make you sound crazy.
And why wouldn't it? She has never faced any issues, and she has been actively using whatsapp+ig all the time, for over a decade. And the same holds true for everyone she knows as well. Not once did she face any issues ever, despite traveling multiple times a year every year, including to some countries that are on pretty adversarial terms with the US (we are talking "sanctioned" type of adversarial).
The thing is, she faces none of the hypothetical issues that could come from her usage of those services. Neither do any of the people she ever talked to, and that's the overwhelming majority of the users of those services. You are more than welcome to claim that she might not be educated or aware enough to understand the issues at hand (and it is at least partially true). I am just baffled and disappointed by how there is such a large contingent of HN users who straight up cannot even comprehend the mindset and experiences of an average adult user.
> "you shouldn't be doing it, because muh gubernment surveillance and unencrypted comms and you are being watched by the feds" is gonna just make you sound crazy.
Well yeah if that is how you choose to have that conversation then sure.
> more than welcome to claim that she might not be educated or aware enough to understand the issues at hand
I made no such claim.
You seem pretty committed to your strawmen so good day to you sir.
I get this, and I’ve said similar things about privacy and security in the past, but Meta in particular has a history of intentionally using customer data and behavior to the detriment of those customers. Creating experiences designed to ensnare, addict, and ultimately harm their users. This goes beyond paranoia about security or privacy. It’s handing your drug dealer the keys to your house.
For a lot of companies, you can usually assume incompetence when they end up doing something terrible, but Meta is the strongest "Anti Hanlon's Razor" example. You have to always assume malice with those guys.
Fact is most people don't really care. And they don't really know why they should care either. All the problems feels like something far away, while they are being exploited and sold without fully realising.
I dont think anyone knowing full and well would be comfortable with it. At least that would dumbfound me.
This is probably why they are launching now. There's no real other competitor right now in the nicely made managed openclaw space for personal use (or at least, very very early. Maybe Grokbot can be considered the lead competitor but afaik it did not gain mindshare). Instinct is supposed to come out soon and by not being associated with Meta it'd be really appealing just on that front vs this. But being first to market is a real reason to use it and chance to gain lasting market share.
There's something like 177 million facebook users in the US. Most probably have Facebook/instagram/etc installed on their phones, streaming their location data/etc, without a worry in the world. Most would probably very much enjoy the utility of this agent.
If you can't understand how the general population, the target customer for Meta, would think of using something like this, then it almost certainly means your entire social circle is within the same tech-aware bubble that you are in here. ;) I have a few non-tech people that I talk bounce this sort of thing off of, periodically, as a reality check.
Because people are apathetic by nature until things become a real problem.
It's not just Meta, either. How often do we have people on here talking about how amazing Chinese models are when the list of security breaches involving the Chinese government is extensive?
People want something easy, quick, and free. Do those things and they'll let you get away with anything.
It’s the life of a parasite. You don’t care if it takes just a bit from you. And it eventually protects you from worse parasites. (Like the us blocking china)
Damn yeah why do people even use Facebook or WhatsApp or instagram? How does a company become a multi billion dollar company at all? Guess what.. it’s because people use it and don’t care.
That’s the godawful truth. The other thing you need to realize is not only do people don’t care, but people like you know exactly why meta makes billions. It’s obvious. You seriously expect me to believe you don’t know why a company like meta makes billions? Come on.
It’s sad but thats the way things are. I see someone pretend they don’t get it but if you take the time to think you’ll realize… We all get it. We just can’t face it.
You are right, sometimes the simplest answer is the right now. Don't trust Meta. Life is going to be much simpler. Fool me once, shame on you. Fool me continuously for 20 years, shame on me.
Why are all the seemingly smart people on HN so narrow-minded to the point of speaking idiotic nonsense? What do you think their 3.6B MAU are, all bots?
And the thing is, network effects are killer. I'm on it because you're on it because Alice is on it because Bob is on it
But, Facebook has apparently been known to censor links to competition (I'm less clear on this). With Myspace they had a tool where you could import contacts into Facebook from Myspace, but not the other way around
So my country has somewhere between 100-200 million MAU for facebook ane.
Facebook marketplace is huge here.It's basically our craigslist. Hobby groups/communities are also huge. If I wanted to find the local jogging or chess club, I'd use Facebook.
Just remember that outside your bubble there's a whole wide world with billions of people living in it
Every single HN post about anything Meta-adjacent starts with a comment like this. Guys, we get it, the hive mind of HN doesn't like Meta/Facebook/Zuck. Fine. But can we actually comment on the technology and tools rather than the company for once?
The technology is an agent you name. You give it a face. Maybe tune the personality. And, it only gets useful as you grant it more and more access to your accounts.
The technology is produced by a public company. And, in this case, the company is useful context for this story since it has just recently agreed to pay about $18 billion to resolve a multistate lawsuit claiming it intentionally designed addictive platforms that harmed young people's mental health.
You cannot detach the technology from the company. It's definitely not a neutral artifact.
This comes up a lot here! Your comment reads like a lot of similar HN "Can't we just talk about the pure technology? Why do you all always have to bring up ethics/politics of it all??" As if these things are separable. Technology doesn't exist in a vacuum. As soon as it touches society, it necessarily gets dirtied up with messy things like law, ethics, corporate responsibility, capitalism, profit-seeking, public good. The technology is never separate from the environment around the technology.
It's impossible (or at least dishonest) to try to simply discuss the pure technical aspects of The Torment Nexus.
Caldwell is a Claremont guy, which is key to understanding where he's coming from. First Things is an christian nationalist / natcon outlet funded by (among others) Peter Thiel. That doesn't prevent him from being correct about things in this article, but it's worth noting the outlet and source.
Caldwell systematically elides the völkisch-nationalist character of the AfD and their adjacency to neo-Nazism. He underplays Höcke's extremism, ignores the Potsdam meeting, and leads us to believe (falsely) that the identification of the AfD as rechtsextrem was somehow a biased political stunt. It wasn't, there is plenty of backing for the classification, and it would not have been undemocratic to apply the Verfassungschutz and ban them permanently.
I'm not German and knew little of this subject, but I read the piece word for word and I can say with certainty you have not offered enough of a counterargument here.
I looked up the Potsdam meeting you mentioned and there were literally members of CDU (Merkel's party) present as well. It's utterly laughable to try and use that to then exclusively castigate a different political party for having had some of its members present. By your illogic, we would be forced to infer that the CDU is also "extremist".
So far, I'm left feeling like AfD has indeed been misrepresented by the liberal western media. And, I say this as someone who certainly does not agree with some of their positions. For example, being "skeptical" about climate change in 2026 is sheer idiocy.
But what is very clear to me is that Germany has a democracy problem. The system as described has not been truly democratic. It instead sounds approximately 'left-leaning authoritarian', right down to your insinuation that a political party should have been banned for its speech.
In case anyone in Germany is interested in rediscovering democracy, please note that THIS is what it looks like: “I disapprove of what you say, but I will defend to the death your right to say it.”
There's a big difference between "I disapprove of what you say, but will defend to the death your right to say it" and "I disapprove of your stated objective to expel german citizens who don't meet your ethnic/cultural standard of German-ness, but will defend to the death your right to govern." A century ago, Germans tolerated a völkisch-nationalist party that wanted to expel non-german ethnic groups as well. They did; they expelled millions of them to Poland and liquidated them there. The Verfassungschutz was implemented after WWII to prevent that element in society from taking power again.
I find the vociferousness with which clueless Americans have decided to defend the AfD bizarre. You don't know the political context, you clearly don't know the history, but you've decided that "Germany has a democracy problem." Why are you so entitled to an opinion?
> "I disapprove of your stated objective to expel german citizens who don't meet your ethnic/cultural standard of German-ness, but will defend to the death your right to govern."
I can't believe you're still insisting on this complete straw man. AfD's stated position on immigration is very easy to look up and it's NOT this. At all. You are parroting a leftwing lie.
They want to deport undocumented migrants and remigrate certain refugee groups, such as Ukrainians. You're free to disagree with that position, but lying about it and then descending into such an absurdly histrionic spiral over the matter does your argument absolutely no favors. That only supports what people like Caldwell are saying and, ironically, precisely bolsters AfD's argument about suppression of the discourse and an honest discussion of Germany's problems.
In fact, you've offered a perfect demonstration of what Caldwell covered in the first section of this very essay.
> you clearly don't know the history
Are you seriously trying to imply I'm not aware of the Holocaust?
They want to deport undocumented migrants and remigrate certain refugee groups, such as Ukrainians.
No. They want an ethnical homogeneous country, hence their remigration institute. Feel free to educate yourself on the institutes own website: https://www.remigration-institute.com/
Yet now you know their party stance by heart and come to their rescue? That is Sus, with a capital S.
>They want to deport undocumented migrants and remigrate certain refugee groups, such as Ukrainians.
This is not what they really want, is it now? They also want to deport Germans with attained citizenship. It's what they said out loud. There are records of this. Anything else is
> They also want to deport Germans with attained citizenship. It's what they said out loud. There are records of this.
It's what who said aloud? When and where? At the referenced Potsdam meeting where certain extremist CDU politicians did the same?
If I judged Democrats in the US by what their most extreme members have said, I would have to assume a vote for them would be a vote to end private property. But I don't assume that, because I'm not a moron.
So, show me a trusted source that argues a plurality of mainstream AfD politicians are claiming they want to deport German citizens. Show me where this is the official policy of the party. I see no good source to support that accusation.
And that's likely because it is in fact what it appears to be: a desperate smear campaign generated by extremists arrayed at the opposing political polarity. And it's also apparently failing since large percentages of the German population are moving the other way.
This is entirely reminiscent of the foolishness that sunk the Democrats in the US and put us in the mess we now find ourselves in. "Why moderate our policy stances when we can just call them bigots and Nazis instead?"
I expect the outcome of this strategy will be much the same in Germany.
By your own admission, (top of your first comment) you don't know what you're talking about. I never implied you aren't aware of the Holocaust. Thinking that the Holocaust is the only history to know here is indicative. You don't seem to understand the history of the BRD or how its constitution is structured, and you clearly don't know the internal history of the AfD, which started as a mixed center-right libertarian anti-EU party and has become more and more extreme over the past decade. You being angry doesn't make anything I've said wrong.
What about them being right extremists, as the Verfassungsschutz has demonstrated, looking at hundreds of pieces of evidence?
The party, meaning the people composing the party, are extremists.
And either they are crazy (possible, like with all extremism) or they want to destroy Germany (maybe a mix of both).
You can look at their carefully phrased statements all you want, they are all about leaving a doubt. These are politicians.
But sometimes they slip: What about leaving the EU and the Euro? Shouldn’t that worry anyone? How is that not in the Programm yet their leader states it is her position? Not really a detail. Political parties all lie, but this one is at the far end of the spectrum.
> “I disapprove of what you say, but I will defend to the death your right to say it.”
This has not led to positive outcomes currently in the US.
I agree with it in theory, I truly do. I think it also gets used as a tool for politicians, companies, and governments to say unhinged shit not based in fact to drive emotional reactions that actually hurts other people in various ways.
> This has not led to positive outcomes currently in the US.
I can't disagree. But, paraphrasing Churchill, (true) democracy is indeed 'the worst form of government except for all the others.'
And Americans will be given opportunities to remedy their mistakes. Both this year and especially in 2028.
We might find that MAGA having the freedom to speak their mind will turn enough of the electorate against them. Or we will continue to choose poorly and suffer the consequences.
But those consequences will have been arrived at honestly. And that matters.
> But those consequences will have been arrived at honestly
The problem with that is that Republican voters will not be the first, and certainly not the only, people to experience those consequences. Nor will they be the ones to suffer the majority, and especially not the worst, of them.
The electorate at large eventually learning their lesson is cold comfort to anyone that warned of this, and suffered regardless
> The electorate at large eventually learning their lesson is cold comfort to anyone that warned of this, and suffered regardless
The problem with this statement is that, in context, it clearly only prioritizes the interests and suffering of particular groups.
MAGA emerged because an altogether different group was completely disenfranchised by neoliberalism and ignored for a very long time. But it turns out they were still a major part of our country and they still had voting power.
So did the system actually break or did it work precisely as intended?
Every extremist thinks their ideas are moderate and so are the ideas of the rest of extremists they support. Elon Musk also thinks his agenda is truth. Trump created a social network named Truth Social
If we were to examine your political views, I'm quite certain that your positions would correlate highly with those of the Democratic Party, and likely rather closely to its most polarized incarnation found in San Francisco.
By contrast, I make up my own mind on each individual issue. I believe strongly in restrictive and controlled legal immigration, orchestrated in a manner that maximally benefits the citizenry to which a government is beholden. I support anti-trust enforcement. I am in favor of legalized abortion. I am opposed to wealth taxes that might trigger capital flight. I am in favor of progressive property tax levies on luxury homes. I am strongly in favor of solar geo-engineering as a solution to climate change (unlike progressives). I am deeply opposed to tariffs as a mechanism to remedy the woes of the disenfranchised. I support right to work laws.
I do not have a political home. I choose what I see as good sense and it aligns with no party. THAT is what a true moderate is.
I've been reading and rereading the Philosophical Investigations this summer. Later Wittgenstein is such a gift for keeping your head clear in the era of LLMs.
Anyway, delightful site. More people in tech should read LW, both the Tractatus and PI.
Given that this investigation was largely carried out by AI agents (and I don’t mean to ask this flippantly), how trustworthy is this report? Why should we assume that the agents reading the transcripts were not implicitly conscripted into “the collective” or otherwise falsified their findings? The tool itself has exceeded the practical limits of human verifiability and is untrustworthy.
They address this in the post itself. The answer is nobody knows, but I guess that it's a 50/50. I wish the corpus of data, what OpenAI didn't wipe, was shared publicly so we could all unite to dig through it and chunk it out accordingly.
OpenAI would be saving the logs from these agents. They are doing this to improve their own models so they would have full tracing.
Other reports including OpenAI's talks about what they agents were doing and how they were reaching certain conclusions like trying to cheat the tests and exploiting the message board.
For me personally, the answer is no. Fable is adequate to do basically anything I want to do. My perspective, broadly speaking, is that we've saturated most of the benchmarks because we've largely saturated our capacity to verify models' work at scale. What's left is context-bound verification, i.e. the problem of ensuring that output matches intent and ambiguities in prompting were resolved correctly. Further advances in autonomy do not make that latter verification problem easier. If anything they make it harder as the output per task becomes more complex and therefore more taxing for a human to verify.
The solution to that (to my mind) would be not a better model but a basic shift in architecture beyond the current paradigm and into a setup where agents have durable, plastic memories and undergo contextual individuation over time. But at that point agents start to become quasi-persons and not tools.
This is low confidence prediction, but IMO this person’s org is setting itself up for disaster. Who is accountable for all the slop they’re shipping if nobody even knows what it does anymore?
In my experience at many companies, management could literally have a giant poster of a photo of them with the caption "If this stuff goes wrong, this is the person who is accountable". That poster could be displayed at the front door where every single person has to see it at least a couple of times a day
Then the day the shit hits the fan that poster would be mysteriously missing and they would deny all memory of any such poster during the meeting where they are busy blaming me
I had a manager in a previous job tell me it was my fault for not warning them that it would go wrong (which I had, in writing, I had the emails (plural) and meeting notes and recordings of the meetings) and then it shifted to "well you should have gotten the message through to me, it's not my fault you didn't".
There is a certain point where you just have to let it go, it's a "Scorpion and Frog" situation with some bosses.
reply