Worth reposting a follow-up tweet from the author Tom Zahavy [1] after this made the rounds on X/Twitter recently:
> A few reflections on my "LLMs Can’t Jump" paper:
> My position paper recently got some traction here, so I wanted to share a few thoughts and clarify a few things.
> First things first: some people are framing this as "DeepMind is throwing cold water on AI for science" or claiming the paper argues LLMs can never make real scientific discoveries. This is NOT the case.
> This is a personal position paper, not the company's view on AI for science. This is also not my position. As a core contributor to AlphaProof (the first AI system to win an IMO medal), I know firsthand that my colleagues at DeepMind, other frontier labs, and academia have made amazing discoveries with LLMs and will continue to do so. This paper is NOT an "LLMs are a dead end" kind of thing.
> Rather, the paper is the result of a deep dive I took to study the invention of General Relativity. I wanted to explore what it would take for a modern AI system to make that exact kind of jump. Specifically, I focused on the equivalence principle—a key axiom that Einstein formulated through thought experiments grounded in his physical intuition. I was trying to figure out what it would take to give modern AI systems that sort of thinking.
> Giving AI this specific capability isn't necessarily the most urgent thing to do next. It is very likely that improving our current recipes will lead to many exciting discoveries in the near future. In fact, that is what I am personally working on these days (sorry to disappoint you!). It is also quite possible that I am wrong, and that simply scaling our current systems will lead to new inventions in physics and elsewhere.
> Nevertheless, this was my position last winter when I wrote the paper, and I'm sticking to it. I think that there are a few interesting ideas to explore in this space which could influence the next generation of AI systems. I was very lucky to receive a lot of interesting feedback about this position—thank you for all the messages!
>> Specifically, I focused on the equivalence principle—a key axiom that Einstein formulated through thought experiments grounded in his physical intuition.
It's weird because the equivalence principle is very unintuitive. Aristotle's Mechanics does not have it. It took almost two thousand years to discover inertia that is the most simple version of the equivalence principle. Einstein understood the idea of the the equivalence principle because he had a physics degree, not because he feel that in real life.
Moreover, if you ever have to study or teach Quantum Mechanics, physical intuition gets in the way. A lot of properties contradict the physical intuition but after a while you get use to them. If we continue with Einstein, the photoelectric effect does not aperar in real life.
But he didn't have one :) I guess unless you use the raw images, I think the cameras fix a lot of the noise caused by the Poisson distribution. I imagine a movie where Einstein get bored in the patent office and decide to open a photography shop, perhaps with very grainy underexposed photography is possible to notice the quantization of photons, and then he gets a Nobel prize.
Another bad idea for a movie is a watermillpunk universe, where during a practice on a hot day the best ever curling player discover inercia.
There were already prior reported cases of the phenomenon. Einstein provided the mathematical foundation explaining it, in particular that the discharge is quantized, indicating that it came from an electron. Millikan proved it a couple of years later with his oil drop experiment and by measuring the elementary charge of an electron, garnered his own Nobel:
I agree, I agree, but I think the guy only read about it and if lucky measured it once in a lab class. I'm trying to imagine a fake world were he could get enough experience in real life to justify the article claim "grounded in his physical intuition".
That’s funny, when I first learned about the equivalence principle, my first thought was “of course!” I have always found it to be very intuitive. The great leap is being able to frame it that way.
I worked on the early iPod scroll wheel, before there were advertisements for it and it was in common use. I found the UI interaction odd and unintuitive, particularly the "menu" button and the dead end you hit when "playing". Of course by the time the demonstrator ads came out and everyone was talking about how easy it was use, I'd already spent 10s of hours on it, and it WAS second nature. Being first and embedded in the culture of the time has huge UI advantages.
(Note that the HP Chipmunk 9836 also had a scroll selector wheel in 84)
Most people are UI bigots. Once they get used to a first something, they expect everything to work that way and hate learning anew. They get stuck on keyboards, mice, trackpoint nubs, trackpads, trackballs, scroll wheels, or touchscreens and refuse to move on. Of course there are 'objective' performance tests for each including Fitt's test of accuracy and latency, as well as, cognitive load. I guess once you have a hammer, every screw looks like a nail.
So I'd be a bit careful with the "of course!". It may well be obvious only because that is the first mental model you latch onto.
If your glib comment is referring to me as a techbro and doing the annoying worst thing, then maybe you should explain how I am comfortably abusing the goalpost fallacy given the explosion of agentic AI capability.
I simply point out here, that fully accepting the paper’s premise, the paper’s conclusion isn’t limiting on frontier AI reasoning agents. The paper posits the necessity of multimodal world models and the limitations of LLMs. Frontier agents aren’t simply LLMs and do increasingly integrate increasingly capable multimodal models.
> Be kind. Don't be snarky. Converse curiously; don't cross-examine. Edit out swipes.
> Comments should get more thoughtful and substantive, not less, as a topic gets more divisive.
> When disagreeing, please reply to the argument instead of calling names. "That is idiotic; 1 + 1 is 2, not 3" can be shortened to "1 + 1 is 2, not 3."
I mean Einstein had help, he was networked with the best scientific minds of the planet and his discoveries were grounded in experimental results that contradicted existing theories (at least for specialized relativity), and without Riemann’s work he wouldn’t have been able to formulate his theory either. So not sure if AI couldn’t do that if you kept feeding it with new research results and let it correspond with top human scientists. Einstein was a genius but I don’t think his thought process is beyond what an LLM could simulate. And again this is probably the most impressive scientific achievement in theoretical physics in the 20. century so maybe it’s hanging the bar a bit high for LLMs.
Exactly. The post read to me as another variation of the denial that people with expertise are reaching for right now. My sense is that we as programmers went through it over a year ago already (perhaps not all of us, but at least anyone paying attention), and so it's easy to overlook that it's still new to people who do other forms of "knowledge work," i.e. people whose identity is bound up with their expertise.
My working theory at the moment is that for programmers it was relatively "clean" and took the form of an inside-out transformation of the work, where AIs directly produced the central work product more or less adequately and relatively early on, but for other forms of work it will appear as some mixture of inside-out (in which case it will appear similarly first as a tool, then as something more than mere tool) and outside-in (the things surrounding their work and the supports their work processes rely on will be progressively automated). This is going to give rise to all sorts of pathologies in the white collar world, we'll get all kinds of variations on denial/negotiation, and so on, until it fully transforms the division of labor.
One interesting point of reference here: Yuval Harari gave a talk recently about the radical changes that will take place relatively quickly, in which he noted the AIs are not quite as good at writing as he is yet, although he expects they will be relatively soon. He then gave the timeline for what he considered "soon": 10 years! So we find the denial ("I still have time, they're not as good as me yet, maybe in 10 years...") even among the most vocal "prophets," among those supposedly most wised-up to what's going on and where the capability frontier lies.
Did you maybe consider that the 10 year timeline is not denial, but actually well educated reasoning based on Yuval’s experience and understanding of the problem space? You shouldn’t dismiss people’s thoughts as denial just because they don’t match your perspective.
I would guess if asked Harari would actually revise that lower. The point was: even those arguing most strongly that AI is an autonomous historical force can slide back into the mere abstract recognition of it. In context, Harari was saying "I'm speaking to you as a writer now, but my own standpoint will eventually be undermined." I'm saying "yes, and your 10-year timeline suggests you along with all of us are not taking your own thesis seriously enough."
Have you considered the possibility that you're simply not as good as the experts, and that your experience of LLMs being capable of performing your work up to your standards doesn't imply that experts are necessarily in denial?
Being fairly at expert level in a "solved" domain (for some definition of solved) but also having near-expert level proficiency in a non-technical as-yet "unsolved" domain and watching the process repeat there (and watching how people react as inroads are made progressively deeper) is basically my own standpoint. But I am intentionally using "solved" and "unsolved" very loosely here: you can get stuck in the thicket of arguments about verifiable domains, what it means for a domain to be solved, whether a set of evals can tell us something has been solved or not, and so on. One way to avoid that (as my original comment pointed to) and focus on what matters for us is to look instead at the effects being produced in the work process and on the division of labor as a whole.
He's saying: "Expertise improves AI-assisted work, experts can steer better and get more out of models."
The obvious conclusion for anyone is "therefore experts will remain the indispensable and specially rewarded center of the production process."
That this is appearing exactly now, and in this form, strikes me as extremely suspicious. I don't doubt the author's sincerity on the surface. What I suspect is that anxiety over the possibility that the (unstated) conclusion might be false (!) motivates the argument in the first place.
I'm asking the question, "Why is this argument appearing now?" At least one reason seems to me to be, "because we're afraid of what the world could look like if it's not true."
However, I personally agree with the author and I don't think his argument is necessarily motivated out of an anxious fear. On the contrary I think it may be motivated out of a sense of extreme exhilaration and empowerment.
Because experts (like myself as a programmer for 15+ years) who are using AI in many fields are suddenly empowered and much more useful than we were before AI. My employability and value has gone up and not down, precisely because of being able to apply my expertise with AI, which people without expertise simply cannot do. I am a professional programmer and also owner of my own startup.
Let me give you a concrete example that I am dealing with at my startup. I'm a small business owner. Before AI if i wanted to produce production quality video for marketing it would taken such a huge budget and such a large team of people (or an expensive agency) that I wouldn't even have considered it due to the enormous cost. I'm talking about Apple quality video production which takes millions of dollars to produce.
Not anymore. A single competent person with AI can replace an entire marketing video production department or agency. But expertise is key here: knowledge of film terminology to be able to describe the effect you want, and ability to use video editing tools effectively, as well aesthetic taste. I as a programmer with no filmmaking experience don't even know how to write the prompt which makes the video that i want because I don't even have the terminology. But a person with that expertise has suddenly become more employable and more valuable to my business because I as a small business now have the capability to create Apple quality marketing videos.
So AI actually created a new job for an expert that would have otherwise not existed because it was outside the budget of small businesses. Previously somebody like that would have been employable to only a few large production studios but now they become employable by almost any small business.
> A single competent person with AI can replace an entire marketing video production department or agency.
This is a good example of what I was pointing to with "fully transforms the division of labor."
I'm personally in a position similar to yours, but as I watch the different moves the labs make, I see the edge we've been handed (for now) also constantly under attack from different angles. This is why it seems to me that so many who can make the most of things for the moment also feel the clock is ticking.
I've seen this sentiment so much in the past month that it's starting to make me worried that the labs will silently make this the default output to satisfy users. I do recognize the outputs can be so information-dense that it strains reading for many but they absolutely aren't nonsense or mere affectation. What people are asking for now amounts to asking for less intelligence per token only a few months after everyone was demanding more intelligence per token.
Yeah I worry about this too. I have a remote agent that writes PRs for my job, and at the request of other devs I implemented a simplified writing skill that does the technical english thing mentioned elsewhere in these comments. It's definitely easier to read and the other devs seem to like it, so I'm okay with keeping it, but I haven't changed my personal claude settings because I never had a problem with the density of the output myself. Probably a side effect from having a multi-disciplinary background.
How long does a train track last? Does a fiber optic cable last? Both are greater than 30 years, both will persist (relatively well) without use, and the benefits of scrapping or removing them are minimal. This allowed future companies to take advantage of them. Even if GPUs running at high load last five years, if the data center they are in goes bankrupt (because the AI bubble bursts), it’s likely they’ll be stripped and sold to make way for more productive CPU-based uses and to recover some of the cost of the bankruptcy. The surrounding buildings and infrastructure will have longer use, but it doesn’t translate to a net future benefit with AI.
That reminds me of saying the early Google was doomed because they put all their money into cheap pcs acting as servers - how long do those last? But the enduring value was Google dominating search which was worth billions/trillions, not the heaps of pcs.
Same here - the main value is in dominating AI or something like that, not in the stack of hardware.
We're going to see the battle intensify here because "spikes" of ASI are emerging that can't be ignored. Models are now better than humans in certain domains or for certain tasks, which means capital as a moat is being eroded. This is why you see people like Jamie Dimon sounding the alarm. Most of the talk until now has been about how the models would be a serious problem for labor, but if that was true how could it not also be an issue for capital?
This is where we need the hardware companies and neoclouds to start speaking up. The labs want to elevate matters from the level of civil society (basically, competing firms) to the State (enclosure), and as always, in the name of security. But other actors in the same ecosystem have strictly opposed interests here, and are equally if not more credible as far as the State is concerned. If players like Nebius, Baseten, Fireworks, etc. among many others including obviously Nvidia, Dell, AMD, and so on don't get ahead of this they will be sacrificing trillions.
Exactly, it's about taking this stuff off the open market where anyone can judge it and there's competition, into government contracts where competence to judge the offer is scarce or absent, and they can ask much higher prices. And with this much investment at stake, any lie that sells the narrative will serve.
Yes, and given the nature of the current administration, whose actors are not inclined to see themselves as independent competing capitals among others, but rather as privileged capitals, and therefore more inclined to move towards taking an interest in the process of enclosure, ensuring it includes them, the hope seems to lie with companies at the hardware layer who have an interest in seeing the diffusion of intelligence play out freely at all levels of society.
The labs have to be told NO--the problem they're dealing with, that model outputs give the game away, and that in turn the distiller becomes the distilled, is a fundamental problem they have to figure out how to deal with without going to the State.
It's a form of denial. We're getting another "de-thronement of man" on the order of Copernicus and Darwin. Some get excited, others turn away in horror. Negation is the outward expression of the desire to keep human intelligence wrapped in its mystical veil.
One popular idea is that these systems will asymptotically approximate human intelligence because they're trained on mostly human-written texts. Not only is that untrue, it's also directly contradicted by our experience with previous RL-systems, where they seem to breeze right by human ability without even the slightest hiccup.
Most human systems are much, much, much more complicated than most closed world games (which is where RL approaches have seen massive success, mostly through self-play).
Like LLMs are great, but I honestly can't see us getting actual general intelligence out of them.
They already possess general intelligence by many metrics. Sure they miss a few, but that's nitpicky and goalpost shifty - lots of humans make errors of all sorts as well, or incur brain damage limiting them in one or a few areas of intelligence - we don't then say they are not general intelligence anymore.
I think maybe you mean superintelligence, which is a more fair critique.
> They already possess general intelligence by many metrics.
Can you share the metrics you are using for this assessment?
They are really powerful tools, but a quick glance at their thinking tokens (which is a bad name, tbh) rapidly disabuses me of the notion that they are general intelligences.
They possess large amounts of crystallised intelligence (i.e. they have a lot of knowledge), but their fluid intelligence is definitely lower than the human median.
My take is that fluid intelligence (quite a bit) below human median is still general intelligence (that is, general intelligence doesn't mean no gaps).
It feels like we have collectively goal-post shifted the definition of AGI to be closer to that of ASI.
One of those "what do you call a doctor who graduated at the bottom of their class" type things. I think despite their genuine deficits (and there are many), frontier LLMs have basically cleared the minimum AGI bar.
Though I realize a lot of people don't agree with this take :)
Approximate as a limit and not surpass, where the hidden hope actually seems to be that they don't surpass. There are lots of variations floating around, from simple metaphors like "LLM as librarian that speaks to you," to even Sutton's remark that rather than being a case where we've taken the "bitter lesson" to heart, LLMs may be a yet another case where we're limited by it.
A justification for this would be: states choose these providers to accurately present building codes (keeping up with the right versions/revisions is a lot of work, etc) because there's risk in simply publishing the PDFs and leaving it to basically anyone to curate. But the providers (like ICC) also tend to lock these deals in with state laws (not sure if they lobby for it, but I would imagine they do). When you think about who needs to have access to the correct and current building code or electric code or whatever for regular reference, it's really your lawyer, plumber, electrician, architect, builder, etc. and in that context $170 or $130 isn't a bad deal.
It also locks up things for homeowners that want to DIY a solution fixing a house. Obviously people fix things without looking at the codes (and there are plenty of horror stories out there), but if we opened up house codes for people to actually look at and refer to, homeowners could potentially better find how to do things properly, especially with AI.
This would potentially lead to a whole lot of injured people. There are already DIY jobs that are safe for homeowners to do. But some really should be done by a professional.
It's one thing to make the POP3 standard free; worst case your mail gets lost. It's another to make the standard for the electrical code free, so that people can incorrectly implement its quite complicated rules, and result in things like fires and electrocution.
This is a patently ridiculous take. People are still DIYing the very things you are concerned about, just not knowing whether what they did was up to code or not. Just look up DIY on YouTube and see.
Having access to the code won't change that. As laypeople it is too complex for them to understand. Often they simply ignore the code if they personally believe the code is unnecessary. It takes years of experience to understand why you should follow instructions you'd rather not. That's part of why the apprentice period is so long.
I have never seen a YouTube video where anyone even mentioned the code unless they were a professional or engineer. Most people are not very smart. Making the code free isn't going to make them smarter. I'm the one guy on subreddits telling people not to modify random beams in their house unless they know how it ties into the rest of the structure to determine the structural impact. 99% of people reply that I'm over-reacting. That's how they act towards the code. Slap it real hard and if it seems solid it's good to go.
> As laypeople it is too complex for them to understand. Often they simply ignore the code if they personally believe the code is unnecessary. It takes years of experience to understand why you should follow instructions you'd rather not.
Same is true about anything from cooking to crocheting to brushing teeth. Documentation isn't written for those who know, but for those want to know, and those who know better and would rather not follow it tend to not read it in the first place. It's as true of ISO standards as it is of instruction manual to your induction stove, or electric toothbrush.
> That's part of why the apprentice period is so long.
I'm going to bet that technically, post-COVID, it's "$300 + few hours of videos and a quiz" long.
(At least that seems to be the case for basic electrical work over here, in Poland, according to the electrician who did lights in my apartment the other day.)
> I have never seen a YouTube video where anyone even mentioned the code unless they were a professional or engineer.
Perhaps because they have no access to it without spending unreasonable amounts of money on it, unless they were a professional or engineer? YouTube videos are bargain bin education. A random video you referred to costed less to make than getting a bootleg copy of the code would. Probably whole channel did.
> Most people are not very smart. Making the code free isn't going to make them smarter.
Nothing will help "most people". But the rest would benefit.
BTW. most contractors are in the "not very smart" group too, which is a possible reason why you won't get many answers from them - they're either unable to, or plain unwilling to entertain people asking "why".
> I'm the one guy on subreddits telling people not to modify random beams in their house unless they know how it ties into the rest of the structure to determine the structural impact. 99% of people reply that I'm over-reacting. That's how they act towards the code. Slap it real hard and if it seems solid it's good to go.
It's a separate subject, but thing is, they're probably right.
In my experience, the difference between a load-bearing wall and a regular one boils down, for most people, to the question of whether they need a hammer drill or will regular cordless drill suffice. The hole is happening either way. I used to be the guy worrying about it a lot, until I noticed that even contractors don't care. Unless I'm literally asking them to cut a new entryway through the load-bearing part, they don't even parse the question.
Was I right to worry? Are they wrong? Well, I presume no to both, or else apartment building collapse would be daily news. Then again, I can't tell for sure, because of people in power thinking "it's too complex for [laypeople] to understand" and gate-keeping standards and codes.
> > That's part of why the apprentice period is so long.
> I'm going to bet that technically, post-COVID, it's "$300 + few hours of videos and a quiz" long.
Try 5 years.
> BTW. most contractors are in the "not very smart" group too
And that's why there's a code, and inspections, and penalties. They know that if they mess it up they could be fined, sued, or have their license pulled. Even with all their training they can still mess up. Which is why just giving a homeowner the electrical code and saying "good luck" is even less likely to work out. Many contractors learn just enough to do one specific kind of job, and then just keep doing that job. But they do learn the right way to do that one job.
> Was I right to worry? Are they wrong? Well, I presume no to both, or else apartment building collapse would be daily news
Buildings are usually engineered so that if a single structural member fails the whole thing doesn't come crashing down, because sometimes shit happens. That doesn't mean it's a great idea to weaken the structural integrity of your house.
When it comes to putting a hole in a beam or column, it depends on the member, and there are specific sizes the hole can be, locations, etc. Sometimes contractors don't go into the details and say "it'll be fine", because the way they do it is fine, because they learned the right way a long time ago and don't tell you that part. They might even be dumb enough to not realize that you would do it a different/less-safe way. But they'd probably do it the right way themselves, because that's how they learned it.
However, buildings do fail on a daily basis. They just aren't news, or aren't catastrophic / life-threatening. A lot of bad stuff happens on a daily basis and almost none of it makes it into the news.