> LLMs interpret (so, "execute" in a way) natural language
You could say the same thing about human programmers, but I've never heard anyone say they think that programmers "execute" Jira tickets.
> they have internal logic that assigns to the sequence of tokens in context a next token.
I don't think this means what you think it means, because it has almost no information content relevant to what we're discussing. The probabilities that are most relevant at the level we're discussing are satisfying a reward function from post-training, which approximates to:
"What is the likelihood the solution the agent is pursuing will be marked correct by the automated grader based on the full prompt and other context provided?"
It still has to predict the next token but that isn't based on a likelihood of that token appearing in a corpus of internet text consumed in pretraining. That was eons ago. Every predicted token is shaped by the probabilities of the predicted solution, which must already be very specific and shaped completely by the request and associated context that is built during investigation of the same.
"You could say the same thing about human programmers, but I've never heard anyone say they think that programmers "execute" Jira tickets."
Yes I could. We have "executives", for starters. And first "computers" were actual humans.
"That was eons ago."
Yes, technically I should call them LRMs (large reasoning models) not LLMs. But that doesn't seem relevant here, to my point they encode some logic (which we want to be close to classical logic, i.e. behavior of words like "true", "and", "not" and so on matches).
That isn't what executives do, and I'm not talking about semantics. If you believe that the term "token predictor" has any relevance when discussing the capabilities of these models, you are misinformed. That is the inner, inner loop and it is simply the substrate through which reasoning and action is expressed.
I agree with the 2nd sentence onwards, and I think I expressed it in the other comments I made here. I am not really sure what was your point about "execution", though. Humans might have inner interpreter loops as well.
You could say the same thing about human programmers, but I've never heard anyone say they think that programmers "execute" Jira tickets.
> they have internal logic that assigns to the sequence of tokens in context a next token.
I don't think this means what you think it means, because it has almost no information content relevant to what we're discussing. The probabilities that are most relevant at the level we're discussing are satisfying a reward function from post-training, which approximates to:
"What is the likelihood the solution the agent is pursuing will be marked correct by the automated grader based on the full prompt and other context provided?"
It still has to predict the next token but that isn't based on a likelihood of that token appearing in a corpus of internet text consumed in pretraining. That was eons ago. Every predicted token is shaped by the probabilities of the predicted solution, which must already be very specific and shaped completely by the request and associated context that is built during investigation of the same.