Has industry "cracked the mystery of consciousness"? Has anyone at OpenAI even made a plausible start at explaining what the hell is going on inside these LLMs that enables them to produce such human-like comprehension? To me these are the really exciting questions.
From what I can tell, the team at OpenAI (following after Google/Deepmind etc.) are simply mashing the pedal to the floor to get bigger and better models from their existing techniques, and then tuning the resulting black box to make it produce more "useful" answers. And that's fine! That's precisely what an industry lab is expected to do: they have the resources to do the training and the need to get products in front of paying customers as quickly as possible to justify it. And frankly with top AI engineers getting paid millions of dollars and Google/Meta tight behind you, emphasizing results is the most viable strategy. If "turn the needle on the box to the right" is giving you good answers, why would you waste a $1-$5m-salary engineer on academic questions like "why does the box do that?"
And yet, asking questions like "why does the box do that?" is the reason technology didn't stop at the steam engine. I suspect that finding the answer to those questions won't immediately sell enterprise licenses, but will be very important. And the answers will probably fall to someone who's making $30k/year in a graduate program.
ETA: Of course, it may turn out that "turn the needle on the box to the right" is enough to obsolete all human researchers, in which case I'll be wrong about this. But it'll hardly matter in that case ;)
> Has anyone at OpenAI even made a plausible start at explaining what the hell is going on inside these LLMs that enables them to produce such human-like comprehension?
I'm only an amateur in the field, so my uneducated high-level understanding is that, in a sufficiently high-dimensional latent space, there's more than enough dimensions to assign to any single semantic relationship people ever thought of, which is what the training process effectively does, which reduces an important part of thinking - working with concepts and their relationships - entirely to vector adjacency search.
I'm only beginning to study the details, and I don't know how much of specific understanding of this exists, but at the very least this high-level model explains why scaling makes qualitative difference here.
Now, I agree they have strong commercial incentives to push their models as far as possible as fast as possible, but honestly, if I were a researcher working on these models, even if I was somehow unconcerned with any kind of commercial viability and had access to more compute, I'd absolutely keep scaling those models up and up, all the way until I hit the limit of available compute, or the models stop qualitatively improving with scale.
Basically, there's no reason[0] to stop now and try to fully comprehend how GPT-2 works, when GPT-3 was a qualitative jump, and GPT-4 even more so, and GPT-5 is around the corner, and GPT-6 might be a year away from now. All those steps yield important new insights into how the whole architecture works, and if at some point the scaling breaks, that would be even more important knowledge to have. And this doesn't even take into account the fact that, starting with GPT-3, those models are increasingly useful in accelerating both research and scaling alike.
----
[0] - Except, of course, that if transformer models are the road to generic AI, then we'll just blindly race straight into a point of no return.
From what I can tell, the team at OpenAI (following after Google/Deepmind etc.) are simply mashing the pedal to the floor to get bigger and better models from their existing techniques, and then tuning the resulting black box to make it produce more "useful" answers. And that's fine! That's precisely what an industry lab is expected to do: they have the resources to do the training and the need to get products in front of paying customers as quickly as possible to justify it. And frankly with top AI engineers getting paid millions of dollars and Google/Meta tight behind you, emphasizing results is the most viable strategy. If "turn the needle on the box to the right" is giving you good answers, why would you waste a $1-$5m-salary engineer on academic questions like "why does the box do that?"
And yet, asking questions like "why does the box do that?" is the reason technology didn't stop at the steam engine. I suspect that finding the answer to those questions won't immediately sell enterprise licenses, but will be very important. And the answers will probably fall to someone who's making $30k/year in a graduate program.
ETA: Of course, it may turn out that "turn the needle on the box to the right" is enough to obsolete all human researchers, in which case I'll be wrong about this. But it'll hardly matter in that case ;)