The state of the art model isn't even on that list.
Okay, you've at least given me numbers. You're still not answering my question. Which number signals agi ?
Let's look at the top model in that list (which again isn't close to the best performing LLM) that you say isn't agi. so are you telling me that every human can do those tests and perform better than every model on that list. Is that what you are saying ? because i can tell you right now you're wrong.
> so are you telling me that every human can do those tests and perform better than every model on that list
This is a strange bar to require clearing. Every human is not even capable of reading, or speaking, or thinking at all. But those humans who can't are not the benchmark of general human intelligence.
Let's make this simple - the average human IQ is around 100. Will an LLM ever have an IQ of 100? No, I don't believe they'll ever even be close.
How can we gauge this abstract reasoning ability? Lots of ways. Try to teach them mathematics. Try to teach them physics. Try to explain to an LLM the rules to a card game and have it actually play the game with you accurately.
>Every human is not even capable of reading, or speaking, or thinking at all. But those humans who can't are not the benchmark of general human intelligence.
It's simple. If you're going to disqualify a potential general intelligence because of a particular task solving ability, it better be something every general intelligence can do.
Now let's forget every comparison with the blind, disabled and so on. Let's compare with humans who according to you are "capable of reading, or speaking, or thinking at all ", what's the magic number? Do you really think every human capable of doing all three will hit 86% or more average?, because again, you are wrong. Numerous replies and you still won't tell me what number will signify agi. That says a lot.
>Let's make this simple - the average human IQ is around 100. Will an LLM ever have an IQ of 100? No, I don't believe they'll ever even be close.
No, it's actually exceptionally easy to quantify. Take a look at the current leaderboard for LLM reasoning capabilities: https://huggingface.co/spaces/HuggingFaceH4/open_llm_leaderb...