Hacker Newsnew | past | comments | ask | show | jobs | submitlogin

> or is this just a vague "i'll know it when i see it" assertion

No, it's actually exceptionally easy to quantify. Take a look at the current leaderboard for LLM reasoning capabilities: https://huggingface.co/spaces/HuggingFaceH4/open_llm_leaderb...



The state of the art model isn't even on that list.

Okay, you've at least given me numbers. You're still not answering my question. Which number signals agi ?

Let's look at the top model in that list (which again isn't close to the best performing LLM) that you say isn't agi. so are you telling me that every human can do those tests and perform better than every model on that list. Is that what you are saying ? because i can tell you right now you're wrong.

Now lets see how GPT-4 performs.

ARC - 96.3%, MMLU - 86.4%, HellaSwag - 95.3%, WinoGrande - 87.5%, GSM-8K - 92.0%, TruthfulQA - 60%

Ave - 86.25%

Is this worse than every human that can take these tests. Is this even worse than most ? I can tell you it's not. So again, why is GPT-4 not agi ?


> so are you telling me that every human can do those tests and perform better than every model on that list

This is a strange bar to require clearing. Every human is not even capable of reading, or speaking, or thinking at all. But those humans who can't are not the benchmark of general human intelligence.

Let's make this simple - the average human IQ is around 100. Will an LLM ever have an IQ of 100? No, I don't believe they'll ever even be close.

How can we gauge this abstract reasoning ability? Lots of ways. Try to teach them mathematics. Try to teach them physics. Try to explain to an LLM the rules to a card game and have it actually play the game with you accurately.


>Every human is not even capable of reading, or speaking, or thinking at all. But those humans who can't are not the benchmark of general human intelligence.

It's simple. If you're going to disqualify a potential general intelligence because of a particular task solving ability, it better be something every general intelligence can do.

Now let's forget every comparison with the blind, disabled and so on. Let's compare with humans who according to you are "capable of reading, or speaking, or thinking at all ", what's the magic number? Do you really think every human capable of doing all three will hit 86% or more average?, because again, you are wrong. Numerous replies and you still won't tell me what number will signify agi. That says a lot.

>Let's make this simple - the average human IQ is around 100. Will an LLM ever have an IQ of 100? No, I don't believe they'll ever even be close.

Well you are wrong

https://arxiv.org/abs/2212.09196

What's the next post ?

>Try to teach them mathematics.

Ok. https://arxiv.org/abs/2211.09066

https://arxiv.org/abs/2308.00304

>Try to explain to an LLM the rules to a card game and have it actually play the game with you accurately.

Have you tried to do this with GPT-4 ?


> Is this worse than every human that can take these tests. Is this even worse than most ? I can tell you it's not. So again, why is GPT-4 not agi ?

https://chat.openai.com/share/4a92c752-b5bb-4a07-beed-f57786...


Love it when people can only demonstrate tasks handicapped by Tokenization. Shows how far we've come.

That said, By that test, I guess some dyslexics aren't general intelligences then.


Tokenization is not the problem in this case.

https://chat.openai.com/share/90c7a9e2-77d0-4049-8eff-4312ca...

You see the difference?




Guidelines | FAQ | Lists | API | Security | Legal | Apply to YC | Contact

Search: