I think that is a sort of a reverse scaling fallacy. Given the right resources and environments, many small models can function together in an emergent way. I’ve been on the lookout for an SLM version of Conway’s game of life. SLM always reminds me of slime molds, which demonstrate a form of intelligence which is remarkable.
> Given the right resources and environments, many small models can function together in an emergent way.
What is this based on? Every researcher I've heard talk about this says it's exactly not true, as an uncontested rule, because the larger models will more effectively contain the smaller models, and use them together in ways that the connections between the smaller models can't. Remember, even MOE is to save compute/memory, not to help performance/parameter.
That is my understanding as well. Thousands of monkeys do not equal or surpass a man, intellectually, even if working together. There is some intrinsic super linear scaling in intelligence.
This is not a good comparison, because the brain doesn't only do "being smart" - it has to do things like innervate muscle and other tissue through the body, elephants will require more neurons given their larger size, just to be able to *walk*
You're comparing completely different training data, harness overhead, and cost function. You're also comparing brains, which aren't really related to this discussion at all.
But, it depends on what you're measuring. By spatial/navigational memory, yes, elephants are far far better. Reasoning, no. It would be interesting to see what an elephant or whale eugenics program could result in, since humans have that pesky (or maybe instrumental?) birth canal problem.