I think it's a pretty well established in learning theory that the combination of neural network topology + the training procedure sets up an implicit prior (it establishes a bias towards the type of things the network will learn more easily from the input data). In fact, this is a corollary of the No Free Lunch Theorem.
The problem is that these are extremely high-dimensional spaces, and even narrowing down the space within which the prior might be defined (much less finding the prior itself) is mathematically difficult. Hence the progress is slow and halting—it's consisted of discussions like what the representations fully-trained neural networks find within the early activation layers, which means the question is being investigated empirically, without the benefit of a fully-fledged theory.
Initeresting, the 'nature vs. nurture' debate can be viewed as basically a discussion of learning priors in the context of human beings. Turns out the truth is pretty subtle—for many problem domains, humans (and animals) can be made to learn a very large set of things with repeated training, but we also have strong priors, so some things can be taught much more easily and quickly than others.
Our brain networks are primed to learn certain things, e.g. ducklings have a prior that the first animate object they see is their mother. It's possible to force a duckling to unlearn its maternal imprinting, but only with repeated and large amounts of negative conditioning. It seems pretty likely that this prior is embedded in both the topology and neurochemical functioning of a normally developing brain.
> the combination of neural network topology + the training procedure sets up an implicit prior [...]
That's what I got from the koan. Sussman thinks that a randomly wired NN has no prior, but that's false. Same as Minsky "thinks" that the room is empty when he closes his eyes.
EDIT: actually if you read the source (spetharrific.tumblr.com) it spells this out.
That's certainly the moral. I think what's changed with the new deep learning paradigm is a willingness to accept that implicit prior, not so much because it's good as because we don't know how to improve it consistently.
The problem is that these are extremely high-dimensional spaces, and even narrowing down the space within which the prior might be defined (much less finding the prior itself) is mathematically difficult. Hence the progress is slow and halting—it's consisted of discussions like what the representations fully-trained neural networks find within the early activation layers, which means the question is being investigated empirically, without the benefit of a fully-fledged theory.
Initeresting, the 'nature vs. nurture' debate can be viewed as basically a discussion of learning priors in the context of human beings. Turns out the truth is pretty subtle—for many problem domains, humans (and animals) can be made to learn a very large set of things with repeated training, but we also have strong priors, so some things can be taught much more easily and quickly than others.
Our brain networks are primed to learn certain things, e.g. ducklings have a prior that the first animate object they see is their mother. It's possible to force a duckling to unlearn its maternal imprinting, but only with repeated and large amounts of negative conditioning. It seems pretty likely that this prior is embedded in both the topology and neurochemical functioning of a normally developing brain.