Psychology dealt with a very similar — maybe identical issue — with measurement and prediction in the 60s and 70s. The tension was between "empirically keyed" and "content" approaches to tests, the former being use of tests entirely based on large black box predictive item pools akin to large DL models, and tests based on some items selected based on some theory of test content.
Many of the issues were similar. In the end I think they both lost out to an approach in which tests were empirically validated, but with items retained on their ability to meet various internal structural criteria.
The "discovery" versus "justification" distinction reminds me of this in some ways. It would be akin to if some set of criteria were developed, not based on target prediction criteria, that DL model components would have to meet as constraints. Or, alternatively, you might formalize in some quantitative model what characterizes "justification" characteristics, in the sense of how to interpret a given DL model.
This paper makes we wonder whether there are ways to automate the search for the most plausible posits (see Figure 1). If scientists could choose the right combinations of data, then could a lot the discovery process could be automated?
TLDR: while the way deep learning reaches conclusions can be opaque, those conclusions can supercharge human intuition, leading to mathematical and scientific conjectures that can then be justified or proven with rigorous, transparent, logical arguments and data.
Some famous mathematician, whose name I've forgotten, said something like "the trouble is not the proofs, but knowing what to prove." Humans have always relied on heuristics to discover what might be true before confirming it rigorously. If deep learning provides powerful heuristics, it can be a tremendous aid to scientific progress.
This reminds me of the concept (parallel reconstruction I think) where law enforcement gets information from a covert informant but can't "blow" them (or or maybe can't use it in court) and so finds some other way to show how they got to the result.
Many of the issues were similar. In the end I think they both lost out to an approach in which tests were empirically validated, but with items retained on their ability to meet various internal structural criteria.
The "discovery" versus "justification" distinction reminds me of this in some ways. It would be akin to if some set of criteria were developed, not based on target prediction criteria, that DL model components would have to meet as constraints. Or, alternatively, you might formalize in some quantitative model what characterizes "justification" characteristics, in the sense of how to interpret a given DL model.