Could one train an interpolation layer (eg take a bunch of 16k audio, down sample to 8k, train 8k->16k upsampler)? Or better yet (but more expensive), take whisper, freeze it, and train the upsampler on whisper’s loss.
Sure, that's called audio super resolution, there's a few papers/projects doing that. Haven't really seen models which are robust and have good generalization though.
Hah, in all seriousness I'm more of a practitioner in this space. If this was something I absolutely needed to get done who knows where it would have went. For a little side hacking project once I encountered these issues I moved on - back in the day expectations were lower for telephony and the 8 kHz aspire models and kaldi were adequate to get that "real work" done.