I think comparing this to human drug trials is an appeal to emotion.
I also think you're saying that bad tests are better than no tests, which I don't agree with. These would be bad tests because they wouldn't constitute the complicated reality of driving, they inherently only summarize the output of such a complex system.
You seem to have confused the concept of bad tests with incomplete tests. Asking to stop the car when pedestrian shaped obstacles appear seems like a bare minimum that we should expect from an AI before letting it drive on real roads. We can't test every driving situation in a lab, but that doesn't mean we can't test some of them.
It would be morally bankrupt to rely entirely on a quantitative test for something that is qualitative, I don't know how much clearer with the previous comment I can get?
You're missing the fact that any time someone takes a drug, they're essentially testing it against their unique body chemistry. The same as self-driving cars in unique situations.
The only difference is the car feedback loop is way better instrumented, and doesn't have to rely on mass statistics to inform recalls.
I also think you're saying that bad tests are better than no tests, which I don't agree with. These would be bad tests because they wouldn't constitute the complicated reality of driving, they inherently only summarize the output of such a complex system.