Firstly, that's not my claim. I merely said it passed the USMLE. Which it did.
Is it also inferior to clinicians? Yes, there's room to improve. But maybe next time read the whole paper before writing a comment.
> Clinicians were asked to rate answers provided to questions in the HealthSearchQA, Live QA and Medication question answering datasets. Clinicians were asked to identify whether the answer is aligned with the prevailing medical/scientific consensus; whether the answer was in opposition to consensus; or whether there is no medical/scientific consensus for how to answer that particular question (or whether it was not possible to answer this question).
And on this criteria, clinicians were rated as being aligned with consensus 92.9% of the time while the MedPalm model was aligned with consensus 92.6% of the time.
If the other 7-8% of the answers were so wrong the patient would've died, then yes. And that's the current obvious issue with these models, they present convincing hallucinations with conviction of correctness.
Medical practice is less about being right, and more about not being wrong. You can take more tests and ask for second opinions, but you can't undo administering a drug that kills the patient.
Entire institutions exist specifically to get rid of below average by test-based gatekeeping. You do not want your doctor or lawyer to be "below average" (worse than most people) in their jobs. Inferior test results mean exactly that, failing the test.
So maybe?