Hacker Newsnew | past | comments | ask | show | jobs | submitlogin

If people and especially doctors were serious about improving healthcare, they should strive to build reliable data collection means. That's the prerequisite to efficient 'AI', and would be a far better use of funding than all the present half-baked 'improvements' and star-system research. We could really do science, then. A man can dream...


There is a lot of data collection. With everyone using EMRs massive datasets are available. They just aren't available to people outside the walled gardens for PHI/privacy reasons.

At Kaiser Permanente, I regularly train models on 10s/100s M patient/dr encounters. Our transformer language models are fine tuned on multi-billion word corpuses. Some of our models do real time inference on millions of patient notes per month.

The thing is, from data governance standpoint our org, and most other health orgs, just aren't comfortable sharing this data with outside businesses or even each other. And of course orgs like KP strongly don't believe in licensing internally developed products to other health orgs.


There is a huge governance issue as you note (honestly, most of this should probably belong to the patients but disentangling that from the parts that are clearly the providers is tough). A lot of the data generated isn't at the right granularity either, because it assumes another human in the loop to interpret. This can somewhat work for AI, but also be an impediment.


I personally generate several Gbs of healthcare data per day.

I know we generate a lot of data. I also know it's data that's so unreliable that its business value does not lie in its real-world use for improving healthcare pathways. It's very valuable politically and from a managerial standpoint, though. Unfortunately.


I mean it depends. There is a lot of fairly reliable discrete data in medicine: medications/labs/imaging studies/procedures/flowsheet/etc. ICD10 diagnoses are discrete and fairly reliable. The progress notes have lots of copy/paste/smart-phrase/macro-generated trash, but at least Epic saves a lot of rtf markup about the source of the text data. I am pretty optimistic about data quality, it is availability to 3rd parties that I think is limiting the AI boom.


I remain unconvinced, but let's hope you're right.


> They should strive to build reliable data collection means. That's the prerequisite to efficient 'AI',

It's necessary but not sufficient. The biggest problem in most applications is labeling, which either doesn't exist at all or is insufficient in most clinical workflows.




Guidelines | FAQ | Lists | API | Security | Legal | Apply to YC | Contact

Search: