The NOHARM—Numerous Options Harm Assessment for Risk in Medicine—study evaluated 45 large language models and four specialized clinical systems. Researchers tracked 12,747 expert annotations across 4,249 clinical actions. While the results remain non-peer-reviewed, they provide a rare look at revealed preference, a metric that measures which tools clinicians actually trust when facing real-world cases.
Physicians are notoriously selective regarding medical technology, often abandoning tools that fail to meet high evidentiary standards. Unlike platforms that rely on internal validation, OpenEvidence mandates that every answer be directly sourced and cited from peer-reviewed medical literature. This commitment to transparency recently expanded with the introduction of EvidenceGrade, a feature that assesses the quality of underlying evidence in real time using established frameworks like those employed by the World Health Organization.





Comments (0)
No comments yet. Be the first!