Bedi, Cui, Fuentes et al.; Nature Medicine
Holistic evaluation of large language models for medical tasks with MedHELM ↗
Final journal publication; 37-benchmark analytical snapshot.
Version: 20 January 2026
Evidence locator: Table 1; Extended Data Table 1; Performance metrics
www.nature.com