What do MedHELM’s 37 benchmarks and 121 tasks represent?
Distinguish the clinical taxonomy from the measured benchmark inventory in MedHELM’s final journal publication.
Original analysis / Healthcare Bench
MedHELM supplies a structured view of clinical language-model tasks. This independent publication examines the final journal paper’s 37 evaluations, the evidence available for each task family and the aggregation behind its model rankings. Our analysis and coverage explorer make the unit of measurement visible. Published results belong to the cited study; we have not rerun its models or accessed its private datasets. The current community project is linked separately from this fixed paper snapshot.
Distinguish the clinical taxonomy from the measured benchmark inventory in MedHELM’s final journal publication.
Understand relative ranking, equal benchmark weighting and the three-model jury before interpreting a MedHELM headline.
Compare the 2025 preprint, 2026 journal snapshot and evolving project without mixing their inventories or results.