Original analysis / Healthcare Bench

Read the benchmark closely.

MedHELM supplies a structured view of clinical language-model tasks. This independent publication examines the final journal paper’s 37 evaluations, the evidence available for each task family and the aggregation behind its model rankings. Our analysis and coverage explorer make the unit of measurement visible. Published results belong to the cited study; we have not rerun its models or accessed its private datasets. The current community project is linked separately from this fixed paper snapshot.

Put it into practice ↗