{"publication":"Healthcare Bench","url":"https://healthcarebench.com","publisher":"Arcophos","updated":"2026-09-28","provenance":"Independent analytical publication. Benchmark creation and experimental results belong to their cited authors. Reported results are source-version snapshots, not new Arcophos runs or a live leaderboard.","benchmarks":[{"slug":"medhelm","name":"MedHELM","shortName":"MedHELM","version":"Final journal paper: 37-benchmark snapshot","creators":"Bedi, Cui, Fuentes et al.; Stanford-led collaboration","paperDate":"2026-01-20","headline":"Coverage, access and scoring all shape a hospital benchmark.","summary":"MedHELM organizes clinical language-model evaluation around the work a health system needs performed. We analyze the final journal paper’s 37-benchmark snapshot, its task taxonomy and its combination of deterministic scoring with an LLM jury. The central question is what an overall rank compresses: different task families, different dataset access conditions and different measures of success. Our original contribution is a coverage map and an interpretation of the aggregation choices. We keep the final paper separate from both its earlier preprint and the evolving community leaderboard, because their inventories and comparison sets differ.","task":{"input":"Task-specific records, questions, conversations or research requests.","output":"Task-specific classification, calculation, generated text, code or structured answer.","unit":"A benchmark-specific evaluation instance; 37 benchmark scores feed the aggregate.","setting":"Nine historical models under the published evaluation pipeline."},"dataOrigin":"A mixture of existing public benchmarks, gated clinical datasets and private clinical datasets; no single source population describes the entire suite.","facts":[{"label":"Evaluations","value":"37","detail":"Final-paper benchmark inventory.","sourceIds":["medhelm-full"],"locator":"Extended Data Table 1"},{"label":"Task taxonomy","value":"121 tasks","detail":"Organized into five categories and 22 subcategories; not 121 independently measured datasets.","sourceIds":["medhelm-full"],"locator":"Clinician validation of the taxonomy"},{"label":"Public / gated / private","value":"16 / 7 / 14","detail":"Access categories in the final-paper inventory.","sourceIds":["medhelm-full"],"locator":"Overview of the benchmark suite"},{"label":"Compared models","value":"9","detail":"Historical model snapshots listed in Table 1.","sourceIds":["medhelm-full"],"locator":"Table 1"},{"label":"Jury members","value":"3","detail":"GPT-4o, Claude 3.7 Sonnet and LLaMA 3.3 70B.","sourceIds":["medhelm-full"],"locator":"Evaluation using LLM-jury"}],"metric":{"name":"Pairwise win rate and macro-average","description":"Pairwise wins average comparisons over benchmarks and opponents; macro-average gives every benchmark equal weight.","formula":"Win = 1[score_m,b ≥ score_r,b]; macro = Σ normalized benchmark score / 37","direction":"Higher is better within the fixed suite and opponent set.","comparability":"A win rate depends on the other models. Normalized scores mix different measures; neither aggregate is a patient-outcome percentage.","sourceIds":["medhelm-full"]},"workflow":[{"label":"Locate the clinical task","detail":"Map the task to the clinician-validated taxonomy.","sourceIds":["medhelm-full"]},{"label":"Specify the benchmark","detail":"Choose an input, reference and metric for that task.","sourceIds":["medhelm-full"]},{"label":"Score the output","detail":"Apply a deterministic task metric or the three-model jury.","sourceIds":["medhelm-full"]},{"label":"Aggregate with the denominator visible","detail":"Preserve task-level results before averaging over the 37 benchmark scores.","sourceIds":["medhelm-full"]}],"slices":[{"label":"Clinical decision support","value":12,"unit":"benchmarks","detail":"Final-paper category count.","sourceIds":["medhelm-full"]},{"label":"Patient communication","value":8,"unit":"benchmarks","detail":"Includes education and messaging.","sourceIds":["medhelm-full"]},{"label":"Clinical note generation","value":6,"unit":"benchmarks","detail":"Documentation tasks.","sourceIds":["medhelm-full"]},{"label":"Medical research assistance","value":6,"unit":"benchmarks","detail":"Research tasks.","sourceIds":["medhelm-full"]},{"label":"Administration and workflow","value":5,"unit":"benchmarks","detail":"Operational tasks.","sourceIds":["medhelm-full"]}],"sliceTitle":"Final-paper task coverage","sliceNote":"Counts sum to 37. The paper’s open/closed-ended prose split is internally inconsistent with this total, so it is not charted.","results":[{"id":"mh-win","title":"Final-paper pairwise win rate","metric":"Pairwise win rate","unit":"0–1","lower":0,"upper":1,"scope":"Nature Medicine 2026 final paper, Table 1; 37 benchmarks and nine historical models.","sourceIds":["medhelm-full"],"locator":"Table 1","rows":[{"label":"DeepSeek R1","value":0.66,"display":"0.66","detail":"Win SD 0.11; macro-average 0.75."},{"label":"o3-mini (2025-01-31)","value":0.66,"display":"0.66","detail":"Win SD 0.15; macro-average 0.77."},{"label":"Claude 3.5 Sonnet (20241022)","value":0.63,"display":"0.63","detail":"Win SD 0.15; macro-average 0.74."},{"label":"Claude 3.7 Sonnet (20250219)","value":0.63,"display":"0.63","detail":"Win SD 0.15; macro-average 0.73."},{"label":"GPT-4o (2024-05-13)","value":0.58,"display":"0.58","detail":"Win SD 0.17; macro-average 0.74."}],"note":"SD is variation across the paper’s comparisons, not a confidence interval. Ties count as wins under the stated method. Table 1 takes precedence over a swapped Claude macro-average sentence."}],"analysis":[{"heading":"Taxonomy coverage is not equal evidence density","evidence":"37 benchmarks span 22 subcategories and 121 named tasks.","interpretation":"Coverage of a category establishes representation; it does not establish equally deep testing of every named task.","sourceIds":["medhelm-full"]},{"heading":"A public rerun changes the scope","evidence":"The final inventory contains 16 public, seven gated and 14 private benchmarks.","interpretation":"A public-only result answers a narrower question. It should not inherit the full-suite headline without a matching denominator.","sourceIds":["medhelm-full"]},{"heading":"Rank and average answer different questions","evidence":"DeepSeek R1 and o3-mini both have a 0.66 win rate but macro-averages 0.75 and 0.77.","interpretation":"A tie in relative rank can coexist with a difference in average normalized magnitude. Neither eliminates the need to inspect the intended workflow.","sourceIds":["medhelm-full"]},{"heading":"Version changes affect interpretation","evidence":"The preprint has 35 benchmarks; the final paper has 37.","interpretation":"Even an unchanged model can move when evaluation membership changes. Version the suite alongside the model.","sourceIds":["medhelm-pre","medhelm-full"]}],"limitations":[{"title":"Mixed measurement scales","detail":"Exact matches, task-specific metrics and jury scores are normalized for aggregation.","sourceIds":["medhelm-full"]},{"title":"Private data affects reproducibility","detail":"Some full-suite results cannot be independently rerun from public assets alone.","sourceIds":["medhelm-full"]},{"title":"Jury validation has a scope","detail":"Agreement with clinicians is assessed on selected outputs and does not validate every possible specialty or failure type.","sourceIds":["medhelm-full"]},{"title":"Evolving inventories","detail":"The current community project describes a separate dataset inventory; it is not silently substituted for the paper’s 37 evaluations.","sourceIds":["medhelm-project"]}],"access":{"status":"Public, gated and private benchmark components","license":"Evaluation software Apache 2.0; each dataset has its own terms.","restrictions":"Access to code does not grant access to private benchmarks or waive source-data restrictions.","url":"https://medhelm.org/","sourceIds":["medhelm-project","medhelm-full"]},"sourceIds":["medhelm","medhelm-full","medhelm-pre","medhelm-project"]}],"explorer":{"kind":"coverage","title":"Map MedHELM to the work being evaluated","intro":"Explore concrete benchmarks inside the final 37-evaluation paper. Inputs and endpoints matter more than a broad category name.","caution":"This is an independently annotated coverage map, not a local validation plan or a reproduction of private datasets.","sourceIds":["medhelm-full"],"parameters":[],"rows":[{"label":"Medical calculation","category":"Clinical decision support","input":"Patient note and requested calculation","output":"Numeric or categorical answer","metric":"MedCalc accuracy","constraint":"Exact or tolerance-based scoring varies by calculation.","benchmarkSlug":"medhelm","sourceIds":["medhelm-full"]},{"label":"Clinical error detection","category":"Clinical decision support","input":"Clinical narrative","output":"Error detection or correction","metric":"Task-specific error metric","constraint":"Correctness of detection differs from usefulness of a revised note.","benchmarkSlug":"medhelm","sourceIds":["medhelm-full"]},{"label":"Visit documentation","category":"Clinical note generation","input":"Patient–doctor conversation","output":"Structured clinical note","metric":"LLM-jury assessment","constraint":"A jury assesses outputs; actual user correction is outside this endpoint.","benchmarkSlug":"medhelm","sourceIds":["medhelm-full"]},{"label":"Radiology summarization","category":"Clinical note generation","input":"Findings text","output":"Impression summary","metric":"LLM-jury assessment","constraint":"Text summarization does not establish image interpretation.","benchmarkSlug":"medhelm","sourceIds":["medhelm-full"]},{"label":"Medication questions","category":"Patient communication","input":"Consumer medication query","output":"Long-form answer","metric":"LLM-jury assessment","constraint":"Reference and judge setup shape answer ratings.","benchmarkSlug":"medhelm","sourceIds":["medhelm-full"]},{"label":"Personalized instructions","category":"Patient communication","input":"Procedure and case context","output":"Patient instructions","metric":"LLM-jury assessment","constraint":"Patient comprehension is not directly observed.","benchmarkSlug":"medhelm","sourceIds":["medhelm-full"]},{"label":"Clinical research SQL","category":"Medical research assistance","input":"Natural-language data request","output":"SQL query","metric":"Execution accuracy","constraint":"Correctness is bound to the supplied schema and target.","benchmarkSlug":"medhelm","sourceIds":["medhelm-full"]},{"label":"Billing code assignment","category":"Administration and workflow","input":"Medical note","output":"ICD-10 code set","metric":"F1 score","constraint":"F1 is distinct from raw exact-match correctness.","benchmarkSlug":"medhelm","sourceIds":["medhelm-full"]},{"label":"Referral routing","category":"Administration and workflow","input":"Referral information","output":"Task-defined referral decision","metric":"Task-specific accuracy","constraint":"Private operational data limits independent reruns.","benchmarkSlug":"medhelm","sourceIds":["medhelm-full"]}]},"references":[{"id":"medhelm","title":"Holistic evaluation of large language models for medical tasks with MedHELM","organization":"Bedi, Cui, Fuentes et al.; Nature Medicine","url":"https://www.nature.com/articles/s41591-025-04151-2","note":"Final journal publication; 37-benchmark analytical snapshot.","locator":"Table 1; Extended Data Table 1; Performance metrics","version":"20 January 2026"},{"id":"medhelm-full","title":"MedHELM accepted manuscript with full tables","organization":"Bedi et al.; PubMed Central","url":"https://pmc.ncbi.nlm.nih.gov/articles/PMC13267972/","note":"Full manuscript used to verify final-paper counts, scoring and selected results.","locator":"Table 1; Extended Data Tables 1, 4; Methods","version":"Nature Medicine 2026 accepted manuscript"},{"id":"medhelm-pre","title":"MedHELM original preprint","organization":"Bedi et al.","url":"https://arxiv.org/html/2505.23802v1","note":"Earlier 35-benchmark version; retained for explicit version comparisons.","locator":"Table 1 and Section 2.2","version":"26 May 2025"},{"id":"medhelm-project","title":"MedHELM community project and documentation","organization":"MedHELM community","url":"https://medhelm.org/","note":"Current project site, separate from the final-paper snapshot; software Apache 2.0 does not license underlying datasets.","locator":"Origins and stewardship; Cite and License","version":"Accessed 28 September 2026"}]}