Independent benchmark analysis / Final journal paper: 37-benchmark snapshot

MedHELM

Coverage, access and scoring all shape a hospital benchmark.

MedHELM organizes clinical language-model evaluation around the work a health system needs performed. We analyze the final journal paper’s 37-benchmark snapshot, its task taxonomy and its combination of deterministic scoring with an LLM jury. The central question is what an overall rank compresses: different task families, different dataset access conditions and different measures of success. Our original contribution is a coverage map and an interpretation of the aggregation choices. We keep the final paper separate from both its earlier preprint and the evolving community leaderboard, because their inventories and comparison sets differ.

01 / What is being tested?

The task, before the score.

input
Task-specific records, questions, conversations or research requests.
output
Task-specific classification, calculation, generated text, code or structured answer.
unit
A benchmark-specific evaluation instance; 37 benchmark scores feed the aggregate.
setting
Nine historical models under the published evaluation pipeline.

Data origin. A mixture of existing public benchmarks, gated clinical datasets and private clinical datasets; no single source population describes the entire suite. [1][2][3][4]

Evaluations
37

Final-paper benchmark inventory.

Extended Data Table 1 [2]
Task taxonomy
121 tasks

Organized into five categories and 22 subcategories; not 121 independently measured datasets.

Clinician validation of the taxonomy [2]
Public / gated / private
16 / 7 / 14

Access categories in the final-paper inventory.

Overview of the benchmark suite [2]
Compared models
9

Historical model snapshots listed in Table 1.

Table 1 [2]
Jury members
3

GPT-4o, Claude 3.7 Sonnet and LLaMA 3.3 70B.

Evaluation using LLM-jury [2]
  1. 01

    Locate the clinical task

    Map the task to the clinician-validated taxonomy.

    [2]
  2. 02

    Specify the benchmark

    Choose an input, reference and metric for that task.

    [2]
  3. 03

    Score the output

    Apply a deterministic task metric or the three-model jury.

    [2]
  4. 04

    Aggregate with the denominator visible

    Preserve task-level results before averaging over the 37 benchmark scores.

    [2]

Dataset anatomy

Final-paper task coverage

Clinical decision support

Final-paper category count.

12 benchmarks[2]
Patient communication

Includes education and messaging.

8 benchmarks[2]
Clinical note generation

Documentation tasks.

6 benchmarks[2]
Medical research assistance

Research tasks.

6 benchmarks[2]
Administration and workflow

Operational tasks.

5 benchmarks[2]

Counts sum to 37. The paper’s open/closed-ended prose split is internally inconsistent with this total, so it is not charted. Bar lengths use the largest listed count as their reference; they are not percentages of a shared population.

02 / Measurement

Pairwise win rate and macro-average

Higher is better within the fixed suite and opponent set.

Pairwise wins average comparisons over benchmarks and opponents; macro-average gives every benchmark equal weight.

Scoring definition

Win = 1[score_m,b ≥ score_r,b]; macro = Σ normalized benchmark score / 37

A win rate depends on the other models. Normalized scores mix different measures; neither aggregate is a patient-outcome percentage. [2]

03 / Measured evidence

Results, with their conditions attached.

Paper-reported results / selected rows

Final-paper pairwise win rate

Nature Medicine 2026 final paper, Table 1; 37 benchmarks and nine historical models.

Pairwise win rate · 0–1
00.51
Reported
DeepSeek R1Win SD 0.11; macro-average 0.75.
0.66
o3-mini (2025-01-31)Win SD 0.15; macro-average 0.77.
0.66
Claude 3.5 Sonnet (20241022)Win SD 0.15; macro-average 0.74.
0.63
Claude 3.7 Sonnet (20250219)Win SD 0.15; macro-average 0.73.
0.63
GPT-4o (2024-05-13)Win SD 0.17; macro-average 0.74.
0.58

SD is variation across the paper’s comparisons, not a confidence interval. Ties count as wins under the stated method. Table 1 takes precedence over a swapped Claude macro-average sentence.

Source: Table 1 [2]

04 / Our original analysis

What follows from the design?

01

Taxonomy coverage is not equal evidence density

Published evidence

37 benchmarks span 22 subcategories and 121 named tasks. [2]

Our interpretation

Coverage of a category establishes representation; it does not establish equally deep testing of every named task.

02

A public rerun changes the scope

Published evidence

The final inventory contains 16 public, seven gated and 14 private benchmarks. [2]

Our interpretation

A public-only result answers a narrower question. It should not inherit the full-suite headline without a matching denominator.

03

Rank and average answer different questions

Published evidence

DeepSeek R1 and o3-mini both have a 0.66 win rate but macro-averages 0.75 and 0.77. [2]

Our interpretation

A tie in relative rank can coexist with a difference in average normalized magnitude. Neither eliminates the need to inspect the intended workflow.

04

Version changes affect interpretation

Published evidence

The preprint has 35 benchmarks; the final paper has 37. [3][2]

Our interpretation

Even an unchanged model can move when evaluation membership changes. Version the suite alongside the model.

05 / Scope of the evidence

Where this benchmark stops.

Mixed measurement scales

Exact matches, task-specific metrics and jury scores are normalized for aggregation. [2]

Private data affects reproducibility

Some full-suite results cannot be independently rerun from public assets alone. [2]

Jury validation has a scope

Agreement with clinicians is assessed on selected outputs and does not validate every possible specialty or failure type. [2]

Evolving inventories

The current community project describes a separate dataset inventory; it is not silently substituted for the paper’s 37 evaluations. [4]

06 / Working with the benchmark

Access & reuse.

Open the author’s resource ↗
Availability
Public, gated and private benchmark components
License
Evaluation software Apache 2.0; each dataset has its own terms.
Conditions
Access to code does not grant access to private benchmarks or waive source-data restrictions.
[4][2]

Evidence trail

Read the originals.

  1. Holistic evaluation of large language models for medical tasks with MedHELM ↗

    Bedi, Cui, Fuentes et al.; Nature Medicine. Final journal publication; 37-benchmark analytical snapshot.

  2. MedHELM accepted manuscript with full tables ↗

    Bedi et al.; PubMed Central. Full manuscript used to verify final-paper counts, scoring and selected results.

  3. MedHELM original preprint ↗

    Bedi et al.. Earlier 35-benchmark version; retained for explicit version comparisons.

  4. MedHELM community project and documentation ↗

    MedHELM community. Current project site, separate from the final-paper snapshot; software Apache 2.0 does not license underlying datasets.

Benchmark authors retain authorship of their work. This publication provides independent analysis; published rows are not Arcophos evaluation runs. Editorial method.

Explore the assumptions ↗