Taxonomy coverage is not equal evidence density
37 benchmarks span 22 subcategories and 121 named tasks. [2]
Coverage of a category establishes representation; it does not establish equally deep testing of every named task.
Independent benchmark analysis / Final journal paper: 37-benchmark snapshot
MedHELM organizes clinical language-model evaluation around the work a health system needs performed. We analyze the final journal paper’s 37-benchmark snapshot, its task taxonomy and its combination of deterministic scoring with an LLM jury. The central question is what an overall rank compresses: different task families, different dataset access conditions and different measures of success. Our original contribution is a coverage map and an interpretation of the aggregation choices. We keep the final paper separate from both its earlier preprint and the evolving community leaderboard, because their inventories and comparison sets differ.
01 / What is being tested?
Data origin. A mixture of existing public benchmarks, gated clinical datasets and private clinical datasets; no single source population describes the entire suite. [1][2][3][4]
Final-paper benchmark inventory.
Extended Data Table 1 [2]Organized into five categories and 22 subcategories; not 121 independently measured datasets.
Clinician validation of the taxonomy [2]Access categories in the final-paper inventory.
Overview of the benchmark suite [2]Historical model snapshots listed in Table 1.
Table 1 [2]GPT-4o, Claude 3.7 Sonnet and LLaMA 3.3 70B.
Evaluation using LLM-jury [2]Map the task to the clinician-validated taxonomy.
[2]Choose an input, reference and metric for that task.
[2]Apply a deterministic task metric or the three-model jury.
[2]Preserve task-level results before averaging over the 37 benchmark scores.
[2]Dataset anatomy
Final-paper category count.
Includes education and messaging.
Documentation tasks.
Research tasks.
Operational tasks.
Counts sum to 37. The paper’s open/closed-ended prose split is internally inconsistent with this total, so it is not charted. Bar lengths use the largest listed count as their reference; they are not percentages of a shared population.
02 / Measurement
Higher is better within the fixed suite and opponent set.
Pairwise wins average comparisons over benchmarks and opponents; macro-average gives every benchmark equal weight.
Win = 1[score_m,b ≥ score_r,b]; macro = Σ normalized benchmark score / 37
A win rate depends on the other models. Normalized scores mix different measures; neither aggregate is a patient-outcome percentage. [2]
03 / Measured evidence
Paper-reported results / selected rows
Nature Medicine 2026 final paper, Table 1; 37 benchmarks and nine historical models.
SD is variation across the paper’s comparisons, not a confidence interval. Ties count as wins under the stated method. Table 1 takes precedence over a swapped Claude macro-average sentence.
Source: Table 1 [2]
04 / Our original analysis
37 benchmarks span 22 subcategories and 121 named tasks. [2]
Coverage of a category establishes representation; it does not establish equally deep testing of every named task.
The final inventory contains 16 public, seven gated and 14 private benchmarks. [2]
A public-only result answers a narrower question. It should not inherit the full-suite headline without a matching denominator.
DeepSeek R1 and o3-mini both have a 0.66 win rate but macro-averages 0.75 and 0.77. [2]
A tie in relative rank can coexist with a difference in average normalized magnitude. Neither eliminates the need to inspect the intended workflow.
05 / Scope of the evidence
Exact matches, task-specific metrics and jury scores are normalized for aggregation. [2]
Some full-suite results cannot be independently rerun from public assets alone. [2]
Agreement with clinicians is assessed on selected outputs and does not validate every possible specialty or failure type. [2]
The current community project describes a separate dataset inventory; it is not silently substituted for the paper’s 37 evaluations. [4]
Evidence trail
Bedi, Cui, Fuentes et al.; Nature Medicine. Final journal publication; 37-benchmark analytical snapshot.
Bedi et al.; PubMed Central. Full manuscript used to verify final-paper counts, scoring and selected results.
Bedi et al.. Earlier 35-benchmark version; retained for explicit version comparisons.
MedHELM community. Current project site, separate from the final-paper snapshot; software Apache 2.0 does not license underlying datasets.
Benchmark authors retain authorship of their work. This publication provides independent analysis; published rows are not Arcophos evaluation runs. Editorial method.