The short answer
MedHELM reports more than one view of performance. Pairwise win rate asks how often a model’s benchmark score meets or exceeds a rival’s; macro-average summarizes normalized scores across the suite. Open-ended tasks add another layer because a jury of language models supplies ratings. Our analysis keeps those layers separate. A high relative rank, a strong average score and agreement between automated and human raters are related pieces of evidence, but they do not answer the same question.
A win rate depends on the comparison set
The final paper compares nine historical models across its 37-benchmark inventory. Its stated win rule includes ties. A reader should therefore understand a win as the paper’s defined comparison outcome rather than assume it means strict superiority in every pairing.
Our deduction is that a win rate can change even if one model’s raw task results stay fixed: adding or removing opponents changes the comparisons being averaged. This makes a win rate useful within a fixed leaderboard snapshot but insufficient as a timeless property of the model. Keep the opponent set and benchmark membership with the reported value.
Macro-averaging assigns equal weight to benchmarks
The paper’s macro-average gives each benchmark equal influence regardless of its number of evaluated instances. That is an explicit aggregation choice. It avoids letting the largest dataset dominate, but it also means that an evaluation with fewer cases can affect the average as much as a much larger one.
Neither equal benchmark weighting nor pooled instance weighting is automatically the right measure for a hospital. The appropriate emphasis depends on the intended tasks and consequences. Our proposed interpretation is to use the suite aggregate as an index for exploration, then inspect the relevant task scores before making a narrower system comparison.
The two aggregate measures can disagree
The final table gives DeepSeek R1 and o3-mini the same reported win rate while their macro-averages differ. This is not a contradiction: one summary tracks pairwise order and the other tracks the magnitude of normalized scores. The same distinction can appear within a user’s own evaluation portfolio.
When two summaries point in different directions, identify the contributing tasks rather than selecting whichever supports a preferred model. A model can win narrowly on more tasks while another makes larger gains on fewer. Without task-level inspection, the aggregate alone cannot reveal which pattern matters to the application.
A jury is an evaluator with its own configuration
MedHELM’s final-paper jury contains three model families and rates open-ended responses on specified dimensions. Accuracy, completeness and clarity are central axes; a task-specific adjustment replaces completeness with structure for NoteExtract. The final response rating averages the judgments described in the method.
Our interpretation is that the evaluator must be versioned just as carefully as the candidate. A different jury, rubric or reference policy changes the measurement. Strong writing can also be a different property from correct clinical content, so preserving the separate rubric dimensions can be useful even when the published summary averages them.
Agreement evidence has a bounded meaning
The paper compares jury judgments with clinician ratings on selected outputs. That validation helps establish whether the automated ratings track the chosen human assessment under those conditions. It does not show that every possible error, specialty or deployment outcome has been adequately measured.
For a new use case, identify the important failure types and inspect whether the rubric can express them. Keep human disagreement visible, too: a reference judgment is not made infallible by calling it a gold standard. A transparent report states the candidate configuration, evaluator configuration, aggregation rule and unanswered application question together.
References & further reading
These original sources support the methods discussed here. Our suggested planning steps are editorial guidance, not an endorsement by the source authors.
- MedHELM accepted manuscript with full tables ↗Bedi et al.; PubMed Central. Full manuscript used to verify final-paper counts, scoring and selected results.
- MedHELM community project and documentation ↗MedHELM community. Current project site, separate from the final-paper snapshot; software Apache 2.0 does not license underlying datasets.
Published by Arcophos. Educational material, not clinical advice or a claim of regulatory compliance. Read our editorial method.