Independent benchmark analysis / Final npj Digital Medicine article, 2025

MedS-Bench

Clinical tasks need different output contracts and different metrics.

MedS-Bench assembles medical language tasks that extend beyond selecting an examination answer. Our analysis uses the final journal version and keeps the evaluation collection separate from MedS-Ins, its associated training dataset. The central question is what each metric can actually support: exact or parsed label agreement, entity extraction, or similarity to reference wording. The task atlas exposes those distinctions before any model comparison. Published NER results are shown within one metric family; we do not combine unrelated measures into an invented overall ranking or treat a high closed-set score as evidence of a complete clinical workflow.

01 / What is being tested?

The task, before the score.

input
Task instructions plus medical text or a structured question; input differs by task family.
output
A label, extracted entity set, or generated text depending on the task.
unit
Task-specific evaluation example
setting
Final paper generally uses three-shot prompting; MCQA uses zero-shot. Selected subsets are separately released.

Data origin. A collection of existing biomedical and clinical datasets, including examinations, text records and literature-derived resources; provenance and access vary by constituent dataset. [1][2][3][4][5][6]

Source datasets
28

Final paper and official data card; README retains an earlier count of 39.

Figure 1; data-card Introduction [1][2][3]
Distinct tasks
52

Final journal figure caption.

Figure 1 [1]
High-level categories
11

Data-card taxonomy groups explanation and rationale together.

Introduction [2]
General prompting
3-shot

MCQA is the explicitly separate zero-shot condition.

Results: evaluation settings [1]
Training set ≠ test set
MedS-Ins

The associated training collection’s instance count is not MedS-Bench evaluation size.

Introduction [3]
  1. 01

    Select a named task

    Identify source dataset, task definition and required output format.

    [2]
  2. 02

    Load the reproduction split

    Use the authors’ selected cases where reproducing reported results.

    [2][4]
  3. 03

    Apply the prompt protocol

    Preserve the study’s shot setting and instruction format.

    [1]
  4. 04

    Use the task-specific scorer

    Retain output parsing and metric definitions with the result.

    [1][5][6]

02 / Measurement

Task-specific accuracy, F1, BLEU and ROUGE

Higher is better within a fixed metric and task.

Label tasks use accuracy, entity extraction uses F1, and generated text uses reference-overlap measures. Keep these families separate.

Scoring definition

accuracy = correct/N; entity F1 = 2TP/(2TP + FP + FN)

Match the selected split, task, prompt examples and output parser. The same numeric value has different meanings across metric families. [1][6][5]

03 / Measured evidence

Results, with their conditions attached.

Paper-reported results / selected rows

Selected results within one metric family

Final 2025 journal paper, named-entity recognition tasks; few-shot evaluation.

Mean task F1 × 100 · points
050100
Reported
GPT-4Average NER F1 reported in the paper
59.52
InternLM 2Average NER F1 reported in the paper
45.69
Llama 3Average NER F1 reported in the paper
23.62
MMedIns-Llama 3Instruction-tuned author model; average NER F1
79.29

Paper-reported historical results. These averages concern the NER task set only, not the full benchmark. No numeric uncertainty intervals are supplied.

Source: Table 6 [1]

04 / Our original analysis

What follows from the design?

01

A task name can overstate the output space.

Published evidence

The paper describes SEER treatment choices as eight categories and clinical outcome tasks as binary labels. [1]

Our interpretation

Interpret the score as agreement within that output space; it does not evaluate a complete individualized plan.

02

The parser is part of the measurement.

Published evidence

The pinned information-extraction code tests whether normalized reference text occurs in the response. [5]

Our interpretation

Additional material can leave this metric unchanged. Review contradictions separately when the intended application requires that distinction.

03

A benchmark total can hide unlike units.

Published evidence

The benchmark mixes accuracy, F1 and reference-overlap metrics. [1][6]

Our interpretation

An unqualified average would combine measurements of different things. Read task-level profiles instead.

04

Reproduction begins with the selected cases.

Published evidence

The official card links a separately released reproduction split. [2][4]

Our interpretation

Sampling another subset changes the evaluation condition even when the source dataset name remains constant.

05 / Scope of the evidence

Where this benchmark stops.

Version differences

Final paper/data-card composition differs from the current README; this page fixes the final journal version. [1][3][2]

Proxy metrics

Reference overlap and label agreement do not independently establish medical safety or utility. [1]

Mixed provenance and access

Source resources have different origins and conditions; the collection should not be described as one uniform real-patient cohort. [1][2]

Current code versus paper protocol

Pinned public scripts are observable implementations; this analysis does not claim every historical result used the identical commit. [5][6]

06 / Working with the benchmark

Access & reuse.

Open the author’s resource ↗
Availability
Official data collection, reproduction split and evaluation code are published.
License
Repository LICENSE states CC BY-SA without a version; a uniform constituent-data license is not verified.
Conditions
Consult each source dataset’s access terms. This site publishes aggregate metadata, not records.
[3][2][4]

Evidence trail

Read the originals.

  1. Towards evaluating and building versatile large language models for medicine ↗

    Chaoyi Wu and colleagues · npj Digital Medicine. Final journal version: benchmark composition, evaluation settings, task-specific metrics and results.

  2. MedS-Bench official data card ↗

    MAGIC / Henrychur. Official task taxonomy and distinction between full source data and the reproduction split.

  3. MedS-Ins and MedS-Bench author repository ↗

    MAGIC: Medical Artificial General Intelligence Consortium. Author code; distinguishes training instances from benchmark data. README retains an earlier source-dataset count.

  4. MedS-Bench reproduction split ↗

    MAGIC / Henrychur. Published selected examples used for reproduction; do not silently substitute fresh random samples.

  5. MedS-Bench information-extraction scorer ↗

    MAGIC. Current public scorer normalizes text and checks reference containment in the response.

  6. MedS-Bench entity-recognition scorer ↗

    MAGIC. Current public script aggregates entity TP, FP and FN within tasks and averages task F1 values.

Benchmark authors retain authorship of their work. This publication provides independent analysis; published rows are not Arcophos evaluation runs. Editorial method.

Explore the assumptions ↗