Independent benchmark analysis / Final npj Digital Medicine article, 2025
MedS-Bench
Clinical tasks need different output contracts and different metrics.
MedS-Bench assembles medical language tasks that extend beyond selecting an examination answer. Our analysis uses the final journal version and keeps the evaluation collection separate from MedS-Ins, its associated training dataset. The central question is what each metric can actually support: exact or parsed label agreement, entity extraction, or similarity to reference wording. The task atlas exposes those distinctions before any model comparison. Published NER results are shown within one metric family; we do not combine unrelated measures into an invented overall ranking or treat a high closed-set score as evidence of a complete clinical workflow.
Task instructions plus medical text or a structured question; input differs by task family.
output
A label, extracted entity set, or generated text depending on the task.
unit
Task-specific evaluation example
setting
Final paper generally uses three-shot prompting; MCQA uses zero-shot. Selected subsets are separately released.
Data origin. A collection of existing biomedical and clinical datasets, including examinations, text records and literature-derived resources; provenance and access vary by constituent dataset. [1][2][3][4][5][6]
Source datasets
28
Final paper and official data card; README retains an earlier count of 39.
Match the selected split, task, prompt examples and output parser. The same numeric value has different meanings across metric families. [1][6][5]
03 / Measured evidence
Results, with their conditions attached.
Paper-reported results / selected rows
Selected results within one metric family
Final 2025 journal paper, named-entity recognition tasks; few-shot evaluation.
Mean task F1 × 100 · points
050100
Reported
GPT-4Average NER F1 reported in the paper
59.52
InternLM 2Average NER F1 reported in the paper
45.69
Llama 3Average NER F1 reported in the paper
23.62
MMedIns-Llama 3Instruction-tuned author model; average NER F1
79.29
Paper-reported historical results. These averages concern the NER task set only, not the full benchmark. No numeric uncertainty intervals are supplied.
MAGIC: Medical Artificial General Intelligence Consortium. Author code; distinguishes training instances from benchmark data. README retains an earlier source-dataset count.
MAGIC. Current public script aggregates entity TP, FP and FN within tasks and averages task F1 values.
Benchmark authors retain authorship of their work. This publication provides independent analysis; published rows are not Arcophos evaluation runs. Editorial method.