{"publication":"Clinical Eval","url":"https://clinicaleval.ai","publisher":"Arcophos","updated":"2026-09-28","provenance":"Independent analytical publication. Benchmark creation and experimental results belong to their cited authors. Reported results are source-version snapshots, not new Arcophos runs or a live leaderboard.","benchmarks":[{"slug":"meds-bench","name":"MedS-Bench","shortName":"MedS-Bench","version":"Final npj Digital Medicine article, 2025","creators":"Chaoyi Wu, Pengcheng Qiu, Jinxin Liu, Hongfei Gu, Na Li, Ya Zhang, Yanfeng Wang and Weidi Xie","paperDate":"2025-01-27","headline":"Clinical tasks need different output contracts and different metrics.","summary":"MedS-Bench assembles medical language tasks that extend beyond selecting an examination answer. Our analysis uses the final journal version and keeps the evaluation collection separate from MedS-Ins, its associated training dataset. The central question is what each metric can actually support: exact or parsed label agreement, entity extraction, or similarity to reference wording. The task atlas exposes those distinctions before any model comparison. Published NER results are shown within one metric family; we do not combine unrelated measures into an invented overall ranking or treat a high closed-set score as evidence of a complete clinical workflow.","task":{"input":"Task instructions plus medical text or a structured question; input differs by task family.","output":"A label, extracted entity set, or generated text depending on the task.","unit":"Task-specific evaluation example","setting":"Final paper generally uses three-shot prompting; MCQA uses zero-shot. Selected subsets are separately released."},"dataOrigin":"A collection of existing biomedical and clinical datasets, including examinations, text records and literature-derived resources; provenance and access vary by constituent dataset.","facts":[{"label":"Source datasets","value":"28","detail":"Final paper and official data card; README retains an earlier count of 39.","sourceIds":["meds-paper","meds-card","meds-repo"],"locator":"Figure 1; data-card Introduction"},{"label":"Distinct tasks","value":"52","detail":"Final journal figure caption.","sourceIds":["meds-paper"],"locator":"Figure 1"},{"label":"High-level categories","value":"11","detail":"Data-card taxonomy groups explanation and rationale together.","sourceIds":["meds-card"],"locator":"Introduction"},{"label":"General prompting","value":"3-shot","detail":"MCQA is the explicitly separate zero-shot condition.","sourceIds":["meds-paper"],"locator":"Results: evaluation settings"},{"label":"Training set ≠ test set","value":"MedS-Ins","detail":"The associated training collection’s instance count is not MedS-Bench evaluation size.","sourceIds":["meds-repo"],"locator":"Introduction"}],"metric":{"name":"Task-specific accuracy, F1, BLEU and ROUGE","description":"Label tasks use accuracy, entity extraction uses F1, and generated text uses reference-overlap measures. Keep these families separate.","formula":"accuracy = correct/N; entity F1 = 2TP/(2TP + FP + FN)","direction":"Higher is better within a fixed metric and task.","comparability":"Match the selected split, task, prompt examples and output parser. The same numeric value has different meanings across metric families.","sourceIds":["meds-paper","meds-ner","meds-extraction"]},"workflow":[{"label":"Select a named task","detail":"Identify source dataset, task definition and required output format.","sourceIds":["meds-card"]},{"label":"Load the reproduction split","detail":"Use the authors’ selected cases where reproducing reported results.","sourceIds":["meds-card","meds-split"]},{"label":"Apply the prompt protocol","detail":"Preserve the study’s shot setting and instruction format.","sourceIds":["meds-paper"]},{"label":"Use the task-specific scorer","detail":"Retain output parsing and metric definitions with the result.","sourceIds":["meds-paper","meds-extraction","meds-ner"]}],"slices":[],"sliceTitle":"Coverage is organized by task, not invented sample counts","sliceNote":"The atlas below is a qualitative mapping. No verified all-task test-example denominator is asserted here.","results":[{"id":"meds-ner","title":"Selected results within one metric family","metric":"Mean task F1 × 100","unit":"points","lower":0,"upper":100,"scope":"Final 2025 journal paper, named-entity recognition tasks; few-shot evaluation.","sourceIds":["meds-paper"],"locator":"Table 6","rows":[{"label":"GPT-4","value":59.52,"display":"59.52","detail":"Average NER F1 reported in the paper"},{"label":"InternLM 2","value":45.69,"display":"45.69","detail":"Average NER F1 reported in the paper"},{"label":"Llama 3","value":23.62,"display":"23.62","detail":"Average NER F1 reported in the paper"},{"label":"MMedIns-Llama 3","value":79.29,"display":"79.29","detail":"Instruction-tuned author model; average NER F1"}],"note":"Paper-reported historical results. These averages concern the NER task set only, not the full benchmark. No numeric uncertainty intervals are supplied."}],"analysis":[{"heading":"A task name can overstate the output space.","evidence":"The paper describes SEER treatment choices as eight categories and clinical outcome tasks as binary labels.","interpretation":"Interpret the score as agreement within that output space; it does not evaluate a complete individualized plan.","sourceIds":["meds-paper"]},{"heading":"The parser is part of the measurement.","evidence":"The pinned information-extraction code tests whether normalized reference text occurs in the response.","interpretation":"Additional material can leave this metric unchanged. Review contradictions separately when the intended application requires that distinction.","sourceIds":["meds-extraction"]},{"heading":"A benchmark total can hide unlike units.","evidence":"The benchmark mixes accuracy, F1 and reference-overlap metrics.","interpretation":"An unqualified average would combine measurements of different things. Read task-level profiles instead.","sourceIds":["meds-paper","meds-ner"]},{"heading":"Reproduction begins with the selected cases.","evidence":"The official card links a separately released reproduction split.","interpretation":"Sampling another subset changes the evaluation condition even when the source dataset name remains constant.","sourceIds":["meds-card","meds-split"]}],"limitations":[{"title":"Version differences","detail":"Final paper/data-card composition differs from the current README; this page fixes the final journal version.","sourceIds":["meds-paper","meds-repo","meds-card"]},{"title":"Proxy metrics","detail":"Reference overlap and label agreement do not independently establish medical safety or utility.","sourceIds":["meds-paper"]},{"title":"Mixed provenance and access","detail":"Source resources have different origins and conditions; the collection should not be described as one uniform real-patient cohort.","sourceIds":["meds-paper","meds-card"]},{"title":"Current code versus paper protocol","detail":"Pinned public scripts are observable implementations; this analysis does not claim every historical result used the identical commit.","sourceIds":["meds-extraction","meds-ner"]}],"access":{"status":"Official data collection, reproduction split and evaluation code are published.","license":"Repository LICENSE states CC BY-SA without a version; a uniform constituent-data license is not verified.","restrictions":"Consult each source dataset’s access terms. This site publishes aggregate metadata, not records.","url":"https://huggingface.co/datasets/Henrychur/MedS-Bench","sourceIds":["meds-repo","meds-card","meds-split"]},"sourceIds":["meds-paper","meds-card","meds-repo","meds-split","meds-extraction","meds-ner"]}],"explorer":{"kind":"coverage","title":"MedS-Bench task and metric atlas","intro":"Filter the published task families by the kind of output a system must produce. Categories here are Arcophos editorial groupings, not measured performance or an official ranking.","caution":"A task-family label does not certify a clinical workflow. Counts and scores remain specific to the underlying dataset, selected split and scorer.","sourceIds":["meds-paper","meds-card"],"rows":[{"label":"Medical MCQA","category":"Select","input":"Exam question and options","output":"Selected option","metric":"Accuracy","constraint":"Correct selection does not score the safety of an explanation.","benchmarkSlug":"meds-bench","sourceIds":["meds-paper","meds-card"]},{"label":"Text summarization","category":"Generate","input":"Medical source text","output":"Condensed text","metric":"BLEU / ROUGE","constraint":"Reference overlap is not a direct factuality judgment.","benchmarkSlug":"meds-bench","sourceIds":["meds-paper","meds-card"]},{"label":"Information extraction","category":"Extract","input":"Text plus extraction instruction","output":"Requested fields","metric":"Parsed accuracy","constraint":"Current script checks normalized reference containment.","benchmarkSlug":"meds-bench","sourceIds":["meds-paper","meds-card"]},{"label":"Explanation and rationale","category":"Generate","input":"Concept or answer to explain","output":"Explanatory text","metric":"BLEU / ROUGE","constraint":"Lexical similarity cannot establish every reasoning step.","benchmarkSlug":"meds-bench","sourceIds":["meds-paper","meds-card"]},{"label":"Named entity recognition","category":"Extract","input":"Biomedical text","output":"Entity set","metric":"F1","constraint":"The set parser and normalization are part of the metric.","benchmarkSlug":"meds-bench","sourceIds":["meds-paper","meds-card"]},{"label":"Diagnosis","category":"Select","input":"DDXPlus case information","output":"Diagnosis from defined labels","metric":"Accuracy","constraint":"Closed disease vocabulary bounds the task.","benchmarkSlug":"meds-bench","sourceIds":["meds-paper","meds-card"]},{"label":"Treatment planning","category":"Select","input":"SEER case information","output":"Treatment category","metric":"Accuracy","constraint":"Eight high-level categories do not constitute a complete plan.","benchmarkSlug":"meds-bench","sourceIds":["meds-paper","meds-card"]},{"label":"Clinical outcome prediction","category":"Select","input":"MIMIC4ED case information","output":"Binary outcome label","metric":"Accuracy","constraint":"The task scores an outcome label, not a care action.","benchmarkSlug":"meds-bench","sourceIds":["meds-paper","meds-card"]},{"label":"Text classification","category":"Select","input":"Text plus candidate labels","output":"One or more labels","metric":"Precision / recall / F1","constraint":"Multilabel errors need the task-specific aggregation.","benchmarkSlug":"meds-bench","sourceIds":["meds-paper","meds-card"]},{"label":"Fact verification","category":"Mixed","input":"Claim or answer with task context","output":"Label or generated justification","metric":"Accuracy or BLEU / ROUGE","constraint":"Different subtask outputs use different metrics.","benchmarkSlug":"meds-bench","sourceIds":["meds-paper","meds-card"]},{"label":"Natural language inference","category":"Mixed","input":"Premise and hypothesis","output":"Entailment label or explanation","metric":"Accuracy or BLEU / ROUGE","constraint":"Label and generative versions are separate measurements.","benchmarkSlug":"meds-bench","sourceIds":["meds-paper","meds-card"]}],"parameters":[]},"references":[{"id":"meds-paper","title":"Towards evaluating and building versatile large language models for medicine","organization":"Chaoyi Wu and colleagues · npj Digital Medicine","url":"https://www.nature.com/articles/s41746-024-01390-4","note":"Final journal version: benchmark composition, evaluation settings, task-specific metrics and results.","locator":"Figure 1; Results; Tables 6–7","version":"Published 2025-01-27"},{"id":"meds-card","title":"MedS-Bench official data card","organization":"MAGIC / Henrychur","url":"https://huggingface.co/datasets/Henrychur/MedS-Bench","note":"Official task taxonomy and distinction between full source data and the reproduction split.","locator":"Introduction; sampling note","version":"Accessed 2026-09-28"},{"id":"meds-repo","title":"MedS-Ins and MedS-Bench author repository","organization":"MAGIC: Medical Artificial General Intelligence Consortium","url":"https://github.com/MAGIC-AI4Med/MedS-Ins","note":"Author code; distinguishes training instances from benchmark data. README retains an earlier source-dataset count.","locator":"README; LICENSE","version":"Commit 185cfc04a27023983e48b357a6f2fd525049b9af"},{"id":"meds-split","title":"MedS-Bench reproduction split","organization":"MAGIC / Henrychur","url":"https://huggingface.co/datasets/Henrychur/MedS-Bench-SPLIT","note":"Published selected examples used for reproduction; do not silently substitute fresh random samples.","locator":"Files and versions","version":"Accessed 2026-09-28"},{"id":"meds-extraction","title":"MedS-Bench information-extraction scorer","organization":"MAGIC","url":"https://github.com/MAGIC-AI4Med/MedS-Ins/blob/185cfc04a27023983e48b357a6f2fd525049b9af/Metrics/InformationExtraction.py","note":"Current public scorer normalizes text and checks reference containment in the response.","locator":"information_extraction_accuracy; evaluate_information_extraction","version":"Pinned 2025-05-05 commit"},{"id":"meds-ner","title":"MedS-Bench entity-recognition scorer","organization":"MAGIC","url":"https://github.com/MAGIC-AI4Med/MedS-Ins/blob/185cfc04a27023983e48b357a6f2fd525049b9af/Metrics/NamedEntityRecognition.py","note":"Current public script aggregates entity TP, FP and FN within tasks and averages task F1 values.","locator":"evaluate_hard_ner","version":"Pinned 2025-05-05 commit"}]}