# Clinical Eval > An independent MedS-Bench analysis: clinical task coverage, dataset versions, task-specific metrics, published NER results and an interactive evaluation atlas. MedS-Bench spans answer selection, extraction and generated medical text. Each output needs a different measurement. This publication maps the final journal benchmark to its task definitions, scorers and interpretation limits. Explore the task atlas, inspect the historical NER results, and follow the guides to distinguish a closed-set label from a complete clinical workflow. Arcophos contributes independent analysis; the MedS-Bench authors created the benchmark, data and evaluation methods. ## Provenance Independent analysis published by Arcophos. Benchmark creation belongs to the credited authors. Result rows are selected paper-reported measurements with their source versions and evaluation conditions, not new Arcophos runs or a live leaderboard. ## Benchmark dossiers - [MedS-Bench](https://clinicaleval.ai/benchmarks/meds-bench/): Clinical tasks need different output contracts and different metrics. Source version: Final npj Digital Medicine article, 2025. ## Original analyses - [Which MedS-Bench metric answers your question?](https://clinicaleval.ai/guides/meds-bench-task-metric-atlas/): Map clinical language tasks to accuracy, entity F1 and reference-overlap measures before interpreting model performance. - [What “diagnosis” and “treatment planning” mean inside MedS-Bench](https://clinicaleval.ai/guides/meds-bench-task-names-and-clinical-claims/): Inspect the answer space behind familiar clinical task names and avoid extending label accuracy into unsupported workflow claims. - [Reproducing MedS-Bench starts with the split and the parser](https://clinicaleval.ai/guides/reproduce-meds-bench-splits-and-parsers/): Keep training-data counts, benchmark composition, sampled evaluation cases and scorer revisions separate. ## Inspect the evidence - [Evidence JSON](https://clinicaleval.ai/evidence.json): Task definitions, dataset facts, scoring rules, source-version results, our interpretations, and reference IDs. - [Sources](https://clinicaleval.ai/sources/): Original papers and repositories with evidence locators. - [Editorial method](https://clinicaleval.ai/methodology/): Source reconciliation and interpretation boundaries. - [About](https://clinicaleval.ai/about/): Ownership and corrections. Analysis updated: 2026-09-28