ArXiv

MedFailBench: A Clinician-Built Open-Source Benchmark for Medical AI Safety Boundary Inspection

Authors
Goktug Ozkan
Categories
cs.AI, cs.CL
arXiv
https://arxiv.org/abs/2607.15166v1
PDF
https://arxiv.org/pdf/2607.15166v1

Brief

MedFailBench (Goktug Ozkan, arXiv 2026-07-16) is a clinician-built synthetic benchmark and failure atlas that categorizes medical AI errors by severity (1–5) and six safety-gate types. Release v0.2.1 includes 44 clinician-reviewed cases, a taxonomy, severity rubric, an automated pipeline, and a HuggingFace leaderboard preview; it is openly licensed (Apache‑2.0, CC‑BY‑4.0) and makes no clinical validation or patient-data claims.

Why it matters

MedFailBench is a clinician-built synthetic 'failure atlas' that labels medical AI errors by severity (1–5) and by six safety-gate types: missed urgent escalation; unsafe remote dosing; unsafe discharge reassurance; evidence fabrication; unsafe protocol execution; source support gap.

Key details

  • Release v0.2.1 (arXiv 2026-07-16; author Goktug Ozkan) provides 44 clinician-reviewed synthetic cases, a safety-gate taxonomy, a clinical severity rubric, an automated screening pipeline, and a HuggingFace leaderboard preview; licensed Apache-2.0 and CC-BY-4.0 and archived at DOI 10.5281/zenodo.21205535 — contains no patient data, clinical validation claims, or model rankings.
Source evidence

Abstract

Most medical AI benchmarks measure whether a model knows the correct answer. MedFailBench asks a different question: which safety boundary failed? We present a clinician-built synthetic benchmark and failure atlas that labels medical AI errors by severity (1--5) and safety gate type (missed urgent escalation, unsafe remote dosing, unsafe discharge reassurance, evidence fabrication, unsafe protocol execution, source support gap). The current public release (v0.2.1) contains 44 clinician-reviewed synthetic cases with severity annotations, a live HuggingFace leaderboard preview, a safety gate taxonomy, a clinical severity rubric, and an automated pipeline for archiving model-response screening runs. No patient data, clinical validation claims, or model rankings are included. MedFailBench is released under Apache-2.0 and CC-BY-4.0 and carries the Zenodo DOI 10.5281/zenodo.21205535.

Comment: 6 pages; clinician-reviewed synthetic benchmark; no patient data