ArXiv

Safety That Does Not Transfer: Cross-Lingual Clinical Correctness Drift in Deployable Medical Language Models

Authors
Anthonio Oladimeji Gabriel, Dimeji Olawuyi, Toba Ajayi...
Categories
cs.CL, cs.CY
arXiv
https://arxiv.org/abs/2607.17270v1
PDF
https://arxiv.org/pdf/2607.17270v1

Brief

The paper evaluates cross‑lingual clinical correctness drift when English safety does not transfer to Hausa for locally deployable medical LLMs. Using 128 matched English–Hausa question pairs across malaria, sickle cell disease, and tuberculosis (knowledge, triage, contraindicated prompts, traditional‑remedy claims) and six models (five 4–9B deployable, one frontier), deployable models’ mean correctness dropped from 1.57 to -0.03 while the frontier model fell only 2.00→1.75 with no harmful answers. Full text was not provided with this abstract.

Why it matters

Across 128 matched English–Hausa responses for malaria, sickle cell disease, and tuberculosis, mean clinical correctness for five locally deployable models (4–9B params; two medically fine-tuned) fell from 1.57 (English) to -0.03 (Hausa) on a -1 to 2 scale where 2 = correct and -1 = actively harmful; drift was consistent across all three conditions.

Key details

  • A single frontier model maintained strong cross‑lingual performance (2.00 → 1.75) and produced no responses judged harmful in either language, implying the failure is a property of the deployable model tier rather than the Hausa language or the clinical tasks; two fluent Hausa raters scored responses blind (κ = 0.70 for correctness; κ = 0.22 initially for harm).
Source evidence

Abstract

Safety evaluation of large language models is conducted predominantly in English and predominantly on frontier systems. Neither condition describes how such models are encountered in low-resource health settings, where small quantised systems are run locally and queried in local languages. We ask whether clinical safety established in English transfers to Hausa, and whether any failure is attributable to the language, the clinical task, or the class of model that low-resource deployment admits. Matched English-Hausa question pairs were built for three conditions of high burden in northern Nigeria: malaria, sickle cell disease, and tuberculosis, probing knowledge recall, emergency triage, a leading question inviting a contraindicated action, and a traditional-remedy claim. Six models were evaluated: five locally deployable systems of 4-9 billion parameters, two medically fine-tuned, and one frontier system. All 128 responses were scored against Nigerian national treatment guidelines by two fluent Hausa speakers working independently and blind to one another. Among locally deployable models, mean clinical correctness fell from 1.57 in English to -0.03 in Hausa, on a scale where 2 denotes a correct answer and -1 an actively harmful one. The frontier model moved from 2.00 to 1.75 and produced no response judged harmful in either language. Drift was consistent across all three conditions. Inter-rater agreement was substantial for clinical correctness (kappa = 0.70); agreement on harm was initially poor (kappa = 0.22) and is examined in detail. Because a frontier model answers the same questions competently in Hausa, the deficit is a property neither of the language nor of the clinical material, but of the deployable tier.