ArXiv

Improving the Realism of Synthetic Clinical Benchmarks Under Utility Constraints

Authors
Omid Bazgir, Md Nasir, Jacob Hoffman...
Categories
cs.AI, cs.DB, cs.LG
arXiv
https://arxiv.org/abs/2608.06265v1
PDF
https://arxiv.org/pdf/2608.06265v1

Brief

The paper introduces utility‑constrained realism improvement for synthetic clinical benchmarks and tests it on a Synthea‑derived care‑gap dataset run through demonstration EHR workflows. Realism is quantified by missingness structure, simplicity, structural plausibility, and population alignment. Two deterministic revisions increased realism without dropping below an operational utility floor, whereas naive densification kept unrealistic templating, highlighting realism as a separate optimization target.

Why it matters

Formulates 'utility‑constrained realism improvement' and applies it to a Synthea‑derived care‑gap benchmark; baseline diagnostics show sampled‑pair missingness = 79.44%, only 12.75% of rows are actionable, 38.94% of patients have zero actionable measures, and top‑three token concentration = 100.0%.

Key details

  • Two deterministic revisions improved realism (missingness, actionability, structural plausibility) while remaining above the operational utility floor; a naive densification control preserved templated/unrealistic structure, and internal benchmark realism is distinct from aggregate source fidelity.
Source evidence

Abstract

Synthetic clinical benchmarks for enterprise AI agents can pass existing utility checks and still remain structurally unrealistic, especially in privacy-sensitive healthcare settings where operational data are hard to access. We study how to improve such benchmarks without breaking the downstream utility checks already used in practice. We formulate benchmark revision as utility-constrained realism improvement: dataset changes should increase realism while staying above an operational utility floor. We instantiate this idea on a care-gap benchmark derived from Synthea-generated patients exercised through demonstration electronic health record workflows and then processed by the same downstream pipeline as operational data. Realism is measured through missingness structure, simplicity, structural plausibility, and population alignment. The baseline benchmark is extremely thin: sampled-pair missingness is 79.44%, only 12.75% of rows are actionable, 38.94% of patients have zero actionable measures, and top-three token concentration reaches 100.0%. Two deterministic revisions improve these panels while remaining above the current utility floor, whereas a naive densification control preserves unrealistic templating. We further show that internal benchmark realism and source fidelity to an aggregate operational reference are related but distinct objectives. These results suggest that synthetic benchmark quality should be optimized explicitly, with utility treated as one constraint rather than as sufficient evidence of realism.