ArXiv

When Does Synthetic Data Augmentation Improve Score-Based Imbalanced Classification?

Authors
Zhengchi Ma, Pengfei Lyu, Anru R. Zhang
Categories
stat.ML, cs.LG
arXiv
https://arxiv.org/abs/2606.26053v1
PDF
https://arxiv.org/pdf/2606.26053v1

Brief

Synthetic minority-data augmentation: Ma, Lyu, and Zhang (2026) develop a theoretical framework separating augmentation effects into effective class-weight changes and synthetic-vs-true minority distribution discrepancy. They show that under well-specified score models the raw estimator already yields population-optimal likelihood-ratio ranking (so augmentation can only reduce variance or add bias), while under misspecification augmentation can correct ranking errors via reweighting; explicit bounds and simulations validate these claims.

Why it matters

For AUROC, AUPRC, best-threshold balanced accuracy, and best-threshold F1, the paper proves that under well-specified score models the empirical (raw) estimator already attains the population-optimal likelihood-ratio ordering; synthetic minority augmentation cannot improve population-level metrics beyond possible finite-sample variance reduction and can introduce bias from synthetic-vs-true minority distribution discrepancy, and minimax lower bounds show the raw estimator achieves the optimal metric-regret rate.

Key details

  • Under model misspecification, augmentation can change the effective class weighting and therefore the restricted-class projection, enabling correction of ranking errors caused by the imbalanced objective; the authors derive explicit improvement bounds that decompose gains into approximation error, finite-sample estimation error, and synthetic distributional error.
  • Simulation studies (Ma, Lyu, Zhang; arXiv:2606.26053v1, posted 2026-06-24) corroborate the theory: only limited gains appear under well-specification, while misspecified settings yield nontrivial but nonmonotone improvements from synthetic minority augmentation.
Source evidence

Abstract

Synthetic data augmentation is widely used to mitigate class imbalance, but its theoretical effects on score-based classification remain poorly understood. This paper develops a framework for characterizing when synthetic minority augmentation can improve threshold-integrated and threshold-optimized metrics, including AUROC, AUPRC, best-threshold balanced accuracy, and best-threshold (\F_1) score. We separate the effect of augmentation into two components: a change in effective class weighting and a discrepancy between the synthetic and true minority distributions. Under well-specified score models, the raw estimator already targets the likelihood-ratio ordering, which is population-optimal for the metrics considered. Consequently, augmentation cannot provide a fundamental population-level improvement beyond possible finite-sample variance reduction, and may introduce additional bias through synthetic distributional error. We further establish minimax lower bounds showing that the raw estimator already achieves the optimal metric-regret rate in the well-specified regime. Under misspecification, however, augmentation can play a qualitatively different role: by changing the effective class balance, it can alter the restricted-class projection and correct ranking errors induced by the raw imbalanced objective. We provide explicit improvement bounds quantifying the roles of approximation error, finite-sample estimation error, and synthetic distributional error. Simulation studies corroborate the theory, demonstrating limited gains under well-specification and nontrivial but nonmonotone improvements under misspecification.