ArXiv

Private Generative Bootstrap via Blocking

Authors
Jinwon Sohn, Veronika Ročková
Categories
stat.ML, cs.LG, stat.ME
arXiv
https://arxiv.org/abs/2608.02480v1
PDF
https://arxiv.org/pdf/2608.02480v1

Brief

The paper introduces PGBB, a differentially private, likelihood-free Bayesian bootstrap that groups records and assigns block weights to hide individual contributions, learns a noise-calibrated push-forward map via amortized inference, and provides posterior draws without further privacy cost. The authors prove DP, convergence and dispersion-restoring tuning, and demonstrate improved private UQ on Census and natality examples (abstract-based summary).

Why it matters

Sohn and Ročková (2026) propose the Private Generative Bayesian Bootstrap (PGBB): a blocked Bayesian bootstrap that randomly groups individuals, assigns one weight per group, and uses amortized inference with a privately trained push-forward map (noise added during training) so subsequent posterior draws incur no additional privacy or computation cost.

Key details

  • The paper proves a differential privacy guarantee, analyzes convergence of PGBB to the non-private blocked-bootstrap target, quantifies the discrepancy versus the ordinary Bayesian bootstrap, and gives a data-free tuning rule for the block Dirichlet concentration to asymptotically restore posterior dispersion.
  • Empirically, PGBB delivers competitive private uncertainty quantification in simulations and in applications to U.S. Census returns-to-schooling and U.S. natality birthweight quantiles, improving over private Bayesian alternatives that require a specified data-generating model.
Source evidence

Abstract

With AI systems gaining more access to individuals' information, it is important to protect privacy when reporting statistical answers. Equally important is to privatize the reporting of uncertainty in such answers. To this end, we adopt a Bayesian likelihood-free framework and make simulation from the posterior private. In particular, we propose a new private instantiation of the Bayesian bootstrap using a blocking strategy. Rather than assigning idiosyncratic random weights to each individual, we randomly group individuals and assign a single weight to each group. By concealing individuals' contributions within a group, we fortify differential privacy gates. We harness amortized inference that decouples private learning from posterior sampling. A push-forward map from observation weights to posterior samples is learned privately by adding calibrated noise during training. Subsequent posterior draws require no additional privacy and computation budget. We call the resulting method the Private Generative Bayesian Bootstrap (PGBB). We establish a differential privacy guarantee, analyze convergence to the non-private blocked-bootstrap target, and quantify the discrepancy between the ordinary and blocked Bayesian-bootstrap posteriors. In addition, we derive data-free tuning of the block Dirichlet concentration parameter that restores posterior dispersion asymptotically. We also show a single fit of PGBB can support a family of loss-based decision rules simultaneously without additional privacy cost. In simulations and in applications to U.S. Census returns to schooling and U.S. natality birthweight quantiles, PGBB gives competitive private uncertainty quantification and improves over private Bayesian alternatives that require a specified data-generating model in common settings.