ArXiv

Marginal conformal prediction severely under-covers rare costly…

Authors
Manpreet Singh, Akshatha Srikantha, Shyamal Lakhanpal
Categories
cs.LG, cs.AI
arXiv
https://arxiv.org/abs/2607.27143v1
PDF
https://arxiv.org/pdf/2607.27143v1

Brief

Cost-sensitive conformal prediction and human-in-the-loop abstention address uncertainty quantification for imbalanced, high-stakes classification. The authors benchmark marginal CP, Mondrian (class-conditional) CP, and cost-controlled abstention across 15 real-world tabular datasets, 7 models, 3 calibration techniques and 3,150 runs, showing marginal CP can under-cover minorities (down to 0.5%). Mondrian CP raises minority coverage by 61.7 percentage points (p < 1e-80), and its combination with cost-aware abstention lowers expected decision cost while identifying dataset-specific break-even thresholds for human deferral.

Why it matters

Marginal conformal prediction severely under-covers rare costly classes—minority-class coverage dropped to as low as 0.5% on some datasets.

Key details

  • In a large benchmark (15 real-world imbalanced tabular datasets, 7 models, 3 calibration methods, 10 seeds → 3,150 runs), Mondrian (class-conditional) CP restored valid minority-class coverage, improving average minority coverage by 61.7 percentage points over marginal CP (p < 1e-80); combining Mondrian CP with cost-controlled abstention also significantly reduced expected decision cost versus standard decision boundaries, confidence-based rejectors, and risk-controlled rejectors and yielded dataset-specific human-review break-even thresholds.
Source evidence

Abstract

High-stakes decision systems in credit scoring, fraud detection, healthcare, and industrial safety require reliable uncertainty quantification under severe class imbalance and asymmetric error costs. Standard marginal conformal prediction (CP) provides valid overall coverage guarantees; however, we show that it severely under-covers rare, costly minority classes, with minority-class coverage dropping to as low as 0.5% on certain datasets. To characterize and address this limitation, we conduct a comprehensive benchmark comparing marginal CP, class-conditional (Mondrian) CP, and cost-controlled abstention mechanisms across 15 real-world imbalanced tabular datasets, 7 classification models, 3 probability calibration techniques, and 10 random seeds, resulting in 3,150 experimental runs. Our results show that Mondrian CP restores valid minority-class coverage, achieving an average minority-coverage improvement of 61.7 percentage points over marginal CP (p < 1e-80). Furthermore, combining Mondrian CP with cost-controlled abstention significantly reduces expected decision cost compared with standard decision boundaries, confidence-based rejectors, and risk-controlled rejectors under realistic human review budgets. We further quantify dataset-specific break-even thresholds at which deferring ambiguous instances to human experts becomes cost-effective. These findings provide practical guidance for deploying distribution-free, cost-aware uncertainty quantification in high-stakes decision support systems.