ArXiv

QQWorld: Quantile-Quantile Matching for World Model Regularization

Authors
Zhoushun Yu, Xiaoyu Hu, Xiangyu Xu
Categories
cs.LG, cs.AI, cs.CV, cs.MM, cs.RO
arXiv
https://arxiv.org/abs/2607.28415v1
PDF
https://arxiv.org/pdf/2607.28415v1

Brief

QQWorld proposes replacing the Epps–Pulley regularizer used in LeWorldModel with a quantile–quantile matching objective that directly aligns projected latent samples to rank-matched Gaussian quantiles, maintaining corrective gradients in distribution tails. The paper also develops cross-batch QQ (using detached past-batch samples) and analyzes its bias–variance trade-off. On four control environments the method improves LeWM planning success and yields tighter Gaussian alignment and thinner latent tails; full text is available on arXiv (2026-07-30).

Why it matters

QQWorld (Zhoushun Yu, Xiaoyu Hu, Xiangyu Xu; arXiv 2026-07-30) replaces LeWorldModel's Epps-Pulley (EP) latent regularizer with a quantile–quantile (QQ) matching objective that aligns projected latent samples to rank-matched Gaussian quantiles, preserving corrective gradients in tail samples and addressing EP's vanishing-gradient problem for isolated tails.

Key details

  • The authors introduce cross-batch QQ, which enlarges the effective ranking pool by using detached samples from previous batches and characterize its bias–variance trade-off; empirically, across four control environments QQWorld raises LeWM's average planning success rate while producing better Gaussian alignment and thinner latent tails.
Source evidence

Abstract

Latent world models enable efficient planning by predicting future states in a compact representation space, but their performance depends critically on the quality of the learned latent distribution. LeWorldModel (LeWM) regularizes its latents toward an isotropic Gaussian using the Epps-Pulley (EP) objective. We show that the corrective gradients of EP rapidly vanish for isolated tail samples, leaving heavy-tailed deviations insufficiently controlled. To address this limitation, we propose QQWorld, which replaces EP with a quantile-quantile matching objective that directly aligns projected latent samples with rank-matched Gaussian quantiles, thereby maintaining effective corrective gradients in the tails. We further develop cross-batch QQ, which enlarges the effective ranking pool using detached samples from previous batches, and characterize its bias-variance trade-off. Across four control environments, QQWorld effectively improves the average planning success rate of LeWM, while consistently yielding better Gaussian alignment and thinner latent tails.