ArXiv

One Future, Every Robot: Label-Efficient Collective-State Prediction with Decentralized JEPA

Authors
Alan-Barsag Gazzaev, Alexey Garvilov, Sergey Muravyov
Categories
cs.RO
arXiv
https://arxiv.org/abs/2607.28443v1
PDF
https://arxiv.org/pdf/2607.28443v1

Brief

Collective-State JEPA (CS-JEPA) is a decentralized recurrent joint-embedding predictor that enables each robot to produce a common future token field from 16-frame local histories and a 64-float per-edge recurrent message. Pretrained without labels and probed on 6/12/24 global episodes, CS-JEPA outperforms a raw-future baseline (+9,607 params) on prediction error and inter-robot agreement up to 108 robots, and cuts branch-value MSE by 45.5% while raising candidate-score Pearson correlation by 0.1291.

Why it matters

Introduces Collective-State JEPA (CS-JEPA), a decentralized recurrent joint-embedding predictive architecture where each robot uses a 16-frame local history and one 64-float recurrent message per directed edge to output a common future token field; design omits global pooling, target encoder, episode clock, and recorded future action.

Key details

  • After unsupervised pretraining and ridge-probe evaluation on 6, 12, or 24 globally labeled episodes, CS-JEPA improves prediction-error and inter-robot-agreement label-budget AUC versus a raw-future reconstruction baseline (which used 9,607 extra training-only parameters) across in-distribution, ring, mutual-kNN, and unseen-size families up to 108 robots; effects favored CS-JEPA in 5/5 outer seeds in a registered follow-up.
  • In an action-conditioned sealed eight-seed follow-up (predictors given four-step plans), CS-JEPA reduced branch-value MSE by 45.5% and increased within-context candidate-score Pearson correlation by 0.1291, with both effects favorable in 8/8 seeds including at unseen N=32.
Source evidence

Abstract

Can every robot in a swarm predict the same future collective state from only local observations and bandwidth-limited messages? We formulate this as decentralized shared-state prediction and introduce Collective-State JEPA (CS-JEPA), a recurrent joint-embedding predictive architecture whose output at every robot represents one common future token field. At deployment, each robot uses a 16-frame local history and one 64-float recurrent message per directed edge; there is no global pooling, target encoder, episode clock, or recorded future action. After pretraining without downstream collective labels, frozen representations are evaluated with ridge probes fitted on 6, 12, or 24 globally labeled episodes. Against raw-future reconstruction with the same receiver anchor and deployment capacity but 9,607 additional training-only parameters, a prospectively registered five-seed follow-up improves prediction-error and inter-robot-agreement label-budget AUC on in-distribution, ring, mutual-kNN, and unseen-size families up to 108 robots. Every effect favors CS-JEPA in 5/5 outer seeds. In a separate sealed eight-seed follow-up, matched action-conditioned predictors receive each candidate four-step plan before producing receiver-local predictive representations. CS-JEPA reduces branch-value MSE by 45.5% and improves within-context candidate-score Pearson correlation by 0.1291, with both effects favorable in 8/8 seeds, including at unseen N=32. These results support common-future JEPA targets as a label-efficient primitive for decentralized swarm prediction under topology and size shift, with additional evidence of planning-relevant value estimation.

Comment: Submitted to IEEE ICRA 2027