Twitter/X

@randall_balestr: Oops, SIGReg did it again! Large scale (CC12M->Datacomp-L) vision-language JEPA pretraining beats...

Oops, SIGReg did it again! Large scale (CC12M->Datacomp-L) vision-language JEPA pretraining beats CLIP and SigLIP objectives! Thanks to SIGReg, our LeVLJEPA has no collapse, no EMA, no stop-gradient, no negatives, no problem! Checkpoints/demo are live: levljepa.github.io

Lukas Kuhn (@lukaskuhn77)

🔥 We introduce LeVLJEPA: the first fully non-contrastive end-to-end vision-language pretraining method competitive with CLIP & SigLIP 💪🏼

👀 No negatives. No temperature. No momentum encoder. No teacher-student.

TL;DR: LeVLJEPA learns image to text structure by prediction: each modality predicts the other's embedding, while SIGReg keeps each embedding isotropic Gaussian. 🧵

📄 arxiv.org/abs/2607.00784

Video

— https://nitter.net/lukaskuhn77/status/2072685678572269701#m