Can regularization based JEPA (e.g. SIGReg) scale and compete with SOTA foundation models (DINO)? Here is the answer: yes and with 10x less data.
VISReg (slight variation of SIGReg) competes with DINOv2-LVD142M while only training on inet22k.
Try it out: huggingface.co/BooBooWu/visr…
Haiyu Wu (@HaiyuWu1)
Working on world model or SSL? You definitely need to try our new work: VISReg!
What does it achieve?
💪 Strong collapse prevention: High gradient when embedding collapse
⚡ Friendly to scale training: Linear complexity to scaling factors
🧩 Easy to train: Similar to LeJEPA, it is a heuristic-free method
🏆 Best OOD performance: Achieving the best accuracy on 6 OOD datasets
📉 Data efficiency: Achieving a similar OOD average accuracy to DINOv2 with 90% less data
🧬 Robust to low-quality datasets: It is robust to long-tailed and sparse datasets
Our results also indicate that SIGReg type methods can scale up, filling in the missing piece in @ylecun's great talk piped.video/watch?v=72Xj8k5W….
A big thanks to my co-author @randall_balestr and my manager @DrMorganLevine. Also, huge gratitude to @ylecun for connecting us to make this project happen! 🤝
SelfSupervisedLearning #JEPA #WorldModel
Video
— https://nitter.net/HaiyuWu1/status/2070866534462087626#m