🚀🚀🚀 Excited to introduce AFUN, our first step toward an affordance foundation model for functionality understanding in robotics, led by my PhD student @ZhaoningEricW
Affordance understanding bridges visual perception and physical action, serving as an explainable interface for robot manipulation in open and unstructured real-world environments. Yet, the challenges of building an affordance foundation model are (i) limited datasets to reflect task/environment diversity in the real world, (ii) precise task-conditioned mask understanding indicating where to interact, and (iii) actionable motion to indicate how to interact for robot manipulation. To tackle these, we (i) build a large-scale standardized data pipeline for "affordance" extraction, (ii) leverage a unified model that connects VLMs and SAM3 through MetaQuery, and (iii) use a Bézier Curve representation for motion prediction. The output can be directly deployed on a real robot for execution with zero fine-tuning.
For affordance segmentation, AFUN outperforms all baselines by a large margin across 8 test sets from 4 benchmarks; for contact-point prediction, it predicts substantially more accurate points, with a 12.7–61.3% hit-rate gain over the best baseline; and for 3D motion, it achieves the best performance on all three test sets.
Code & Weights are public. You're welcome to try it out!
Paper: arxiv.org/abs/2606.02551
Project page: zhaoningwang.com/AFUN
Code: github.com/EricWang12/AFUN
Robotics #Affordance #EmbodiedAI #Manipulation #AI
Video