Twitter/X

T-Rex unifies vision, language, and tactile sensing so robots can respond to…

Brief

T-Rex, from researchers at UC Berkeley, NVIDIA, and Stanford, integrates vision, language, and tactile inputs to enable real-time robotic responses to contact. The work is grounded in a 100-hour tactile-synchronized teleoperation dataset across 200+ objects and 22 motor primitives; finger motion was captured with Manus Meta gloves and retargeted to Sharpa Robotics Wave bimanual hands. Code and video are available at tactile-rex.github.io.

Why it matters

T-Rex unifies vision, language, and tactile sensing so robots can respond to physical contact in real time, moving beyond vision-only control.

Key details

  • The project is built on a 100-hour tactile-synchronized teleoperation dataset covering 200+ everyday objects and 22 motor primitives; data were recorded with Manus Meta gloves and retargeted to Sharpa Robotics Wave dexterous hands. Code: tactile-rex.github.io
Source evidence

Researchers from @UCBerkeley, @nvidia, and @Stanford introduce T-Rex, a framework that unifies vision, language, and tactile sensing so robots can respond to physical contact in real time rather than relying on vision alone.

The foundation is a 100-hour tactile-synchronized teleoperation dataset spanning 200+ everyday objects and 22 motor primitives. During data collection, researchers wore @ManusMeta gloves to capture precise finger motion, which was then retargeted onto @SharpaRobotics Wave dexterous hands for bimanual teleoperation.

Code: tactile-rex.github.io

Video