ArXiv

TurboVLA: Real-Time Vision-Language-Action Model at 32 Hz on an RTX 4090 with <1 GB VRAM

Authors
Hengyi Xie, Chenfei Yao, Xianjin Wu...
Categories
cs.CV, cs.RO
arXiv
https://arxiv.org/abs/2607.27205v1
PDF
https://arxiv.org/pdf/2607.27205v1

Brief

TurboVLA presents a lightweight vision-language-action policy that replaces costly LLM-centred pipelines with a direct V+L→A design: separate visual and language encoders, bidirectional interaction, and a compact decoder that outputs continuous action chunks. On LIBERO it reaches 97.7% success with 0.2B parameters, 31.2 ms latency, and 0.9 GB VRAM on an RTX 4090, enabling real-time (≈32 Hz) robot control. Full paper available as an arXiv abstract (full text not provided here); code repository is linked by the authors.

Why it matters

TurboVLA discards the LLM-centric V→L→A pipeline in favor of a direct V+L→A mapping: it independently encodes vision and language, uses lightweight bidirectional vision-language interaction, and predicts continuous action chunks with a compact decoder to cut inference compute and memory.

Key details

  • On the LIBERO benchmark TurboVLA achieves 97.7% average success while using only 0.2B parameters, 31.2 ms inference latency (~32 Hz), and 0.9 GB VRAM on a consumer RTX 4090, matching or outperforming substantially larger VLA policies; code is on GitHub.
Source evidence

Abstract

Vision-language-action (VLA) models commonly adopt an LLM-centric $V \to L \to A$ pathway, where visual observations are projected into the representation space of a large language model before being decoded into robot actions. Although effective, this design incurs substantial computation and memory overhead at every policy invocation. In this work, we introduce TurboVLA, a new VLA paradigm that reformulates the conventional $V \to L \to A$ pathway as a direct $V + L \to A$ mapping. Instead of using a large language model as the central interface between perception and action, TurboVLA independently encodes visual observations and language instructions, directly exchanges information between them through lightweight bidirectional vision-language interaction, and predicts continuous action chunks with a compact decoder. This simple design constructs task-conditioned representations directly from visual and linguistic features, significantly reducing the computational and memory costs of VLA inference. On LIBERO, TurboVLA achieves 97.7% average success with only 0.2B parameters, 31.2 ms inference latency, and 0.9 GB inference VRAM on a consumer-grade RTX 4090, matching or outperforming substantially larger VLA policies. These results establish TurboVLA as a simple and effective alternative to the prevailing LLM-centric VLA paradigm, offering a new perspective on how vision, language, and action can be connected for efficient robotic manipulation. Code is available at https://github.com/H-EmbodVis/TurboVLA.

Comment: Code is available at https://github.com/H-EmbodVis/TurboVLA