ArXiv

Patch Policy: Efficient Embodied Control via Dense Visual Representations

Authors
Gaoyue Zhou, Zichen Jeff Cui, Ada Langford...
Categories
cs.RO, cs.LG
arXiv
https://arxiv.org/abs/2607.18236v1
PDF
https://arxiv.org/pdf/2607.18236v1

Brief

Patch Policy presents a lightweight transformer extension that feeds dense pretrained ViT patch tokens to control policies via a block-causal attention mask, preserving temporal causality and reducing compute compared to full VLMs. According to the abstract, it achieves a 40% relative improvement over global-pooled baselines and an 18% win versus OpenVLA-OFT at ~0.7% of parameters. Summary based on the abstract only.

Why it matters

Patch Policy introduces a block-causal attention mask that lets transformer-based robot policies consume dense, pre-trained ViT patch tokens directly while preserving temporal causality and avoiding the computational cost of full vision-language models.

Key details

  • Empirical gains: across four simulated and three real-world environment suites, Patch Policy yields a 40% relative improvement over policies using state-of-the-art global-pooled representations and outperforms fine-tuned OpenVLA-OFT by 18% while using roughly 0.7% of the parameters.
Source evidence

Abstract

Pretrained dense visual features from Vision Transformers (ViTs) are powerful yet have been underutilized in robot learning. Modern robot policies either compress each observation into a single global token, or rely on visual backbones trained from scratch, sacrificing both fine-grained spatial detail and the benefits of large-scale visual pre-training. While there exist policies that do operate on dense patch features like large vision-language-action models (VLAs), they tend to be heavy and slow, inheriting the full cost of a billion-parameter vision-language model (VLM) backbone. We close this gap with Patch Policy, a minimal architectural extension that enables transformer-based policies to consume dense pre-trained patch tokens directly without the computational overhead of a full VLM. At its core is a block-causal attention mask that preserves the temporal causality of standard policies while letting the model attend over many patch tokens per observation, alongside other state information. Patch Policy is lightweight, fast, and highly effective. Across four simulated and three real-world environment suites, our method achieves a 40% relative improvement over policies using state-of-the-art global-pooled representations. Furthermore, it surpasses fine-tuned OpenVLA-OFT by 18% while using roughly 0.7% of the parameters. We believe Patch Policy provides a pipeline for the robotics community to readily leverage continuing progress in visual representation learning, without sacrificing the training efficiency or inference speed required for high-frequency, reactive control. Videos can be viewed at https://patch-policy.github.io