ArXiv

RouterVLA: Turning Smoke Tests into Supervision for Heterogeneous VLA Selection

Authors
Xingyu Ren, Chugang Yi, Ge Ma...
Categories
cs.RO
arXiv
https://arxiv.org/abs/2606.27355v1
PDF
https://arxiv.org/pdf/2606.27355v1

Brief

RouterVLA evaluates whether pre-deployment smoke-test rollouts can supervise selection among heterogeneous vision–language–action (VLA) policies by using outcome-disjoint cross-fitting: one set of probes builds frozen-expert profiles while separate scored trials estimate held-out performance. On 34,752 LIBERO-Plus rollouts the approach raises held-out success from 0.4686 to 0.6149 (+14.64 pp). Learned scalar scorers matched a transparent probe-success rule, and reusing the scored trial overstates benefits by 1.87×. Full text was not available; results suggest commissioning-aware routing adds system-level value beyond per-model scaling.

Why it matters

RouterVLA uses outcome-disjoint cross-fitting on 34,752 LIBERO-Plus rollout records to build probe profiles for frozen VLA experts and raises held-out success from 0.4686 to 0.6149 (a +14.64 percentage-point gain) using a transparent probe-success rule.

Key details

  • Under the scalar-only profiles studied, learned scorers are statistically indistinguishable from the simple rule; reusing the scored trial inflates measured gain by 1.87×, indicating routing/commissioning drives system gains beyond mere model scaling.
Source evidence

Abstract

We study whether pre-deployment evaluation rollouts can be reused to supervise policy selection. Robot teams routinely smoke test candidate vision-language-action (VLA) policies, then compress those trials into a global winner. RouterVLA evaluates this idea with outcome-disjoint cross-fitting: recorded probes build a profile for each frozen expert, and a separate trial scores the selected expert without entering its profile. Across 34,752 LIBERO-Plus rollout records, a transparent probe-success rule raises held-out success from 0.4686 to 0.6149, a +14.64pp gain. Under the scalar-only profiles studied here, learned scorers are statistically indistinguishable from this rule, showing that commissioning carries the routing value while extra scalar scorer capacity does not create it. Reusing the scored trial inflates the measured gain by $1.87\times$, so credible ledger routing needs outcome separation; model scaling improves individual policies, while commissioning-aware routing improves the system built from them.