ChainVLA: Chaining Vision-Language-Action Queries through a Unified Execution State for Long-Horizon Manipulation
Preprint, 2026
Abstract
ChainVLA is a 1.2B-parameter VLA policy that connects successive action queries through a joint, revisable execution state, carrying both task progress and unfinished motion across long-horizon manipulation.
- We introduce ChainVLA, which combines Progress Context with Motion Tail to preserve observation-derived task progress and the preceding prediction's unexecuted continuation across VLA queries. ChainVLA reaches 62.8% average success on RMBench and 98.8% across four LIBERO suites.