Vela: 適応的行動曲線パラメータ化による視覚-言語-行動モデルのスケーリング
Vela: Scaling Vision-Language-Action Models with Adaptive Action Curve Parametrization
未来のロボット動作を連続軌道として表し、スプライン表現と動作依存の時間サポートを組み合わせることで、固定出力予算でも時間解像度を適応させる視覚-言語-行動基盤モデルVelaを提案。大規模マルチ身体データで事前学習し、シミュレーションと実世界の長期的タスクで有効性を示した。
著者: Yifan Li, Jiaxu Wang, Dongming Wu, Yicheng Jiang, Ryan Ji, Xiangyu Yue, Yanwei Fu
分類: cs.RO, cs.AI
原文アブストラクト
Most vision-language-action models represent future motion as fixed-rate action chunks, tying temporal resolution and prediction horizon to a fixed output budget. This pointwise representation wastes capacity on highly correlated neighboring actions, leaves temporal continuity and smoothness to be learned implicitly, and forces a tradeoff between long-horizon coverage and the local precision required for contact-rich manipulation. To address these limitations, we introduce Vela, a vision-language-action foundation model that represents future robot behavior as continuous trajectories. Vela combines a compact spline-based action representation with motion-dependent temporal support and a shared action interface for heterogeneous embodiments, allowing a fixed output budget to adapt its temporal resolution across motions. We pretrain Vela on large-scale multi-embodiment robot data and evaluate it on LIBERO-X, EBench, and two real-world long-horizon tasks, egg-cake cooking and potato shredding, obtaining promising results across simulation and physical manipulation. These results highlight the potential of continuous action representations as a foundation for future embodied foundation models. Project page and more results: https://clementine24.github.io/Vela/ .