GT-VLA: 汎化可能なロボットマニピュレーションのための目標条件付きトレース誘導
GT-VLA: Target-Conditioned Trace Guidance for Generalizable Robotic Manipulation
汎用VLMが生成する2D視覚トレースを条件としてVLAの行動生成を誘導し、未見タスクや長期的操作への汎化性能を高めるフレームワークを提案。
著者: Ninghan Zhong, Jing-Chen Peng, Sriram Vishwanath
分類: cs.RO, cs.LG
原文アブストラクト
Vision-Language-Action (VLA) models have shown strong performance on robotic manipulation, but they often struggle to generalize to unseen tasks, configurations, and long-horizon settings. A key challenge is that VLAs overfit to training scenes and fail to follow novel language instructions. Off-the-shelf vision-language models (VLMs) often provide stronger generalization, but cannot directly control robot actions. To combine the common sense of VLMs with VLA control, we propose Guided Trace VLA (GT-VLA), a steerable framework that accepts guidance from an external generalist VLM through trace-conditioned action generation. GT-VLA uses a generalist model to identify semantic guidance for the current skill, converts this guidance into a 2D visual trace, and conditions its action policy on the resulting trace-rendered observation. This design separates semantic target acquisition, trace generation, and low-level action execution, allowing high-level guidance to propagate to robot actions. GT-VLA uses a Mixture-of-Experts architecture with skill-specific trace and action modules for robust execution. We evaluate GT-VLA on LIBERO and a physical robot platform, showing improved generalization over recent VLA baselines in both settings. The code and additional supplemental materials are available on our project website at https://ivaniz.github.io/gt-vla/.