TempoBridge: 視覚言語行動ポリシーのための言語誘導テンポ制御
TempoBridge: Language-Guided Tempo Control for Vision-Language-Action Policies
凍結したVLA表現を利用し、指示中のテンポ手がかりに応じて動作を調整する軽量フレームワークを提案。追加のテンポ条件付き実演や微調整なしで、LIBEROタスクのテンポ成功率を52.6%から89.7%に改善した。
詳しい要約
1. どんなもの?
2. 先行研究と比べてどこがすごい?
3. 技術・手法の肝は?
4. どうやって有効だと検証した?
5. 議論はある?
6. 次に読むべき論文は?
※ AIが要旨から生成した要約です。正確性は原文をご確認ください。
著者: Yeonseo Lee, Hyosup Shin, Guebin Hwang, Sungho Jo
分類: cs.RO
原文アブストラクト
Vision-Language-Action (VLA) models are effective at understanding what task to perform, but provide limited control over how it should be executed, such as moving quickly or slowly. We introduce TempoBridge, a lightweight framework that uses frozen VLA representations to modulate actions according to tempo cues in the instruction at each task phase, without additional tempo-conditioned robot demonstrations or tempo-specific base-policy fine-tuning. TempoBridge extracts tempo cues from contextual VLM representations, aligns them with task progress through a causal phase router, and modulates nominal motion commands during execution. Across LIBERO tasks, TempoBridge improves Tempo Success Rate from 52.6% to 89.7% under canonical tempo instructions while retaining high task success. It also preserves near-baseline performance when no tempo cue is present and generalizes to unseen tempo expressions without additional training. Experiments on a physical robot further demonstrate language-conditioned tempo modulation in real-world manipulation.