日本フィジカルAI新聞

世界のフィジカルAIを、日本語で。

週刊ニュースレター購読
VLAarXiv:2609.30833

高速計画と忠実な行動:階層型視覚言語行動モデルにおける計画・実行ギャップの解消

Fast Plans, Faithful Actions: Closing the Planning-Execution Gap in Hierarchical Vision-Language-Action Models

シェア:XThreadsFacebookLINEはてブBluesky

階層型VLAの計画器と実行器の不一致を分析し、ブロック自己回帰デコーディングと正規化ゴール変調で計画遅延を8.7倍削減しつつ計画の活用を改善した。

詳しい要約

1. どんなもの?

階層型 vision-language-action (VLA) システムの planning-execution gap を扱う研究。 - 対象は $π_{0.5}$ を改変した waypoint hierarchy pipeline。 - 高レベル vision-language planner と低レベル action expert から成る。 - 課題は planner の計画生成が実時間制御に間に合わないことと、計画が action 生成に十分寄与しないこと。 - 提案は waypoint-aligned block-autoregressive decoding (Block-AR) と normalized goal modulation (NGM)。

2. 先行研究と比べてどこがすごい?

既存の階層型 VLA ベースラインの欠点を明示し改善。 - baseline は token-level autoregressive decoding (Token-AR) で waypoint plan を生成し、57 回の高コスト VLM forward pass を要した。 - 提案手法は LIBERO で最大 VLM forward pass を 57 から 8 に削減。 - Rokae dual-arm robot で planning latency を 8.7× 削減。 - Block-AR と NGM により LIBERO-Long の成功率を 91.0% から 96.2% に、4 suite 平均を 95.85% から 98.45% に向上。

3. 技術・手法の肝は?

2つの技術が肝。 - waypoint-aligned block-autoregressive decoding (Block-AR): waypoint に整列した block 単位の自己回帰デコードで VLM forward pass を削減。 - normalized goal modulation (NGM): layer-wise goal path を phase gating と anti-shortcut training で制約し、waypoint が action 生成に影響を与えつつ他の信号を維持する。 - これにより planner の過剰に細かい出力粒度と executor の plan 低活用を是正。

4. どうやって有効だと検証した?

LIBERO と実機で検証。 - LIBERO で最大 VLM forward pass を 57 から 8 に削減(prefix prefill 1回を含む)。 - Rokae dual-arm robot で planning latency 8.7× 削減。 - Block-AR + NGM + anti-shortcut training で LIBERO-Long 成功率 91.0%→96.2%、4 suite 平均 95.85%→98.45%。 - 同ロボットの3つの bimanual task では手法間で成功率が同等。

5. 議論はある?

planner-executor の misalignment を指摘。 - planner の出力粒度が過剰に細かい。 - executor が plan を制御条件として十分使っていない。 - waypoint endpoints を消去しても task success への影響が小さいことを発見。 - ただし bimanual task では成功率が同等に留まる点など、限界や一般化可能性は要旨からは不明。

6. 次に読むべき論文は?

要旨で参照・比較されている研究や関連手法。 - $π_{0.5}$ を基にした waypoint hierarchy pipeline。 - token-level autoregressive decoding (Token-AR)。 - waypoint-aligned block-autoregressive decoding (Block-AR)。 - normalized goal modulation (NGM)。 - LIBERO benchmark。 - Rokae dual-arm robot。

※ AIが要旨から生成した要約です。正確性は原文をご確認ください。

著者: Chuanliang Xie, Boyu Ma, Gen Li, Yizhou Liu, Houwang Chen, Xinyu Zhou, Jianfei Yang

分類: cs.RO

原文アブストラクト

Hierarchical vision-language-action (VLA) systems consist of a high-level vision-language planner and a low-level action expert that generates continuous actions. This hierarchical design has practical value only if the planner can generate plans fast enough to meet real-time control requirements, and the resulting plans actually contribute to the generation of action. We study one such system, a waypoint hierarchy pipeline adapted from $π_{0.5}$, and find that neither requirement is satisfied. This baseline relies on token-level autoregressive decoding (Token-AR) to generate a waypoint plan, requiring 57 very expensive vision-language model (VLM) forward passes. However, we find that erasing the waypoint endpoints has little effect on task success. Two findings reveal the misalignment of planner-executor: the planner generates outputs at an excessively fine granularity, and the executor underuses plans as a control condition. We address the latency issue with waypoint-aligned block-autoregressive decoding (Block-AR), and plan underuse issue with normalized goal modulation (NGM), a layer-wise goal path constrained by phase gating and anti-shortcut training so that the waypoint influences action generation maintaining other signals. Our method reduces the maximum number of VLM forward passes from 57 to 8 on LIBERO, including one prefix prefill, and achieves an $8.7\times$ reduction in planning latency on a Rokae dual-arm robot. With normalized goal modulation and anti-shortcut training, Block-AR's success rate on LIBERO-Long increases from 91.0% to 96.2%, while its average success rate across the four suites increases from 95.85% to 98.45%. On three bimanual tasks with this robot, success rates remain comparable across methods.

関連論文

PR本紙発行元 EmplifAI