日本フィジカルAI新聞

世界のフィジカルAIを、日本語で。

週刊ニュースレター購読
VLAarXiv:2609.15005

IMPACT-VLA: 視覚言語行動ポリシーのための反事実軌道による相互作用対応マルチモーダル伝播帰属

IMPACT-VLA: Interaction-aware Multimodal Propagation Attribution via Counterfactual Trajectories for Vision-Language-Action Policies

シェア:XThreadsFacebookLINEはてブBluesky

VLAポリシーにおいて、各モダリティがタスク成功に寄与する実行段階を明らかにするため、反事実的な再実行に基づく帰属手法を提案し、LIBEROタスクで有効性を示した。

詳しい要約

1. どんなもの?

- Vision-Language-Action (VLA) policiesのマルチモーダル入力(視覚、固有受容感覚、言語指示)が、実行段階ごとに最終タスク成功にどう寄与するかを明らかにする研究。 - 既存の帰属手法は局所感度や時間集約重要度に限定され、段階依存の寄与や段階間依存を捉えられない。 - 提案手法IMPACT-VLAは、成功参照ロールアウトの行動段階をポリシークエリ境界に整列し、段階-モダリティブロックを帰属単位として定義。 - 閉ループ反事実再実行により各ブロックの最終タスク成功への寄与を定量化する。

2. 先行研究と比べてどこがすごい?

- 既存の帰属アプローチは局所感度や時間集約重要度を測るため、段階依存の寄与や段階間依存を捉えられない。 - IMPACT-VLAは行動段階をポリシークエリ境界に整列し、段階-モダリティブロックを帰属単位とすることで、段階依存の寄与を評価可能。 - 閉ループ反事実再実行により、入力介入が後続の状態・観測・行動に伝播する効果を定量化。 - 30のLIBEROタスクでOpenVLA-OFTを用い、25タスク(83.3%)で支配的モダリティ遷移を検出し、閉ループ帰属がStatic Action Perturbationより忠実にタスククリティカル情報を特定。

3. 技術・手法の肝は?

- 成功参照ロールアウトの行動遷移から行動段階を構築し、ポリシークエリ境界に整列。 - 段階-モダリティブロックを帰属単位として定義。 - 閉ループ反事実再実行を実施し、各ブロックの最終タスク成功への寄与を定量化。 - 段階間の非加法的相互作用と軌道伝播を分析し、行動的回復と機能的回復を区別。

4. どうやって有効だと検証した?

- 30のLIBEROロボットマニピュレーションタスクでOpenVLA-OFTを使用。 - 25タスク(83.3%)で支配的モダリティ遷移が発生。 - 閉ループ帰属がStatic Action Perturbationより忠実にタスククリティカル情報を特定。 - 負に相互作用するペアの後続ブロックの限界利得が、早期段階入力置換下で約3.3倍増加。 - 行動的回復なしに機能的回復が起こり得ることを示した。

5. 議論はある?

- マルチモーダル入力がいつタスク成功を支えるか、閉ループ実行中に寄与が条件付きで結合する仕組みを明らかに。 - 段階間の非加法的相互作用と軌道伝播の分析から、行動的回復と機能的回復の区別が重要。 - 限界や制限についての具体的な議論は要旨からは不明。

6. 次に読むべき論文は?

- 要旨で参照/比較されている研究: Static Action Perturbation。 - 関連手法: OpenVLA-OFT、LIBERO。 - 同分野の定番: Vision-Language-Action (VLA) policies、マルチモーダル帰属手法。

※ AIが要旨から生成した要約です。正確性は原文をご確認ください。

著者: Jinwoong Kim, Sangjin Park

分類: cs.RO, cs.AI

原文アブストラクト

Vision-Language-Action (VLA) policies perform robot manipulation tasks using multimodal inputs such as visual observations, proprioceptive states, and language instructions. However, it remains unclear at which execution stages each modality contributes to final task success and how input interventions propagate through subsequent states, observations, and actions. Existing attribution approaches primarily measure local sensitivity or temporally aggregated importance, limiting their ability to capture phase-dependent contributions and cross-phase dependencies. We propose Interaction-aware Multimodal Propagation Attribution via Counterfactual Trajectories for Vision-Language-Action Policies (IMPACT-VLA). IMPACT-VLA constructs behavioral phases from action transitions in a successful reference rollout, aligns them with policy query boundaries, and defines phase-modality blocks as attribution units. It then performs closed-loop counterfactual re-execution to quantify each block's contribution to final task success. We further analyze cross-phase non-additive interactions and trajectory propagation while distinguishing behavioral from functional recovery. Across 30 LIBERO robot manipulation tasks using OpenVLA-OFT, dominant-modality transitions occurred in 25 tasks (83.3%), and closed-loop attribution identified task-critical information more faithfully than Static Action Perturbation. Later-block marginal gains for negatively interacting pairs increased by approximately 3.3x under early-phase input replacement, while functional recovery could occur without behavioral recovery. These results reveal when multimodal inputs support task success and how their contributions become conditionally coupled during closed-loop execution.

関連論文