日本フィジカルAI新聞

世界のフィジカルAIを、日本語で。

週刊ニュースレター購読
VLAarXiv:2610.01741

ATI-VLA: 行動中心の予測型視覚-言語-行動モデル

ATI-VLA: Action-Centric Predictive Vision-Language-Action Models via Actionable Alignment Then Adaptive Injection

シェア:XThreadsFacebookLINEはてブBluesky

予測観測と行動の表現を共有コードブックで揃え、予測潜在を行動生成に適応的に注入することで、操作タスクの性能と収束速度を向上させたVLAモデル。

詳しい要約

1. どんなもの?

- 予測的なVision-Language-Action (VLA) モデルは、将来の観測や世界ダイナミクスの予測を通じてロボットマニピュレーションを改善することを目指す。 - しかし、既存のアプローチはこの可能性を実現できず、直接行動予測モデルよりも性能が劣ることが多い。 - この制限は、観測と行動の間のモダリティのミスアラインメントと、行動中心の目的から学習を逸らす共同最適化の競合に起因すると主張。 - そこで、ATI-VLAを導入。これはActionable Alignment Then Adaptive Injectionによる行動中心の予測的VLAフレームワーク。 - 2段階の設計:1) 共有コードブックによる行動可能な表現アラインメント、2) 予測潜在の行動中心適応的注入。 - シミュレーションと実世界のロボットタスクで最先端性能を達成し、収束も速い。

2. 先行研究と比べてどこがすごい?

- 既存の予測的VLAモデルは、将来の観測や世界ダイナミクスを予測するが、直接行動予測モデルに劣ることが多い。 - その原因は、観測と行動のモダリティのミスアラインメントと、共同最適化の競合にあると分析。 - ATI-VLAは、共有コードブックによるアラインメントと適応的注入により、これらの問題を解決。 - 結果として、シミュレーションと実世界タスクで最先端性能を達成し、収束も速い。 - 具体的な比較対象は要旨からは不明。

3. 技術・手法の肝は?

- 2段階の設計:1) Actionable Representation Alignment via a Shared Codebook。 - 予測観測と行動の表現を、統一コードブックを介して共有離散潜在空間にマッピングすることでアラインメント。 - これにより、予測観測潜在が行動生成に利用可能になり、モダリティのミスアラインメントを軽減。 - 2) Action-Centric Adaptive Injection of Predictive Latents。 - 予測観測潜在を、軽量な適応的サイドパスを介して行動デコーディングに明示的な予測事前として注入。 - 単一の行動中心目的の下で適応的な予測ガイダンスを可能にする。

4. どうやって有効だと検証した?

- シミュレーションと実世界のロボットタスクの両方で広範な実験を実施。 - ATI-VLAが最先端性能を達成し、収束が速いことを実証。 - 具体的なタスクやベースライン、評価指標は要旨からは不明。

5. 議論はある?

- 要旨からは不明。

6. 次に読むべき論文は?

- 要旨で参照/比較されている研究は明示されていない。 - 同分野の定番として、Vision-Language-Action (VLA) モデル、予測的VLAモデル、行動予測モデル、ロボットマニピュレーションに関する研究が挙げられる。

※ AIが要旨から生成した要約です。正確性は原文をご確認ください。

著者: Yijie Zhu, Rui Shao, Jie He, Wei Li, Bo Zhao, Yelin Wang, Xiaochen Yuan, Tao Tan, Miao Zhang, Xiaojiang Peng, Zitong Yu

分類: cs.CV

原文アブストラクト

Predictive Vision-Language-Action (VLA) models aim to improve robotic manipulation via future observation or world dynamics forecasting. However, existing approaches often fail to realize this potential and underperform direct action prediction models. We argue that these limitations stem from modality misalignment between observations and actions, together with joint optimization conflicts that drive learning away from an action-centric objective. To this end, we introduce ATI-VLA, an Action-Centric Predictive Vision-Language-Action framework via Actionable Alignment Then Adaptive Injection. Specifically, it follows a two-step design: 1) Actionable Representation Alignment via a Shared Codebook. It aligns predictive observation and action representations by mapping both modalities into a shared discrete latent space via a unified codebook, making predictive observation latents readily usable for action generation and mitigating modality misalignment. 2) Action-Centric Adaptive Injection of Predictive Latents. Building upon this, it then injects predictive observation latents into action decoding as explicit predictive priors via a lightweight adaptive side-path, enabling adaptive predictive guidance under a single action-centric objective. Extensive experiments on both simulation and real-world robotic tasks demonstrate that ATI-VLA achieves state-of-the-art performance with faster convergence.

関連論文

PR本紙発行元 EmplifAI