日本フィジカルAI新聞

世界のフィジカルAIを、日本語で。

週刊ニュースレター購読
VLAarXiv:2609.34467

三モーダル整合性誘導フロートランスフォーマーによる効率的な視覚-言語-行動ポリシー学習

Alignment-Guided Flow Transformer for Efficient Vision-Language-Action Policy Learning

シェア:XThreadsFacebookLINEはてブBluesky

視覚・言語・行動の三モーダル整合性を明示的に強制する損失を導入し、フローマッチング目的で推論を高速化したVLAモデルAGFTを提案。ベンチマークで成功率向上と低遅延を実現。

詳しい要約

1. どんなもの?

- Vision-Language-Action (VLA) モデルにおける tri-modal misalignment を解決する Alignment-Guided Flow Transformer (AGFT) を提案。 - 視覚・言語・行動の三モーダル間の表現ギャップを埋めるため、専用の alignment loss を導入。 - flow-matching 目的を採用し、diffusion-based policies より少ない推論ステップで高精度を実現。 - 大規模ベンチマークで SOTA ベースラインを上回る成功率と低遅延を達成。

2. 先行研究と比べてどこがすごい?

- 従来研究は主に bi-modal な vision-language alignment に焦点を当てていた。 - 本研究は tri-modal alignment を体系的に定式化し、VLA モデルにおけるその役割を分離・分析。 - アブレーションと分析により、適応性とロバスト性向上への寄与を実証。 - flow-matching により diffusion-based policies より推論ステップを大幅に削減しつつ精度を維持。

3. 技術・手法の肝は?

- tri-modal alignment を明示的に強制する alignment loss を導入。 - flow-matching 目的を採用し、効率的な推論を実現。 - 理論的に tri-modal alignment gap と flow matching の最適化 tightness の定量的関係を確立。 - アブレーションと分析により alignment の役割を分離。

4. どうやって有効だと検証した?

- 広範なベンチマーク実験を実施。 - SOTA ベースラインと比較し、成功率と推論遅延で優位性を確認。 - アブレーションと分析により tri-modal alignment の寄与を検証。 - 理論的定量的関係の確立も検証の一部。

5. 議論はある?

- tri-modal alignment がスケーラブルでロバストな VLA マニピュレーションの鍵であると主張。 - 理論的枠組みと実験結果から有効性を示すが、限界や課題については要旨からは不明。 - 今後の研究方向や実世界応用への議論は要旨からは不明。

6. 次に読むべき論文は?

- 要旨で参照/比較されている研究:diffusion-based policies、bi-modal vision-language alignment 研究。 - 関連手法:Vision-Language-Action (VLA) モデル、flow-matching。 - 同分野の定番:RT-1, RT-2, Octo, OpenVLA など(一般名として)。

※ AIが要旨から生成した要約です。正確性は原文をご確認ください。

著者: Shengchao Hu, Peng Wang, Qiyang Zhou, Guodong Zheng, Yuqi Huang, Li Shen, Ya Zhang, Dacheng Tao

分類: cs.LG

原文アブストラクト

Recent advances in Vision-Language-Action (VLA) models point toward general-purpose robotic intelligence by unifying perception, instruction, and control. Despite impressive progress, existing VLA models often adapt poorly due to \emph{tri-modal misalignment} among vision, language, and action, which weakens action grounding and hurts generalization and fine-tuning efficiency. In this work, we present Alignment-Guided Flow Transformer (AGFT), a novel framework that explicitly enforces tri-modal alignment through a dedicated alignment loss, bridging the representational gap across modalities and enhancing task adaptation. While prior research has predominantly emphasized bi-modal vision--language alignment, we systematically formalize and study tri-modal alignment in VLA models, and provide both ablations and analysis to isolate its role in improving adaptation and robustness. To further accelerate deployment, we adopt a flow-matching objective, enabling substantially fewer inference steps than diffusion-based policies while maintaining accuracy. Theoretically, we establish a quantitative connection between the tri-modal alignment gap and the optimization tightness of flow matching; empirically, experiments on the extensive benchmark show that AGFT achieves superior success rates and lower inference latency compared to SOTA baselines, underscoring tri-modal alignment as a key ingredient for scaling robust VLA manipulation.

関連論文

PR本紙発行元 EmplifAI