日本フィジカルAI新聞

世界のフィジカルAIを、日本語で。

週刊ニュースレター購読
VLAarXiv:2609.21948

GALA: 幾何情報を考慮した潜在行動モデリングによるマルチ身体VLA事前学習

GALA: Geometry-Aware Latent Action Modeling for Vision-Language-Action Model Pretraining across Embodiments

シェア:XThreadsFacebookLINEはてブBluesky

エンドエフェクタの3D幾何運動を統合した潜在行動表現を提案し、多様な身体のデータから細粒度の動作を捉えるVLAモデル事前学習を実現した。

詳しい要約

1. どんなもの?

- 多様なembodimentのデータから大規模VLAモデルを学習する枠組み - 画像ベースのlatent actionに3D end-effectorの幾何運動を追加するGALAを提案 - 手指レベルの微細なarticulationを捉え、embodiment横断の事前学習を狙う - 人間のaction-free ego-centric動画も活用対象に含む

2. 先行研究と比べてどこがすごい?

- 既存のimage-based LAMは微細なend-effector articulation、特に手指の幾何変化を捉えにくい - 単純にpoint cloudを足すと微細表現は得られるが共有semanticsが乏しくcross-embodiment事前学習を阻害 - GALAはUEMRで微細運動を保ちつつembodiment横断の汎化性を改善 - 従来のscene-level dynamicsに加え、共有された微細articulationの監督を提供

3. 技術・手法の肝は?

- Geometry-Aware Latent-Action modelingフレームワークGALAを提案 - Unified End-effector Motion Representation (UEMR)を導入 - UEMRは微細なmotion情報を保持しつつlatent actionのcross-embodiment汎化性を高める - 視覚latent action(scene-level dynamics)と幾何latent action(共有微細articulation)を組み合わせ - これによりmulti-embodimentデータとaction-free ego-centric人間動画からVLA事前学習を監督

4. どうやって有効だと検証した?

- fine-grained motion probingで評価 - cross-embodiment retrievalで評価 - 下流VLA評価を実施 - RoboCasa-GR1で68.3%のsuccess rateを達成 - 実世界で75.5%のsuccess rateを達成

5. 議論はある?

- 単純なpoint cloud導入は微細表現の共有semanticsを損なう点を課題として議論 - UEMRが微細運動保持とcross-embodiment汎化の両立に寄与すると主張 - 具体的な限界や失敗事例、計算コスト、データ依存性などの議論は要旨からは不明

6. 次に読むべき論文は?

- 既存のlatent action models (LAMs) - image-based LAM - 関連するVLA事前学習手法 - 具体的な参照論文名は要旨からは不明

※ AIが要旨から生成した要約です。正確性は原文をご確認ください。

著者: Yichen Liu, Puzhen Yuan, Xiang Zhu, Yanjiang Guo, Jianyu Chen

分類: cs.RO, cs.CV

原文アブストラクト

Learning large-scale vision-language-action (VLA) models from multi-embodiment datasets remains challenging due to heterogeneous action spaces across end effectors. Although latent action models (LAMs) can learn embodiment-agnostic action representations from diverse video data, existing image-based LAMs often fail to capture fine-grained end-effector articulation, particularly finger-level geometric changes in human and dexterous robot hands. To address this limitation, we propose GALA, a Geometry-Aware Latent-Action modeling framework that augments image-based latent actions with 3D end-effector geometric motion. However, naively incorporating point clouds yields fine-grained action representations with limited shared semantics, hindering cross-embodiment pretraining. To address this issue, we introduce the Unified End-effector Motion Representation (UEMR), which preserves fine-grained motion information while improving the cross-embodiment generalizability of latent actions. Building upon UEMR, GALA combines visual latent actions that capture scene-level dynamics with geometric latent actions that capture shared fine-grained end-effector articulation, providing effective supervision for VLA pretraining from multi-embodiment data, including action-free ego-centric human videos. Experiments on fine-grained motion probing, cross-embodiment retrieval, and downstream VLA evaluation demonstrate GALA's effectiveness in modeling generalizable fine-grained motions across embodiments, achieving 68.3% RoboCasa-GR1 success rate and 75.5% real-world success rate. Code, appendix, and demos are available at https://puzhenyuan.github.io/GALA-website/.

関連論文

PR本紙発行元 EmplifAI