日本フィジカルAI新聞

世界のフィジカルAIを、日本語で。

週刊ニュースレター購読
VLAarXiv:2609.15870

WLA³: 意味・動力学・運動学のための世界潜在行動モデリング

WLA$^3$: World Latent Action Modeling for Semantics, Dynamics, and Kinematics

シェア:XThreadsFacebookLINEはてブBluesky

世界状態の変化から潜在行動を学習し、意味・動力学・運動学を統合した汎用ポリシーモデルを構築するフレームワークを提案。

詳しい要約

1. どんなもの?

- 異種データで汎用policy modelを学習する枠組み WLA$^3$ を提案。 - World Latent Action Model (WLAM) が学習する表現を中心に構成。 - WLAMは局所区間でのマルチモーダル世界状態の変化を学習。 - 同期したcamera viewとembodiment-state変化を入力。 - コンパクトなlocal latent actionと遷移特徴に符号化。 - 表現をsemantics, dynamics, kinematicsで再利用。 - local latent actionは物理dynamicsモデリングに寄与。 - segment-level特徴はSemantic Latent Aggregate (SLA)経由でVLMを監督。 - action expertがlatent actionとembodiment固有のrobot制御を同時予測。

2. 先行研究と比べてどこがすごい?

- 異種データでの汎用policy学習は統一された低ノイズなaction supervision不足が制約。 - 一人称human videoは豊富だが高品質なhand-actionラベル付きは僅少。 - 観測されたworld transitionをaction関連supervisionの共通源として活用。 - human videoでスケーラブルなtransition supervisionを得る。 - robot trajectoryで共有表現を実行可能なnative controlに接地。 - 結果としてLARYBenchで最終32D latent actionが平均67.89%分類精度。 - 実ロボット6タスクで平均成功率81.9%、$\pi_{0.5}$の66.2%を上回る。

3. 技術・手法の肝は?

- World Latent Action Model (WLAM) を中核に据える。 - 局所区間のマルチモーダル世界状態変化を学習。 - 同期camera viewと利用可能なembodiment-state変化を入力。 - コンパクトなlocal latent actionとより豊かなtransition featureに符号化。 - 部分モダリティからの再構成と重複window間の一貫性で頑健な遷移表現を促す。 - WLA$^3$は表現をsemantics, dynamics, kinematicsで再利用。 - local latent actionがaction-sensitiveな物理dynamicsモデリングを支援。 - segment-level特徴がSemantic Latent Aggregate (SLA)を通じVLMを直接監督。 - action expertがlatent actionとembodiment固有robot制御を同時予測。

4. どうやって有効だと検証した?

- LARYBenchで最終32D latent actionが平均67.89%の分類精度。 - 実ロボット6タスクで平均成功率81.9%。 - 比較対象 $\pi_{0.5}$ は66.2%。 - 汎用policy modelのmid-trainingデータ規模拡大に伴い性能が向上。 - human videoがhuman-to-robot転移を支えることを確認。 - 詳細な実験設定やablationは要旨からは不明。

5. 議論はある?

- human videoはスケーラブルなtransition supervisionを提供。 - robot trajectoryは共有表現を実行可能なnative controlに接地。 - データ規模拡大で性能向上し、human-to-robot転移も支持。 - 限界、失敗事例、計算コスト、倫理面の議論は要旨からは不明。

6. 次に読むべき論文は?

- 比較対象として $\pi_{0.5}$ が挙げられている。 - 関連手法としてWorld Latent Action Model (WLAM)、Semantic Latent Aggregate (SLA)、VLM、action expertが要旨で言及。 - ベンチマークとしてLARYBenchが参照されている。 - その他の具体的な先行研究や次に読むべき論文は要旨からは不明。

※ AIが要旨から生成した要約です。正確性は原文をご確認ください。

著者: Peidong Liu, Zhiyuan Xiang, Mingyang Li, Wenhao Li, Jiale Zhang, Jiahao Sun, Jiawei Li

分類: cs.RO

原文アブストラクト

Scaling generalist policy models with heterogeneous data is limited by the lack of unified, low-noise action supervision. Human egocentric videos are abundant, but only a small fraction comes with high-quality hand-action labels. Observed world transitions offer a common source of action-related supervision across data sources. We introduce WLA$^3$ (World Latent Action Modeling for Semantics, Dynamics, and Kinematics), a unified generalist policy model framework built around representations learned by a World Latent Action Model (WLAM). WLAM first learns how multimodal world states change over a local interval, encoding synchronized camera views and available embodiment-state changes into a compact local latent action and a richer transition feature. Reconstruction from partial modalities and consistency across overlapping windows encourage robust transition representations. WLA$^3$ reuses them across semantics, dynamics, and kinematics: local latent actions support action-sensitive physical-dynamics modeling, segment-level features directly supervise the VLM through a Semantic Latent Aggregate (SLA), and an action expert jointly predicts latent actions together with embodiment-specific robot controls. Human videos provide scalable transition supervision, while robot trajectories ground the shared representation in executable native controls. On LARYBench, the final 32D latent action reaches 67.89\% average classification accuracy. WLA$^3$ achieves 81.9% average success across six real-robot tasks versus 66.2% for $π_{0.5}$. Performance improves as generalist policy model mid-training data scales, and human videos support human-to-robot transfer. Project page can be found at https://wla-3.github.io/.

関連論文