日本フィジカルAI新聞

世界のフィジカルAIを、日本語で。

週刊ニュースレター購読
VLAarXiv:2609.18462

CSWAM: 世界行動モデルの分布外汎化のための因果意味表現の改善

CSWAM: Better Causal Semantic Representations for Out-of-Distribution Generalization in World Action Models

シェア:XThreadsFacebookLINEはてブBluesky

V-JEPA 2.1をベースにした因果意味エキスパートをFastWAMに追加し、見たことのない環境や物体でもロボット行動を汎化させる手法を提案。

詳しい要約

1. どんなもの?

- FastWAM 系の world action model を拡張した CSWAM を提案。 - 外観依存の再構成表現を因果意味表現に置き換え、視覚分布シフト下の汎化を改善。 - 推論時は action-only の効率的推論を維持。

2. 先行研究と比べてどこがすごい?

- FastWAM は action-only 推論が可能だが視覚分布シフトに弱い。 - 再構成志向の表現が外観固有の詳細に偏り、未知シーン・物体への汎化を制限。 - 観測履歴が無く、未知視覚条件下で状態変化・運動を頑健に同定できない。 - CSWAM は V-JEPA 2.1 ベースの因果意味 expert でこれらを補う。

3. 技術・手法の肝は?

- FastWAM に V-JEPA 2.1 上の causal semantic expert を追加。 - V-JEPA は外観依存を抑えた時間的に接地された意味的状態変化・運動表現を提供。 - expert は現在と過去の疎な観測履歴から将来の進展を学習。 - 履歴由来の context を causal attention で video と action の両ストリームに共有。 - 推論時は現在の video 状態と観測意味履歴に条件付け、action denoising を実行。

4. どうやって有効だと検証した?

- シミュレーションと実ロボット実験で分布シフト下の汎化を評価。 - embodied pretraining により RoboTwin 2.0 の Clean-to-Randomized transfer で Randomized success を 10.16% から 45.18% へ改善(FastWAM 比 +35.02pt)。 - 実ロボット 2 タスク・OOD 難易度 3 水準で平均成功率を 27.5% から 70.0% へ改善(+42.5pt)。

5. 議論はある?

- 要旨からは不明。 - 限界や失敗事例、計算コスト、V-JEPA 依存性などの議論は要旨に記載なし。

6. 次に読むべき論文は?

- FastWAM(比較対象の基盤モデル)。 - V-JEPA 2.1(因果意味 expert の基盤表現)。 - RoboTwin 2.0(評価ベンチマーク)。 - world action model 分野の関連手法(一般名)。

※ AIが要旨から生成した要約です。正確性は原文をご確認ください。

著者: Tianbin Liu, Jian Zhu, Taiyi Su, Jianjun Zhang, Chong Ma, Zitai Huang, Yi Xu

分類: cs.CV, cs.AI

原文アブストラクト

FastWAM-style world action models enable efficient action-only inference, but generalize poorly under visual distribution shifts. Their reconstruction-oriented representations emphasize appearance-specific details, limiting generalization to unseen scenes and objects. Without observation history, the model also lacks temporal evidence for robustly identifying task-relevant state changes and motion in unfamiliar visual conditions. To address these limitations, we present the Causal Semantic World Action Model (CSWAM), which augments FastWAM with a causal semantic expert built on V-JEPA 2.1. V-JEPA provides temporally grounded representations of semantic state changes and motion with less dependence on appearance-specific details. The expert learns their future evolution from a sparse history of current and past observations and shares the history-derived context with both the video and action streams through causal attention. At inference, CSWAM conditions action denoising on the current video state and observed semantic history, retaining efficient action-only inference. We conduct simulation and real-robot experiments to evaluate generalization under distribution shifts. With embodied pretraining, CSWAM raises Randomized success on RoboTwin 2.0 Clean-to-Randomized transfer from 10.16% to 45.18%, a gain of 35.02 percentage points over FastWAM. Across two real-robot tasks and three OOD difficulty levels, CSWAM improves average success over FastWAM by 42.5 percentage points, from 27.5% to 70.0%.

関連論文

PR本紙発行元 EmplifAI