日本フィジカルAI新聞

世界のフィジカルAIを、日本語で。

週刊ニュースレター購読
VLAarXiv:2609.39973

EWAM:統合身体モデルにおける創発的な深さ方向の専門化——意味理解から視覚的先読み、行動まで

EWAM: Emergent Depth-Wise Specialization in a Unified Embodied Model -- From Semantic Understanding through Visual Foresight to Action

シェア:XThreadsFacebookLINEはてブBluesky

行動中心の統合身体モデルEWAMを提案し、層ごとの監督なしに浅い層で意味、中間層で予測未来、深い層で行動に注意が向く創発的な深さ方向の専門化が生じることを示した。

詳しい要約

1. どんなもの?

EWAMは、action-centricなunified embodied modelであり、asymmetric joint attentionによりaction tokensが各層でsemantic、current-visual、predicted-future、action情報を読む一方、perceptual expertsは役割を維持する。層ごとのsupervisionなしに、emergent depth-wise specializationが生じる。

2. 先行研究と比べてどこがすごい?

VLA policiesはsemantic understandingを重視し、WAMsは予測表現を学習する。両者を統合してもaction computationが単一expertに集中する既存システムに対し、EWAMはaction tokensが層ごとに異なる情報源へ注意を移すemergent specializationを示し、VLA、WAM、hybrid baselinesを上回る。

3. 技術・手法の肝は?

asymmetric joint attentionにより、action tokensが各層でsemantic、current-visual、predicted-future、action情報を読み、perceptual expertsは役割を保持する。層別supervisionなしで、shallow層ではvision-language features、intermediate層ではpredicted future frames、deep層ではaction tokensへ注意が移る。

4. どうやって有効だと検証した?

simulationとreal-robot experimentsでVLA、WAM、hybrid baselinesを上回る。checkpoint trackingとcausal interventionsにより、depth-wise specializationが学習され、action generationがそれに依存することを示す。human egocentric dataがcross-embodiment transferとreal-robot robustnessを改善し、subtask-phase supervisionがlong-horizon completionを改善する。

5. 議論はある?

emergent depth-wise specializationはタスク間で再現し、denoising steps間で安定する。unified embodied learningがsemantic understandingからvisual foresightを経てaction formationへ至る順序的な内部進行を誘導しうることを示唆する。

6. 次に読むべき論文は?

要旨で参照/比較されている研究はVLA policies、WAMs、hybrid baselines。関連手法としてcross-embodiment robot trajectories、human egocentric video、subtask-phase supervisionが挙げられる。

※ AIが要旨から生成した要約です。正確性は原文をご確認ください。

著者: Hao Wang, Jiajun Wen, Jingzhi Liu, Shuoshuo Xue, Zhiliang Chen, Min Lin, Yicheng Chang, Xiaoyu Guo, Yukang Zhuo, Zheng Chong, Yunshuang Nie, Jian Zhang, Weijia Liufu, Qingman Wu, Heming Xu, Bingchang Song, Dantong Wu, Zhiyuan Wang, Hang Xu, Jianhua Han, Bokui Chen, Shen Zhao, Rui Li, Xiaodan Liang

分類: cs.RO

原文アブストラクト

Vision-language-action (VLA) policies emphasize semantic understanding, whereas world-action models (WAMs) learn predictive representations of environment dynamics. Systems that expose a policy to both sources often still concentrate action computation on a single expert. We present EWAM, an action-centric unified embodied model whose asymmetric joint attention lets action tokens read semantic, current-visual, predicted-future, and action information at every layer while the perceptual experts retain their distinct roles. Without layer-wise supervision, EWAM develops an emergent depth-wise specialization: action queries attend mainly to vision-language features in shallow layers, to predicted future frames in intermediate layers, and to action tokens themselves in deep layers. This handoff replicates across tasks and is stable across denoising steps. Checkpoint tracking and causal interventions show that it is learned and that action generation depends on it. EWAM is pretrained in two separate regimes, one on cross-embodiment robot trajectories and one on human egocentric video. In simulation and real-robot experiments, it surpasses existing VLA, WAM, and hybrid baselines. Human egocentric data improve both cross-embodiment transfer and real-robot robustness, and subtask-phase supervision improves long-horizon completion. Together, these results suggest that unified embodied learning can induce an ordered internal progression from semantic understanding, through visual foresight, to action formation.

関連論文

PR本紙発行元 EmplifAI