日本フィジカルAI新聞

世界のフィジカルAIを、日本語で。

週刊ニュースレター購読
VLAarXiv:2610.09427

イベント整合型視覚行動推論によるワールドアクションモデル

Event-Aligned Visual Action Reasoning for World Action Models

シェア:XThreadsFacebookLINEはてブBluesky

タスク上重要なインタラクションイベントに視覚予測を整合させ、行動生成を導くフレームワークを提案し、DOMINOで成功率を10.26ポイント改善した。

詳しい要約

1. どんなもの?

- World-Action Models (WAMs) の新しい枠組み - 未来の視覚予測を中間推論として行動生成を導く - 既存WAMsは固定時間間隔で視覚想像を構成 - タスク重要interactionと接続transitionの役割を区別しない - 提案: event-aligned visual action reasoning framework - interaction eventsを中心に視覚-行動予測を組織 - event-aligned visual-action supervisionを導入 - 各imagined rolloutでevent-aligned visual contextを生成 - 重要state changesに重点を置く - execution validity headで有効部分を特定 - chunked inferenceでの冗長行動を回避

2. 先行研究と比べてどこがすごい?

- 既存WAMsはpredefined temporal intervalsで視覚想像を構造化 - task-critical interactionsとconnecting transitionsの役割を明示的に考慮しない - 提案は視覚的先見をtask-relevant interactionsとその推論要求に直接整合 - event-aligned supervisionで重要state changesを強調 - 視覚推論の粒度をinteraction dynamicsに応じて形成 - task-critical events周辺で詳細推論、接続transitionで粗い進行 - execution validity headでchunked inferenceの冗長行動を回避 - DOMINO success rateでbaseline比10.26 percentage point改善 - RoboTwin 2.0で競争力のある性能 - DOMINO Level 1からLevels 2,3へtarget-level adaptationなしで転移

3. 技術・手法の肝は?

- event-aligned visual action reasoning framework - interaction eventsを中心に視覚-行動予測を組織 - event-aligned visual-action supervision - WAMが各imagined rolloutでevent-aligned visual contextを生成 - 重要state changesに重点 - 視覚推論粒度をinteraction dynamicsに応じて形成 - task-critical events周辺で詳細推論 - connecting transitionsで粗い進行 - execution validity headを導入 - 各predicted action sequenceの有効部分を特定 - chunked inferenceでの冗長行動を回避

4. どうやって有効だと検証した?

- DOMINO success rateでbaseline比10.26 percentage point改善 - RoboTwin 2.0で競争力のある性能 - DOMINO Level 1からLevels 2,3へtarget-level adaptationなしで転移 - 詳細な実験設定は要旨からは不明

5. 議論はある?

- 要旨からは不明 - 限界や議論の詳細は記述なし

6. 次に読むべき論文は?

- World-Action Models (WAMs) - DOMINO - RoboTwin 2.0 - 関連手法: event-aligned visual action reasoning, execution validity head

※ AIが要旨から生成した要約です。正確性は原文をご確認ください。

著者: Xiaomeng Yang, Yushu Wu, Yi Gao, Yuhao Lei, Xuan Zhang, Pu Zhao, Yanzhi Wang

分類: cs.CV, cs.RO

原文アブストラクト

World-Action Models (WAMs) utilize future visual prediction as an intermediate reasoning process to guide action generation. However, existing WAMs typically structure visual imagination according to predefined temporal intervals, without explicitly accounting for the different roles of task-critical interactions and connecting transitions. We argue that effective visual foresight should align directly with task-relevant interactions and their corresponding reasoning demands. To this end, we introduce an event-aligned visual action reasoning framework that organizes visual-action prediction around interaction events. Through event-aligned visual-action supervision, WAM learns to generate event-aligned visual context in each imagined rollout, placing greater emphasis on critical state changes that inform action generation. This shapes the visual reasoning granularity according to the underlying interaction dynamics, with detailed reasoning around task-critical events and coarser progression through connecting transitions. Furthermore, we introduce an execution validity head that identifies the valid portion of each predicted action sequence, avoiding redundant actions during chunked inference. Experiments demonstrate a 10.26 percentage point improvement in DOMINO success rate over baseline and competitive performance on RoboTwin 2.0. It also transfers from DOMINO Level 1 to Levels 2 and 3 without target-level adaptation.

関連論文

PR本紙発行元 EmplifAI