日本フィジカルAI新聞

世界のフィジカルAIを、日本語で。

週刊ニュースレター購読
VLAarXiv:2609.40153

Dream4ACT: 複数身体に対応した動画・行動モデリングのための共有視覚行動インターフェース

Dream4ACT: A Shared Visual Action Interface for Multi-Embodiment Video-Action Modeling

シェア:XThreadsFacebookLINEはてブBluesky

関節構成を4つの仮想カメラ映像として描画する「action views」で身体差を吸収し、単一の動画拡散モデルで順動力学・逆動力学・観測行動生成を統合した世界モデルを提案。

詳しい要約

1. どんなもの?

- 複数embodimentにまたがるvideo-action modelingのためのworld model「Dream4ACT」を提案。 - 行動を「action views」と呼ぶ共有visual action interfaceで表現。 - URDF-based forward kinematicsで4つの仮想カメラからtarget joint configurationsを描画。 - 観測と行動が同一のvideo autoencoderとdiffusion transformerを共有可能。 - masked flow-matchingによりforward dynamics、inverse dynamics、joint observation-action generationを単一モデルで実現。

2. 先行研究と比べてどこがすごい?

- 従来のjoint-space action vectorsはimage-space構造を持たず、embodiment間で次元・意味が異なるためVGMの時空間priorを直接活用しにくい。 - end-effector visualizationsは代替になるが、ロボット実行に必要な全関節構成を指定できない。 - 提案のaction viewsはembodiment固有のarticulated geometryを保ちつつ、観測と行動でvideo autoencoderとdiffusion transformerを共有できる点が新しい。 - 学習済みembodiment固有decoderを必要とせず、training-freeなURDF-constrained multiview recoveryで実行可能行動列を復元。

3. 技術・手法の肝は?

- 共有visual action interface「action views」を導入。 - URDF-based forward kinematicsによりtarget joint configurationsを4つの所定仮想カメラからレンダリング。 - 観測と行動の系列が同一のvideo autoencoderとdiffusion transformerを共有。 - masked flow-matchingを採用し、未来系列のどの部分をcorruptするかを変えることでforward dynamics、inverse dynamics、joint observation-action generationを単一のjointly trained modelでサポート。 - 予測されたaction viewsから実行可能な行動列を復元するため、学習不要のURDF-constrained multiview recovery mechanismを提案。

4. どうやって有効だと検証した?

- RoboTwin 2.0で平均成功率88.98%を達成。 - TriWorldBenchで総合スコア65.66を達成。 - 閉ループmanipulationと、action-conditioned multiview predictionが有効であることを示す。 - 具体的な実験設定・比較対象・アブレーションの詳細は要旨からは不明。

5. 議論はある?

- action viewsがembodiment固有のarticulated geometryを保持しつつembodiment間の行動表現を統一できる点を主張。 - 学習済みembodiment固有decoderなしで実行可能行動列を復元できる点を強調。 - 限界・失敗事例・計算コスト・スケーラビリティ・一般化範囲についての議論は要旨からは不明。

6. 次に読むべき論文は?

- 要旨で参照・比較されている個別研究は明示されていない。 - 関連手法としてVideo Generation Models (VGMs)、masked flow-matching、URDF-based forward kinematics、diffusion transformer、RoboTwin 2.0、TriWorldBenchが挙げられる。 - 同分野の定番としてembodied video-action modeling、world models、robot manipulation benchmarksに関する研究を読むとよい。

※ AIが要旨から生成した要約です。正確性は原文をご確認ください。

著者: Xiangyu Zhu, Jin Xu, Yue Guo, Xin Wu, Yifan Sun, Xiancong Ren, Jianxin Sun, Yong Dai, Xiaozhu Ju

分類: cs.RO

原文アブストラクト

Video generation models (VGMs) offer strong spatiotemporal priors for embodied observation--action modeling. However, joint-space action vectors lack explicit image-space structure and vary in dimensionality and semantics across embodiments, making it challenging to directly leverage the rich spatiotemporal priors of VGMs. End-effector visualizations provide an alternative but do not specify the full articulated configuration needed for robot execution. We present Dream4ACT, a world model built for joint video-action modeling across embodiments. To unify action representations across embodiments, we introduce a shared visual action interface, called action views, which render target joint configurations from four prescribed virtual cameras using URDF-based forward kinematics. This shared visual representation preserves embodiment-specific articulated geometry while allowing observation and action sequences to share a video autoencoder and diffusion transformer. Through masked flow-matching, our model supports forward dynamics, inverse dynamics, and joint observation--action generation within a single jointly trained model by varying which future sequences are corrupted. To recover executable action sequences from predicted action views, we propose a training-free, URDF-constrained multiview recovery mechanism, without a learned embodiment-specific decoder. Dream4ACT achieves an average success rate of 88.98\% on RoboTwin~2.0 and an overall score of 65.66 on TriWorldBench, supporting effective closed-loop manipulation and competitive action-conditioned multiview prediction through the visual action interface.

関連論文

PR本紙発行元 EmplifAI