日本フィジカルAI新聞

世界のフィジカルAIを、日本語で。

週刊ニュースレター購読
VLAarXiv:2610.03516

XGenAct: クロスタスク生成による幾何学強化型ワールドアクションモデル

XGenAct: Geometry-Enhanced World Action Models through Cross-Task Generation

シェア:XThreadsFacebookLINEはてブBluesky

RGB観測・行動・深度・法線・機能セグメンテーションを統一的なRGB動画表現に変換し、単一の動画拡散トランスフォーマで予測するワールドアクションモデルを提案。RLBenchで52%の成功率を達成。

詳しい要約

1. どんなもの?

- XGenActは、World Action Models (WAMs) の一種で、ロボット制御のための未来予測モデル。 - RGB観測、ロボット行動、metric depth、surface normals、functional role segmentationをRGBビデオとして決定論的コーデックで表現。 - 単一のビデオ拡散transformerと単一の目的関数で、これらの空間を横断した時間予測を学習。 - モダリティ固有の学習済みヘッドを必要としない。

2. 先行研究と比べてどこがすごい?

- 既存のRGBと行動に基づく未来予測は、ロボットマニピュレーションに必要な空間理解を明示的に扱わない。 - 従来は特殊なヘッドやブランチで限られた空間予測タスクを追加し、空間監督の範囲とモデルアーキテクチャが断片化していた。 - XGenActは、複数の空間表現を統一されたビデオ拡散transformerで扱い、単一目的関数で学習する点が新しい。 - 外部比較5タスクで52%の成功率を達成し、最強ベースラインの26%を上回る。

3. 技術・手法の肝は?

- RGB観測、行動、metric depth、surface normals、functional role segmentationを決定論的コーデックでRGBビデオに変換。 - 訓練中に知覚タスクと行動タスクをサンプリングし、単一のビデオ拡散transformerと単一の目的関数で時間予測を学習。 - モダリティ固有の学習済みヘッドを使用しない。 - これにより、空間監督の範囲を広げつつ、アーキテクチャの断片化を回避。

4. どうやって有効だと検証した?

- 保留されたRLBenchタスクで評価。 - 構造化された知覚訓練が、RGBのみの訓練よりも平均閉ループ成功率を改善。 - 5タスクの外部比較でXGenActは52%の成功率を達成し、最強ベースラインの26%を上回る。 - 未来のdepthとsegmentationの予測精度が、RGBを先に生成してから凍結した知覚エキスパートを適用するパイプラインよりも高い。

5. 議論はある?

- 要旨からは、具体的な議論や限界についての記述は不明。 - 構造化知覚訓練の有効性は示されているが、他のタスクや環境への一般化については言及されていない。 - 計算コストやリアルタイム性に関する議論は要旨からは不明。

6. 次に読むべき論文は?

- 要旨で参照されている研究:RLBench、World Action Models (WAMs)、ビデオ拡散transformer、凍結した知覚エキスパートを用いたパイプライン。 - 関連手法:RGBベースの未来予測、空間予測タスクを追加する特殊ヘッドやブランチ。 - 同分野の定番:ロボットマニピュレーションのためのビデオ予測モデル、拡散モデルを用いた行動生成。

※ AIが要旨から生成した要約です。正確性は原文をご確認ください。

著者: Tingting Du, Ziyao Wang, Guoheng Sun, Ang Li

分類: cs.RO, cs.CV, cs.LG

原文アブストラクト

World action models (WAMs) have advanced robot control by predicting how observations and actions evolve over time. Despite this progress, RGB and action based future prediction does not explicitly address the spatial understanding needed for robot manipulation. Existing efforts often add a limited set of spatial prediction tasks through specialized heads or branches, leaving both the range of spatial supervision and the model architecture fragmented. We introduce XGenAct, a world action model that represents RGB observations, robot actions, metric depth, surface normals, and functional role segmentation as RGB videos through deterministic codecs. By sampling perception and action tasks during training, XGenAct uses one video diffusion transformer and one objective to learn temporal prediction across these spaces without modality specific learned heads. On held out RLBench tasks, structured perception training improves average closed loop success over RGB only training, and XGenAct achieves 52% success in the five task external comparison, versus 26% for the strongest evaluated baselines. It also predicts future depth and segmentation more accurately than the evaluated pipelines that generate RGB first and then apply a frozen perception expert.

関連論文

PR本紙発行元 EmplifAI