XGenAct: クロスタスク生成による幾何学強化型ワールドアクションモデル
XGenAct: Geometry-Enhanced World Action Models through Cross-Task Generation
RGB観測・行動・深度・法線・機能セグメンテーションを統一的なRGB動画表現に変換し、単一の動画拡散トランスフォーマで予測するワールドアクションモデルを提案。RLBenchで52%の成功率を達成。
詳しい要約
1. どんなもの?
2. 先行研究と比べてどこがすごい?
3. 技術・手法の肝は?
4. どうやって有効だと検証した?
5. 議論はある?
6. 次に読むべき論文は?
※ AIが要旨から生成した要約です。正確性は原文をご確認ください。
著者: Tingting Du, Ziyao Wang, Guoheng Sun, Ang Li
分類: cs.RO, cs.CV, cs.LG
原文アブストラクト
World action models (WAMs) have advanced robot control by predicting how observations and actions evolve over time. Despite this progress, RGB and action based future prediction does not explicitly address the spatial understanding needed for robot manipulation. Existing efforts often add a limited set of spatial prediction tasks through specialized heads or branches, leaving both the range of spatial supervision and the model architecture fragmented. We introduce XGenAct, a world action model that represents RGB observations, robot actions, metric depth, surface normals, and functional role segmentation as RGB videos through deterministic codecs. By sampling perception and action tasks during training, XGenAct uses one video diffusion transformer and one objective to learn temporal prediction across these spaces without modality specific learned heads. On held out RLBench tasks, structured perception training improves average closed loop success over RGB only training, and XGenAct achieves 52% success in the five task external comparison, versus 26% for the strongest evaluated baselines. It also predicts future depth and segmentation more accurately than the evaluated pipelines that generate RGB first and then apply a frozen perception expert.