日本フィジカルAI新聞

世界のフィジカルAIを、日本語で。

週刊ニュースレター購読
VLAarXiv:2608.26058v1

一つのポリシー、多様な身体:異種具現化操作のためのカメラ中心の統一アクション幾何学事前学習

One Policy, Many Embodiments: Unified Camera-Centric Action Geometry Pre-training for Heterogeneous Embodied Manipulation

シェア:XThreadsFacebookLINEはてブBluesky

異なるロボット形態やカメラ設定をまたいで操作データを統一するため、カメラで観測可能なアンカー動作を共通のアクション表現として用いる新しいフレームワークUCAG-Pを提案。単一のVLAポリシーで複数のベンチマークを高精度に達成した。

詳しい要約

1. どんなもの?

UCAG-Pは、異種のembodied manipulationデータセットを統一的に扱うためのcamera-centricなaction formulationを導入した、視覚言語行動(VLA)ポリシーの事前学習手法。ロボット固有のコマンドではなく、カメラで観測可能なアンカーの動きを画像・カメラ座標で表現し、ロボットアーム、ヒューマノイド、人間の手を共通のaction schemaの異なるembodimentとして扱う。これにより、単一の共有VLAポリシーが転移可能な操作ジオメトリを学習しつつ、embodiment固有の制御可能性を維持する。

2. 先行研究と比べてどこがすごい?

既存手法は、明示的なaction retargeting、human-to-robot video synthesis、データセット固有のadaptation branchなどを用いて異種データの不一致に対処していたが、これらは統一ポリシーの共同学習を妨げていた。UCAG-Pは、構造的に異種データセットを共通の幾何学的action spaceに整列させることで、この問題を解決し、ベンチマーク固有のfine-tuningなしで単一チェックポイントが複数ベンチマークで高い性能を達成している点が優れている。

3. 技術・手法の肝は?

手法の核は、camera-centricな統一action formulationと、geometry-conditioned action translatorの2つ。まず、操作をカメラで観測可能なアンカーの動きとして画像・カメラ座標で表現し、ロボット固有のコマンドを共有ポリシーのターゲットとしない。次に、geometry-conditioned action translatorが、予測されたモーションとターゲットembodimentのキネマティクスを組み合わせて実行可能な制御コマンドを生成する。この分離アーキテクチャにより、共有VLAポリシーは転移可能な操作ジオメトリを学習し、embodiment固有の制御性を保持する。

4. どうやって有効だと検証した?

UCAG-Pは、4.03K時間のロボット・シミュレーションデータと2.34K時間の人間デモンストレーションで訓練され、単一チェックポイントでLIBERO 98.3%、RoboTwin Easy 88.7%、RoboTwin Hard 89.2%、LIBERO-Plusでゼロショット82.0%、RoboCasa GR-1で62.0%を達成。ベンチマーク固有のfine-tuningなしで検証された。

5. 議論はある?

要旨からは、議論の余地や限界についての詳細は不明。ただし、異種データを統一するアプローチの有効性が示される一方で、実世界の多様な環境や未見のembodimentへの一般化、計算コスト、データの質や多様性の影響などが今後の課題として考えられるが、要旨には明記されていない。

6. 次に読むべき論文は?

要旨で参照されている関連研究として、VLAポリシー、action retargeting、human-to-robot video synthesis、データセット固有のadaptation branchに関する論文が挙げられる。具体的には、Vision-Language-Actionモデル(例:RT-2、OpenVLA)、action retargeting手法、人間デモからの模倣学習などが関連する。

※ AIが要旨から生成した要約です。正確性は原文をご確認ください。

著者: Xiaomi Embodied Intelligence Team, University of Macau, :, Shaoqing Xu, Fang Li, Guozhi Zhan, Zhixiang Duan, Yuhan Wang, Yuechen Luo, Shengyin Jiang, Hanbing Li, Zhiying Du, Longlong Wang, Longmei Jiang, Weixiang Liang, Ying Gong, Yong Pan, Ziping Zhao, Zhiyuan Chen, Yangwei You, Kun Ma, Qinyuan Liu, Hangjun Ye, Zhi-xin Yang

分類: cs.RO

原文アブストラクト

Scaling generalist vision-language-action (VLA) policies is severely bottlenecked by the inherent heterogeneity of embodied data, which spans diverse robot morphologies, camera configurations, and low-level action spaces. Existing paradigms typically address this mismatch through explicit action retargeting, human-to-robot video synthesis, or dataset-specific adaptation branches, fundamentally hindering the joint learning of a unified policy. We introduce UCAG-P, a camera-centric unified action formulation that structurally aligns heterogeneous embodied datasets into a shared geometric action space. Rather than treating robot-specific commands as the shared policy target, UCAG-P represents manipulation through camera-observable anchor motion in image and camera-frame coordinates, treating robot arms, humanoids, and human hands as different embodiments of a common action schema. A geometry-conditioned action translator combines predicted motion with target-embodiment kinematics to produce executable controls. The resulting decoupled architecture allows a shared VLA policy to learn transferable manipulation geometry while retaining embodiment-specific controllability. UCAG-P is trained on 4.03K hours of robot and simulation data and 2.34K hours of human demonstrations. A single checkpoint reaches 98.3% on LIBERO, 88.7% and 89.2% on RoboTwin Easy and Hard, 82.0% zero-shot on LIBERO-Plus, and 62.0% on RoboCasa GR-1, without benchmark-specific fine-tuning.

関連論文