日本フィジカルAI新聞

世界のフィジカルAIを、日本語で。

週刊ニュースレター購読
マニピュレーションarXiv:2610.02840

PointWAM: 3D世界行動モデリングによる器用なロボットマニピュレーション

PointWAM: 3D World Action Modeling for Dexterous Robotic Manipulation

シェア:XThreadsFacebookLINEはてブBluesky

世界をシーンと手に分解し、両者を3D点軌跡として予測することで、人間の動画から事前学習し器用な操作を実現する手法を提案。

詳しい要約

1. どんなもの?

- 3D world action model「PointWAM」を提案 - 世界をscene(環境)とhands(actor)に分解 - 両者を共有時空間座標系の3D point trajectoriesとして予測 - 色付きpoint cloudと言語指示からsceneとhandsの3D共進化を予測 - 予測したhand motionをrobot actionsにretarget - dexterous robotic manipulationを対象

2. 先行研究と比べてどこがすごい?

- 既存手法は世界をRGB framesやlatentで表現 - 行動はend-effector posesやjoint anglesで予測 - 3D spatial structureやcontact geometryの捕捉が苦手 - PointWAMは3D point trajectoriesで明示的・分離表現 - 大規模human demonstration videosでtask-specificなobject/keypoint選択なしに事前学習可能 - DexJoCoでprior SOTAを10タスクで11.7ポイント上回る

3. 技術・手法の肝は?

- 世界をsceneとhandsに分解 - 共有space-time coordinate frame内で3D point trajectoriesをjointly forecast - 入力はcolored point cloudと言語指示 - sceneとhandsの3D空間での共進化を予測 - 予測hand motionをrobot actionsへretarget - human videosでのpre-trainingとscene-trajectory supervisionを活用

4. どうやって有効だと検証した?

- human videosでのpre-trainingがDexJoCo平均成功率を56.9ポイント改善 - scene-trajectory supervisionがhandsのみ予測より10.9ポイント追加 - 両者併用でDexJoCo 10タスクにおいてprior state of the artを11.7ポイント上回る - 実機ロボットで強いVLAsを上回る

5. 議論はある?

- 要旨からは不明

6. 次に読むべき論文は?

- DexJoCo - VLAs - RGB framesやlatent表現を用いる既存world action models - end-effector posesやjoint anglesを予測する既存手法

※ AIが要旨から生成した要約です。正確性は原文をご確認ください。

著者: Chunghyun Park, Beomjun Kim, Seungcheol Park, Heeseung Kwon, Yashu Shukla, Seunghoon Sim, Jinwoo Shin, Minsu Cho

分類: cs.RO, cs.CV

原文アブストラクト

World action models jointly learn to forecast world dynamics and predict robot actions, such that the learned internal world dynamics guide accurate actions. Existing approaches typically represent the world as RGB frames or latent counterparts while predicting actions as end-effector poses or joint angles, but they often struggle to capture the 3D spatial structure and contact geometry central to dexterous manipulation. We introduce Point World Action Model (PointWAM), a 3D world action model that decomposes the world into a scene (i.e., environment) and hands (i.e., actor), and jointly forecasts both as 3D point trajectories within a shared space-time coordinate frame. This explicit, disentangled representation enables effective pre-training on large-scale human demonstration videos without requiring any task-specific object or keypoint selection. Given a colored point cloud and a language instruction, PointWAM predicts how the scene and hands co-evolve in 3D space over time, then retargets the forecast hand motion to robot actions. Pre-training on human videos improves average DexJoCo success by 56.9 percentage points, and scene-trajectory supervision adds 10.9 points over forecasting the hands alone. With both, PointWAM surpasses the prior state of the art on ten DexJoCo tasks by 11.7 points and outperforms strong VLAs on a real robot.

関連論文

PR本紙発行元 EmplifAI