日本フィジカルAI新聞

世界のフィジカルAIを、日本語で。

週刊ニュースレター購読
sim2realarXiv:2610.06137

人間の動画を物理的に整合したロボット操作データに変換

Robotizing Human Videos with Physically Consistent Interactions

シェア:XThreadsFacebookLINEはてブBluesky

人間の手をロボットハンドに置き換える際の接触形状と遮蔽を物理的に整合させ、ロボット実演と共学習することで操作性能を向上させる手法を提案。

著者: Ching-Lam Cheng, Shengfeng He, Bin Zhu

分類: cs.RO

原文アブストラクト

Human videos offer scalable manipulation data, but the embodiment gap between human hands and robot manipulators limits their direct use. Existing video-editing methods replace hands with rendered robots, yet inaccurate interaction reconstruction and compositing can produce inconsistent grasps and implausible robot-object occlusions. We address these failures from two complementary physical aspects: interaction geometry and scene visibility. First, an interaction-aware contact reconstruction module combines hand-object segmentation with mesh-level contact prediction to recover dense 3D contacts, then converts them into temporally stabilized grasps for parallel-jaw grippers. Second, a depth-aware compositing module uses scene and robot depth to enforce physically consistent robot-object occlusions. The resulting videos preserve the interaction structure of human demonstrations in a robot-compatible form and are co-trained with robot demonstrations. Using identical human videos and robot data, we compare against robot-only training and the original Masquerade pipeline. Across four RoboTwin tasks and two Diffusion Policy visual encoders, our method achieves the highest average success rates, with especially strong gains under out-of-distribution scene variation. Real-world deployment further shows that the proposed co-training approach improves robustness to visual distractors when the task geometry is observable, while performance on depth-sensitive grasps remains limited by the single-camera setup.

関連論文

PR本紙発行元 EmplifAI