日本フィジカルAI新聞

世界のフィジカルAIを、日本語で。

週刊ニュースレター購読
マニピュレーションarXiv:2609.33104

VPTwin: 実機-シミュレーション-実機のビデオ予測によるロボットマニピュレーション計画

VPTwin: Real-Sim-Real Video Prediction for Robotic Manipulation Planning

シェア:XThreadsFacebookLINEはてブBluesky

実演からデジタルツインを再構築し、Isaac Simの複数ロールアウトを参照して実世界の未来映像を予測するReal-Sim-Real枠組みを提案し、物理的に妥当な予測でマニピュレーション計画を改善した。

著者: Zhenghao Xiao, Minting Pan, Nantian He, Dongzhan Zhou, Yunbo Wang

分類: cs.RO, cs.AI

原文アブストラクト

While action-conditioned video prediction provides an intuitive world model for robotics, purely data-driven predictors often suffer from compounding errors and physically implausible hallucinations in long-horizon rollouts, severely undermining downstream action planning. We propose VPTwin, a Real-Sim-Real video prediction framework that anchors real-world future prediction using real-synchronized simulation twins. For a target manipulation task, a VLM reconstructs an executable digital twin from a real demonstration episode. To accommodate the ill-posed estimation of unobserved physical properties, Isaac Sim simulates multiple forward dynamic rollouts across randomized physical configurations under candidate action trajectories. Using these rollouts as in-context references, VPTwin harmonizes both domains, using simulation dynamics to enforce physical plausibility while capturing unmodeled contact interactions from real video. Furthermore, we establish a predictive planning loop using VPTwin to visually verify VLM-proposed actions and guide reliable real-world execution. Evaluations show substantial reductions in physical hallucinations during video prediction and marked improvements in manipulation planning performance.

関連論文

PR本紙発行元 EmplifAI