日本フィジカルAI新聞

世界のフィジカルAIを、日本語で。

週刊ニュースレター購読
行動予測arXiv:2507.14809

軽量未来予測:InstructPix2Pixによるマルチモーダル行動フレーム予測

Light Future: Multimodal Action Frame Prediction via InstructPix2Pix

シェア:XThreadsFacebookLINEはてブBluesky

InstructPix2Pixを微調整し、現在の画像とテキスト指示からロボットが10秒後に見る視覚フレームを予測する軽量なマルチモーダル手法を提案。

著者: Zesen Zhong, Duomin Zhang, Yijia Li

分類: cs.CV, cs.MM, cs.RO

原文アブストラクト

Predicting future motion trajectories is a critical capability across domains such as robotics, autonomous systems, and human activity forecasting, enabling safer and more intelligent decision-making. This paper proposes a novel, efficient, and lightweight approach for robot action prediction, offering significantly reduced computational cost and inference latency compared to conventional video prediction models. Importantly, it pioneers the adaptation of the InstructPix2Pix model for forecasting future visual frames in robotic tasks, extending its utility beyond static image editing. We implement a deep learning-based visual prediction framework that forecasts what a robot will observe 100 frames (10 seconds) into the future, given a current image and a textual instruction. We repurpose and fine-tune the InstructPix2Pix model to accept both visual and textual inputs, enabling multimodal future frame prediction. Experiments on the RoboTWin dataset (generated based on real-world scenarios) demonstrate that our method achieves superior SSIM and PSNR compared to state-of-the-art baselines in robot action prediction tasks. Unlike conventional video prediction models that require multiple input frames, heavy computation, and slow inference latency, our approach only needs a single image and a text prompt as input. This lightweight design enables faster inference, reduced GPU demands, and flexible multimodal control, particularly valuable for applications like robotics and sports motion trajectory analytics, where motion trajectory precision is prioritized over visual fidelity.

関連論文

PR本紙発行元 EmplifAI