日本フィジカルAI新聞

世界のフィジカルAIを、日本語で。

週刊ニュースレター購読
動画生成arXiv:2610.09454

RobotAPO: ロボットマニピュレーション動画生成のための敵対的物理選好最適化

RobotAPO: Adversarial Physics Preference Optimization for Robotic Manipulation Video Generation

シェア:XThreadsFacebookLINEはてブBluesky

ロボット操作動画生成において、物理違反を選好データセットで学習し、敵対的選好最適化で物理整合性を高める手法を提案。下流のロボット実行精度が向上。

詳しい要約

1. どんなもの?

- ロボット操作動画生成のための手法 - 物理的に妥当な操作動画を生成する - 生成動画をembodied agentのvisual planとして利用 - 課題 - 視覚的もっともらしさだけでは物理的整合性が不十分 - 相互作用境界での物理違反が下流実行を無効化 - 提案 - AgiBot-PhysPref: 物理違反を分離した10,000サンプルのpreference dataset - RobotAPO: adversarial physics preference optimizationフレームワーク

2. 先行研究と比べてどこがすごい?

- 従来のsupervised fine-tuning - 局所的な物理違反を直接罰する圧力が欠如 - 提案手法 - 生成器自身のfailure distributionを物理整合性の信号として活用 - 静的なcurated failureの丸暗記を防ぐ - 外部構造条件付け不要でprompt-and-reference推論を維持 - 性能 - 物理整合性: hard score 6.8%、soft score 10.0%改善 - 実機リプレイ: タスク成功率37.4%相対改善

3. 技術・手法の肝は?

- RobotAPO - continuous flow-matching denoising spaceで動作 - adversarial physics preference optimization - 構成要素 - lightweight adversarial counterfactual proposer - 条件依存の物理失敗バイアス方向をdenoising spaceで学習 - モデルに物理相互作用境界の探索と尊重を促す - データ - AgiBot-PhysPref: 条件一致した物理違反を分離したpreference dataset

4. どうやって有効だと検証した?

- 包括的評価を実施 - held-out AgiBot条件で評価 - 物理整合性: hard score 6.8%、soft score 10.0%改善 - 実機リプレイ: タスク成功率37.4%相対改善 - 比較対象 - 最強のcontrolled internal baseline

5. 議論はある?

- 局所的な物理違反の明示的修正が下流ロボット実行を改善 - 物理整合性の向上が実機タスク成功率に翻訳されることを確認 - 限界や議論の詳細は要旨からは不明

6. 次に読むべき論文は?

- AgiBot-PhysPref (提案データセット) - RobotAPO (提案手法) - flow-matchingに基づく動画生成手法 - preference optimizationを用いた生成モデル改善 - ロボット操作動画生成の先行研究

※ AIが要旨から生成した要約です。正確性は原文をご確認ください。

著者: Kerui Li, Zhe Jing, Chenyi Huang, Xiaofeng Wang, Zheng Zhu, Haoming Cui, Huaibo Huang

分類: cs.RO

原文アブストラクト

Robotic manipulation videos are increasingly used as visual plans for embodied agents, but optimizing purely for visual plausibility often fails to capture the fragile physical manifold of real-world interactions. Even minor physics-violating errors at the interaction boundary, such as interpenetration or premature object motion, can completely invalidate the inferred timing and pose needed for downstream execution. Because standard supervised fine-tuning lacks the direct pressure to penalize these localized failures, we introduce AgiBot-PhysPref. This rigorously curated 10,000-sample preference dataset isolates condition-matched physics violations, turning the generator's own failure distribution into a foundational signal for physical consistency. Building upon this, we propose RobotAPO, an adversarial physics preference optimization framework operating in the continuous flow-matching denoising space. To prevent the policy from merely memorizing static curated failures, RobotAPO employs a lightweight adversarial counterfactual proposer that learns a condition-dependent, physical-failure-biased direction in denoising space. This encourages the model to explore and better respect the physical interaction boundary, all while maintaining a pure prompt-and-reference inference interface without requiring external structural conditioning. Comprehensive evaluations demonstrate that explicitly correcting these localized physics violations improves downstream robot execution from generated videos. On held-out AgiBot conditions, RobotAPO outperforms the strongest controlled internal baseline in physical consistency by 6.8% hard score and 10.0% soft score. Crucially, in real-robot replay, it translates these physical-consistency gains into a 37.4% relative improvement in task success over the strongest controlled internal baseline.

関連論文

PR本紙発行元 EmplifAI