IronMan: 情報制約付き動画行動学習によるロボットマニピュレーション
IronMan: Information-Constrained Video-Action Learning for Robot Manipulation
情報ボトルネック原理に基づき、行動に無関係な視覚情報を抑制しつつ行動関連の動的手がかりを保持する動画行動学習フレームワークIronManを提案。LIBEROで99.0%、RoboTwinで79.4%の成功率を達成し、分布外シフト下でも高い頑健性を示した。
詳しい要約
1. どんなもの?
2. 先行研究と比べてどこがすごい?
3. 技術・手法の肝は?
4. どうやって有効だと検証した?
5. 議論はある?
6. 次に読むべき論文は?
※ AIが要旨から生成した要約です。正確性は原文をご確認ください。
著者: Yuanshuo Zhang, Wenzhe Zhao, Zixing Lei, Bin Chen, Siheng Chen
分類: cs.RO
原文アブストラクト
Video Action Models (VAMs) couple visual dynamics modeling with action generation for robot manipulation. However, video representations are not naturally suited to action generation, as exposing the action policy to excessive visual detail can impair its generalization ability. Therefore, we introduce IronMan (Information-constRained videO-actioN learning for robot MANipulation), a robust video-action learning framework built on the information bottleneck principle. The core principle of this framework is to impose information constraints that suppress irrelevant visual information while preserving action-relevant dynamics cues. IronMan employs a dynamics-aware bottleneck that distills noisy, entangled one-step video features into compact world representations. Extensive simulation and real-world experiments demonstrate strong in-distribution (ID) performance and out-of-distribution (OOD) robustness while maintaining efficient inference. IronMan achieves success rates of 99.0% on LIBERO and 79.4% on RoboTwin clean2clean, outperforming all the evaluated baselines. Under OOD shifts, IronMan achieves a success rate of 79.1% on LIBERO-Plus, exceeding the strongest baseline by 10.4 percentage points. Project page: https://youngsoul0731.github.io/ironman-project-page/
関連論文
- 密度関数を用いた安全なマルチロボット協調搬送マニピュレーション
- コンパクトなロボットポリシーに必要なのは細粒度の視覚表現マニピュレーション
- ずれた座標系を見抜く:視覚・力覚精密組立のための特権ノイズ蒸留マニピュレーション
- 把持後における物体再配向によるロボット挿入の運動学的修復マニピュレーション
- 未来が一致する間だけコミット:ロボットマニピュレーションのための結果認識型適応アクションチャンキングマニピュレーション
- TacZero: 触覚フィードバックと汎用視覚言語モデルによる訓練不要のペグ挿入マニピュレーション