日本フィジカルAI新聞

世界のフィジカルAIを、日本語で。

週刊ニュースレター購読
マニピュレーションarXiv:2610.07961

IronMan: 情報制約付き動画行動学習によるロボットマニピュレーション

IronMan: Information-Constrained Video-Action Learning for Robot Manipulation

シェア:XThreadsFacebookLINEはてブBluesky

情報ボトルネック原理に基づき、行動に無関係な視覚情報を抑制しつつ行動関連の動的手がかりを保持する動画行動学習フレームワークIronManを提案。LIBEROで99.0%、RoboTwinで79.4%の成功率を達成し、分布外シフト下でも高い頑健性を示した。

詳しい要約

1. どんなもの?

- Video Action Models (VAMs) は視覚ダイナミクスと行動生成を統合するロボットマニピュレーション手法。 - しかし視覚表現は行動生成に適さず、過剰な視覚詳細が汎化を損なう問題がある。 - 本研究は Information Bottleneck 原理に基づく IronMan を提案。 - 不要な視覚情報を抑制し、行動に関連するダイナミクス手がかりを保持する。 - シミュレーションと実世界実験で ID 性能と OOD ロバスト性を検証。

2. 先行研究と比べてどこがすごい?

- 従来の VAMs は視覚詳細をそのまま行動ポリシーに露出し、汎化が劣化。 - IronMan は情報制約により無関係な視覚情報を抑制し、行動関連のダイナミクスを保持。 - LIBERO で 99.0%、RoboTwin clean2clean で 79.4% の成功率を達成し、評価した全ベースラインを上回る。 - OOD シフト下の LIBERO-Plus で 79.1% を達成し、最強ベースラインを 10.4 ポイント上回る。 - 効率的な推論を維持しつつ ID 性能と OOD ロバスト性を両立。

3. 技術・手法の肝は?

- Information Bottleneck 原理に基づく video-action 学習フレームワーク。 - dynamics-aware bottleneck を採用。 - ノイズが多く絡み合った one-step video features をコンパクトな world representations に蒸留。 - 無関係な視覚情報を抑制し、行動に関連するダイナミクス手がかりを保持。 - これにより行動ポリシーの汎化能力を向上。

4. どうやって有効だと検証した?

- 広範なシミュレーションと実世界実験を実施。 - LIBERO で成功率 99.0%、RoboTwin clean2clean で 79.4% を達成。 - 評価した全ベースラインを上回る。 - OOD シフト下の LIBERO-Plus で成功率 79.1% を達成。 - 最強ベースラインを 10.4 パーセントポイント上回る。 - 効率的な推論を維持。

5. 議論はある?

- 要旨からは不明。 - 情報制約の強さや設計選択に関する議論は要旨に記載なし。 - 実世界実験の詳細や失敗事例の分析も要旨からは不明。

6. 次に読むべき論文は?

- Video Action Models (VAMs) に関する研究。 - Information Bottleneck 原理をロボット学習に適用した研究。 - LIBERO、RoboTwin、LIBERO-Plus ベンチマークを用いた研究。 - 視覚表現と行動生成の結合に関する研究。 - 要旨で具体的な参照論文は明示されていないため、同分野の定番として VAMs や Information Bottleneck 関連の論文を挙げる。

※ AIが要旨から生成した要約です。正確性は原文をご確認ください。

著者: Yuanshuo Zhang, Wenzhe Zhao, Zixing Lei, Bin Chen, Siheng Chen

分類: cs.RO

原文アブストラクト

Video Action Models (VAMs) couple visual dynamics modeling with action generation for robot manipulation. However, video representations are not naturally suited to action generation, as exposing the action policy to excessive visual detail can impair its generalization ability. Therefore, we introduce IronMan (Information-constRained videO-actioN learning for robot MANipulation), a robust video-action learning framework built on the information bottleneck principle. The core principle of this framework is to impose information constraints that suppress irrelevant visual information while preserving action-relevant dynamics cues. IronMan employs a dynamics-aware bottleneck that distills noisy, entangled one-step video features into compact world representations. Extensive simulation and real-world experiments demonstrate strong in-distribution (ID) performance and out-of-distribution (OOD) robustness while maintaining efficient inference. IronMan achieves success rates of 99.0% on LIBERO and 79.4% on RoboTwin clean2clean, outperforming all the evaluated baselines. Under OOD shifts, IronMan achieves a success rate of 79.1% on LIBERO-Plus, exceeding the strongest baseline by 10.4 percentage points. Project page: https://youngsoul0731.github.io/ironman-project-page/

関連論文

PR本紙発行元 EmplifAI