日本フィジカルAI新聞

世界のフィジカルAIを、日本語で。

週刊ニュースレター購読
マニピュレーションarXiv:2609.37583

RoboHarn-Evo: 階層的物理知識を進化させ自己改善するロボットマニピュレーション

RoboHarn-Evo: Evolving Hierarchical Physical Knowledge for Self-Improving Robotic Manipulation

シェア:XThreadsFacebookLINEはてブBluesky

物理的な試行錯誤からタスク知識と行動知識を階層的に蓄積・修正し、基盤モデルを更新せずにロボット操作の成功率を高めるフレームワークを提案。

詳しい要約

1. どんなもの?

- Vision-language models (VLM) を用いた長期的ロボットマニピュレーションの改善手法 - RoboHarn-Evo という dual-loop harness を提案 - 物理的経験から Hierarchical Physical Knowledge (HPK) を進化させる - HPK は Task Knowledge と Action Knowledge の2階層からなる - ベースモデルを更新せずに反復相互作用で能力向上を目指す

2. 先行研究と比べてどこがすごい?

- 従来の VLM ベース手法は局所的な物理相互作用の成否に依存し、成功推論が不安定 - ベースモデルをファインチューニングせずに物理経験を再利用可能な知識として蓄積 - 異なるエージェントモデル間で平均成功率を最大 24.2 ポイント改善 - 80 回の相互作用ロールアウトで held-out 成功率が GPT-5.5 で 48.3%→75.0%、GPT-6 で 70.0%→88.3% に向上 - 歴史的知識エラーの 83% 以上を解消しつつ、有効知識の 95.8% を保持 - RMBench から RoboDojo への zero-shot 転移で 35.0 および 25.0 ポイントの向上

3. 技術・手法の肝は?

- dual-loop harness により物理経験から HPK を進化 - HPK は2階層の再利用可能知識: - Task Knowledge: どのサブタスクをいつ実行すべきか、いつ完了かを捉える - Action Knowledge: 物体相対の幾何学的戦略とその物理的効果を捉える - 実行時、エージェントは対応する意思決定レベルで知識を検索し、タスク目標の下で現在のシーンに接地 - エピソード間で物理フィードバックを用いて履歴知識を修正し、適用可能性を更新し、再利用可能なエントリを整理

4. どうやって有効だと検証した?

- RMBench での実験 - 異なるエージェントモデルで平均成功率を最大 24.2 ポイント改善 - 80 回の相互作用ロールアウトで held-out 成功率を評価 - 歴史的知識エラーの解消率と有効知識の保持率を測定 - RMBench から RoboDojo への zero-shot 転移を評価

5. 議論はある?

- 物理的相互作用を再利用可能な知識として蓄積することが有効であることを示唆 - ベースモデルを更新せずに性能向上が可能 - 知識の修正と適用可能性の更新メカニズムが重要 - 限界や失敗ケース、計算コスト、スケーラビリティに関する議論は要旨からは不明

6. 次に読むべき論文は?

- 要旨で参照/比較されている研究: 明示的な参照はなし - 関連手法: Vision-language models (VLM) を用いたロボットマニピュレーション、RMBench、RoboDojo - 同分野の定番: 階層的強化学習、知識転移、ロボットマニピュレーションのベンチマーク

※ AIが要旨から生成した要約です。正確性は原文をご確認ください。

著者: Shifeng Bao, Fanding Huang, Yihan Lin, Youhe Feng, Guanlin Li, Chen Zhao, Yang Li, Jiawei He, Cheng Chi, Jing Zhang

分類: cs.RO

原文アブストラクト

Vision-language models can coordinate long-horizon robot manipulation, yet successful task reasoning still depends on whether local physical interactions produce the intended effects. We study how repeated interaction can improve this capability without updating the base model. We introduce RoboHarn-Evo, a dual-loop harness that evolves Hierarchical Physical Knowledge (HPK) from physical experience. HPK couples two levels of reusable knowledge: Task Knowledge captures which subtask should be executed and when it is complete, while Action Knowledge captures object-relative geometric strategies and their physical effects. During execution, the agent retrieves knowledge at the corresponding decision level and grounds it in the current scene under the task goal. Across episodes, physical feedback is used to revise historical knowledge, update its applicability, and organize reusable entries for subsequent retrieval. Experiments on RMBench show that HPK improves average success by up to 24.2 percentage points across different agent models. With 80 interaction rollouts, held-out success rises from 48.3% to 75.0% for GPT-5.5 and from 70.0% to 88.3% for GPT-6. RoboHarn-Evo also resolves over 83% of historical knowledge errors while retaining 95.8% of valid knowledge, and transfers zero-shot from RMBench to RoboDojo with gains of 35.0 and 25.0 percentage points. These results demonstrate that physical interaction can be accumulated into reusable knowledge for improving subsequent manipulation.

関連論文

PR本紙発行元 EmplifAI