日本フィジカルAI新聞

世界のフィジカルAIを、日本語で。

週刊ニュースレター購読
マニピュレーションarXiv:2609.29091

受動的実行から能動的探索へ:実環境におけるエージェント型身体性マニピュレーション

From Passive Execution to Active Exploration: Agentic Embodied Manipulation in Realistic Environments

シェア:XThreadsFacebookLINEはてブBluesky

ロボットが指示をただ実行するのではなく、計画・知覚・実行の3モジュールを連携させて環境と動的に相互作用し、隠れた対象物を能動的に探索して操作するフレームワークを提案した。

詳しい要約

1. どんなもの?

- エージェントベースの能動的探索フレームワークを提案。 - ロボットが環境と動的に相互作用し、事前定義された指示を実行するだけでなく、能動的に情報を取得。 - 3つの協調モジュール:計画(高レベルタスク推論)、知覚(視覚シーン理解)、実行(低レベル操作)。 - 部分観測下での操作タスクを完了可能。 - 現実的なFind-and-Placeタスクで評価。

2. 先行研究と比べてどこがすごい?

- 既存のエージェントシステムは受動的実行パラダイムが多く、テキスト意味的手がかり、妨害物、初期不可視ターゲットを含む実世界シナリオへの適用が限定的。 - 提案手法は能動的探索により、これらの課題に対処。 - 環境フィードバックに基づく行動適応を実現。

3. 技術・手法の肝は?

- 計画・知覚・実行の3モジュール協調。 - 細粒度の知覚-実行インターリービング戦略:視覚フィードバックとスキル実行を密結合。 - これにより探索ロバスト性を向上。 - 能動的にタスク関連情報を取得し、部分観測下で操作を完了。

4. どうやって有効だと検証した?

- 現実的なFind-and-Placeタスクで評価。 - ターゲットオブジェクトを能動的に発見してから操作する必要がある挑戦的環境で有効性を実証。

5. 議論はある?

- 要旨からは不明。

6. 次に読むべき論文は?

- 要旨で参照/比較されている研究は明示されていない。 - 同分野の関連手法として、embodied manipulation、agentic systems、active exploration、visual scene understanding、planning for manipulationなどが挙げられる。

※ AIが要旨から生成した要約です。正確性は原文をご確認ください。

著者: Shilin Ma, Chubin Zhang, Xulong Bai, Zifeng Gao, Shiyi Zhang, Yansong Tang

分類: cs.RO

原文アブストラクト

Recent advances in agentic systems have substantially enhanced the long-horizon capability of embodied manipulation. However, many existing frameworks still follow a passive execution paradigm, which limits their applicability to real-world scenarios involving textual semantic cues, distractors, and initially invisible targets. To bridge this gap, we propose an agent-based active exploration framework that enables robots to dynamically interact with the environment rather than merely execute predefined instructions. Specifically, our framework consists of three collaborative modules: a planning module for high-level task reasoning, a perception module for visual scene understanding, and an execution module for low-level manipulation. This design allows the robot to actively acquire task-relevant information, adapt its behavior based on environmental feedback, and complete manipulation tasks under partial observability. Furthermore, we introduce a fine-grained perception-execution interleaving strategy, which tightly couples visual feedback with skill execution to improve exploration robustness. We evaluate our method on a realistic Find-and-Place task, demonstrating its effectiveness in challenging environments where target objects must be actively discovered before manipulation.

関連論文

PR本紙発行元 EmplifAI