日本フィジカルAI新聞

世界のフィジカルAIを、日本語で。

週刊ニュースレター購読
水中ロボティクスarXiv:2609.31211

CoralPlan: 水中ロボット検査のための観測スキル選択と実行

CoralPlan: Observation Skill Selection and Execution for Underwater Robotic Inspection

シェア:XThreadsFacebookLINEはてブBluesky

視覚言語モデルを用いて、サンゴ礁の水中検査においてカメラ画像とタスク文から適切な観測動作(周回・パッチ・調査)を選択・実行するシステムを提案し、シミュレーションと実機で評価した。

詳しい要約

1. どんなもの?

- 水中ロボット検査のための vision-language system『CoralPlan』 - 現在の camera image と episode manifest の task text から observation skill を選択 - 共有 motion interface で orbit, patch, survey を target-relative trajectory として実行 - 残りの plan fields は operator guidance を提供 - 観測完了には target keeping と primitive-specific coverage が必要 - joint success には選択が記録された reference と一致することも必要

2. 先行研究と比べてどこがすごい?

- 従来は対象認識が出発点だったが、本手法は検査タスクに適した viewing motion の選択と実行まで扱う - 構造的に複雑な coral colony に対し、認識だけでなく観測 skill 選択を統合 - 選択と実行を分離せず、共有 motion interface で結びつけた点が新しい - 要旨からは具体的な先行研究との定量比較は不明

3. 技術・手法の肝は?

- vision-language system が camera image と task text から observation skill を選択 - 共有 motion interface が orbit, patch, survey を target-relative trajectory として実行 - episode manifest が task text と reference を提供 - 観測完了判定は target keeping と primitive-specific coverage に基づく - joint success は skill selection が recorded reference と一致することも要求

4. どうやって有効だと検証した?

- 144 simulated episodes で評価 - 36 matched simulation-hardware pairs で評価 - clear-water pool と external target-reference poses を使用 - hardware observation completion は 77.8% - hardware joint success は 63.9% - reference-mismatched completions と、matching skill selection 後の incomplete observations を同定

5. 議論はある?

- observation-skill choice と測定可能な underwater execution outcomes を結びつけた - task-directed acquisition が成功/失敗する箇所を特定 - reference-mismatched completions と incomplete observations の両方が課題として残る - 要旨からは失敗原因の詳細や改善策の議論は不明

6. 次に読むべき論文は?

- 要旨で参照/比較されている研究は明示されていない - 同分野の定番として underwater robotic inspection, vision-language navigation, target-relative trajectory planning, coral reef monitoring に関する研究が候補 - 具体的な論文名は要旨からは不明

※ AIが要旨から生成した要約です。正確性は原文をご確認ください。

著者: Yuer Gao, Yu Zhao, Yi Cai

分類: cs.RO

原文アブストラクト

Underwater robotic inspection depends on acquiring views that reveal task-relevant structure. For a structurally complex coral colony, recognising the target is only the starting point: the robot must select and execute a viewing motion suited to the inspection task. We present CoralPlan, a vision-language system that selects an observation skill from a current camera image and task text supplied by an episode manifest. A shared motion interface executes orbit, patch, or survey as target-relative trajectories; the remaining plan fields provide operator guidance. Observation completion requires target keeping and primitive-specific coverage, while joint success also requires selection to match the recorded reference. We evaluate this interface in 144 simulated episodes and 36 matched simulation-hardware pairs. In a clear-water pool with external target-reference poses, hardware observation completion reaches 77.8% and joint success reaches 63.9%. The experiments identify both reference-mismatched completions and incomplete observations after a matching skill selection. These results connect observation-skill choice to measurable underwater execution outcomes and identify where task-directed acquisition succeeds or fails.

関連論文

PR本紙発行元 EmplifAI