日本フィジカルAI新聞

世界のフィジカルAIを、日本語で。

週刊ニュースレター購読
物体探索arXiv:2609.20330

RoboFind: 視覚障害者のためのマルチエージェント個人向け物体探索

RoboFind: Multi-Agent Personalized Object Search for People Who Are Blind or Have Low Vision

シェア:XThreadsFacebookLINEはてブBluesky

スマホで対象物を教示し、四足ロボットが探索・照合するマルチエージェントシステムを構築し、実機実験で高い成功率と誤検出の低減を示した。

詳しい要約

1. どんなもの?

- 視覚障害・低視力のユーザーが『自分の特定の持ち物』を探すためのマルチエージェント型ロボット探索フレームワーク - スマートフォンで対象を教示し、quadruped robot が探索を実行 - Target Teaching Agent、Navigation Agent、Verification Agent、Coordination and Recovery Agent で構成 - 32回の実機ミッションで評価

2. 先行研究と比べてどこがすごい?

- 再構成した sequential first-stop baseline は20試行・10対象で成功率25.0%、RoboFind は85.0% - false success を75.0%から5.0%へ削減 - 6つの共有対象で10/12成功、独立実行の GPT-6 Astra-only 12試行は5/12 - 完了宣言前に対象同一性を検証する点が信頼性を高める

3. 技術・手法の肝は?

- Target Teaching Agent が guided smartphone recordings を semantic target profile と再利用可能な multi-view reference bank に変換 - AR guidance、speech、haptic feedback、screen-reader support を備えた accessible capture flow - 実行時は Navigation Agent が環境を探索し候補を提案 - Verification Agent が候補を保存参照と照合 - Coordination and Recovery Agent がミッション完了または recovery と探索継続を担当

4. どうやって有効だと検証した?

- 32回の real-robot missions で評価 - sequential first-stop baseline との比較(20試行、10対象) - 6つの shared targets で10/12成功 - 12回の independently executed GPT-6 Astra-only trials と比較し5/12 - false success 率の比較も実施

5. 議論はある?

- マルチエージェント設計が personalized object search の要求に適合することを示す - 完了宣言前の object identity 検証がユーザーが信頼できる結果につながる - 限界や失敗事例、ユーザビリティの詳細は要旨からは不明

6. 次に読むべき論文は?

- 要旨で参照されている sequential first-stop baseline - GPT-6 Astra-only - 同分野の定番として personalized object search、semantic target profile、multi-view reference bank、quadruped robot navigation に関する研究

※ AIが要旨から生成した要約です。正確性は原文をご確認ください。

著者: Ruiping Liu, Shaofang Quan, Qian Yin, Jingqi Zhang, Junwei Zheng, Yufan Chen, Di Wen, Weijia Fan, Kailun Yang, M. Saquib Sarfraz, Tamim Asfour, Kunyu Peng, Rainer Stiefelhagen

分類: cs.RO

原文アブストラクト

Blind and low-vision users often need to locate a specific personal object rather than an arbitrary instance of the same category. The task calls for a robot that can move through the space and reach viewpoints the user cannot, and for an accessible interface where the user says which object is meant and learns whether the right one was found. We present RoboFind, a multi-agent framework in which a smartphone teaches the target and a quadruped robot carries out the search. A Target Teaching Agent converts guided smartphone recordings into a semantic target profile and a reusable multi-view reference bank through an accessible capture flow with AR guidance, speech and haptic feedback, and screen-reader support, so later missions refer to a stored object without repeating the teaching process. At runtime, a Navigation Agent explores the environment and proposes candidate targets, a Verification Agent checks each candidate against the stored references, and a Coordination and Recovery Agent completes the mission or triggers recovery and continued search. Across 32 real-robot missions, RoboFind reaches 85.0% success against 25.0% for a reconstructed sequential first-stop baseline over 20 trials with ten targets, and reduces false success from 75.0% to 5.0%. On six shared targets it succeeds in 10/12 trials, against 5/12 for 12 independently executed GPT-6 Astra-only trials. These results show that the multi-agent design fits the demands of personalized object search, where verifying object identity before declaring completion is what makes the outcome something a user can rely on.

関連論文

PR本紙発行元 EmplifAI