日本フィジカルAI新聞

世界のフィジカルAIを、日本語で。

週刊ニュースレター購読
人間ロボット協働arXiv:2609.24055

人間参加型ロボット失敗回復に向けて:人間とロボットの協働におけるコミュニケーションギャップの橋渡し

Toward Human-in-the-Loop Robot Failure Recovery: Bridging Communication Gaps in Human-Robot Collaboration

シェア:XThreadsFacebookLINEはてブBluesky

ロボットが失敗から回復するために周囲の人に助けを求める際、聞き手の知識差を考慮したコミュニケーションの重要性を示し、その差を評価するゲーム・データセット・ベンチマークLD-HRIを提案した。

詳しい要約

1. どんなもの?

- ロボットが失敗から回復するために周囲の人に助けを求める際の、人間とロボットの協働におけるコミュニケーションギャップを扱う研究。 - Listener Differences in Human-Robot Interaction (LD-HRI) というゲーム、データセット、ベンチマークを導入。 - 話者の性能を人間の聞き手のパフォーマンスで評価する。 - 聞き手の情報の違いを制御した下で、要求の特性、LLM話者、inverse-semantics要求選択アルゴリズムを評価。 - コーパスは446の人間作成要求と1,302の聞き手試行を含む。 - さらに24の凍結LLM作成要求を70人の聞き手で560試行評価。

2. 先行研究と比べてどこがすごい?

- 先行のinverse-semantics研究は単一の聞き手モデルで要求を生成し、聞き手の知識の違いを検証していなかった。 - 本研究は聞き手の知識の違いを制御して評価する点が新しい。 - 人間の聞き手のパフォーマンスを通じて話者を評価するベンチマークを提供。 - これにより、専門家と初心者の間のギャップを測定可能にする。

3. 技術・手法の肝は?

- LD-HRIゲーム、データセット、ベンチマークを設計。 - 聞き手の情報の違いを制御した実験設定。 - 要求の特性、LLM話者、inverse-semantics要求選択アルゴリズムを評価。 - 人間作成要求とLLM作成要求を比較。 - 聞き手の成功を指標として話者の性能を評価。

4. どうやって有効だと検証した?

- 446の人間作成要求と1,302の聞き手試行を含むコーパスで評価。 - 24の凍結LLM作成要求を70人の聞き手で560試行評価。 - 初心者の成功率はモデル作成要求の方が全4タスクで記述的に高い。 - しかし両要求源とも専門家と初心者の間に大きなギャップが残り、LLM要求では16パーセントポイントのギャップ。

5. 議論はある?

- 聞き手の知識の違いがコミュニケーションの有効性に影響することを示す。 - モデル作成要求でも専門家と初心者のギャップが大きい。 - LD-HRIはこれらのギャップを測定可能にし、より堅牢なコミュニケーション設計の基盤を提供。 - 具体的な議論の詳細は要旨からは不明。

6. 次に読むべき論文は?

- 要旨で参照/比較されている研究: inverse-semantics work。 - 関連手法: large language model (LLM) speakers, inverse-semantics request-selection algorithms。 - 同分野の定番: human-robot collaboration, human-in-the-loop robot failure recovery。

※ AIが要旨から生成した要約です。正確性は原文をご確認ください。

著者: Promise Ekpo, Teju Vijay, Dhruv Mandalik, Tisha Jain, Arman Ibrayeva, Sunishka Sil, Stefanie A. Tellex, Angelique Taylor

分類: cs.RO, cs.HC

原文アブストラクト

Robots can recover from failures by asking bystanders for help, but effective human-in-the-loop recovery requires communication that accounts for differences in people's knowledge. Prior inverse-semantics work generates requests using a single listener model, leaving differences in listener knowledge untested. We introduce Listener Differences in Human-Robot Interaction (LD-HRI), a game, dataset, and benchmark that evaluates speakers through human listener performance. Our evaluation examines request properties, large language model (LLM) speakers, and inverse-semantics request-selection algorithms under controlled differences in listener information. The corpus contains 446 human-written requests and 1{,}302 listener trials. We additionally evaluated 24 frozen LLM-written requests with 70 human listeners across 560 trials. Novice success is descriptively higher with model-written requests across all four tasks, yet both request sources leave substantial expert--novice gaps, including 16 percentage points for LLM requests. LD-HRI makes these gaps measurable, providing a foundation for designing more robust communication in human-robot and human-agent interaction.

関連論文

PR本紙発行元 EmplifAI