日本フィジカルAI新聞

世界のフィジカルAIを、日本語で。

週刊ニュースレター購読
ヒューマンロボット対話arXiv:2609.21942

失敗したロボットはいつ人間に尋ねるべきか:監査済みセンサ証拠に基づく修正対話の開始

When Should a Failing Robot Ask? Initiating Corrective Human-Robot Dialogue from Audited Sensor Evidence

シェア:XThreadsFacebookLINEはてブBluesky

失敗原因の診断精度をセンサごとに測定し、自己診断・追加センサ確認・人間への質問という3択の最適方針を導出。視覚言語モデルは証拠ではなくプロンプト表面に従うことを示した。

詳しい要約

1. どんなもの?

ロボットがタスク失敗時に、自己診断で行動するか、別のオンボードセンサを参照するか、人間に問い合わせるかを選ぶ問題を扱う。失敗原因を注入して真値が既知のシミュレーションベンチマークを構築し、各センサが原因をどれだけ明らかにするか、データリークを明示的に検査して測定する。カメラ画像で診断可能な失敗と、力データでのみ診断可能な失敗がある(力データで0.99、画像手法は0.55超なし)。6つのopen vision-languageモデルをテストし、その挙動は証拠ではなくプロンプトの表面に従うことを示す。

2. 先行研究と比べてどこがすごい?

先行研究との具体的比較は要旨からは不明。ただし、失敗原因を注入して真値が既知のシミュレーションベンチマークを構築し、データリークを明示的に検査する点、センサごとの診断可能性を定量化する点、open vision-languageモデルの挙動をプロンプト変種や信頼度とともに体系的に評価する点が特徴として示される。

3. 技術・手法の肝は?

失敗原因を注入したシミュレーションベンチマークで、各センサが原因をどれだけ明らかにするかをデータリーク検査付きで測定。6つのopen vision-languageモデルを、拒否オプションの位置変更、worked examplesの有無、信頼度表明などのプロンプト変種で評価。力データを10行のテキストとして与える条件も比較。act・consult own sensors・ask a humanの3行動決定問題として定式化し、測定精度から最適方策を導出。

4. どうやって有効だと検証した?

力データでは0.99、画像手法では0.55超なしという診断可能性を測定。6モデル中3つのモデル・ファミリーペアで、拒否オプションを最後から最初に移すと拒否率が78-100%から0-6%に崩壊。フレームからの精度は全プロンプト変種で多数クラスベースライン以下、worked examplesの有無を問わず、表明信頼度は正しさの情報を持たない。力データを10行テキストで与えると6モデル中4つで初のベースライン超え診断。人間への1質問でベースラインから回答者自身の信頼性程度(質問時0.70-0.81)に上昇。

5. 議論はある?

モデルの挙動は証拠ではなくプロンプトの表面に従い、拒否率がオプション位置で崩壊する。精度は多数クラスベースライン以下で、信頼度は正しさと無関係。失敗の多くは能力不足ではなくセンサデータ欠如を反映する。モデルは測定精度から導かれる最適方策に従わず、質問コストの4倍変化を無視する。質問する決定はモデルの信頼度ではなく、測定精度と明示的コストに結び付けるべきと議論。

6. 次に読むべき論文は?

要旨で参照・比較されている具体的な先行研究は明示されていない。関連手法としてopen vision-languageモデル、vision-language-actionモデル、human-robot dialogue、active perception、sensor fusion、calibration、decision-theoretic askingなどが挙げられるが、要旨からは不明。

※ AIが要旨から生成した要約です。正確性は原文をご確認ください。

著者: Eshika Pathak, Leela Krishna

分類: cs.RO, cs.AI, cs.HC

原文アブストラクト

A robot that fails at a task faces the first decision in corrective dialogue: act on its own diagnosis, consult another onboard sensor, or interrupt a person. Choosing well requires knowing how much the robot's sensors reveal about the cause and how reliable the robot's own diagnosis is. We build a simulated benchmark in which every failure's true cause is known, because we injected it, and measure what each sensor reveals, with explicit checks against data leakage. Some failures are diagnosable from camera images; others only from the robot's force data (0.99 from force data, no image method above 0.55). We then test six open vision-language models. Their behavior tracks the surface of the prompt, not the evidence: moving the refusal option from last to first in the answer list collapses refusal rates from 78-100% to 0-6% in three of the six swept model-and-family pairs. Accuracy from frames stays at or below a majority-class baseline under every prompt variant, with or without worked examples, and stated confidence carries no information about correctness. Handing the same models the force data as ten lines of text produces the first above-baseline diagnoses, in four of the six models: much of the failure reflects missing sensor data, not missing ability. We pose the choice as a three-action decision problem, act, consult your own sensors, or ask a human, whose optimal policy follows from measured accuracy. The models do not follow it, and their ask rates ignore a fourfold change in question cost. One question to a human still lifts them from that baseline to roughly the answerer's own reliability (0.70-0.81 when they ask). The decision to ask should be tied to measured accuracy and stated costs, not to the model's confidence.

PR本紙発行元 EmplifAI