日本フィジカルAI新聞

世界のフィジカルAIを、日本語で。

週刊ニュースレター購読
共有制御arXiv:2609.25369

能力を考慮した意味意図に基づく共有制御のための調停

Capability-Aware Arbitration for Semantic Intent-Based Shared Control

シェア:XThreadsFacebookLINEはてブBluesky

VLMによる人間意図の確信度とVLAの実行能力の確信度を組み合わせてロボットの権限を調停する共有制御フレームワークを提案し、12人での実験で成功率92%を達成した。

詳しい要約

1. どんなもの?

- 共有制御において、VLMが人間の意図とその信頼度を推定し、VLAポリシーが自律行動を生成する枠組み。 - VLAの能力信頼度を確率的行動軌跡の分散と局所不安定性からオンライン推定。 - ベイズフィルタで平滑化した意図信頼度とVLA能力信頼度をシグモイド写像で統合し、ロボットの権限を適応的に調停する。 - 12名の参加者によるpick-and-placeと双方向スタッキングの実験で有効性を検証。

2. 先行研究と比べてどこがすごい?

- 従来の共有制御は人間意図の信頼度のみに基づいて権限を配分し、自律実行の信頼性を仮定していた。 - この仮定が崩れると、高い意図信頼度が過剰支援(over-helping)を引き起こす問題があった。 - 提案手法はVLAの能力信頼度を組み込むことで、過剰支援を緩和し、タスク成功率を向上させる点が新しい。

3. 技術・手法の肝は?

- VLMが人間の意図とsemantic-intent confidenceを推定。 - VLAポリシーが自律行動を生成し、そのstochastic action trajectoriesのdispersionとlocal instabilityからVLA capability confidenceをオンライン推定。 - ベイズフィルタで平滑化した意図信頼度とVLA能力信頼度を、非線形なシグモイド写像で統合するarbitration policyを設計。 - これによりロボットの権限を適応的に調停する。

4. どうやって有効だと検証した?

- VLM/VLAの信頼度評価と、12名の参加者によるpick-and-placeおよび双方向スタッキングの実験を組み合わせて評価。 - in-distributionとout-of-distributionの条件下で検証。 - 提案手法はタスク成功率92%を達成し、manual teleoperation(83%)、intent-only arbitration(44%)、fixed equal-weight blending(10%)を上回った。 - また、制御のfriendlinessが高く、authority-weighted disagreementが両共有制御ベースラインより低かった。

5. 議論はある?

- 結果は、VLA capabilityを権限配分に組み込むことでover-helpingを緩和し、共有制御性能を改善する利点を示す。 - ただし、要旨からは具体的な議論や限界、今後の課題については不明。

6. 次に読むべき論文は?

- 要旨で参照/比較されている研究:intent-only arbitration、fixed equal-weight blending、manual teleoperation。 - 関連手法:VLM、VLA、Bayesian filtering、sigmoid mapping。 - 同分野の定番:shared control、human-robot collaboration、intent inference、capability-aware arbitration。

※ AIが要旨から生成した要約です。正確性は原文をご確認ください。

著者: Zhaoda Du, Michael Bowman, Xiaoli Zhang

分類: cs.RO

原文アブストラクト

Shared control often allocates robot authority based on confidence in inferred human intent, assuming reliable autonomous execution. When this assumption fails, high intent confidence can cause over-helping. We present a capability-aware shared-control framework in which a vision-language model (VLM) infers human intent and provides semantic-intent confidence, while a vision-language-action (VLA) policy generates autonomous actions. VLA capability confidence is estimated online from the dispersion and local instability of stochastic action trajectories. We design a nonlinear arbitration policy that combines Bayesian-filtered semantic-intent confidence with VLA capability confidence through a sigmoid mapping to adapt robot authority. Our evaluation combined VLM/VLA confidence assessment with a study involving 12 participants performing pick-and-place and bidirectional stacking under in-distribution and out-of-distribution conditions. The proposed method achieved the highest task success rate (92%), compared with manual teleoperation (83%), intent-only arbitration (44%), and fixed equal-weight blending (10%). It also achieved higher control friendliness and lower authority-weighted disagreement than both shared-control baselines. These results demonstrate the benefit of incorporating VLA capability into authority allocation to mitigate over-helping and improve shared-control performance.

関連論文

PR本紙発行元 EmplifAI