日本フィジカルAI新聞

世界のフィジカルAIを、日本語で。

週刊ニュースレター購読
ハンドオーバーarXiv:2610.08003

生成的仮説選択による反応的なタスク指向ロボット-人間間ハンドオーバー

Reactive Task-Oriented Robot-Human Handovers via Generative Hypothesis Selection

シェア:XThreadsFacebookLINEはてブBluesky

VLMの画像生成でタスクに応じた手と物体のインタラクション仮説を生成し、観察した人間の手姿勢とリアルタイムに照合することで、未知の物体-タスクの組み合わせでも適切な受け渡し方を推論する手法を提案した。

詳しい要約

1. どんなもの?

- タスク指向のロボットから人間への物体受け渡し(handover)を実現する手法「GENESIS-Handover」を提案。 - VLM(Vision-Language Model)の画像生成能力を活用し、タスクに応じた手と物体のインタラクション仮説を複数生成。 - 生成された仮説を実時間で観測された人間の手姿勢と照合し、最適な受け渡し構成を推論する。 - 未知の物体とタスクの組み合わせに対しても、タスク条件付きの受け渡し戦略を生成可能。

2. 先行研究と比べてどこがすごい?

- 従来の最先端手法は物体の幾何形状からアフォーダンスへと進化してきたが、人間が物体を利用する際に選択する明示的でタスク固有の手姿勢を予測しないことが多い。 - 多くの物体は複数のインタラクションモダリティ(例: claw hammer で打つか引くか)を支持するため、この変動性をモデル化する必要がある。 - 提案手法はVLMを人間と物体の相互作用の事前分布として活用し、未知の物体-タスクペアに対してもタスク条件付きの受け渡し戦略を生成できる点が優れている。

3. 技術・手法の肝は?

- VLMの画像生成を用いて、タスク固有の手と物体のインタラクション仮説を多様に生成する。 - 生成された仮説を実時間で観測された人間の手姿勢とマッチングし、最も適切な受け渡し構成を推論する。 - VLMを実現可能な手と物体の相互作用の事前分布として利用することで、未知の物体-タスクペアに対してもタスク条件付きの戦略を生成する。

4. どうやって有効だと検証した?

- 単体のインタラクション提案モジュールを評価した後、モバイルマニピュレータ上でフルシステムを展開。 - 12人の参加者によるユーザスタディを5つのタスク-物体ペアで実施。 - 83.3%の参加者が、提案手法は従来の最先端手法よりもタスク理解が優れていると認識した。

5. 議論はある?

- 要旨からは不明。

6. 次に読むべき論文は?

- 要旨で参照されている先行研究:物体の幾何形状からアフォーダンスへと進化したタスク指向のロボット-人間受け渡しの最先端手法。 - 関連手法:VLM(Vision-Language Model)を用いた画像生成、タスク指向のhandover、アフォーダンスベースの手法。

※ AIが要旨から生成した要約です。正確性は原文をご確認ください。

著者: Carmen Scheidemann, Andreea Tulbure, Pascal Burkhardt, Marco Hutter

分類: cs.RO

原文アブストラクト

When humans hand each other objects, they incorporate both geometric and semantic information into this process. For example, passing a knife with the handle towards the recipient, rather than the blade, is both more ergonomic and safer. Recent state-of-the-art methods for task-oriented robot-human handovers have progressed from modeling object geometry to incorporating object affordances. However, they often forgo predicting the explicit, task-specific hand poses a human selects to utilize an object. Since many objects support multiple interaction modalities, e.g., a claw hammer used to strike or pull nails, this variability must be modeled to achieve robust task-oriented handovers. To tackle this, we propose a novel approach, GENESIS-Handover (GENErative HypotheSIS), which leverages VLM image generation to produce a variety of task-specific hand-object interaction hypotheses. These hypotheses are matched in real time to the observed human hand pose, enabling inference of the most suitable handover configuration. By leveraging VLMs as priors of plausible hand-object interactions, the method produces task-conditioned handover strategies for previously unseen object-task pairs. We evaluate the standalone interaction proposal module before deploying the full system on a mobile manipulator. In a user study with 12 participants across five task-object pairs, 83.3% perceived our method to have better task understanding than the previous state of the art.

関連論文

PR本紙発行元 EmplifAI