日本フィジカルAI新聞

世界のフィジカルAIを、日本語で。

週刊ニュースレター購読
ビデオ生成/ベンチマークarXiv:2608.13049v1

H2R-Bench: 世界モデルにおける人間からロボットへの操作ビデオ生成のベンチマーク

H2R-Bench: Benchmarking Human-to-Robot Manipulation Video Generation in World Models

シェア:XThreadsFacebookLINEはてブBluesky

人間の操作ビデオをロボットの操作ビデオに変換する能力を評価するベンチマークを提案し、既存のビデオ生成モデルの限界を明らかにした。

詳しい要約

1. どんなもの?

H2R-Benchは、人間の操作ビデオをロボット操作ビデオに変換するクロスエンボディメント(人間からロボットへの)操作ビデオ生成を評価するためのベンチマークである。各インスタンスは、人間のデモンストレーションビデオ、ターゲットのエンボディメント制約、およびタスク目標、アクションイベント、機能的接触、オブジェクト応答を含むソース接地アノテーションで構成される。生成されたビデオは、目標状態完了、アクションイベント完了、機能的接触転送、エンボディメント正確性、一般的なビデオ品質の5つの次元で評価される。

2. 先行研究と比べてどこがすごい?

先行研究では、ビデオワールドモデルによるロボット中心の操作ビデオ合成の可能性が示唆されているが、クロスエンボディメント転送能力はほとんど未探求であった。H2R-Benchは、このギャップを埋めるために、人間の観察をロボット操作ビデオに変換するタスクを体系的に評価する初めてのベンチマークである。

3. 技術・手法の肝は?

H2R-Benchは、人間のデモンストレーションビデオとターゲットエンボディメント制約を入力とし、指定されたエンボディメントのロボット操作ビデオを生成する。評価は、5つの次元(目標状態完了、アクションイベント完了、機能的接触転送、エンボディメント正確性、一般的なビデオ品質)に基づいて行われる。ベンチマークには、6つの操作ファミリーと2つのロボットエンボディメントが含まれる。

4. どうやって有効だと検証した?

H2R-Benchは、11の最先端ビデオ生成モデルを6つの操作ファミリーと2つのロボットエンボディメントにわたって評価した。その結果、現在のビデオワールドモデルは人間からロボットへの操作転送において限定的であり、主要なモデルでもエンボディメント一貫性、機能的相互作用、タスク実行に失敗することが多いことが明らかになった。

5. 議論はある?

要旨からは、現在のビデオワールドモデルが人間とロボットのエンボディメントギャップを埋めるには不十分であることが示唆される。H2R-Benchは、このギャップを診断するための体系的なフレームワークを提供するが、具体的な議論や限界については要旨からは不明である。

6. 次に読むべき論文は?

要旨で参照されている関連研究は、ビデオワールドモデルとクロスエンボディメント転送に関するものである。具体的な論文名は挙げられていないが、次に読むべき論文としては、ビデオ生成モデル(例えば、Video Diffusion Models)やロボット学習のためのデータセット(例えば、RoboNet)などが考えられる。

※ AIが要旨から生成した要約です。正確性は原文をご確認ください。

著者: Dingyi Rong, Yue Shi, Chaofan Ma, Jiezhang Cao, Zongrui Wang, Zeyu Zhang, Yao Mu, Guangtao Zhai, Ning Liu

分類: cs.RO, cs.CV

原文アブストラクト

Large-scale manipulation data is essential for robot learning, yet collecting robot demonstrations remains expensive and difficult to scale. Meanwhile, abundant egocentric human manipulation videos provide rich behavioral experiences, but transferring them across embodiments remains challenging due to differences between human hands and robotic end-effectors. Recent advances in video world models offer a promising pathway to synthesize robot-centric manipulation videos from human observations, while their cross-embodiment transfer capability remains largely unexplored. Therefore, we introduce H2R-Bench, a benchmark for evaluating cross-embodiment human-to-robot manipulation video generation, where models transform egocentric human demonstrations into robot manipulation videos under specified embodiments. Each benchmark instance contains a human demonstration video, target embodiment constraints, and source-grounded annotations covering task goals, action events, functional contacts, and object responses. H2R-Bench evaluates generated videos through five dimensions, including goal-state completion, action-event completion, functional contact transfer, embodiment correctness, and general video quality. We benchmark eleven state-of-the-art video generation models across six manipulation families and two robot embodiments. Our evaluation reveals that current video world models remain limited in human-to-robot manipulation transfer: even leading models often fail in embodiment consistency, functional interaction, and task execution. H2R-Bench provides a systematic diagnostic framework for evaluating whether video world models can bridge the human-to-robot embodiment gap and convert human manipulation observations into robot-centric training resources.