日本フィジカルAI新聞

世界のフィジカルAIを、日本語で。

週刊ニュースレター購読
画像編集/ロボティクスarXiv:2608.12122v1

HandEdit: 身体性を考慮した人間からロボットへの器用な手の画像編集のための統一ベンチマーク

HandEdit: A Unified Benchmark for Egocentric Human-to-Robot Dexterous Hand Image Editing

シェア:XThreadsFacebookLINEはてブBluesky

人間の手の画像を様々なロボットハンドに変換する大規模な画像編集データセットとベンチマークを構築し、既存の編集モデルの評価と身体性を考慮した編集モデルの発展を促進する。

詳しい要約

1. どんなもの?

HandEditは、人間の手と腕を多様なロボットの器用な身体(dexterous robotic embodiment)に変換する、大規模で統一されたembodiment-aware画像編集データセットおよびベンチマークである。5つの異なるソースデータセットから2億以上の編集インスタンスを提供し、26種類のURDF(13の手のみ、13の手と腕の構成)をカバーする。Hand-onlyとHand-Armの2つのトラックを持つ統一ベンチマークプロトコルを確立し、URDF条件付き評価をサポートする。

2. 先行研究と比べてどこがすごい?

既存の一般的な画像編集モデルは強力な能力を持つが、embodiment固有の事前知識が欠如しており、人間とロボットデータ間の見た目、関節構造、カメラ視点の大きな差異を埋めることができない。HandEditは、このギャップを埋めるために設計された最初の大規模なembodiment-aware画像編集データセットであり、ロボット操作学習のためのスケーラブルなデータソースを提供する点で先行研究より優れている。

3. 技術・手法の肝は?

HandEditは、5つの多様なソースデータセットから編集インスタンスを生成し、26の異なるURDF(手のみ13、手と腕13)をカバーする。ベンチマークはHand-onlyとHand-Armの2つのトラックで構成され、URDF条件付き評価を可能にする。評価には、一般的な類似度メトリクス、VLMベースの判断、embodiment-awareメトリクスを含む多次元メトリクススイートを使用する。

4. どうやって有効だと検証した?

HandEditの有効性は、11の代表的な画像編集ベースラインを、多次元メトリクススイート(一般的な類似度メトリクス、VLMベースの判断、embodiment-awareメトリクス)を用いて広範に評価することで検証された。

5. 議論はある?

要旨からは、HandEditがembodiment-aware編集モデルを前進させ、豊富な人間のビデオデータからスケーラブルな器用なロボット学習を可能にする一方で、データセットのバイアスや編集品質の限界、実ロボットへの適用可能性などについての議論は明示されていない。

6. 次に読むべき論文は?

要旨で参照されている研究は、一般的な画像編集モデル(例:InstructPix2Pix、ControlNetなど)や、ロボット操作のためのデータセット(例:DexYCB、EgoDex等)が関連する。次に読むべき論文としては、これらの基盤となる画像編集モデルや、embodiment-awareロボット学習の研究が挙げられる。

※ AIが要旨から生成した要約です。正確性は原文をご確認ください。

著者: Zhenjie Yang, Xingyu Jiao, Guopeng Zhong, Shuzhe Yang, Shi Che, Chao Wu, Chenyu Jiang, Dongjie Zhang, Yideng Zhang, Zheng Zhang, Muyun Jiang, Haisheng Su, Shuang Jin, Donghang Zhang, Chao Yang, Li Chen, Hongyang Li, Zuxuan Wu, Yu-Gang Jiang, Xiaosong Jia, Junchi Yan

分類: cs.RO, cs.CV

原文アブストラクト

Robotic manipulation with dexterous hands is a cornerstone of Embodied AI, yet its progress is stifled by the high cost of collecting embodiment-aware teleoperation data. While abundant egocentric videos of human hands offer a scalable alternative, the profound discrepancies in appearance, articulation, and camera viewpoints between human and robotic data raise significant challenges for co-training. Though existing general image-editing models demonstrate strong capabilities, they lack necessary embodiment-specific priors to fully bridge this gap. In this work, we present HandEdit, a unified large-scale embodiment-aware image-editing dataset and benchmark specifically designed to transform human hands and arms into various dexterous robotic embodiments within egocentric frames. HandEdit comprises over 200M editing instances derived from five diverse source datasets, covering 26 distinct URDFs, including 13 hand-only and 13 hand-arm configurations. Alongside the dataset, we establish a unified benchmark protocol with two tracks: Hand-only and Hand-Arm, supporting URDF-conditioned evaluation. We conduct extensive evaluations of 11 representative image-editing baselines using a multi-dimensional metric suite, including generic similarity metrics, VLM-based judgment, and embodiment-aware metrics. HandEdit serves as a critical resource at the intersection of image editing and robotics: it advances embodiment-aware editing models while enabling scalable dexterous robotic learning from abundant human video data, paving the way for more generalizable Embodied AI.