日本フィジカルAI新聞

世界のフィジカルAIを、日本語で。

週刊ニュースレター購読
ロボット学習arXiv:2610.09455

RLHND: 動画基盤モデルを物理的に接地したハンドトラッカーとしてロボット学習に活用

RLHND: Video Foundation Models as Physically Grounded Hand Trackers for Robot Learning

シェア:XThreadsFacebookLINEはてブBluesky

単眼エゴセントリック動画から手の姿勢と接触・力の触覚情報を同時推定する動画基盤モデルベースのハンドトラッカーを提案し、ロボット方策学習のための人間動画活用を促進する。

詳しい要約

1. どんなもの?

- 単眼 egocentric 動画から手の pose と触覚情報(contact/force)を同時推定する video foundation model ベースの hand tracker。 - 名称は RLHND。 - 目的は human video を robot policy 学習に活用すること。 - 手の pose と物理的に整合する tactile 情報を出力する。

2. 先行研究と比べてどこがすごい?

- 既存 hand tracker は cropped frame から pose を回帰し、手の動きや物体との相互作用の prior が乏しく、不正確で物理的に不整合な推定になりがち。 - また contact や force などの物理的手がかりが欠如し、human video の robot policy 学習への利用が制限されていた。 - RLHND は video foundation model の prior を活用し、pose と tactile を同時に推定する点が新しい。

3. 技術・手法の肝は?

- 事前学習済み Cosmos 3 video diffusion backbone を clean-latent conditioning により deterministic な clip-level feature extractor に変換。 - 手の動きと hand-object interaction の学習済み prior を tracking に持ち込む。 - pose 推定では解剖学的に妥当な joint angle を予測し、shape parameter への任意条件付けで同一動画内・同一 actor の動画間で hand shape を一貫させる。 - tactile 推定では pose stream を凍結して学習した別の tactile expert stream が手表面の dense contact と force を予測。 - LBS-based feature spreading により高コストな per-vertex attention なしで vertex-wise 特徴抽出を実現。

4. どうやって有効だと検証した?

- 複数の benchmark dataset で pose estimation の state-of-the-art 性能を達成。 - contact と force estimation でも state-of-the-art 性能を達成。 - retargeting 結果と実世界 robot 実験を通じて robot learning への有用性を実証。

5. 議論はある?

- 要旨からは不明。

6. 次に読むべき論文は?

- Cosmos 3 - LBS-based feature spreading - 関連する hand tracking 手法や human video を用いた robot policy 学習の研究

※ AIが要旨から生成した要約です。正確性は原文をご確認ください。

著者: Seungjun Moon, Subin Jeon, Sangwoo Kim, Hanbyul Joo, Jinwoo Shin

分類: cs.CV, cs.RO

原文アブストラクト

Recently, approaches that leverage human video datasets for robot policy training have become increasingly prevalent. However, most existing hand trackers regress pose from cropped frames with limited priors on hand motion and object interaction, resulting in inaccurate and physically inconsistent estimates. Moreover, the lack of physical cues, e.g., contact and force, limits the use of human videos for robot policy training. To this end, we propose RLHND, a video foundation model-based hand tracking model that jointly estimates hand pose and realistic tactile information from monocular egocentric videos. RLHND turns the pre-trained Cosmos 3 video diffusion backbone into a deterministic clip-level feature extractor via clean-latent conditioning, carrying its learned priors on hand motion and hand-object interaction into tracking. For pose estimation, RLHND (i) predicts hand poses with anatomically plausible joint angles and (ii) enables optional conditioning on the shape parameter to maintain consistent hand shape within the same video and even across videos recorded by the same actor. For tactile estimation, a separate tactile expert stream, trained with the pose stream frozen, predicts dense contact and force over the hand surface. We further adopt LBS-based feature spreading to enable vertex-wise feature extraction without costly per-vertex attention. RLHND achieves state-of-the-art performance across various benchmark datasets for pose estimation, while also achieving state-of-the-art performance in contact and force estimation. Moreover, we demonstrate the utility of RLHND for robot learning through retargeting results and real-world robot experiments. The code will be publicly available at https://seungjun-moon.github.io/rlhnd/.

関連論文

PR本紙発行元 EmplifAI