日本フィジカルAI新聞

世界のフィジカルAIを、日本語で。

週刊ニュースレター購読
一人称視点手再構成arXiv:2610.01210

EgoFound3R: 世界座標空間における一人称視点手再構成と点単位のインタラクション属性推定

EgoFound3R: End-to-End Egocentric Hand Reconstruction in World Space with Point-Wise Interaction Attributes

シェア:XThreadsFacebookLINEはてブBluesky

一人称視点動画から、シーンと共通のメートルスケールで世界座標空間の手形状を再構成し、可視性・接触・距離といった点単位のインタラクション属性を単一パスで推定する統合エンドツーエンドモデルを提案。

詳しい要約

1. どんなもの?

- 一人称視点(egocentric)動画から、世界座標系での手の3D形状を推定する統合モデル - カメラモーションと手のオクルージョンに対処し、シーンと共有されるメートルスケールで手のジオメトリを復元 - 点ごとのinteraction attributes(visibility, contact, distance)も同時に予測 - 従来は手とシーンの推定が分離され、属性は別タスクモデルに委ねられていた - 大規模アノテーションのボトルネックとなるスループット問題にも取り組む

2. 先行研究と比べてどこがすごい?

- 従来の再構成パイプラインは手とシーンの推定を分離し、interaction attributesを別のタスク特化モデルに任せ、動画ごとに複数モデルを呼び出していた - 先行研究ではこれらの属性を推定する再構成モデルは存在しなかった - EgoFound3Rは単一のend-to-endモデルで世界座標系の手ジオメトリと点ごとの属性を一度に予測 - OakInk-v2, TACO, HOI4DでMPJPEをそれぞれ43.2%, 22.4%, 11.6%削減 - 約6倍のスループットを達成

3. 技術・手法の肝は?

- 3つの設計を統合: - (i) structured hand prompts: 事前学習済みの幾何学的priorを世界座標系の手再構成に転移 - (ii) explicit hand representation: 手のジオメトリとinteraction attributesをデコード - (iii) shared-parameter multi-rate design: 推論コストを低減 - これらにより1パスで手のジオメトリと点ごとの属性を予測

4. どうやって有効だと検証した?

- OakInk-v2, TACO, HOI4Dの3つのデータセットで評価 - 平均関節位置誤差(MPJPE)を従来手法比で43.2%, 22.4%, 11.6%削減 - ジオメトリと同時に点ごとのcontactとdistanceを予測 - 約6倍のスループットを達成

5. 議論はある?

- 要旨からは不明

6. 次に読むべき論文は?

- OakInk-v2, TACO, HOI4Dのデータセットを用いた研究 - 一人称視点の手再構成に関する先行研究 - 具体的な論文名は要旨からは不明

※ AIが要旨から生成した要約です。正確性は原文をご確認ください。

著者: Hongming Fu, Jingcheng Shi, Wenjia Wang, Binhua Zuo, Bo Zhao

分類: cs.CV

原文アブストラクト

Egocentric video has become a primary source of supervision for embodied models, and its value rests on recovering hand motion in world coordinates, which camera motion and hand occlusion make difficult. Existing reconstruction pipelines typically separate hand and scene estimation, leave interaction attributes to separate task-specific models, and invoke several models per video, so no prior reconstruction model estimates these attributes and throughput becomes a practical constraint on large-scale annotation. We therefore introduce EgoFound3R, a unified end-to-end model that estimates world-space hand geometry in a metric scale shared with the scene, and predicts point-wise interaction attributes, including visibility, contact, and distance. The model integrates three designs: (i) structured hand prompts that transfer pretrained geometric priors to world-space hand reconstruction; (ii) an explicit hand representation that decodes hand geometry and interaction attributes; and (iii) a shared-parameter multi-rate design that lowers inference cost. Together, these designs predict hand geometry and point-wise attributes in one pass. On OakInk-v2, TACO, and HOI4D, EgoFound3R reduces the mean per-joint position error (MPJPE) by 43.2%, 22.4%, and 11.6% over previous methods and predicts point-wise contact and distance alongside the geometry in the same pass, while attaining approximately 6x higher throughput.

PR本紙発行元 EmplifAI