日本フィジカルAI新聞

世界のフィジカルAIを、日本語で。

週刊ニュースレター購読
手姿勢推定arXiv:2609.34817

ESTHER: 実環境での自己中心ステレオ手推定と再構成

ESTHER: Egocentric Stereo Hand Estimation and Reconstruction in the Wild

シェア:XThreadsFacebookLINEはてブBluesky

ウェアラブルな自己中心ステレオ映像からメートル尺度の3D手形状を再構成するモデルESTHERを提案し、実環境データセットESTHER3Dを構築した。

詳しい要約

1. どんなもの?

- 一人称視点のステレオ映像から手のメトリック3D再構成を行うモデルESTHERを提案。 - ウェアラブルなegocentric stereo向けに設計。 - ステレオ幾何、時間推論、出力表現を統合。 - キャリブレーション済みラベリングパイプラインによる疑似ラベルで学習。 - ベンチマークESTHER3Dも構築。大規模in-the-wild訓練セットとmotion captureによる真のメトリックGTテストセットを含む。

2. 先行研究と比べてどこがすごい?

- 従来、egocentric stereoからのメトリック3D手再構成にはend-to-endモデルもin-the-wildベンチマークも存在しなかった。 - ESTHERは最先端精度を達成。 - 外部汎化性能に優れる。 - 実環境の欠損視点、フレーム落ち、照明・モーションブラー極値に対して頑健。 - 既存手法が破綻する条件でも動作。 - ステレオ指導により見かけの手のスケールをメトリック深度に結びつけ、単眼視に崩しても真のメトリックスケールを保持。

3. 技術・手法の肝は?

- ステレオ幾何、時間推論、出力表現をウェアラブルegocentric stereo向けに設計。 - キャリブレーション済みラベリングパイプラインから疑似ラベルを生成し学習。 - そのモデルがESTHER3Dを構築。大規模in-the-wild訓練セット(モデル生成ラベル)とmotion captureテストセット(真のメトリックGT)をペアリング。 - ステレオ指導が手の見かけスケールをメトリック深度に結びつける。 - 異なるステレオリグやモダリティに最小限のファインチューニングで適応。

4. どうやって有効だと検証した?

- 実験でstate-of-the-art精度を確認。 - 外部汎化性能が優れることを検証。 - 欠損視点、フレーム落ち、照明・モーションブラー極値に対する頑健性を検証。 - 異なるステレオリグやモダリティへの適応を最小限のファインチューニングで確認。 - 単眼視に崩した後も真のメトリックスケールを保持することを確認。

5. 議論はある?

- ステレオ指導が単なる緩やかな劣化ではなく、見かけの手のスケールをメトリック深度に結びつける深い効果を持つと議論。 - 異なるステレオリグやモダリティへの適応、単眼視でのメトリックスケール保持という驚くべき性質を示す。 - その他の限界や議論は要旨からは不明。

6. 次に読むべき論文は?

- 要旨で参照/比較されている研究は明示されていない。 - 同分野の定番として、egocentric hand pose estimation、stereo 3D hand reconstruction、monocular 3D hand reconstruction、motion captureベースの評価手法などが次に読むべき候補。 - 具体的な論文名は要旨からは不明。

※ AIが要旨から生成した要約です。正確性は原文をご確認ください。

著者: Hongyu Ma, Hairong Qu, Shiqi Zhao, Yongsong Yang, Peng Yin

分類: cs.CV, cs.AI

原文アブストラクト

Human dexterity is guided by two eyes watching two hands: binocular vision supplies the metric 3D structure that fine-grained manipulation consumes. Egocentric stereo is therefore the natural perceptual interface for robots, AR, and VR-yet metric 3D hand reconstruction from this very signal still has neither an end-to-end model nor an in-the-wild benchmark. We propose ESTHER, a model whose stereo geometry, temporal reasoning, and output representation are designed for wearable egocentric stereo. It is trained on pseudo-labels from a calibrated labeling pipeline and in turn assembles our benchmark ESTHER3D, an egocentric stereo hand dataset pairing a large in-the-wild training set of model-generated labels with a motion capture test set of true metric ground truth. Experiments show state-of-the-art accu?racy, superior external generalization, and robustness to the missing views, dropped frames, and lighting and motion blur extremes of real egocentric capture that break existing meth?ods. This robustness runs deeper than graceful degradation: stereo guidance teaches the model to bind apparent hand scale to metric depth, so it not only adapts to different stereo rigs and modalities with minimal fine-tuning, but more strikingly preserves true metric scale even after collapsing to a single monocular view.

関連論文

PR本紙発行元 EmplifAI