日本フィジカルAI新聞

世界のフィジカルAIを、日本語で。

週刊ニュースレター購読
一人称視点/力推定arXiv:2610.11347

EgoPhys: 一人称視点の操作動画からピーク接触力と機械的仕事を推定

EgoPhys: Estimating Peak Contact Force and Mechanical Work from Egocentric Manipulation Video

シェア:XThreadsFacebookLINEはてブBluesky

一人称視点のRGB動画のみから、物体操作時の最大接触力と機械的仕事を推定するフレームワークEgoPhysを提案し、Hoi!データセットで高精度を達成した。

詳しい要約

1. どんなもの?

- 一人称視点の操作動画から、接触時のピーク接触力と力学的仕事を推定するRGBのみのフレームワークEgoPhysを提案。 - ピーク力は短い接触イベント、仕事は接触期間全体の力と運動の結合に依存する点が課題。 - Hoi!データセットのテスト分割で評価し、MAE 5.205±0.584 N(力)と0.894±0.081 J(仕事)を達成。

2. 先行研究と比べてどこがすごい?

- 一人称動画からの物理量推定は、接触の手がかりが局所的・間接的で困難。 - ピーク力と仕事は性質が異なるため、単一モデルでの同時推定が難しい。 - EgoPhysはこれらを別々に扱う専用設計で、先行研究と比べ大幅に予測精度を改善(具体的な比較対象は要旨からは不明)。

3. 技術・手法の肝は?

- Contact-Aware Spatial Aggregation (CASA):見た目と幾何特徴を統合し、相互作用に関連する手がかりを強調。 - Target-Specific Multi-Expert Temporal Routing (TMTR):意味・イベント・運動の手がかりを専門の時間エキスパートでモデル化し、力と仕事の予測に別々にルーティング。 - RGBのみを入力とし、物理量を直接推定する。

4. どうやって有効だと検証した?

- Hoi!データセットのテスト分割で評価。 - ピーク力のMAE 5.205±0.584 N、力学的仕事のMAE 0.894±0.081 Jを達成し、予測精度の向上を示した。 - 具体的なベースラインやアブレーションは要旨からは不明。

5. 議論はある?

- 要旨からは不明。 - 限界や失敗事例、一般化可能性についての議論は記載されていない。

6. 次に読むべき論文は?

- 要旨で参照・比較されている研究は明記されていない。 - 同分野の関連手法として、一人称動画からの物理量推定や接触力推定に関する研究(例:Hoi!データセットを用いた手法)を挙げる。 - 具体的な論文名は要旨からは不明。

※ AIが要旨から生成した要約です。正確性は原文をご確認ください。

著者: Zhuo Dong, Jianhua Yang, Haohao Li, Yumeng Zhao, Keji He, Yan Huang, Liang Wang

分類: cs.CV

原文アブストラクト

Physically grounded manipulation of articulated objects requires understanding both the maximum forces encountered during contact and the work performed as their parts move. Peak contact force and mechanical work quantify these complementary aspects, but estimating them from egocentric video is challenging because physical interaction cues are local and indirect. Moreover, peak force is associated with brief contact events, whereas mechanical work depends on force-motion coupling throughout the contact duration. To address these challenges, we propose EgoPhys, an RGB-only framework comprising Contact-Aware Spatial Aggregation (CASA) and Target-Specific Multi-Expert Temporal Routing (TMTR). CASA integrates appearance and geometry features to emphasize interaction-relevant cues, while TMTR models semantic, event, and motion cues with specialized temporal experts and routes them separately for force and work prediction. On the test split from Hoi! dataset, EgoPhys substantially improves predictions of peak force and mechanical work, achieving MAEs of \(5.205 \pm 0.584\) $N$ and $0.894 \pm 0.081$ $J$, respectively.

PR本紙発行元 EmplifAI