日本フィジカルAI新聞

世界のフィジカルAIを、日本語で。

週刊ニュースレター購読
一人称視点/動画理解/ベンチマークarXiv:2609.37938

局所的な動画理解は遭遇をまたいで転移するか?EgoGearsベンチマーク

Does Local Video Understanding Transfer Across Encounters? The EgoGears Benchmark

シェア:XThreadsFacebookLINEはてブBluesky

一人称視点の動画を単一・複数本で比較するベンチマークEgoGearsを提案し、複数動画間での対応付けや証拠統合が必要になるとMLLMの性能が平均22.5ポイント低下することを示した。

詳しい要約

1. どんなもの?

- 一人称視点の動画理解が、別の遭遇(encounter)間で転移するかを診断するベンチマーク EgoGears を提案。 - 単一動画と複数動画の両方を含み、126 の人間収集 egocentric recordings(39 の屋外ルート)から 567 の単一動画質問と 1,487 の複数動画質問を構築。 - 移動速度や照明条件を変えた繰り返し走行により、共有物理環境での比較を可能にし、531 問は独立記録間の alignment を要求。 - 単一動画質問は局所的な視覚・空間・運動 evidence を、複数動画質問は evidence が正しい observation に結びつき route 関係として構成できるかを測る。

2. 先行研究と比べてどこがすごい?

- 従来の cross-video accuracy は局所知覚の失敗と、observation identity の保持・correspondence 確立・evidence 構成の失敗を混同していた。 - EgoGears は単一動画と複数動画を相補的に分離し、局所 video understanding が実際に転移するかを診断できる点が新しい。 - 同一物理環境での繰り返し走行を基盤にし、独立記録間の alignment を要求する質問を含む点で、単なる動画本数や記録境界の影響を切り分けられる。

3. 技術・手法の肝は?

- 単一動画質問と複数動画質問を設計し、局所 evidence の有無と、それを正しい observation に束縛し route 状態として順序付け・統合する能力を分けて評価。 - 複数動画質問では独立記録間の alignment を要求し、observation–evidence binding と ordered route-state tracking を主なボトルネックとして検証。 - 6 つの model families にわたり、単一動画 29 構成・複数動画 31 構成の MLLM configurations を main leaderboard で報告。

4. どうやって有効だと検証した?

- 20 の構成を両 split で同等に評価したところ、全モデルが複数動画質問で性能低下し、平均 22.5 percentage points の減少を確認。 - 回答形式と scoring を固定しても gap が持続することを確認。 - gap は単に動画本数が増えることや recording boundaries では説明できないと報告。

5. 議論はある?

- 中心的なボトルネックは observation–evidence binding と ordered route-state tracking であると議論。 - 集約的な cross-video accuracy では局所知覚の失敗と転移・対応・構成の失敗が混同されるという問題意識を提示。 - その他の議論や限界については要旨からは不明。

6. 次に読むべき論文は?

- 要旨で参照・比較されている個別研究は明示されていない。 - 関連手法として MLLM(multimodal large language model)ベースの video understanding、egocentric video 理解、cross-video / multi-video reasoning、video alignment の定番研究を次に読むべき。 - 公開コードとベンチマーク(https://github.com/lei-qi-233/EgoGears)も参照。

※ AIが要旨から生成した要約です。正確性は原文をご確認ください。

著者: Yuedong Tan, Lei Qi, Yu Liu, Di Wen, Ruiping Liu, Xiaoye Wang, Yufan Chen, Junwei Zheng, Chengzhi Wu, Chen Zhang, Zhihang Chen, Haiwen Sun, Zongwei Wu, Radu Timofte, Danda Pani Paudel, Kunyu Peng

分類: cs.CV, cs.AI, cs.RO

原文アブストラクト

Embodied systems must make knowledge acquired during one encounter usable in another despite changes in viewpoint, motion, and illumination. Yet aggregate cross-video accuracy conflates failures of local perception with failures to preserve observation identity, establish correspondence, and compose evidence, obscuring whether local video understanding actually transfers. We introduce EgoGears, a complementary single- and multi-video benchmark designed to diagnose this transition. It contains 567 single-video and 1,487 multi-video questions derived from 126 human-collected egocentric recordings covering 39 outdoor routes. Repeated traversals across movement speeds and lighting conditions ground comparisons in shared physical environments; 531 questions require alignment across independent recordings. Single-video questions measure the local visual, spatial, and motion evidence available to a model, while multi-video questions test whether evidence remains bound to the correct observation and can be composed into consistent route relationships. We report 29 single-video and 31 multi-video MLLM configurations across six model families in the main leaderboard. Among the 20 configurations evaluated comparably on both splits, every model performs worse on multi-video questions, with a mean decrease of 22.5 percentage points, and the gap persists when answer format and scoring are held fixed. The gap is not explained simply by additional videos or recording boundaries. The central bottlenecks are observation--evidence binding and ordered route-state tracking. The code and benchmark are publicly available at https://github.com/lei-qi-233/EgoGears.

PR本紙発行元 EmplifAI