日本フィジカルAI新聞

世界のフィジカルAIを、日本語で。

週刊ニュースレター購読
データ生成/両眼視覚arXiv:2609.19881

BinoGen: 身体性視覚知覚と学習のための大規模自己中心両眼データ生成

BinoGen: Scaling egocentric binocular data for embodied visual perception and learning

シェア:XThreadsFacebookLINEはてブBluesky

屋内環境で身体性を考慮した自己中心両眼視覚体験を自動生成するフレームワークBinoGenを提案し、2000万枚以上の注釈付き画像データセットを構築して、実世界の深度推定・物体検出・追跡性能を向上させた。

詳しい要約

1. どんなもの?

- 大規模なegocentric binocular視覚体験を自動生成するフレームワークBinoGenを提案。 - 屋内環境でembodiment-awareな視覚データを生成し、depth maps, optical flow, surface normals, semantic maps, object coordinates, camera posesなどの密なマルチモーダル教師信号を提供。 - 2000万枚以上の注釈付き画像を含むデータセットを構築し、実世界の視覚知覚タスクの改善に利用。

2. 先行研究と比べてどこがすごい?

- 従来は大規模なegocentric binocular観測と密な注釈の収集が高コストで困難だった。 - BinoGenは環境と観察者の変動(視点高さ、視野、両眼幾何、動き)を同時にモデル化し、自動生成する点が新しい。 - 生成データが実世界のdepth estimation, object detection, video object trackingを一貫して改善することを示した。

3. 技術・手法の肝は?

- generative scene synthesis, probabilistic object instantiation, appearance randomization, stochastic trajectory generation, configurable binocular camera setupsを組み合わせる。 - 環境と観察者の変動を同時にモデル化し、同期されたbinocular videosと密なマルチモーダル監督を生成。 - 人間とマウスにインスパイアされた観察を同一環境からペアで生成し、embodimentの影響を制御的に調査可能。

4. どうやって有効だと検証した?

- BinoGenデータを組み込むことで、実世界のdepth estimation, object detection, video object trackingが一貫して改善することを示した。 - 同一環境からの人間とマウスにインスパイアされた観察ペアを用い、embodiment-specific adaptationが性能を大幅に向上させ、joint trainingで単一モデルが両embodimentで競争力を持つことを確認。

5. 議論はある?

- 大規模で制御可能な視覚体験がembodied perceptionを改善することを示唆。 - 観察者のembodimentが知覚学習に与える影響を制御的に調査できる可能性を提示。 - 具体的な限界や議論の詳細は要旨からは不明。

6. 次に読むべき論文は?

- 要旨で参照/比較されている研究は明示されていない。 - 関連手法として、egocentric vision, binocular vision, embodied AI, generative scene synthesis, depth estimation, object detection, video object trackingなどの定番分野が挙げられる。

※ AIが要旨から生成した要約です。正確性は原文をご確認ください。

著者: Chunpeng Li, Ya-tang Li

分類: cs.CV, cs.MM

原文アブストラクト

Embodied visual perception relies on temporally coherent visual experience accumulated through continuous engagement with the environment. However, collecting large-scale egocentric binocular observations together with dense annotations remains costly and difficult. Moreover, visual experience is shaped not only by the environment but also by the embodiment of the observer, including viewing height, field of view, binocular geometry, and motion through the scene. To address these challenges, we present BinoGen, an automated framework for generating large-scale, embodiment-aware egocentric binocular visual experiences in indoor environments. BinoGen jointly models environmental and observer variation through generative scene synthesis, probabilistic object instantiation, appearance randomization, stochastic trajectory generation, and configurable binocular camera setups. The framework produces synchronized binocular videos together with dense multimodal supervision, including depth maps, optical flow, surface normals, semantic maps, object coordinates, and camera poses. Using BinoGen, we construct a dataset comprising more than 20 million annotated images for supervised learning. We demonstrate two complementary utilities of BinoGen. First, incorporating BinoGen data consistently improves real-world visual perception, including depth estimation, object detection, and video object tracking. Second, paired human-inspired and mouse-inspired observations from the same environments enable controlled investigation of how observer embodiment affects perceptual learning. Embodiment-specific adaptation substantially improves performance, while joint training enables a single model to perform competitively across both embodiments. Together, these results demonstrate that large-scale, controllable visual experience can improve embodied perception...

PR本紙発行元 EmplifAI