日本フィジカルAI新聞

世界のフィジカルAIを、日本語で。

週刊ニュースレター購読
VLAarXiv:2609.18514

ActiveScale:モデル・データ・ハードウェアを横断したロボットの能動的知覚のスケーリング

ActiveScale: Scaling Active Perception for Robots across Model, Data, and Hardware

シェア:XThreadsFacebookLINEはてブBluesky

視点変更を伴う操作タスクのための能動的知覚フレームワークを提案し、カメラ姿勢を考慮したVLAモデル、1000時間の一人称視点データによる中間学習、単一操作者で遠隔操作可能な移動マニピュレーションプラットフォームを統合した。

詳しい要約

1. どんなもの?

- VLAモデルに能動的知覚を導入するフレームワークActiveScaleを提案。 - モデル・データ・ハードウェアを統合的に設計。 - モデルは履歴動画とカメラ姿勢監督を活用。 - データは1000時間の一人称視点とロボットデータで中間訓練。 - ハードウェアは単一操作者遠隔操作のAMPプラットフォーム。

2. 先行研究と比べてどこがすごい?

- 従来のVLAは固定視点が前提で、視点変化や能動的観測が困難。 - 履歴動画とカメラ姿勢を明示的に監督する点が新しい。 - 人間活動のカメラ運動を活用するスケーラブルな中間訓練を導入。 - 能動知覚と移動操作を両立するAMPを開発。 - これらを統合した基盤を提供。

3. 技術・手法の肝は?

- VLAに履歴動画入力を追加し、フレームごとのpose tokenと軽量予測ヘッドでカメラ姿勢を監督。 - 視点間の観測を関連付け、シーン理解を一貫させる。 - 1000時間のegocentricおよびロボットデータで中間訓練し、時間入力と姿勢監督に適応。 - AMPは単一操作者遠隔操作で視点変更と操作を調整するデモを収集可能。

4. どうやって有効だと検証した?

- 能動知覚タスクで成功率の向上を実験的に示した。 - アブレーション研究でカメラ姿勢認識モデリングとegocentric中間訓練の貢献を検証。 - 具体的なタスクやベースラインは要旨からは不明。

5. 議論はある?

- 能動知覚研究の統合基盤を提供すると主張。 - 限界や課題、今後の方向性は要旨からは不明。

6. 次に読むべき論文は?

- 要旨で参照/比較されている研究は明示されていない。 - 関連手法としてVLAモデル、egocentricデータ、能動知覚、移動操作の研究が挙げられる。 - 具体的な論文名は要旨からは不明。

※ AIが要旨から生成した要約です。正確性は原文をご確認ください。

著者: Shuai Zhou, Kaisheng Pang, Wenxuan Song, Wenjie Zhang, Xinhu Zheng, Haoang Li

分類: cs.RO, cs.LG

原文アブストラクト

Active perception is essential for robotic manipulation when fixed viewpoints leave task-relevant information occluded or unobserved. However, enabling vision-language-action (VLA) models to reason across changing viewpoints and actively acquire informative observations remains challenging. We present ActiveScale, a framework that advances active perception through coordinated model, data, and hardware designs. Our model augments a VLA with historical video observations and explicit camera-pose supervision, using per-frame pose tokens and a lightweight prediction head to associate observations across viewpoints and support a coherent understanding of the scene. To learn from the camera motion naturally present in human activity, we introduce a scalable human--robot mid-training recipe using 1000 hours of egocentric and robotic data, adapting the model to temporal inputs and pose supervision. We further introduce Active-perception Mobile-manipulation Platform (AMP), a robotic platform that supports active perception and mobile manipulation through single-operator teleoperation, enabling scalable collection of demonstrations that coordinate viewpoint changes and manipulation. Experiments demonstrate improved success rates on active-perception tasks, while ablation studies validate the contributions of camera-pose-aware modeling and egocentric mid-training. Together, these components provide an integrated foundation for studying and developing active perception in robotic manipulation.

関連論文

PR本紙発行元 EmplifAI