ROMA: 実世界の物体中心マルチセンサ能動知覚のためのLLMシステム
ROMA: LLM System for Real-World Object-Centric Multi-Sensory Active Perception
視覚・音・触覚・力覚を統合し、不足情報を能動的に獲得するLLMベースのロボット知覚システムROMAを提案。約2000物体のマルチセンサデータセットROMI-2Kと評価ベンチマークも構築した。
詳しい要約
1. どんなもの?
2. 先行研究と比べてどこがすごい?
3. 技術・手法の肝は?
4. どうやって有効だと検証した?
5. 議論はある?
6. 次に読むべき論文は?
※ AIが要旨から生成した要約です。正確性は原文をご確認ください。
著者: Ruoxuan Feng, Yutong Chen, Ruihua Song, Huan Yang, Zhongyuan Wang, Guocai Yao, Di Hu
分類: cs.RO, cs.CV
原文アブストラクト
Humans inherently understand the physical world through an active process. When sensory evidence is insufficient to infer physical properties, we naturally interact with the environment by deciding what information is missing, how to acquire it, and when sufficient evidence has been obtained. In stark contrast, existing multi-sensory robot systems mainly integrate sensory inputs rather than actively acquiring missing evidence through interactions. In this work, we introduce ROMA, an LLM-based system for Real-World Object-Centric Multi-Sensory Active Perception. ROMA integrates vision, audio, tactile, and force sensing into a reasoning-interaction-feedback loop. The model identifies missing evidence and determines the target objects, interactions, and modalities, while a physical interface executes the selected interactions and collects the multi-sensory feedback. To support this capability, we construct ROMI-2K, a large-scale real-world multi-sensory object interaction dataset covering nearly 2,000 objects and 6 atomic interactions with synchronized sensory feedback. Building on these data, we develop a two-stage training framework that aligns sensory modalities and equips the LLM to assess evidence sufficiency, select informative interactions, and reason over the multi-sensory feedback. We further characterize active perception as perception chains, where acquired evidence guides subsequent interactions and reasoning, and establish ROMA Bench to evaluate single-attribute, long-horizon multi-attribute, and intent-driven active perception. Experiments show that ROMA can actively acquire missing evidence and solve complex, long-chain multi-sensory perception tasks that existing methods struggle to handle, laying a strong perceptual foundation for active multi-sensory embodied agents.
関連論文
- ReSPEC: 動的環境におけるオンライン多スペクトルセンサ再構成フレームワークマルチモーダル知覚
- 深層学習による視覚・超音波ロボットシステム:製造現場におけるガス漏れとアーク放電の検知マルチモーダル知覚
- VTD: ドライバー状態と行動知覚のための視覚・触覚データベースマルチモーダル知覚
- 実環境向けマルチモーダル知覚システムマルチモーダル知覚
- 支援ロボティクスのための堅牢な知覚に向けて:RGB-イベント-LiDARデータセットとマルチモーダル検出パイプラインマルチモーダル知覚
- 都市ロボットのための音声・視覚による信号機状態検出マルチモーダル知覚