FUSE: 適応的意味・幾何学的証拠獲得による能動的機能的アフォーダンス接地
FUSE: Active Functional Affordance Grounding through Adaptive Semantic-Geometric Evidence Acquisition
エージェントが機能に基づいて物体を識別・接地するため、不確実性駆動の探索と学習型プランナーを組み合わせて視点を選択する能動的機能的アフォーダンス接地フレームワークFUSEを提案し、Habitatベースのベンチマークで評価した。
詳しい要約
1. どんなもの?
2. 先行研究と比べてどこがすごい?
3. 技術・手法の肝は?
4. どうやって有効だと検証した?
5. 議論はある?
6. 次に読むべき論文は?
※ AIが要旨から生成した要約です。正確性は原文をご確認ください。
著者: Zhou Chen, Sathyanarayanan N. Aakur
分類: cs.RO, cs.CV
原文アブストラクト
Embodied agents must often identify and interact with objects based on their function rather than their identity, requiring them to actively acquire observations that reveal discriminative functional evidence. Existing affordance grounding methods operate from fixed viewpoints and lack mechanisms for deciding where to look when functional cues are occluded or incomplete. We introduce Active Functional Affordance Grounding, a new task in which an agent sequentially explores a scene to identify and spatially ground an object satisfying a functional query. To address this problem, we propose FUSE, an adaptive semantic-geometric evidence acquisition framework that combines explicit uncertainty-driven exploration with a learned amortized planner to efficiently select informative viewpoints. We further introduce a Habitat-based benchmark for evaluating active functional grounding. Experiments show that FUSE achieves the highest observed non-oracle grounding performance while reducing computation by 1.33x relative to fully explicit exploration, and remains effective across multiple affordance knowledge sources.