日本フィジカルAI新聞

世界のフィジカルAIを、日本語で。

週刊ニュースレター購読
探索/シーン表現arXiv:2606.24068v1

ObsGraph: 身体化推論と探索のための階層的観察表現

ObsGraph: Hierarchical Observation Representation for Embodied Reasoning and Exploration

シェア:XThreadsFacebookLINEはてブBluesky

ロボットが複雑な環境でタスクを遂行するために、観察中心の階層的シーングラフを提案し、シーン表現・検索・探索を統合して効率的な情報収集を実現した。

著者: Taekbeom Lee, Youngseok Jang, Jeonghwa Heo, Jeongjun Choi, H. Jin Kim

分類: cs.CV, cs.RO

原文アブストラクト

Embodied reasoning and exploration are increasingly considered crucial abilities for robots operating in complex and unfamiliar environments. To accomplish tasks in such settings, an agent must identify and acquire the information necessary for the task through exploration. We propose ObsGraph, an observation-centric hierarchical scene graph that unifies scene representation, retrieval, and exploration. It retains visual evidence and organizes it into room-view-object layers: rooms provide coarse semantic anchors, views preserve contextual object covisibility, and objects store fine-grained details. On top of this representation, we perform coarse-to-fine hierarchical retrieval under a bounded budget, and crucially use retrieval outcomes to structure the exploration candidate space--activating room-level exploration, view refinement, or frontier exploration--thereby tightly coupling representation, retrieval, and adaptive multi-scale exploration. Experiments across embodied reasoning and exploration benchmarks demonstrate improved success and efficiency, highlighting the benefits of structured scene representation and more targeted information gathering driven by identified evidence gaps.