日本フィジカルAI新聞

世界のフィジカルAIを、日本語で。

週刊ニュースレター購読
タスク計画arXiv:2610.07649

OntoPlan: オントロジーに基づくシーン表現とエージェントフレームワークによるスケーラブルなロボットタスク計画

OntoPlan: An Ontology-Grounded Scene Representation and Agentic Framework for Scalable Robot Task Planning

シェア:XThreadsFacebookLINEはてブBluesky

オントロジーで物体・空間・関係・状態を統一的に表現し、LLMエージェントが指示解釈から計画生成までを行うことで、大規模環境での長期的タスク計画の成功率とトークン効率を大幅に改善した。

詳しい要約

1. どんなもの?

- LLMベースのロボットタスク計画はopen-endedな指示追従に有望だが、大規模環境でのlong-horizonタスクで性能が低下する。 - 空間情報をテキストでLLMに伝えると空間的文脈を捉えられず、tokenコストが環境規模とともに増大する。 - LLMで直接action sequenceを生成すると、現在のworld stateやaction preconditionsを満たすのが難しい。 - 本研究はontology-groundedなscene representationと、OntoPlanというagentic frameworkを提案する。 - 物体・空間・関係・状態を共有symbolic vocabularyで整合させ、指示解釈・選択的検索・目標と制約の形式化・実行可能plan生成を行う。

2. 先行研究と比べてどこがすごい?

- 5つのindoor環境・3つのscene scaleにわたる150のgeneralタスクで、OntoPlanは平均task success 0.89を達成。 - 最強baselineは0.27であり、大幅に上回る。 - 1タスクあたり平均18.1k total tokensで、最も効率的なbaselineより約5.6倍少ない。 - scene scaleが増大しても優位性が持続し、prior methodsはsuccessがより急激に低下しtokenコストもはるかに高い。 - 曖昧または実行不可能な指示に対し、follow-up questionsやinsufficient informationの報告で適切に応答する点も先行研究と異なる。

3. 技術・手法の肝は?

- ontology-grounded scene representationにより、objects・spaces・relations・statesを共有symbolic vocabularyで整合させる。 - これによりspatial reasoningとtask planningを支援する。 - OntoPlanはagentic frameworkとして、指示を解釈し、タスク関連情報を選択的にretrieveする。 - 目標と制約を形式化し、実行可能なplanを生成する。 - 空間情報をテキストで全て渡すのではなく、選択的検索でtokenコスト増大を抑える設計と推測される。

4. どうやって有効だと検証した?

- 5つのindoor environmentsと3つのscene scalesにまたがる150のgeneral tasksで評価。 - 平均task success 0.89を達成し、最強baselineの0.27と比較。 - 平均18.1k total tokens per taskで、最も効率的なbaselineより約5.6倍少ないことを示す。 - scene scale増大に対するsuccessとtokenコストの挙動を比較し、prior methodsの劣化がより急激であることを確認。 - 曖昧・実行不可能な指示への応答(follow-up questionsやinsufficient information報告)も検証。

5. 議論はある?

- 大規模環境でのlong-horizonタスクにおけるLLMベース計画の限界(空間文脈の欠如、tokenコスト増大、world stateとpreconditionsの不整合)を議論。 - ontology-grounded representationとagentic frameworkがこれらの課題を緩和することを示す。 - scene scale増大時の優位性持続を主張。 - 曖昧・実行不可能な指示への適切な応答も議論。 - 具体的な限界や失敗事例、計算コストの詳細は要旨からは不明。

6. 次に読むべき論文は?

- 要旨で参照/比較されている具体的なbaseline名は不明。 - 関連手法としてLLM-based robot task planning、ontology-grounded scene representation、agentic framework、spatial reasoning、task planningが挙げられる。 - 同分野の定番としてLLM-based planning、scene graph、ontology、task and motion planning (TAMP) 関連の研究を次に読むべき。 - コードはhttps://github.com/namhyeongwoo/OntoPlanで公開。

※ AIが要旨から生成した要約です。正確性は原文をご確認ください。

著者: Hyeongwoo Nam, Woongje Cho, Juwon Kim, Jongeun Choi

分類: cs.RO

原文アブストラクト

Large language model (LLM)-based robot task planning is promising for open-ended instruction following, but degrades on long-horizon tasks in large environments. When spatial information is conveyed to the LLM through text, the model can fail to capture spatial context, and token cost grows with environment size. Generating action sequences directly with an LLM also makes it difficult to satisfy the current world state and action preconditions. We address this with an ontology-grounded scene representation that aligns objects, spaces, relations, and states in a shared symbolic vocabulary for spatial reasoning and task planning, and with OntoPlan, an agentic framework that interprets instructions, selectively retrieves task-relevant information, formalizes goals and constraints, and produces executable plans. Across 150 general tasks spanning five indoor environments and three scene scales, OntoPlan achieves 0.89 average task success, compared with 0.27 for the strongest baseline, while using 18.1k total tokens per task on average, about 5.6$\times$ fewer than the most efficient baseline. These advantages persist as scene scale increases, whereas prior methods degrade more sharply in success and remain far more costly in tokens. OntoPlan also responds appropriately to ambiguous or infeasible instructions by asking follow-up questions or reporting insufficient information rather than committing to invalid plans. Code available at https://github.com/namhyeongwoo/OntoPlan.

関連論文

PR本紙発行元 EmplifAI