日本フィジカルAI新聞

世界のフィジカルAIを、日本語で。

週刊ニュースレター購読
VLAarXiv:2609.30428

自然言語の曖昧なタスクにおける文脈的不確実性の能動的解決

Actively Resolving Contextual Uncertainty for Underspecified Tasks in Natural Language

シェア:XThreadsFacebookLINEはてブBluesky

LLMとオンライン構築した言語埋め込みマップを組み合わせ、閉ループ環境相互作用を通じて曖昧な指示の文脈的不確実性を能動的に解消するフレームワークCLUEを提案し、実機Spotで検証した。

詳しい要約

1. どんなもの?

- 自然言語で与えられた underspecified なタスクを、不慣れな環境で実行するロボットのための枠組み CLUE を提案。 - タスク成功条件・関連情報・その情報の所在が不明な「contextual uncertainty」を、閉ループの環境相互作用を通じて能動的に解消する。 - LLM 由来の policy で仮説と計画を生成し、オンライン構築される language-embedded map で接地する。 - Boston Dynamics Spot に実装し、屋内・屋外3環境・15タスクで評価。

2. 先行研究と比べてどこがすごい?

- 多くの language-conditioned policy は目標が well-specified で、事前 map によりタスク関連情報が与えられると仮定する。 - CLUE は underspecified なタスクと未知環境を前提に、成功条件・関連情報・情報の所在を同時に推論する。 - 閉ループフィードバックのない LLM-enabled planner を 4 倍の成功率で上回る。 - oracle policy との差は 7 パーセントポイント以内。 - 単に language-enriched map を構築・照会するだけでは不十分で、成功率は約 3 分の 1、VLM tokens は 10 倍以上必要。

3. 技術・手法の肝は?

- LLM-derived policy がタスク関連概念と候補計画を仮説として生成。 - オンラインで構築される language-embedded map を用いて仮説を行動に接地。 - policy は閉ループ環境相互作用を通じて仮説を逐次評価し、新情報に応じて計画を洗練。 - これにより contextual uncertainty を能動的に解消する。

4. どうやって有効だと検証した?

- Boston Dynamics Spot を用い、屋内・屋外の 3 実環境で 15 タスクを実行。 - タスクは object disambiguation、functional inference、occlusion reasoning を要求。 - CLUE は oracle policy の成功率から 7 パーセントポイント以内を達成。 - 閉ループフィードバックなしの LLM-enabled planner を 4 倍上回る。 - language-enriched map の単純な構築・照会は成功率約 3 分の 1、VLM tokens 10 倍以上であることを示す追加実験。

5. 議論はある?

- 要旨からは、限界や失敗事例、計算コスト、スケーラビリティに関する議論は不明。 - 単純な language-enriched map 照会では複雑な contextual planning を解けないことを示唆。 - 閉ループフィードバックの重要性を主張。

6. 次に読むべき論文は?

- 要旨で比較されている LLM-enabled planner without closed-loop feedback。 - language-enriched map を構築・照会するアプローチ。 - oracle policy。 - 関連手法として language-conditioned policy、VLM、LLM-derived policy が挙げられる。

※ AIが要旨から生成した要約です。正確性は原文をご確認ください。

著者: Zachary Ravichandran, Jonathan Diller, Fernando Cladera, Varun Murali, George J. Pappas, Vijay Kumar

分類: cs.RO, cs.AI

原文アブストラクト

Foundation models provide robots with the ability to interpret natural language and reason about environmental context, yet most language-conditioned policies assume that goals are well-specified and that task-relevant information is provided upfront via a prior map. Operating in unfamiliar environments with underspecified tasks entails high contextual uncertainty: the robot must jointly infer what constitutes task success, what constitutes relevant information, and where (or whether) that information exists. We address these limitations via CLUE (Closed-Loop contextual Uncertainty rEsolution), a framework for actively resolving contextual uncertainty given underspecified tasks in natural language. CLUE uses an LLM-derived policy to hypothesize task-relevant concepts and potential plans. It then uses a language-embedded map, which is constructed online, to ground these hypotheses into actions. The policy sequentially evaluates hypotheses via closed-loop environment interaction and refines its plans as it gathers new information. We deploy CLUE on a Boston Dynamics Spot across three real indoor and outdoor environments spanning 15 tasks that require object disambiguation, functional inference, and occlusion reasoning. CLUE achieves a success rate within 7 percentage points of an oracle policy and outperforms an LLM-enabled planner without closed-loop feedback by a 4x margin. Supporting experiments demonstrate that simply building and then querying a language-enriched map is insufficient to resolve complex contextual planning tasks; these approaches achieve roughly one third the success rate of CLUE while requiring over 10x more VLM tokens. We provide additional information at https://zacravichandran.github.io/CLUE.

関連論文

PR本紙発行元 EmplifAI