日本フィジカルAI新聞

世界のフィジカルAIを、日本語で。

週刊ニュースレター購読
自動運転/言語条件付き計画arXiv:2609.38028

doPlan: 自動運転における多段階言語条件付き計画のための可変ホライズンデータセット

doPlan: A Variable-Horizon Dataset for Multi-Stage Language-Conditioned Planning in Autonomous Driving

シェア:XThreadsFacebookLINEはてブBluesky

乗客の自然言語指示を長期的なタスク文脈として扱う自動運転データセットdoPlanを構築し、既存の言語条件付き運転モデルが指示の方向性に沿った行動をとれないことを示した。

詳しい要約

1. どんなもの?

- 自動運転における乗客の自然言語指示を、持続的なタスク文脈として扱う研究用データセット - 名称は doPlan。nuPlan を基盤に構築 - 5,154件の人手注釈による乗客指示を含む - 累積169.1時間の指示整合文脈、ユニーク走行50.9時間 - 注釈ウィンドウは30.0〜508.8秒 - immediate, deferred, event-conditioned, persistent, multi-stage な意図を収録 - データセット・注釈インターフェース・関連資源を GitHub で公開

2. 先行研究と比べてどこがすごい?

- 既存の言語対応運転データセットは短く局所的な相互作用が中心 - 長期的な乗客意図は比較的未開拓だった - doPlan は初の公開・人手注釈・実世界データセットと主張 - 乗客言語を『持続的タスク文脈』として扱う点が新しい - 複数段階にまたがる意図や将来事象条件付き意図を注釈対象に含む

3. 技術・手法の肝は?

- nuPlan 上に構築し、実走行データに人手注釈を付与 - 注釈ウィンドウを30.0〜508.8秒と長くとる設計 - immediate, deferred, event-conditioned, persistent, multi-stage の意図カテゴリを定義 - 指示と整合する文脈を累積169.1時間分収録 - 言語条件付き運転モデル4種を評価する枠組みを提供 - 将来の maneuver と評価時点の対応を分析する指標を導入

4. どうやって有効だと検証した?

- 言語条件付き運転モデル4種を評価 - 乗客言語への感度が、要求方向と一致する挙動に必ずしも結びつかないことを発見 - 将来 maneuver が対応する2,161例を分析 - 最初の関連 maneuver は評価時点から中央値24.6秒後 - モデルの一般的な5秒予測 horizon 内に収まるのは9.8%のみ - 持続的意図と逐次計画判断の接続が必要と示唆

5. 議論はある?

- 乗客言語への感度と、要求に沿った挙動の一致性は別問題 - 長期的意図はモデルの短期予測 horizon を大きく超える - 未解決ゴールを保持し、変化する scene に接地し、複数段階で追跡する必要 - planner が将来ゴールの現在計画への関連性をいつ判断するかが課題 - データセットの限界や一般化可能性は要旨からは不明

6. 次に読むべき論文は?

- nuPlan を基盤データセットとして参照 - 言語条件付き運転モデル4種(具体名は要旨からは不明) - 同分野の定番として language-conditioned driving, autonomous driving planning, natural language instruction grounding 関連研究

※ AIが要旨から生成した要約です。正確性は原文をご確認ください。

著者: Parthib Roy, Yash Tandon, Marcus Blennemann, Giovanni Tapia Lopez, Angel Martinez-Sanchez, Mohan M. Trivedi, Ross Greer

分類: cs.RO, cs.AI, cs.CV, cs.LG

原文アブストラクト

Autonomous vehicles interacting with passengers through natural language must reason beyond immediate commands. Passenger intent may span multiple stages of behavior, depend on future events, refer to surrounding agents or landmarks, and remain relevant as driving conditions evolve. Existing language-enabled driving datasets largely focus on short, localized interactions, leaving these longer-horizon forms of passenger intent comparatively underexplored. We introduce doPlan, to our knowledge the first publicly available, human-annotated real-world dataset designed to study passenger language as persistent task context. Built on nuPlan, doPlan contains 5,154 human-written passenger instructions spanning 169.1 hours of cumulative instruction-aligned context over 50.9 hours of unique driving, with annotation windows ranging from 30.0 to 508.8 s. The annotations capture immediate, deferred, event-conditioned, persistent, and multi-stage passenger intent. The dataset, annotation interface, and supporting resources are publicly available at https://github.com/Mi3-Lab/doPlan. We evaluate four language-conditioned driving models and find that sensitivity to passenger language does not reliably translate into behavior consistent with the requested direction. More broadly, among 2,161 examples with a matched future maneuver, the first associated maneuver occurs a median of 24.6 s after the evaluation point, and only 9.8% occur within the models' common 5 s prediction horizon. These findings highlight the need to connect persistent passenger intent with successive planning decisions. doPlan provides a setting for studying how unresolved goals can be retained, grounded in evolving scenes, and tracked across multiple stages, including how a planner determines when a future goal becomes relevant to the current plan.

PR本紙発行元 EmplifAI