日本フィジカルAI新聞

世界のフィジカルAIを、日本語で。

週刊ニュースレター購読
自動運転arXiv:2609.31383

終点制約付き軌道最適化によるエンドツーエンド運転モデルの誘導

Guiding End-to-End Driving Models with Endpoint-Constrained Trajectory Optimization

シェア:XThreadsFacebookLINEはてブBluesky

エンドツーエンド運転ポリシーの出力軌道を、予測終点を保ちながら走行履歴に基づいて中間ウェイポイントを再整形する軽量後処理ECOを提案し、閉ループ性能を改善した。

詳しい要約

1. どんなもの?

- 対象は end-to-end driving policies の open-loop 学習と closed-loop 実行の不一致 - 要因として waypoint-based supervision と displacement metrics が中間軌道の物理的整合性や controller 追従性を保証しない点を指摘 - 提案は Endpoint-Constrained Optimization (ECO) - 軽量な postprocessing layer - 車両の実行履歴に軌道を固定 - policy の予測 endpoint を保持 - 中間 waypoint を再整形し feasibility を改善 - map、privileged simulator state、追加学習は不要 - waypoint を出力する幅広い policy と controller の間に挿入可能

2. 先行研究と比べてどこがすごい?

- 従来は open-loop/closed-loop gap の要因として covariate shift と causal confusion が主に研究 - 本研究は補完的要因として waypoint-based supervision と displacement metrics の問題を同定 - 中間 waypoint の不整合が集中し、予測 endpoint は比較的信頼できると観察 - ECO は追加学習や特権情報なしで、既存の生成・回帰ベース policy に後付け可能 - 6つの policy すべてで closed-loop score を改善 - 改善幅は base plan が motion limits に違反する頻度が高いほど増加傾向

3. 技術・手法の肝は?

- Endpoint-Constrained Optimization (ECO) を導入 - 軽量な postprocessing layer として動作 - 軌道を車両の実行履歴に anchor - policy が予測した endpoint を保存 - 中間 waypoint を再整形して feasibility を向上 - map、privileged simulator state、追加 training は不要 - waypoint を出力する policy と controller の間に挿入可能

4. どうやって有効だと検証した?

- 2つの closed-loop simulator で評価 - 6つの generative および regression-based driving policy すべてで aggregate closed-loop score が改善 - HUGSIM では VaVAM を 18.1 から 31.0 HD-Score へ改善 (+71%) - HUGSIM Closed-Loop Driving Challenge で 1st place を獲得 - AlpaSim では VaVAM と DiffusionDrive の scene score をそれぞれ 123% と 22% 改善 - 改善は base plan の motion limits 違反頻度が高いほど大きい傾向

5. 議論はある?

- open-loop/closed-loop gap の新たな補完的要因を提示 - 中間 waypoint の幾何学的修復が closed-loop 性能を大幅に改善し得ることを示す - endpoint を変更せず中間軌道のみを修復するアプローチの有効性を主張 - 幅広い end-to-end driving model に適用可能と示唆 - 限界や失敗事例、計算コスト、リアルタイム性などの詳細は要旨からは不明

6. 次に読むべき論文は?

- VaVAM - DiffusionDrive - HUGSIM - AlpaSim - covariate shift や causal confusion に関する先行研究 - waypoint-based supervision や displacement metrics を用いる end-to-end driving 手法

※ AIが要旨から生成した要約です。正確性は原文をご確認ください。

著者: Brayden Zhang, Mahsa Golchoubian, Igor Gilitschenski, Boris Ivanovic, Kashyap Chitta

分類: cs.RO, cs.AI, cs.CV, cs.LG

原文アブストラクト

End-to-end driving policies are commonly trained through open-loop behavior cloning, yet they must ultimately operate in closed-loop when deployed on a vehicle, creating a fundamental mismatch between training and execution. Beyond the commonly studied effects of covariate shift and causal confusion, we identify a complementary factor for this open-loop/closed-loop gap: waypoint-based supervision and displacement metrics do not ensure that the intermediate trajectory is physically coherent or easy for the controller to track. We observe that these inconsistencies concentrate primarily at intermediate waypoints, while the predicted endpoint remains comparatively reliable. Based on this observation, we introduce Endpoint-Constrained Optimization (ECO), a lightweight postprocessing layer that anchors the trajectory to the vehicle's executed history, preserves the policy's predicted endpoint, and reshapes the intermediate waypoints to improve feasibility. ECO requires no map, privileged simulator state, or additional training, and can be inserted between a broad range of waypoint-emitting policies and their controllers. Across two closed-loop simulators, it improves the aggregate closed-loop score of all six evaluated generative and regression-based driving policies, and the gains tend to increase with how often the base plans violate motion limits. On HUGSIM, ECO improves VaVAM from 18.1 to 31.0 HD-Score (+71%), achieving 1st place on the HUGSIM Closed-Loop Driving Challenge. Similarly, on AlpaSim, ECO increases the scene scores of VaVAM and DiffusionDrive by 123% and 22%, respectively. These results show that for a broad collection of end-to-end driving models, repairing the intermediate geometry of predicted trajectories without changing the policy's predicted endpoint can substantially improve closed-loop performance.

関連論文

PR本紙発行元 EmplifAI