日本フィジカルAI新聞

世界のフィジカルAIを、日本語で。

週刊ニュースレター購読
自動運転計画arXiv:2609.38862

報酬誘導型選好最適化による自動運転のための効率的マルチモーダル計画

Efficient Multi-Modal Planning with Reward-Guided Preference Optimization for Autonomous Driving

シェア:XThreadsFacebookLINEはてブBluesky

スパースアンカーとオフセット精緻化を組み合わせたマルチモーダル軌道計画手法を提案し、報酬誘導型ファインチューニングで安全性を高め、NAVSIMベンチマークで精度と効率の両立を実現した。

著者: Chenglin Chen, Lujia Wang, Xinhu Zheng, Jun Ma, Haoang Li

分類: cs.RO, cs.AI

原文アブストラクト

Safe and efficient trajectory planning is essential in autonomous driving. However, existing end-to-end approaches often fall short in both computational efficiency and safety guarantees. Methods based on imitation learning suffer from causal confusion, while rule-based scoring approaches often incur heavy computational overhead and suffer from objective misalignment. Additionally, preference-based methods rely on strict pairwise annotations, limiting data utilization. To overcome these limitations, we propose EMPlan, an efficient multi-modal trajectory planning method powered by reward-guided fine-tuning. We design a hybrid architecture that combines sparse anchors with an offset refinement module for efficient multi-modal trajectory prediction. Sparse anchors provide coarse trajectory candidates with low latency, which are subsequently refined by the offset module for higher prediction accuracy. To enhance safety without incurring additional inference costs, we adopt a two-stage training paradigm consisting of pretraining and reward-guided fine-tuning. During fine-tuning, we leverage rule-based reward signals and unpaired preference supervision to refine the pretrained policy toward safer trajectory selection. We evaluate EMPlan on the non-reactive NAVSIM benchmark, where it strikes a favorable balance between planning accuracy and efficiency, demonstrating superior performance under real-time constraints.

関連論文

PR本紙発行元 EmplifAI