日本フィジカルAI新聞

世界のフィジカルAIを、日本語で。

週刊ニュースレター購読
自動運転arXiv:2609.33190

FocusDrive:視覚的フォーカスを用いた自動運転の推論

FocusDrive: Reasoning with Visual Focus for Autonomous Driving

シェア:XThreadsFacebookLINEはてブBluesky

自動運転の計画推論に、注視すべき物体とその位置を画像パッチ参照で明示する視覚フォーカス表現を導入し、注視予測とend-to-end計画性能を向上させた研究。

著者: Zhiyuan Liu, Zehong Ke, Yuanxin Tian, Hao Cheng, Jinhao Li, Yining Xing, Yanbo Jiang, Zhenhua Xu, Wenhao Yu, Jianqiang Wang

分類: cs.CV, cs.RO

原文アブストラクト

Driving decisions depend on both where to focus and how to act on what is seen. Effective driving reasoning must establish which objects matter, where they are, and how they inform the intended action. Text-based rationales can describe a driving response while leaving its correspondence to specific visual evidence implicit. Visual focus provides a concrete starting point for this connection by identifying what matters in the scene and where it is. We propose FocusDrive, a structured multimodal reasoning framework that organizes end-to-end planning around explicit visual focus. It pairs descriptions of decision-relevant objects with image-patch references, bringing explicit visual focus into the reasoning that generates driving plans and trajectories. We first assess this focus representation through driver gaze prediction, then investigate its role in planning reasoning using driving-focus annotations within existing NAVSIM training scenes. Experiments on W3DA and NAVSIM demonstrate competitive gaze prediction and end-to-end planning performance, with FocusDrive improving over text-based chain-of-thought. These results support visual focus as an effective link between scene understanding and driving action.

関連論文

PR本紙発行元 EmplifAI