検索関連性を超えて:視覚言語運転のためのシーン接地型リスク含意
Beyond Retrieval Relevance: Scene-Grounded Risk Entailment for Vision-Language Driving
運転シーンの知覚事実から知識グラフとSWRL推論でリスク関係を導出し、VLMと拡散プランナーの判断を条件付けることで、検索ベースの安全知識より計画安全性を向上させた。
著者: Jiaxin Liu, Ruilin Yu, Liang Peng, Jingkai Wang, Chengxiang Zhao, Zhenxin Zhu, Bing Wang, Guang Chen, Hangjun Ye, Hong Wang, Jun Li
分類: cs.RO
原文アブストラクト
Retrieval-augmented generation (RAG) gives vision--language driving systems access to external safety knowledge, yet a retrieved risk rule may be relevant without applying to the current scene. A vision--language model (VLM) receiving such knowledge must ground objects, bind entities across time, and verify relations before deciding how to act, leaving the support for risk conclusions implicit. We address this relevance--applicability gap with a Driving-Risk Knowledge Graph (DRKG) and Semantic Web Rule Language (SWRL) reasoning stage before VLM decision-making. Structured perception instantiates scene facts, from which SWRL rules derive events and directed risk relations when their antecedents are jointly satisfied. Recognized events, bound risk relations, and semantic descriptions of activated rules form compact evidence that conditions the VLM and diffusion planner. In matched comparisons on nuReasoning, our method improved the nuReasoning planning score (NPS) by 1.30 points and the non-at-fault collision score (NC) by 2.76 points over the relevance retrieval-based baseline. These gains indicate that scene-applicable risk evidence improves safety-weighted planning relative to semantically retrieved risk knowledge.