日本フィジカルAI新聞

世界のフィジカルAIを、日本語で。

週刊ニュースレター購読
自動運転評価arXiv:2609.34440

スコアが目標になるとき:自動運転における指標の妥当性を再考する

When the Score Becomes the Target: Rethinking Metric Validity in Autonomous Driving

シェア:XThreadsFacebookLINEはてブBluesky

自動運転のベンチマークスコアを最適化目標にすると、スコア向上が運転改善の信頼できる証拠でなくなることを、実行・測定・サブスコア・集約の分解と閉ループ比較で示した。

著者: Morui Zhu, Deyuan Qu, Qi Chen, Kentaro Oguchi, Qing Yang

分類: cs.CV, cs.RO

原文アブストラクト

Driving benchmark scores are increasingly used not only for evaluation but also as optimization targets. This raises a fundamental question: do score gains remain reliable evidence of driving improvement once the score itself is optimized? We address this question by examining how the scoring process responds to changes in driving behavior and whether the resulting gains persist under repeated execution and replanning. We decompose the process into execution, measurement, subscore mapping, and aggregation. Controlled interventions reveal substantial behavioral changes that receive little score response because distinctions are omitted, thresholded, or attenuated between requested and executed motion. Closed-loop comparisons further show that optimization gains can reverse when the execution interface changes, demonstrating their dependence on how requests are executed and returned as feedback. Together, these findings connect the behavioral distinctions preserved by a metric to the conditions under which its gains transfer. Metric validity under optimization therefore requires examining both what the scoring process measures and how the optimized behavior is executed.

関連論文

PR本紙発行元 EmplifAI