日本フィジカルAI新聞

世界のフィジカルAIを、日本語で。

週刊ニュースレター購読
自動運転セキュリティ/LLMプランニングarXiv:2609.39969

TACTIC: 路側LiDAR攻撃のための時空間・文脈認識型LLM戦術プランニング

TACTIC: Temporal and Context-Aware LLM Tactical Planning for Roadside LiDAR Attacks

シェア:XThreadsFacebookLINEはてブBluesky

マルチモーダルLLMを用いて周囲の交通状況を認識し、路側LiDARへの攻撃戦術を状況に応じて選択・調整するフレームワークを提案。CARLA実験で100%の衝突率を達成した。

詳しい要約

1. どんなもの?

- 道路脇の LiDAR に対する物理攻撃を、周囲の交通状況に応じて適応的に計画する枠組み TACTIC を提案する研究。 - マルチモーダル大規模言語モデル (MLLM) を用い、攻撃者側の路側知覚スタックのみで車両状態と路側画像から交通文脈を推論する。 - 意味的シーングラフを構築し、push-away と phantom-obstacle braking の2つのプリミティブを選択・設定する。 - gray-box 脅威モデルを前提とし、被害者 LiDAR の点群や内部処理にはアクセスしない。

2. 先行研究と比べてどこがすごい?

- 従来の物理 LiDAR 攻撃は固定プリミティブと手動選択パラメータで評価されることが多く、周囲交通への依存を考慮していなかった。 - TACTIC はシーン依存の戦術計画により、固定攻撃ポリシーでは見逃しうる文脈依存の LiDAR 故障モードを顕在化させる。 - 280 回のランダム化 CARLA 試行で、完全ポリシーは衝突率 100% を達成し、固定ルール 35%、ランダム選択 60%、デフォルトパラメータのモード選択を行う制限付き LLM 75% を上回った。

3. 技術・手法の肝は?

- 局所知覚が車両のメートル法状態を提供し、MLLM がそれを路側画像と組み合わせて関係的交通文脈を推論、意味的シーングラフを構築する。 - この表現に基づき push-away (先行車の知覚距離をずらす) と phantom-obstacle braking (障害物注入で緊急制動を誘発) を選択・設定する。 - 計測された交通状態と経験的に較正された制約により、生成戦術を物理的に実行可能な動作領域に接地させる。 - MLLM の遅延に対応するため、推論と実行を非同期に重ね、高頻度の局所知覚がシーン変化を検出して再計画をトリガーする。

4. どうやって有効だと検証した?

- 280 回のランダム化 CARLA 試行で評価し、完全ポリシーの衝突率 100% を確認。 - 比較対象は固定ルール 35%、ランダム選択 60%、制限付き LLM 75%。 - 物理情報と画像の統合入力は成功率 100% で、物理計測のみ 65%、画像のみ 75% を上回った。 - 非同期 Δ refresh によりシーン変化への応答が 7.4 秒から 2.0 秒に短縮された。

5. 議論はある?

- シーン依存の戦術計画が、固定攻撃ポリシーでは見逃しうる文脈依存の LiDAR 故障モードを明らかにできると主張。 - 物理情報と画像の統合、非同期再計画の有効性が示唆される。 - ただし要旨からは、攻撃の倫理的・防御的含意や限界、実世界への一般化可能性についての議論は不明。

6. 次に読むべき論文は?

- 要旨で参照・比較されている研究: 固定ルール、ランダム選択、制限付き LLM (mode selection with default parameters)。 - 関連手法: 物理 LiDAR 攻撃、MLLM を用いた計画、CARLA シミュレーション。 - 同分野の定番として、LiDAR スプーフィング/ジャミング攻撃、敵対的知覚、自動運転の安全評価に関する研究が次に読むべき候補。

※ AIが要旨から生成した要約です。正確性は原文をご確認ください。

著者: Yiming Gao, Shaocheng Luo

分類: cs.RO, cs.AI, eess.SY

原文アブストラクト

Physical LiDAR attacks are often evaluated using fixed primitives and manually selected parameters, despite their strong dependence on surrounding traffic. We present TACTIC, a scene-aware framework that uses a multimodal large language model (MLLM) to coordinate state-adaptive roadside LiDAR attacks. Under a gray-box threat model, TACTIC relies only on an attacker-operated roadside perception stack, without accessing the victim LiDAR's native point clouds or internal processing. Local perception provides metric vehicle states, while the MLLM combines these measurements with roadside imagery to infer relational traffic context and construct a semantic scene graph. Based on this representation, TACTIC selects and configures two complementary primitives: \emph{push-away}, which shifts the perceived range of a lead vehicle, and \emph{phantom-obstacle braking}, which triggers emergency braking through obstacle injection. Measured traffic states and empirically calibrated constraints ground the generated tactics in physically feasible operating regions. To accommodate MLLM latency, TACTIC overlaps reasoning and execution asynchronously while high-rate local perception detects scene changes and triggers replanning. Across 280 randomized CARLA trials, the full policy achieves a 100% collision rate, compared with 35% for a fixed rule, 60% for random selection, and 75% for a restricted LLM using mode selection with default parameters. Joint physical-and-image input achieves 100% success, versus 65% with physical measurements alone and 75% with imagery alone, while asynchronous $Δ$ refresh reduces scene-mutation response from 7.4 s to 2.0 s. These results show that scene-dependent tactical planning can expose context-sensitive LiDAR failure modes that fixed attack policies may miss.

PR本紙発行元 EmplifAI