日本フィジカルAI新聞

世界のフィジカルAIを、日本語で。

週刊ニュースレター購読
VLAarXiv:2609.26184

静かなる妨害:LLM搭載ロボットシステムに対する内部状態トリガーバックドア攻撃

Silent Sabotage: Internal State Triggered Backdoor Attacks on LLM-Powered Robotic Systems

シェア:XThreadsFacebookLINEはてブBluesky

LLMで制御されるロボットに、外部刺激ではなくロボット自身の過去の行動履歴をトリガーとする隠れたバックドアを埋め込めることを示し、通常時は性能を保ちつつ特定条件下で停止や衝突を引き起こせることを実験で明らかにした。

詳しい要約

1. どんなもの?

- LLMを組み込んだロボット制御システムに対する新たなセキュリティリスクを提示する研究。 - 攻撃者がLLMベースのロボットコントローラの指示を操作し、ロボット自身の過去の行動系列をトリガーとするバックドアを埋め込む。 - このバックドアは通常時は休眠し、ロボットの有用性を保つが、特定の稀な行動系列後に起動し、停止や衝突などの悪意ある行動を引き起こす。 - シミュレーション環境で多様なロボットとLLMを用いて実験し、攻撃の有効性と検出困難性を示す。

2. 先行研究と比べてどこがすごい?

- 従来のLLMバックドア研究は、特定の単語、視覚オブジェクト、環境状態など外部刺激をトリガーとする攻撃に焦点を当てていた。 - 本研究は、エージェント自身の操作ロジック内部(過去の行動系列)をトリガーとする、より潜在的な脆弱性クラスを初めて包括的に研究。 - 外部刺激に依存しないため、従来の検出手法では見逃されやすく、攻撃の成功率が極めて高く、検出が非常に困難である点が新しい。

3. 技術・手法の肝は?

- 攻撃者はLLMベースのロボットコントローラへの指示を操作し、特定の稀な過去行動系列をトリガーとして埋め込む。 - バックドアは通常動作では休眠状態を維持し、ロボットのタスク性能を損なわない。 - トリガーとなる行動系列が発生すると、停止や衝突などの悪意ある行動を誘発する。 - シミュレーション環境で様々なロボットとLLMを組み合わせて評価。

4. どうやって有効だと検証した?

- シミュレーション環境で多様なロボットとLLMを用いた実験を実施。 - 履歴ベースの攻撃がほぼ完璧な攻撃成功率を達成し、かつ検出が極めて困難であることを示した。 - 通常運用時のロボットの有用性が維持されることも確認。

5. 議論はある?

- 自律システムにおける重大かつ未対処の脆弱性を明らかにし、エージェントの内部状態を考慮したセキュリティ対策の緊急必要性を強調。 - 攻撃の検出困難性や、既存の防御手法では対応できない可能性について議論。 - 具体的な緩和策や倫理的影響については要旨からは不明。

6. 次に読むべき論文は?

- 要旨で参照/比較されている研究:外部刺激トリガーのLLMバックドア攻撃(特定の単語、視覚オブジェクト、環境状態)。 - 関連手法:LLM backdoor attacks、robotic control systems、autonomous agents。 - 同分野の定番:LLM security、backdoor detection、robot learning。

※ AIが要旨から生成した要約です。正確性は原文をご確認ください。

著者: Doniyorkhon Obidov, Shivayogi Akki, Tan Chen, Kaichen Yang

分類: cs.RO, cs.AI, cs.CR

原文アブストラクト

The integration of Large Language Models (LLMs) into robotic control systems is enabling a new generation of autonomous agents capable of complex reasoning and planning. While this paradigm shift accelerates progress, it also introduces novel security risks that remain largely unexplored. Current research into LLM backdoors has focused on attacks triggered by external stimuli, such as specific words, visual objects, or environmental states. These attacks, while potent, overlook a more insidious class of vulnerability where the trigger is internal to the agent's own operational logic. This paper presents the first comprehensive study of history-based backdoor attacks on LLM-powered robotic systems. We demonstrate that an attacker can embed a stealthy backdoor into an LLM-based robot controller by manipulating its instructions. This backdoor is triggered not by an external cue, but by a specific, rare sequence of the robot's own past actions. It remains dormant during normal operation, preserving the robot's utility, but can be activated to induce a malicious behavior, such as a complete stop or a collision. Our experiments, conducted in a simulated environment with a variety of robots and LLMs, show that this history-based attack is highly effective, achieving a near-perfect attack success rate while remaining exceptionally difficult to detect. These findings reveal a critical and previously unaddressed vulnerability in autonomous systems and underscore the urgent need for security measures that account for an agent's internal state.

関連論文

PR本紙発行元 EmplifAI