日本フィジカルAI新聞

世界のフィジカルAIを、日本語で。

週刊ニュースレター購読
敵対的攻撃arXiv:2610.06814

TAPDreamer: ワールドアクションモデルに対する転移可能な敵対的パッチ

TAPDreamer: Transferable Adversarial Patches for World Action Models

シェア:XThreadsFacebookLINEはてブBluesky

公開エンコーダのみを用いて、タスクや行動アーキテクチャをまたいで転移する固定の敵対的パッチを生成し、ワールドアクションモデルの成功率をほぼゼロに低下させる攻撃手法を提案した。

詳しい要約

1. どんなもの?

World Action Model(環境の将来を予測しロボット制御の基盤となるモデル)に対する敵対的攻撃 TAPDreamer を提案する研究。 - カメラ入力を操作し、タスクや行動ポリシー間で共有される視覚表現を破壊する。 - 単一の固定局所パッチ(入力の約6.5%)を1つ凍結して使う転移可能な攻撃。 - 対象ポリシーへのクエリを必要としない。

2. 先行研究と比べてどこがすごい?

既存の World Action Model への攻撃は、victim の行動や予測未来に対して最適化するため対象モデルの出力へのアクセスが必要だった。 - TAPDreamer は公開 encoder のみを使い、対象ポリシーへのクエリ不要。 - タスクと行動アーキテクチャをまたいで転移する固定パッチを構築できる点が新しい。

3. 技術・手法の肝は?

鍵となる洞察は、パッチが誘起する attention weight の変化と value vector の相互作用が、パッチ領域をはるかに超えてほぼ同一の表現シフトを伝播させ、それがタスク観測間で安定するという点。 - これを利用し、1つの source task の6フレームを用いて clean と patched の encoder 表現間の global L1 distance を最大化する。 - 結果として固定された局所摂動パッチを得る。

4. どうやって有効だと検証した?

closed-loop 評価で検証。 - ベンチマークごとに1つの凍結パッチ(入力の約6.5%)を使用。 - FastWAM の成功率を LIBERO 40タスクで 97.7%→0.0%、RoboTwin 50タスクで 90.8%→0.0% に低下。 - 同条件の random patch では 81.5%、79.2% の成功率が残る。 - 同じパッチで DreamWAM 2構成が 2.1%、0.8%、Motus が 10.0% に低下。

5. 議論はある?

下流の行動生成のみを保護しても不十分だと主張。 - World Action Model の防御は、共有視覚 encoder を持続的な局所摂動から守る必要がある。 - 攻撃の転移性と表現シフトの広域性が防御設計上の課題として議論される。

6. 次に読むべき論文は?

要旨で参照・比較されている研究:既存の World Action Model への攻撃(victim の行動や予測未来に最適化する手法)。 - 評価対象モデル:FastWAM、DreamWAM、Motus。 - ベンチマーク:LIBERO、RoboTwin。 - 関連する基盤:World Model、World Action Model、visual encoder、attention weight、value vector。

※ AIが要旨から生成した要約です。正確性は原文をご確認ください。

著者: Xuanyu Lu, Fengqing Jiang, Kaiyuan Zheng, Yichen Feng, Yaorui Ding, Yuetai Li, Zhen Xiang, Bhaskar Ramasubramanian, Basel Alomair, Luyao Niu, Radha Poovendran

分類: cs.CV, cs.AI, cs.RO

原文アブストラクト

World models learn to predict how their environment will evolve, making them an important foundation for general-purpose robotic control. Yet world action models depend on camera inputs whose manipulation can corrupt the visual representations used across tasks and action policies. Existing attacks on these models optimize against the victim's actions or predicted futures and therefore require access to target-model outputs. In this paper, we propose an attack, TAPDreamer, against world action models that instead uses a public encoder alone to construct a fixed local perturbation that transfers across tasks and action architectures. TAPDreamer requires no target-policy queries. Our key insight is that interactions between patch-induced changes in attention weights and value vectors broadcast a nearly identical representation shift far beyond the patch footprint, and this shift remains stable across task observations. Guided by this insight, TAPDreamer uses six frames from one source task to maximize the global L1 distance between clean and patched encoder representations. In closed-loop evaluation, one frozen patch per benchmark, covering about 6.5% of the input, reduces FastWAM's success rate from 97.7% to 0.0% across 40 LIBERO tasks and from 90.8% to 0.0% across 50 RoboTwin tasks; matched random patches retain 81.5% and 79.2% success. The same patches reduce success to 2.1% and 0.8% on two DreamWAM configurations and to 10.0% on Motus. These results show that protecting downstream action generation alone is insufficient: defenses for world action models must also secure shared visual encoders against persistent local perturbations.

関連論文

PR本紙発行元 EmplifAI