CyberWorld: 自律型サイバー防御のためのサンプル効率の高い世界モデル
CyberWorld: World Models for Sample-Efficient Autonomous Cyber Defense
サイバー防御の強化学習にDreamer型の世界モデルを導入し、ネットワークのグラフ表現などを用いて潜在ダイナミクスを学習することで、モデルフリーPPOよりはるかに少ない環境ステップで攻撃戦略に対応できることを示した。
著者: Ryozo Masukawa, Sanggeon Yun, Raheeb Hassan, Hyunwoo Oh, SungHeon Jeong, Mohsen Imani
分類: cs.LG, cs.AI, cs.CR
原文アブストラクト
Deep reinforcement learning has become a prominent approach to autonomous cyber defense. Existing methods are predominantly model-free and consequently require extensive environment interaction. World models provide an alternative by learning predictive dynamics and optimizing policies through imagined trajectories, yielding substantial gains in sample efficiency in robotics and embodied control. Extending this paradigm to cybersecurity raises a fundamental question: what should constitute the "world" in a cyber world model? We introduce CyberWorld, a Dreamer-style world modeling framework that learns latent cyber dynamics from vector, graph, textual, and multimodal representations of the defended network. Across all four scoreable CyberWheel attack strategies, the graph-based CyberWorld variant exceeds a strategy-agnostic control after 3.6k-15.8k environment steps, compared with millions of steps required by model-free PPO. Across representation choices, graph structure provides greater robustness under topology-dependent attacks, while simpler representations remain competitive in overall performance. Among successful runs, the number of episodes required to reach the control remains approximately constant as network size increases from 15 to 100 hosts. These results establish learned cyber dynamics as a sample-efficient and scalable basis for autonomous defense, and identify world representation as a central design axis for robustness and scalability.
関連論文
- 効率的な一次強化学習のための局所・大域世界モデルの結合強化学習/世界モデル
- 因果関係を考慮した強化学習のためのオブジェクト中心の世界モデル強化学習/世界モデル
- 注目を学ぶ:部分観測強化学習における構造的注意機構による情報履歴の優先強化学習/世界モデル
- DreamerV3-XP:不確実性推定による探索の最適化強化学習/世界モデル
- 非キュレーション�データで世界モデルを導く効率的強化学習強化学習/世界モデル