日本フィジカルAI新聞

世界のフィジカルAIを、日本語で。

週刊ニュースレター購読
強化学習制御arXiv:2609.11815

FDI攻撃と外乱下の入力制約付き未知非線形システムに対する事前定義時間ロバスト積分強化学習:完全データ駆動アプローチ

Predefined-Time Resilient Integral Reinforcement Learning for Input-Constrained Unknown Nonlinear Systems Under FDI Attacks and Disturbances: A Fully Data-Driven Approach

シェア:XThreadsFacebookLINEはてブBluesky

システムの動力学が未知で入力制約・外乱・敵対的信号がある非線形システムに対し、収束時間を事前に指定できる積分強化学習ベースの最適制御法を提案し、2リンクロボットマニピュレータで有効性を示した。

詳しい要約

1. どんなもの?

- 入力制約付き未知非線形システムの最適制御を扱う。 - FDI攻撃と外乱が存在する状況を想定。 - 設計者が事前に収束時間を指定できる学習ベース制御法を提案。 - integral reinforcement learning (IRL) フレームワークを採用。 - システムダイナミクスの正確な知識を必要としない。 - アクチュエータ制約を常に満たす。

2. 先行研究と比べてどこがすごい?

- 従来のIRLはシステムダイナミクスや持続的励起を必要とすることが多い。 - 提案法は現在と記録データを組み合わせ、持続的励起を不要とする。 - 収束時間を事前に指定できる点が特徴。 - 外乱と敵対的チャネル下でも実用的な事前定義時間収束を保証。 - 入力制約を考慮しつつ学習を行う点で先行研究と異なる。

3. 技術・手法の肝は?

- integral reinforcement learning (IRL) を基盤とする。 - 現在と記録データを組み合わせてcriticを訓練。 - 学習ゲインを指定収束期限から直接選択。 - Lyapunov解析により状態とcriticの結合系の実用的事前定義時間収束を証明。 - 入力制約を満たすように制御入力を設計。 - 未知ダイナミクスに対応するためデータ駆動型アプローチを採用。

4. どうやって有効だと検証した?

- 2リンクロボットマニピュレータの安定化制御で有効性を検証。 - シミュレーションにより提案法の性能を確認。 - 外乱とFDI攻撃下での収束を評価。 - 入力制約の遵守を確認。 - 具体的な評価指標は要旨からは不明。

5. 議論はある?

- 外乱と敵対的チャネル下での実用的事前定義時間収束を理論的に保証。 - 持続的励起を必要としない点が利点。 - 入力制約を常に満たすことを強調。 - 計算負荷や実装上の課題については要旨からは不明。 - 他のシステムへの適用可能性は要旨からは不明。

6. 次に読むべき論文は?

- 要旨で参照/比較されている研究は明示されていない。 - 関連手法として integral reinforcement learning (IRL) や Lyapunov解析が挙げられる。 - 同分野の定番として adaptive dynamic programming (ADP) や robust control が考えられる。 - 具体的な論文は要旨からは不明。

※ AIが要旨から生成した要約です。正確性は原文をご確認ください。

著者: Tien Dat Vu, Minh Doan

分類: eess.SY

原文アブストラクト

This paper investigates optimal control for nonlinear systems with unknown dynamics, input constraints, disturbances, and adversarial signals. The objective is to develop a learning-based control method that allows the designer to prescribe the desired convergence time in advance. An integral reinforcement-learning framework is proposed to avoid requiring exact knowledge of the system dynamics while ensuring that the control input always satisfies the actuator constraints. Current and recorded data are combined to train the critic without requiring persistent excitation. The learning gain is selected directly from the prescribed convergence deadline. Lyapunov analysis is then used to establish practical predefined-time convergence of the coupled state-critic system in the presence of disturbances and adversarial channels. The effectiveness of the proposed method is validated through the stabilization control of a two-link robot manipulator.