日本フィジカルAI新聞

世界のフィジカルAIを、日本語で。

週刊ニュースレター購読
強化学習/制御arXiv:2609.28396

時空間チューブ報酬を用いた信号時相論理仕様の全クラスに対する実用的強化学習

Tractable Reinforcement Learning for Full Class of Signal Temporal Logic Specifications Using Spatiotemporal Tube Reward

シェア:XThreadsFacebookLINEはてブBluesky

信号時相論理(STL)で表された複雑な高レベル仕様を満たすため、時空間チューブの幾何学的性質を利用した時間認識型強化学習フレームワークを提案し、入力制約を厳守しつつ履歴不要で効率的に連続制御方策を学習する。

詳しい要約

1. どんなもの?

- 未知のダイナミクスと厳しいアクチュエータ制限下で動作するロボティクスシステム(非ホロノミックおよび劣駆動プラットフォームを含む)の制御問題を扱う。 - 高レベル仕様をSignal Temporal Logic (STL)で表現し、Spatiotemporal Tubes (STTs)の幾何学的性質を活用した時間認識型強化学習 (RL) フレームワークを提案。 - 完全なクラスのSTL仕様を時変幾何学的境界にマッピングし、スカラーロバストネス指標に依存せずに多次元システム状態を直接制約。 - 状態空間に時間を追加し、時間認識型Soft Actor-Critic (SAC) エージェントを連続的で幾何学認識型の報酬関数を用いて訓練。 - 履歴不要で計算効率の良い連続制御ポリシーを学習し、仕様のロバストな満足と入力制約の厳守を両立。

2. 先行研究と比べてどこがすごい?

- 従来の解析的STTコントローラは入力制約の強制が困難であったが、提案手法はこれを克服。 - 既存のRLアプローチはメモリ集約的な状態履歴に依存していたが、提案手法は履歴不要でネイティブにこの制限を克服。 - スカラー・ロバストネス指標に依存せず、時変幾何学的境界を直接制約として用いる点が新しい。 - 実行中に複雑な論理セマンティクスを明示的に評価する必要を排除し、計算効率を向上。

3. 技術・手法の肝は?

- STL仕様の論理的・時間的複雑さを時変幾何学的境界にマッピングし、多次元状態を直接制約。 - 状態空間に時間を追加して時間認識型SACエージェントを訓練。 - 連続的で幾何学認識型の報酬関数を設計し、実行中の論理セマンティクス評価を不要に。 - 履歴不要のアプローチで、入力制約を厳守しながら仕様のロバスト満足を保証する連続制御ポリシーを学習。

4. どうやって有効だと検証した?

- 要旨からは不明(具体的な検証方法や実験結果についての記述がない)。

5. 議論はある?

- 要旨からは不明(議論や限界についての記述がない)。

6. 次に読むべき論文は?

- 要旨で参照/比較されている研究:従来の解析的STTコントローラ、既存のRLアプローチ(状態履歴に依存するもの)。 - 関連手法:Signal Temporal Logic (STL)、Spatiotemporal Tubes (STTs)、Soft Actor-Critic (SAC)。 - 同分野の定番:強化学習を用いたSTL仕様の満足に関する研究(例:ロバストネス指標に基づくRL)。

※ AIが要旨から生成した要約です。正確性は原文をご確認ください。

著者: Vaishnavi Jagabathula, P Sangeerth, Pushpak Jagtap

分類: cs.RO, eess.SY

原文アブストラクト

This paper addresses the control problem for robotic systems, including non-holonomic and underactuated platforms operating under unknown dynamics and strict actuator limits to satisfy complex high-level specifications. We denote these high-level specifications using Signal Temporal Logic (STL) and propose a novel time-aware Reinforcement Learning (RL) framework that leverages the geometric properties of Spatiotemporal Tubes (STTs). While traditional analytical STT controllers often struggle to enforce input constraints, and existing RL approaches rely on memory-intensive state history, our method natively overcomes both limitations. By mapping the logical and temporal complexities of the full class of STL into time-varying geometric boundaries, we directly constrain the multidimensional system state without relying on scalar robustness metrics. Augmenting the state space with time, we train a time-aware Soft Actor-Critic (SAC) agent using a continuous, geometry-aware reward function that eliminates the need to explicitly evaluate complex logical semantics during execution. The proposed framework offers a history-free, computationally efficient approach to learn continuous control policies that ensure robust satisfaction of specifications while strictly adhering to system input constraints.

関連論文

PR本紙発行元 EmplifAI