ポテンシャルベース報酬整形と制御リアプノフ・バリア関数によるUAVナビゲーションのゼロショット安全・時間効率化
Zero-Shot, Safe and Time-Efficient UAV Navigation via Potential-Based Reward Shaping, Control Lyapunov and Barrier Functions
強化学習にポテンシャルベース報酬整形と制御リアプノフ・バリア関数を統合し、UAVのミッション時間短縮と安全性保証を両立するナビゲーション手法を提案した。
著者: Ashik Abrar Naeem, Mohammad Ariful Haque
分類: eess.SY, cs.LG, cs.RO, cs.SY
原文アブストラクト
Autonomous navigation and obstacle avoidance remain a core challenge of modern Unmanned Aerial Vehicles (UAVs). While traditional control methods struggle with the complexity and variability of the environment, reinforcement learning (RL) enables UAVs to learn adaptive behaviors through interaction with the environment. Existing research with RL prioritizes the mission success at the expense of mission time and safety of UAVs. This study integrates Potential Based Reward Shaping (PBRS) with Control Lyapunov Functions (CLF) and Control Barrier Functions (CBF) to simultaneously optimize mission time and ensure formal safety guarantees. An RL model is trained in a generalized simple environment, then used in complex scenarios incorporating a CLF-CBF-QP filter without further training. Experimental results in simulated environments demonstrate a significant reduction in mission time and outstanding performance in complex environment.
関連論文
- RACO: 点検指向UAV視覚言語ナビゲーションのための信頼性を考慮した粗目標最適化UAVナビゲーション
- AirAlign: ジオメトリを考慮したUAV最終メートル航法のための相対ポーズ整列UAVナビゲーション
- 形状適合領域による閉鎖・劣化環境での飛行UAVナビゲーション
- PILOT: 部分観測下での自律UAVのエンドツーエンド動作計画のための特権模倣学習UAVナビゲーション
- FlowPilot: 俊敏なUAVナビゲーションのためのリアルタイム世界行動モデリングUAVナビゲーション
- RASR: 距離認識スケール復元によるメートル単位UAVナビゲーションUAVナビゲーション