日本フィジカルAI新聞

世界のフィジカルAIを、日本語で。

週刊ニュースレター購読
強化学習arXiv:2610.07910

深層強化学習における滑らかな制御のための時間的正則化の再検討

Revisiting Temporal Regularization for Smooth Control in Deep Reinforcement Learning

シェア:XThreadsFacebookLINEはてブBluesky

時間的ペナルティが空間的な滑らかさももたらすことを証明し、線形ランプアップと組み合わせたCATSを提案。シミュレーションと実機で行動振動を抑えつつタスク性能を維持できることを示した。

詳しい要約

1. どんなもの?

- 深層強化学習(Deep Reinforcement Learning)のポリシーが生じる非平滑な行動振動を抑制し、物理ロボットへの展開を容易にする手法。 - 時間的正則化(temporal regularization)に着目し、Conditioning for Action using only Temporal Smoothness (CATS) を提案。 - 時間的ペナルティと線形ランプアップ(linear ramp-up)を組み合わせる。 - シミュレーションと実世界の両方で評価。

2. 先行研究と比べてどこがすごい?

- 既存のアーキテクチャやペナルティベースのアプローチは、状態入力の変化に対する感度を直接減らすことで空間的平滑性(spatial smoothness)を追求するが、強い平滑化はタスク性能を劣化させる。 - 時間的正則化は観測ノイズ下で空間的平滑性を提供できないと考えられてきた。 - 本研究はこの仮定を再検討し、時間的ペナルティが空間的平滑性をもたらすことを証明。 - CATSは、明示的な空間的正則化よりもタスク性能を保持しつつ空間的平滑性を提供できる点が優れている。

3. 技術・手法の肝は?

- 時間的ペナルティが、同じ次状態を共有する現在状態間の期待行動差を制限することを証明。 - この空間的効果が経験的に空間的平滑性に拡張されることを示す。 - CATSは時間的ペナルティと線形ランプアップを組み合わせる。 - 線形ランプアップにより、ポリシーはまず報酬の高い行動を学習し、その後徐々に行動を平滑化する。 - これによりリターン保持と時間的・空間的平滑性の両方を改善。

4. どうやって有効だと検証した?

- シミュレーションと実世界の両方で実験を実施。 - CATSがタスク性能を劣化させることなく行動振動を大幅に低減することを示す。 - 計算オーバーヘッドが小さいことも確認。

5. 議論はある?

- 時間的正則化が空間的平滑性を提供できるという新たな知見を提示。 - 明示的な空間的正則化と比較して、タスク性能をより良く保持できることを強調。 - 線形ランプアップの有効性を議論。 - 計算オーバーヘッドが小さい点も言及。 - ただし、具体的な限界や今後の課題については要旨からは不明。

6. 次に読むべき論文は?

- 要旨で参照/比較されている研究:既存のアーキテクチャベースのアプローチ、ペナルティベースのアプローチ、明示的な空間的正則化。 - 関連手法:temporal regularization、spatial regularization、linear ramp-up。 - 同分野の定番:Deep Reinforcement Learning、smooth control、action oscillation。

※ AIが要旨から生成した要約です。正確性は原文をご確認ください。

著者: SungJae Ahn, Jeong Woon Lee, Kyoleen Kwak, Hyoseok Hwang

分類: cs.LG, cs.RO

原文アブストラクト

Deep Reinforcement Learning policies can produce nonsmooth action oscillations that hinder deployment on physical robots. Existing architectural and penalty-based approaches seek spatial smoothness by directly reducing sensitivity to changes in state inputs, but their broad constraints can degrade task performance as stronger smoothing is pursued. Temporal regularization instead constrains action differences along observed transitions, but has been considered unable to provide the spatial smoothness needed under observation noise. We revisit this assumption by proving that the temporal penalty bounds the expected action differences between current states sharing a next state, revealing a spatial effect that empirically extends to spatial smoothness. Building on this finding, we propose Conditioning for Action using only Temporal Smoothness (CATS), which combines a temporal penalty with linear ramp-up. We highlight temporal regularization's ability to provide spatial smoothness while better preserving task performance than explicit spatial regularization. Through linear ramp-up, CATS allows the policy to learn rewarding behavior before progressively smoothing its actions, improving return preservation and both temporal and spatial smoothness. Experiments in both simulation and the real world show that CATS substantially reduces action oscillation without degrading task performance, with little computational overhead.

関連論文

PR本紙発行元 EmplifAI