日本フィジカルAI新聞

世界のフィジカルAIを、日本語で。

週刊ニュースレター購読
軌道最適化arXiv:2610.08532

パレート最適エントロピー正則化軌道最適化

Pareto-Optimal Entropy-Regularized Trajectory Optimization

シェア:XThreadsFacebookLINEはてブBluesky

エントロピー正則化とパレートフィルタリングを組み合わせ、DDPの探索範囲を広げて非凸な軌道最適化問題の成功率を向上させる手法を提案。

詳しい要約

1. どんなもの?

- 非線形ダイナミクス、アクチュエータ制限、衝突回避制約下での軌道最適化(TO)問題を扱う。 - 特に非凸性が高く困難な問題に対し、Differential Dynamic Programming (DDP) を拡張した手法を提案。 - 提案手法は Pareto-Optimal Entropy-Regularized DDP (PER-DDP) と呼ばれ、エントロピー正則化された母集団フレームワーク。 - 自由エネルギー/相対エントロピー不等式に基づく。 - サンプリングと最適化を組み合わせ、局所解への脆弱性を軽減する。

2. 先行研究と比べてどこがすごい?

- 従来のDDPは二次収束の射撃法だが、局所構造のため劣最適な盆地に陥りやすい。 - サンプリング拡張型の変種は、コストのみに基づいて保持する少数の軌道周辺でのみサンプリングし、探索の広がりが制限される。 - PER-DDPは、事前誘導サンプリングと拡張ロールアウト評価、パレートフィルタリングを組み合わせ、サンプリング努力を最適化母集団サイズから分離。 - これにより、DDPの有効性を支える二次構造を犠牲にせずに探索を広げる。 - 複数のシステムと数百の環境で、最先端のサンプリング拡張TO法よりも高い成功率を達成し、全てのベースラインが到達できない環境でも信頼できる解を見つける。

3. 技術・手法の肝は?

- 自由エネルギー/相対エントロピー不等式から導出されたエントロピー正則化母集団フレームワーク。 - 事前誘導サンプリング:各保持軌道の周りで探索を形成。 - 拡張ロールアウト評価:保持された少数の候補を超えてサンプリングポリシーを調査。 - パレートフィルタリング:反復を通じてタスク制約の代替案を保持。 - これらにより、サンプリング努力を最適化母集団サイズから分離し、探索を広げる。

4. どうやって有効だと検証した?

- 複数のシステムと数百の環境で評価。 - 最先端のサンプリング拡張TO法と比較して高い成功率を達成。 - 全てのベースラインが到達できない環境でも信頼できる解を見つける。 - 具体的な検証方法の詳細は要旨からは不明。

5. 議論はある?

- 要旨からは不明。

6. 次に読むべき論文は?

- Differential Dynamic Programming (DDP) - サンプリング拡張型軌道最適化法(sampling-augmented TO methods) - エントロピー正則化強化学習(entropy-regularized RL) - パレート最適化(Pareto optimization)

※ AIが要旨から生成した要約です。正確性は原文をご確認ください。

著者: Dimitrios S. Georgiou, Augustinos D. Saravanos, Evangelos A. Theodorou

分類: cs.RO, eess.SY

原文アブストラクト

Trajectory optimization (TO) under nonlinear dynamics, actuation limits and collision avoidance constraints is a fundamental problem in robotics, albeit especially challenging due to its highly non-convex nature. For this setting, Differential Dynamic Programming (DDP) is an efficient second-order shooting method, yet its local structure renders it vulnerable to suboptimal basins. Sampling-augmented variants mitigate this susceptibility through stochastic exploration, but often sample only around the few trajectories they retain for reoptimization, based solely on their cost which restricts exploration breadth. We introduce Pareto-Optimal Entropy-Regularized DDP (PER-DDP), an entropy-regularized population framework derived from the free-energy/relative-entropy inequality. Our method combines prior-guided sampling that shapes exploration around each retained trajectory, with expanded rollout evaluations that probe these sampling policies beyond the few retained candidates, and Pareto filtering for preserving task-constraint alternatives across iterations. This decouples sampling effort from the optimization population size and broadens exploration without sacrificing the second-order structure that makes DDP effective. Across multiple systems and hundreds of environments, PER-DDP achieves higher success rates than state-of-the-art sampling-augmented TO methods and finds reliable solutions in environments beyond the reach of all baselines.

関連論文

PR本紙発行元 EmplifAI