日本フィジカルAI新聞

世界のフィジカルAIを、日本語で。

週刊ニュースレター購読
MPCarXiv:2610.03563

Empowermentと予測制御の統合フレームワーク

A Unified Framework for Empowerment and Predictive Control

シェア:XThreadsFacebookLINEはてブBluesky

探査と行動に同一のポリシーを用いてEmpowermentを定式化し、標準的なMPCソルバで最適化可能にすることで、タスクコストとの組み合わせが成功率を向上させることを示した。

詳しい要約

1. どんなもの?

- サンプリングベースのMPCとempowermentを統合するフレームワークを提案。 - empowermentを単一のポリシーで推定し、実行される制御ポリシーと直接リンク。 - 一次導関数のみを必要とし、標準的なMPCソルバーで最適化可能。 - タスク固有のコストと組み合わせて使用できる。

2. 先行研究と比べてどこがすごい?

- 既存のempowermentベースのコントローラは、仮想プロービングポリシーと実行ポリシーが異なり、行動との関連が不明瞭。 - また、動力学の二次導関数を必要とし、標準的なMPCとの互換性が低い。 - 提案手法は単一ポリシーで両方を兼ね、一次導関数のみで済むため、標準MPCと互換性が高い。

3. 技術・手法の肝は?

- empowermentをチャネル容量として定義し、単一のポリシーでプロービングと行動を統合。 - 目的関数は一次導関数のみを必要とし、標準的なMPCソルバーで最適化可能。 - タスク固有のコストと組み合わせて、単独または併用で最適化。

4. どうやって有効だと検証した?

- 古典的な制御タスクで評価。 - empowermentとタスクコストを組み合わせることで、どちらか単独の場合と同等かそれを上回る成功率を達成。

5. 議論はある?

- 提案手法は標準MPCと内在的動機付け制御の実用的な橋渡しを提供。 - 制限や議論の詳細は要旨からは不明。

6. 次に読むべき論文は?

- 要旨で参照されている研究:empowerment-based controllers、sampling-based MPC。 - 関連手法:model predictive control (MPC)、empowerment、intrinsically motivated control。

※ AIが要旨から生成した要約です。正確性は原文をご確認ください。

著者: Wooyoung Chung, Tristan Shah, Volodymyr Makarenko, Stas Tiomkin

分類: cs.RO

原文アブストラクト

Sampling-based model predictive control (MPC) is a powerful approach to trajectory optimization, but its performance depends on an informative cost function that often requires substantial domain knowledge to design. An alternative is to derive objectives directly from the system dynamics. Empowerment, defined as the channel capacity between an agent's inputs and its future state, provides such a goal-agnostic objective and has been shown to produce useful behaviors across a range of domains. However, existing empowerment-based controllers have two key limitations: empowerment is estimated using a virtual probing policy distinct from the executed control policy, obscuring its connection to the resulting behavior; and its computation typically requires second-order derivatives of the dynamics, limiting compatibility with standard MPC methods. We formulate empowerment using a single policy for both probing and acting, directly linking the empowerment objective to the executed behavior. The resulting objective requires only first-order derivatives and can be optimized with standard MPC solvers, either alone or in combination with a task-specific cost. We evaluate the proposed approach on classical control tasks and show that combining empowerment with task costs matches or exceeds the success rate achieved by either objective alone. Our formulation provides a practical bridge between standard MPC and intrinsically motivated control.

関連論文

PR本紙発行元 EmplifAI