日本フィジカルAI新聞

世界のフィジカルAIを、日本語で。

週刊ニュースレター購読
歩行arXiv:2610.12113

ゾノトープを用いた幾何学的アプローチによるSoft Actor-Criticの歩行学習

A Geometric Approach to Soft Actor-Critic with Zonotopes for Locomotion Learning

シェア:XThreadsFacebookLINEはてブBluesky

各Criticがゾノトープを予測し、その幾何学的幅とCritic間の不一致度に応じて悲観性を適応的に調整するGeZo-SACを提案し、MuJoCo歩行タスクで性能と省エネルギー性を向上させた。

詳しい要約

1. どんなもの?

- 強化学習の off-policy actor-critic 手法である GeZo-SAC を提案する研究。 - 各 critic がスカラー値に加えて zonotope を定義する generators を予測する。 - この幾何表現を用いて critic の悲観度を state と action に適応させる。 - 推論時は通常の SAC actor をそのまま用い、generators は訓練時の critic 側のみで使う。 - MuJoCo-v5 の locomotion ベンチマークで評価している。

2. 先行研究と比べてどこがすごい?

- 従来の off-policy actor-critic は二つの critic の最小値を取ることで過大評価バイアスを制御し、critic の不一致に関わらずどこでも同じ集約規則を使う。 - GeZo-SAC は critic の不一致に応じて集約規則を変える点が異なる。 - 不一致が小さいときは width-weighted average に近づき、不一致が大きくなると通常の minimum に移行する。 - これにより critic の悲観度を state と action に適応させられる。

3. 技術・手法の肝は?

- 各 critic はスカラー値とともに zonotope を定義する generators の集合を予測する。 - サンプルした方向に沿って zonotope を調べ、geometric width を得る。 - この width を各 critic 値から悲観的オフセットとして差し引く。 - 二つの critic 間の不一致も測り、log-sum-exp で集約する。 - この不一致が critic の結合方法を制御し、width-weighted average から minimum へと移行させる。 - 推論時は generators を使わず、未変更の SAC actor を展開する。

4. どうやって有効だと検証した?

- 四つの MuJoCo-v5 locomotion ベンチマークと六つの off-policy ベースラインで評価した。 - Ant-v5 と Hopper-v5 で最高の平均リターンを達成した。 - 残りのタスクでも他の手法と競争力がある。 - 評価した手法の中で平均アクチュエータ仕事とメートルあたりの行動努力が最小であることを示した。 - 四環境すべてで測定された過大評価頻度がほぼゼロであることを維持した。

5. 議論はある?

- 要旨からは不明。 - ただし、GeZo-SAC が最低の平均アクチュエータ仕事と行動努力、ほぼゼロの過大評価頻度を達成した点は分析で示されている。

6. 次に読むべき論文は?

- 要旨で参照・比較されている研究として、SAC (Soft Actor-Critic) と off-policy actor-critic の最小値集約を用いる手法が挙げられる。 - また、MuJoCo-v5 locomotion ベンチマークと六つの off-policy ベースラインが比較対象として言及されている。 - 関連手法として zonotope を用いた幾何表現や log-sum-exp 集約も挙げられる。

※ AIが要旨から生成した要約です。正確性は原文をご確認ください。

著者: Panagiotis Roditis, Panagiotis P. Filntisis, Petros Maragos

分類: cs.LG, cs.AI

原文アブストラクト

Off-policy actor--critic methods control overestimation bias by taking the minimum of two critics. This uses the same aggregation rule everywhere, regardless of how the critics disagree. We propose \textbf{GeZo-SAC}, which uses auxiliary geometric representations to adapt critic pessimism to the state and action. Alongside its scalar value, each critic predicts a set of generators defining a zonotope. Probing this zonotope along sampled directions provides a geometric width, "subtracted from each critic value as a pessimistic offset, and a measure of disagreement between the two critics, aggregated with log-sum-exp. This disagreement controls how the critics are combined, moving from a width-weighted average toward the usual minimum as disagreement increases. At inference, the deployed policy is an unmodified SAC actor, since the generators are used only on the critic side during training.Across four MuJoCo-v5 locomotion benchmarks and six off-policy baselines, GeZo-SAC achieves the highest mean return on Ant-v5 and Hopper-v5 and remains competitive with other methods on the remaining tasks. Our analysis further shows that GeZo-SAC achieves the lowest average actuator work and action effort per metre among the evaluated methods, while maintaining near-zero measured overestimation frequency across all four environments.

関連論文

PR本紙発行元 EmplifAI