日本フィジカルAI新聞

世界のフィジカルAIを、日本語で。

週刊ニュースレター購読
強化学習arXiv:2610.10437

スカラー随伴マッチングによるQ学習

Q-Learning with Scalar Adjoint Matching

シェア:XThreadsFacebookLINEはてブBluesky

フローポリシーのファインチューニングにおいて、各ステップのヤコビアン計算を不要にするスカラー随伴を提案し、OGBenchの難タスクで成功率を大幅に改善した。

詳しい要約

1. どんなもの?

- Flow policy を off-policy RL で fine-tuning する手法 SQAM を提案。 - Flow policy は多段の flow step で action を生成するため、value function による fine-tuning が難しい。 - Adjoint matching を scalar 化し、per-step の vector-Jacobian product を不要にした。 - さらに policy 生成 action 上での critic value を制御する value penalty を組み合わせる。 - OGBench と実機 bimanual robot で有効性を検証。

2. 先行研究と比べてどこがすごい?

- 従来の adjoint matching は各 flow step で vector-Jacobian product が必要で、step 数と policy サイズに比例して計算コストが増大。 - 提案手法は pretrained flow policy の batch-averaged velocity Jacobian が対角に集中する観察に基づき、closed-form scalar adjoint を導出。 - これにより per-step の vector-Jacobian product を排除し、計算効率を改善。 - OGBench の最も難しい 4 ドメインで最強 baseline を 18〜35 ポイント上回る。 - 実機 bimanual robot でも supervised fine-tuning を全 3 タスクで改善。

3. 技術・手法の肝は?

- Pretrained flow policy の batch-averaged velocity Jacobian が対角的に集中することを観察。 - この観察から closed-form scalar adjoint を導出し、最終 action の value gradient を flow time でスケール。 - これにより各 flow step での vector-Jacobian product を不要化。 - Scalar adjoint 下では policy 生成 action 上の critic value 制御が重要と判明。 - そこで policy 生成 action に value penalty を課す Q-learning with Scalar Adjoint Matching (SQAM) を構成。

4. どうやって有効だと検証した?

- OGBench の最も難しい 4 ドメインで評価し、各ドメインで最強 baseline を 18〜35 ポイント上回る success rate を達成。 - 大規模 pretrained policy への拡張性を検証するため、実機 bimanual robot 上で vision-language-action policy を fine-tune。 - その結果、3 タスクすべてで supervised fine-tuning を上回る性能を確認。

5. 議論はある?

- 要旨からは不明。 - ただし scalar adjoint の導出根拠である Jacobian の対角集中性や、value penalty の必要性についての議論が含まれる可能性がある。

6. 次に読むべき論文は?

- Adjoint matching に関する元論文。 - Flow policy の off-policy RL fine-tuning に関する研究。 - OGBench をベンチマークとして用いた関連手法。 - Vision-language-action policy の fine-tuning に関する研究。

※ AIが要旨から生成した要約です。正確性は原文をご確認ください。

著者: Yonghoon Dong, Minsung Yoon, Jaehyuk Kim, Jungwoo Park, Changyeon Kim, Jinwoo Shin

分類: cs.LG, cs.AI, cs.RO

原文アブストラクト

Flow policies capture rich and diverse action distributions, and fine-tuning them with off-policy RL to improve beyond the demonstrations has drawn growing interest. However, fine-tuning a flow policy against a learned value function is not trivial, because the policy generates its action over many flow steps. Adjoint matching offers a principled way to update the flow model itself by propagating value information from the final action back to each flow step, but it requires a vector--Jacobian product through the policy at every step, a cost that grows with the number of flow steps and the policy size. We observe that the batch-averaged velocity Jacobian of pretrained flow policies concentrates on its diagonal. Motivated by this finding, we derive a closed-form scalar adjoint that scales the value gradient at the final action by the flow time, eliminating the per-step vector--Jacobian products. We further find that controlling the critic's value at policy-generated actions is particularly important under the scalar adjoint. Based on these findings, we propose Q-learning with Scalar Adjoint Matching (SQAM), which combines the scalar adjoint with a value penalty at those actions. SQAM's gains concentrate on the four hardest OGBench domains, where its success rate exceeds that of the strongest baseline in each domain by 18 to 35 percentage points. To test whether SQAM extends to large pretrained policies, we also fine-tune a vision-language-action policy on a real bimanual robot. SQAM improves over supervised fine-tuning on all three tasks.

関連論文

PR本紙発行元 EmplifAI