日本フィジカルAI新聞

世界のフィジカルAIを、日本語で。

週刊ニュースレター購読
VLAarXiv:2609.34688

統一軌道マッチング方策最適化:多様なT2I生成とVLA汎化

Unified Trajectory Matching Policy Optimization: Diverse T2I Generation and VLA Generalization

シェア:XThreadsFacebookLINEはてブBluesky

拡散・フローポリシーの事後学習において、報酬最大化ではなく軌道分布マッチングを行う統一RLフレームワークUni-TMPOを提案し、T2I生成の多様性とVLAのタスク・シーン汎化を両立させた。

詳しい要約

1. どんなもの?

本論文は、stochastic diffusion および flow policies の post-training のための統一的な reinforcement learning (RL) フレームワーク Uni-TMPO (Unified Trajectory Matching Policy Optimization) を提案する。 - 対象は text-to-image (T2I) 生成と vision-language-action (VLA) モデル。 - 従来の reward-maximizing RL が引き起こす policy mode collapse を解決することを目的とする。 - T2I では多様性の低下と reward hacking、VLA では代替戦略の消失と汎化性能の低下を防ぐ。

2. 先行研究と比べてどこがすごい?

従来の reward-maximizing RL は reference KL や entropy regularization を用いても policy mode collapse を起こし、単一の高報酬モードに収束してしまう。 - これに対し Uni-TMPO は forward Kullback-Leibler optimization により、報酬を目標分布に変換して方策分布とマッチングさせる。 - その結果、T2I ではより高い報酬と reward-diversity-efficiency trade-off を達成し、VLA では held-out タスク・シーンへの汎化性能と ID 成功率で最強ベースラインを上回る。

3. 技術・手法の肝は?

Uni-TMPO の肝は以下の通り。 - 各 trajectory group 内で標準化された報酬を目標分布に変換し、trajectory log probabilities から方策分布を導出する。 - 期待報酬の最大化ではなく、forward Kullback-Leibler optimization で両分布をマッチングさせる。 - T2I 用に progress-conditioned coarse-to-fine scheduler を導入し、効率的に trajectory を構築する。 - VLA 用には feedback-conditioned sampling を用い、更新された observation から trajectory を構築する。

4. どうやって有効だと検証した?

広範な実験により有効性を検証。 - T2I では最強ベースラインより高い報酬を達成。 - VLA では ID 成功率が最強ベースラインを上回る。 - さらに T2I の reward-diversity-efficiency trade-off と、VLA の held-out タスク・シーンへの汎化で最良の結果を示す。 - 実ロボット評価では、より高報酬のターゲットがブロックされた際に複数の行動戦略が有用であることを実証。

5. 議論はある?

本論文は reward-maximizing RL の mode collapse 問題を指摘し、forward KL による分布マッチングが有効であることを示す。 - しかし、目標分布の設計や trajectory group の構成方法、計算コスト、他の RL アルゴリズムとの比較などについては要旨からは不明。 - 実ロボット評価の詳細な設定や被験者数なども要旨からは不明。

6. 次に読むべき論文は?

要旨で参照・比較されている研究として、reward-maximizing reinforcement learning (RL) を用いた stochastic diffusion および flow policies の post-training 手法、reference KL や entropy regularization を適用した既存手法、text-to-image (T2I) 生成のための RL ベース手法、vision-language-action (VLA) モデルが挙げられる。 - 関連手法として、diffusion policy や flow matching、KL 正則化付き RL、reward hacking に関する研究も次に読むべき候補である。

※ AIが要旨から生成した要約です。正確性は原文をご確認ください。

著者: Zhiyuan Ma, Jiaming Li, Lingzhen Li, Yu Liu, Xuekai Zhu, Dingkang Liang, Kaiyan Zhang, Jianjun Li, Bowen Zhou, Xiang Bai

分類: cs.CV

原文アブストラクト

Reward-maximizing reinforcement learning (RL) is widely used to post-train stochastic diffusion and flow policies for text-to-image (T2I) generation. However, reward-maximizing RL causes policy mode collapse even under reference KL or entropy regularization, reducing the policy to a single high-reward mode. In T2I, this produces similar images and reward hacking. When extended to vision-language-action (VLA) models, the same collapse removes alternative successful strategies and weakens task and scene generalization. To address this limitation, we introduce Unified Trajectory Matching Policy Optimization (Uni-TMPO), a unified RL post-training framework for diffusion and flow policies. First, Uni-TMPO converts standardized rewards into a target distribution within each trajectory group and derives the policy distribution from trajectory log probabilities. Then, forward Kullback-Leibler optimization matches the two distributions instead of maximizing expected reward. A progress-conditioned coarse-to-fine scheduler efficiently constructs T2I trajectories. Within the unified framework, feedback-conditioned sampling uses updated observations to construct VLA trajectories. Extensive experiments show that Uni-TMPO achieves higher T2I rewards and VLA ID success rates than the strongest baselines. More importantly, it achieves the best T2I reward-diversity-efficiency trade-off and VLA generalization to held-out tasks and scenes, while real-robot evaluation demonstrates the value of multiple action strategies when the higher-reward target is blocked.

関連論文

PR本紙発行元 EmplifAI