日本フィジカルAI新聞

世界のフィジカルAIを、日本語で。

週刊ニュースレター購読
VLAarXiv:2609.30127

リーマン平均流による行動多様体上の高速視覚運動方策学習

Faster Visuomotor Policy Learning on Action Manifolds via Riemannian MeanFlow

シェア:XThreadsFacebookLINEはてブBluesky

ロボットの行動系列が滑らかな多様体上に定義されることに着目し、リーマン条件付きフローマッチングを基盤としたフローマップ整合性目的で、1回のネットワーク評価で多様体上の行動生成を可能にするRMFPを提案した。

詳しい要約

1. どんなもの?

- 視覚運動ポリシー(visuomotor policies)は、生の感覚観測からロボットの行動系列への直接的なマッピングを学習する。 - DiffusionやFlow Matchingに基づくポリシーは、行動系列の多峰性分布をエンドツーエンドで捉えるが、行動生成に学習されたベクトル場の多段階数値積分が必要で、計算コストが高く、ロボットに必要な高速制御レートを妨げる。 - また、ロボットの行動系列は通常、滑らかで微分可能な多様体(manifold)上で定義され、学習されたポリシーが行動空間の内在的幾何学を尊重する必要がある。 - 本研究では、Riemannian MeanFlow Policy (RMFP) を提案する。これは、ロボット行動多様体上の確率経路の条件付きフローマップを学習する。 - 定式化は、Riemannian Conditional Flow Matchingアンカーによってデータに基づくフローマップ整合性目的を採用する。 - フローマップ整合性条件は訓練が安定しており、学習モデルを有限時間輸送に制約し、わずか1回のネットワーク関数評価で多様体上の行動系列生成を可能に…

2. 先行研究と比べてどこがすごい?

- 先行研究(DiffusionやFlow Matchingに基づくポリシー)と比べて、RMFPは多段階数値積分を必要とせず、1回のネットワーク関数評価で行動生成が可能であり、サンプリングコストを大幅に削減する。 - また、ロボット行動空間の内在的幾何学(多様体構造)を尊重する点で、従来のユークリッド空間での手法とは異なる。 - 具体的な性能比較では、RMFPは先行研究と競争力のある性能をより低いサンプリングコストで達成することを示した。

3. 技術・手法の肝は?

- 技術の肝は、Riemannian MeanFlow Policy (RMFP) の定式化にある。 - ロボット行動多様体上の確率経路の条件付きフローマップを学習する。 - フローマップ整合性目的をRiemannian Conditional Flow Matchingアンカーでデータに基づかせる。 - この整合性条件は訓練が安定しており、有限時間輸送を制約し、1回のネットワーク関数評価での多様体上の行動系列生成を実現する。

4. どうやって有効だと検証した?

- 検証は、球面LASAとPush-Tベンチマーク、RobomimicスイートのTool HangとTransportタスク、多様体制約付き行動生成を伴うFranka Kitchenタスクで行った。 - その結果、RMFPは先行研究と競争力のある性能をより低いサンプリングコストで達成することを示した。 - さらに、実世界のロボットマニピュレーションタスクでRMFPを適用し、不完全なセンサ測定下での高速行動生成を実証した。

5. 議論はある?

- 要旨からは、議論や限界についての明示的な記述はない。 - ただし、実世界タスクでの不完全なセンサ測定下での高速行動生成の実証から、実用性が示唆される。 - 訓練の安定性や有限時間輸送の制約について言及されているが、具体的な議論は要旨からは不明。

6. 次に読むべき論文は?

- 要旨で参照/比較されている研究:Diffusionに基づくポリシー、Flow Matchingに基づくポリシー、Riemannian Conditional Flow Matching。 - 関連手法:MeanFlow、Riemannian Flow Matching。 - 同分野の定番:Diffusion Policy、Flow Matching Policy。

※ AIが要旨から生成した要約です。正確性は原文をご確認ください。

著者: S. Talha Bukhari, Austin Garrett, Yi Wei, Ruiqi Ni, Zachary Kingston, Aniket Bera

分類: cs.RO

原文アブストラクト

Visuomotor policies learn a direct map from raw sensory observations to robot action sequences. Policies based on Diffusion and Flow Matching capture the multimodal distribution over action sequences in an end-to-end manner. This expressivity comes at the cost of multi-step numerical integration of the learned vector field for action generation, which can be expensive and time-consuming, impeding fast control rates required in robotics applications. Furthermore, robot action sequences are usually defined on a smooth, differentiable manifold, requiring that the learned policy respects the intrinsic geometry of the robot's action space. Here, we present Riemannian MeanFlow Policy (RMFP), which learns the conditioned flow map of the probability path on the robot action manifold. Our formulation employs a flow map consistency objective grounded in the data by a Riemannian Conditional Flow Matching anchor. The flow map consistency condition is stable to train and constrains the learned model to finite-time transport, which yields on-manifold action sequence generation with as few as one network function evaluation. We present results on the spherical LASA and Push-T benchmarks, on the Tool Hang and Transport tasks of the Robomimic suite, and on the Franka Kitchen task with manifold-constrained action generation, and demonstrate that RMFP attains performance competitive with prior work at a lower sampling cost. We also employ RMFP on a real-world robotic manipulation task to demonstrate fast action generation under imperfect sensor measurements in the physical world.

関連論文

PR本紙発行元 EmplifAI