日本フィジカルAI新聞

世界のフィジカルAIを、日本語で。

週刊ニュースレター購読
マニピュレーションarXiv:2609.24155

物体中心条件付けによる視覚運動フローマッチング

Object-Centric Conditioning for Visuomotor Flow Matching

シェア:XThreadsFacebookLINEはてブBluesky

シーンを物体ごとの意味特徴と空間手がかりに分離し、フローマッチング方策の条件付けに用いることで、視覚的妨害や位置ずれに頑健なマニピュレーションを実現した。

詳しい要約

1. どんなもの?

- ロボットの visuomotor policy として、autoregressive、diffusion、flow matching などが使われる。 - その中で Action-to-Action (A2A) flow matching は、stochastic noise ではなく過去の action prior から生成を初期化し、推論効率を改善する。 - しかし、古い historical motion pattern と絡み合った global visual representation が、spatial out-of-distribution (OOD) shift や visual distractor 下での robustness を低下させる。 - 本研究では SlotFlow を提案する。object-centric flow matching policy であり、robust な visuomotor manipulation を目指す。 - SlotFlow は scene observation を semantic ("what") feature…

2. 先行研究と比べてどこがすごい?

- 従来の A2A flow matching は推論効率に優れるが、stale な historical motion pattern と entangled な global visual representation により、spatial OOD shift や visual distractor に弱い。 - SlotFlow は object-centric に scene を decouple することで、semantic representation が無関係な背景相関を抑制し、spatial cue が shifted object configuration への適応を改善する。 - これにより、A2A の低ステップ推論効率を保ちつつ、visual distractor や severe spatial perturbation 下での robustness を向上させる。 - 制御された initialization と perception ablation により、object-centric grounding が主要な改善要因であり、有用な histori…

3. 技術・手法の肝は?

- SlotFlow は object-centric flow matching policy である。 - 観測を semantic ("what") feature と軽量な image-plane spatial ("where") cue に分離する。 - semantic representation は無関係な背景相関を抑制する。 - spatial cue は shifted object configuration への適応を改善する。 - これらを policy conditioning と current-state grounding に用いる。 - 生成は A2A と同様に historical action prior から初期化し、低ステップ推論を維持する。

4. どうやって有効だと検証した?

- 広範な simulation と real-world experiment を実施。 - visual distractor と severe spatial perturbation 下での robustness 改善を実証。 - A2A の低ステップ推論効率が維持されることを確認。 - controlled initialization と perception ablation により、object-centric grounding が主要な改善源であることを特定。 - また、object-centric grounding が有用な historical motion prior を置き換えるのではなく補完することを示した。

5. 議論はある?

- 要旨からは、SlotFlow の限界や失敗ケース、計算コスト、一般化可能性についての議論は明示されていない。 - ただし、object-centric grounding が historical motion prior を補完するという知見が得られている。 - 今後の課題や応用範囲については要旨からは不明。

6. 次に読むべき論文は?

- 要旨で直接参照・比較されている研究は Action-to-Action (A2A) flow matching である。 - 関連手法として diffusion-based policy、autoregressive policy、flow matching が挙げられる。 - また、object-centric representation learning や visuomotor manipulation の robustness に関する研究が関連する。 - 具体的な論文名は要旨に記載がないため、同分野の定番として A2A flow matching の原論文や object-centric learning の代表的手法を読むとよい。

※ AIが要旨から生成した要約です。正確性は原文をご確認ください。

著者: Jijie Li, Xu Yang, Junhong Zou, Chunhai Zhao, Chaoyang Zhao, Zhen Lei, Xiangyu Zhu

分類: cs.RO

原文アブストラクト

Robot visuomotor policies are commonly formulated as autoregressive, diffusion-based, or more recently, flow matching models. Among them, Action-to-Action (A2A) flow matching improves inference efficiency by initializing generation from historical action priors rather than stochastic noise. However, stale historical motion patterns and entangled global visual representations can jointly reduce robustness under spatial out-of-distribution (OOD) shifts and visual distractors. In this work, we propose SlotFlow, an object-centric flow matching policy for robust visuomotor manipulation. SlotFlow decouples scene observations into semantic ("what") features and lightweight image-plane spatial ("where") cues to provide object-aware policy conditioning and current-state grounding. The semantic representation suppresses irrelevant background correlations, while the spatial cue improves adaptation to shifted object configurations. Extensive simulation and real-world experiments demonstrate improved robustness under visual distractors and severe spatial perturbations while preserving the low-step inference efficiency of A2A. Controlled initialization and perception ablations further identify object-centric grounding as a major source of the gains and show that it complements, rather than replaces, useful historical motion priors.

関連論文

PR本紙発行元 EmplifAI