日本フィジカルAI新聞

世界のフィジカルAIを、日本語で。

週刊ニュースレター購読
VLA/拡散ポリシーarXiv:2610.09369

非混和拡散ポリシー:ラベルフリーなノイズ割り当てによるマルチモーダルなロボット行動の保持

Immiscible Diffusion Policy: Preserving Multimodal Robot Actions through Label-Free Noise Assignment

シェア:XThreadsFacebookLINEはてブBluesky

拡散ポリシーがマルチモーダルな行動分布を崩壊させる問題に対し、行動とノイズの割り当てを工夫するラベルフリーな学習時追加手法を提案し、モダリティ保持を大幅に改善した。

詳しい要約

1. どんなもの?

本論文は、diffusion policy が multi-modal action distribution を回復できず単一モードに collapse する問題を指摘し、その原因を action-noise pairing の独立性による diffusion path の mixing/crossing にあると分析する。対策として Immiscible Diffusion Policy を提案する。これは label-free な training-time add-on であり、policy architecture や inference procedure を変更せず、action-noise assignment によって比較的 distinct な noise-to-action route を保持する。

2. 先行研究と比べてどこがすごい?

従来の diffusion policy は multi-modal action distribution の回復を期待されていたが、dataset modality の balance や within-batch symmetry を保証しても単一モードに collapse することがある。本手法は architecture や inference を変えずに training 時の action-noise assignment のみで modality 保持を改善する点が異なる。

3. 技術・手法の肝は?

技術の肝は label-free な action-noise assignment である。独立な action-noise pairing が diffusion path の mixing/crossing を増やし、averaged denoising response を生んで modality-specific behavior を抑制するという分析に基づき、比較的 distinct な noise-to-action route を保持するように割り当てる。policy architecture や inference procedure の変更は不要。

4. どうやって有効だと検証した?

5つの simulated と 2つの real-world humanoid manipulation tasks で検証。observation は state, RGB, point-cloud を含む。3つの two-modality tasks で non-dominant modality の割合を 6.0x-14.6x 増加させ、2つの four-modality tasks では vanilla policy rollouts で完全に欠落していた demonstrated modalities を回復した。

5. 議論はある?

diffusion policy の multi-modal action 保持失敗の原因として action-noise pairing の独立性と diffusion path の mixing/crossing を挙げ、dense で low-dimensional な robot planning の action space で特に深刻だと議論している。提案手法は simple yet robust な approach と位置づけられている。

6. 次に読むべき論文は?

要旨で参照/比較されている研究は明示されていない。関連手法として diffusion policy、および multi-modal action distribution を扱う robot learning の定番手法を挙げる。

※ AIが要旨から生成した要約です。正確性は原文をご確認ください。

著者: Xiao Zhang, Yuxin Chen, Zhixuan Liang, Guojian Zhan, Chenran Li, Chenfeng Xu, Masayoshi Tomizuka, Yiheng Li

分類: cs.RO, cs.LG

原文アブストラクト

When diffusion policies were first introduced, they were expected to recover multi-modal action distributions. However, we find this expectation does not always hold, as diffusion policies often collapse to a single modality even when we guarantee the balance of dataset modalities and exact within-batch symmetry. Our analysis indicates that independent action-noise pairing contributes to this failure by increasing mixing and crossing among diffusion paths, which can produce averaged denoising responses and suppress modality-specific behavior. This issue is especially severe in robot planning, where action spaces are dense and low-dimensional, significantly increasing such mixing and crossing. To alleviate this problem, we propose Immiscible Diffusion Policy, a label-free training-time add-on to diffusion policy that uses action-noise assignment to preserve relatively distinct noise-to-action routes without modifying the policy architecture or inference procedure. Across five simulated and two real-world humanoid manipulation tasks spanning state, RGB, and point-cloud observations, our method significantly improves the policy's preservation of action modalities while maintaining strong task performance. It increases the proportion of the non-dominant modality by 6.0x-14.6x across three two-modality tasks and recovers demonstrated modalities that are entirely absent from vanilla policy rollouts on both four-modality tasks. These results demonstrate that Immiscible Diffusion Policy provides a simple yet robust approach to preserving action multi-modality in general robot learning tasks.

PR本紙発行元 EmplifAI