日本フィジカルAI新聞

世界のフィジカルAIを、日本語で。

週刊ニュースレター購読
操作arXiv:2608.02990v1

EmbodiedVAE: 効率的で制御可能な身体化操作のための分離型ビデオVAE

EmbodiedVAE: Disentangled Video VAE for Efficient and Controllable Embodied Manipulation

シェア:XThreadsFacebookLINEはてブBluesky

ロボット操作の世界モデル向けに、ロボットアームの動きと背景を分離して圧縮する新しいビデオVAEを提案し、高圧縮率で高品質な再構成と精密な動作制御を実現した。

詳しい要約

1. どんなもの?

EmbodiedVAEは、ロボット操作の世界モデル構築に用いるLatent Diffusion Models (LDMs)のための、コンパクトかつ制御可能な潜在表現を提供する新しいビデオVAEである。従来の自然シーン向けVAEがロボット操作特有の特性を考慮せず、潜在表現が非コンパクトで制御不能である問題を解決する。

2. 先行研究と比べてどこがすごい?

既存のLDMsは自然シーン向けに最適化されたVAEを使用しており、ロボット操作シナリオの特性を考慮していない。EmbodiedVAEは、デュアルエンコーダ・シングルデコーダ構造と非対称な時空間圧縮モジュールにより、ロボットアームの動きと背景環境を自動的に分離し、コンパクトさと明示的なembodied潜在表現を両立する点が優れている。

3. 技術・手法の肝は?

手法の核心は、デュアルエンコーダ・シングルデコーダ構造と非対称な時空間圧縮モジュールにより、ロボットアームの動きと背景環境を自動的に分離すること。さらに、最適輸送に基づく一貫性モジュールを導入し、学習されたロボット動作潜在表現の時間的一貫性を明示的に保証する。

4. どうやって有効だと検証した?

広範な実験により、再構成品質と高圧縮率の両立を実証し、最先端のビデオVAEと比較して平均2dBのPSNR改善を達成した。また、ロボット操作シナリオにおけるより精密な行動制御を可能にすることを示した。

5. 議論はある?

要旨からは、潜在表現の解釈可能性や、異なるロボット操作タスクへの汎用性、計算コストなどの議論は不明。また、実機での検証や、他の世界モデルへの応用可能性については言及されていない。

6. 次に読むべき論文は?

要旨で参照されているのは、Latent Diffusion Models (LDMs)とVariational Autoencoders (VAEs)である。次に読むべき論文としては、これらの基礎となる論文や、ロボット操作の世界モデルに関する関連研究が挙げられる。

※ AIが要旨から生成した要約です。正確性は原文をご確認ください。

著者: Jiayi Luo, Hanxin Zhu, Chen Gao, Jiankun Wang, Cong Wang, Tianyu He, Jianxin Li, Zhibo Chen

分類: cs.RO

原文アブストラクト

Latent diffusion models (LDMs) have recently significantly advanced embodied learning in constructing powerful embodied manipulation world models. However, despite the remarkable performance, existing LDMs predominantly rely on Variational Autoencoders (VAEs) optimized for natural scenes while failing to account for the unique characteristics of embodied manipulation scenarios, yielding latent representations that are neither compact nor controllable, thereby hindering efficient training of LDMs and precise robotic control. To solve this problem, we present EmbodiedVAE, a novel video VAE that provides compact yet controllable latent representations tailored for the robotic manipulation world models. Specifically, EmbodiedVAE adopts a dual-encoder, single-decoder architecture with an asymmetric spatio-temporal compression module, which automatically disentangles the robot arm's motion from background environment, resulting in overall compactness while providing explicit embodied latent to support fine-grained action control. To further preserve the temporal consistency of learned robotic motion latent, we introduce an optimal-transport-based consistency module that explicitly enforces motion fidelity and inter-frame coherence. Extensive experiments demonstrate that our proposed EmbodiedVAE achieves superior reconstruction quality with high compression rate, while enabling more precise action control in robotic manipulation scenarios with an average of 2dB PSNR improvement over state-of-the-art video VAEs.