日本フィジカルAI新聞

世界のフィジカルAIを、日本語で。

週刊ニュースレター購読
マニピュレーションarXiv:2609.28431

LiMA: 非同期拡散による長期想像からリアルタイム巧みな操作への橋渡し

LiMA: Bridging Long-term Imagination to Real-time Dexterous Manipulation via Asynchronous Diffusion

シェア:XThreadsFacebookLINEはてブBluesky

低速な長期意図生成と高速な動作精緻化を非同期に分離した二重システム拡散フレームワークで、両腕巧みな操作をリアルタイムに実行する。

詳しい要約

1. どんなもの?

- 長期的な意図計画とリアルタイムの反応的実行を非同期に分離した dual-system 生成フレームワーク LiMA を提案。 - slow system が疎な long-horizon 時空間意図を生成し、fast system が高頻度の動作 refinement を担う。 - bimanual dexterous manipulation タスクを対象とし、複数 horizon にわたる6タスクで評価。

2. 先行研究と比べてどこがすごい?

- VLA モデルは高レベル推論に優れるが物理動力学・空間知覚の細粒度理解が不足。 - WAM は反復生成により推論遅延が大きく、意図が急速な接触変化に適応できない temporal misalignment が問題。 - LiMA は非同期 decoupling により Cosmos-Policy 比で推論遅延を45.8%削減。

3. 技術・手法の肝は?

- slow system と fast system からなる multi-scale hierarchy で計算を構成。 - 疎な意図予測と密な行動軌跡を整合させる Latent Schrödinger Bridge Coupling を導入。 - refinement を entropy-regularized probabilistic transport process として定式化。

4. どうやって有効だと検証した?

- 複数 horizon にまたがる6つの bimanual dexterous manipulation タスクで評価。 - 全体成功率70.8%、平均 subtask 成功率78.9%を達成。 - 未見シナリオでも性能を維持。

5. 議論はある?

- 要旨からは不明。

6. 次に読むべき論文は?

- Cosmos-Policy、VLA モデル、World-Action Models (WAMs) が参照・比較されている。

※ AIが要旨から生成した要約です。正確性は原文をご確認ください。

著者: Ning Chen, Yankai Fu, Junkai Zhao, Qianpu Sun, Guocai Yao, Pengwei Wang, Zhongyuan Wang, Shanghang Zhang

分類: cs.RO

原文アブストラクト

Dexterous manipulation demands long-term foresight and rapid reactive control. Vision-Language-Action (VLA) models, while proficient in high-level reasoning, often lack a fine-grained understanding of physical dynamics and spatial perception. Conversely, World-Action Models (WAMs) typically suffer from high inference latency due to iterative generation. These deficiencies result in a critical temporal misalignment where the model's intent fails to adapt to rapid physical contact changes. To overcome this fundamental bottleneck, we propose LiMA, an asynchronous dual-system generative framework that systematically decouples intent planning from reactive execution. LiMA organizes computation into a multi-scale hierarchy: a slow system handles sparse long-horizon spatiotemporal intent generation, while a fast system focuses on dense high-frequency motion refinement. To align sparse intent predictions with dense action trajectories, we introduce a Latent Schrödinger Bridge Coupling mechanism that formulates refinement as an entropy-regularized probabilistic transport process. LiMA reduces inference latency by 45.8% compared with Cosmos-Policy via asynchronous decoupling. Evaluated across six bimanual dexterous manipulation tasks spanning multiple horizons, LiMA achieves an overall success rate of 70.8% and an average subtask success rate of 78.9%, while maintaining performance in unseen scenarios. The project website is available at https://ccdcs.github.io/LiMA_repo/

関連論文

PR本紙発行元 EmplifAI