日本フィジカルAI新聞

世界のフィジカルAIを、日本語で。

週刊ニュースレター購読
マニピュレーションarXiv:2610.07917

PACE: 進捗整合コンテキストによる段階一貫性のある長期的ロボットマニピュレーション

PACE: Stage-Consistent Long-Horizon Robot Manipulation via Progress-Aligned Context for Execution

シェア:XThreadsFacebookLINEはてブBluesky

実演を段階構造を持つトークン列に圧縮し、実行中の行動・観測履歴を記憶して進捗に応じて参照することで、段階の混同を防ぎ長期的な操作タスクの成功率を高める手法を提案。

詳しい要約

1. どんなもの?

- 長期的なロボットマニピュレーションを、デモンストレーション条件付きポリシーで実行する手法。 - 視覚的に類似した状態が異なる段階で再出現する場合や、デモと実行の速度が異なる場合に生じる『stage confusion』を特定。 - これを解決するため、実行の進捗に応じて完全なデモを継続的に再解釈するstatefulな手法『Progress-Aligned Context for Execution (PACE)』を提案。 - テスト時の段階ラベルや段階別ポリシーを必要とせず、統一的なdiffusion action expertで動作する。

2. 先行研究と比べてどこがすごい?

- 従来のdemonstration-conditioned policiesは、視覚的に類似した状態が異なる段階で現れる場合や、デモと実行の速度差がある場合にstage confusionを起こしやすい。 - PACEは、デモを順序付きマルチモーダルprompt tokensに圧縮し、訓練時のみのdual-edge attention supervisionで潜在的な段階構造を露出。 - 実行時にはepisode-local fast-weight memoryが実現されたaction-observation遷移を因果的に符号化し、prompt cross-attentionを調節。 - これにより、テスト時の段階ラベルや段階別ポリシーなしで進捗に整合したcontextを生成し、stage confusionを軽減する点が新しい。

3. 技術・手法の肝は?

- デモンストレーションを順序付きマルチモーダルprompt tokensに圧縮。 - 訓練時のみdual-edge attention supervisionを適用し、デモの潜在的な段階構造を明示的に学習。 - 実行時、episode-local fast-weight memoryが因果的にaction-observation遷移を符号化。 - このmemoryがprompt cross-attentionを調節し、進捗に整合したcontextを生成。 - 統一的なdiffusion action expertがそのcontextを用いて行動を生成。テスト時の段階ラベルや段階別ポリシーは不要。

4. どうやって有効だと検証した?

- LIBERO-Gen Goal Chainで成功率が88.9%から94.0%に向上。 - Spatial Combinationで79.1%から83.3%に向上。 - two-step Block Routingタスクで33.3%から73.3%に向上。 - 失敗分析により、構造化されたデモンストレーション整合と因果的実行メモリが共同でstage confusionを軽減することを示唆。

5. 議論はある?

- 失敗分析から、構造化されたデモンストレーション整合と因果的実行メモリがstage confusionの軽減に寄与することが示唆された。 - ただし、stage confusionの定量的な定義や、他の失敗モードとの比較、計算コストやスケーラビリティに関する議論は要旨からは不明。

6. 次に読むべき論文は?

- 要旨で参照/比較されている研究は明示されていない。 - 関連手法として、demonstration-conditioned policies、diffusion policy、fast-weight memory、cross-attention modulation、LIBERO-Genベンチマークなどが挙げられる。 - 同分野の定番として、Behavior Cloning、Imitation Learning、Transformer-based policies、Long-horizon manipulation benchmarks(例:LIBERO)を次に読むべき論文として推奨。

※ AIが要旨から生成した要約です。正確性は原文をご確認ください。

著者: Yenan Chen, Junjie Shi, Lu Chen, Zhongxiang Zhou, Rong Xiong

分類: cs.RO

原文アブストラクト

Demonstration-conditioned policies provide a natural interface for specifying robot behavior, yet long-horizon manipulation remains difficult when visually similar states recur across different stages or when demonstrations and executions proceed at different speeds. We identify the resulting failure mode as stage confusion and introduce Progress-Aligned Context for Execution (PACE), a stateful method that continually reinterprets a complete demonstration according to realized execution progress. PACE compresses the demonstration into ordered multimodal prompt tokens and uses training-only dual-edge attention supervision to expose its latent stage structure. During execution, an episode-local fast-weight memory causally encodes realized action-observation transitions and modulates prompt cross-attention, producing a progress-aligned context for a unified diffusion action expert without test-time stage labels or stage-specific policies. PACE improves success from 88.9% to 94.0% on LIBERO-Gen Goal Chain, from 79.1% to 83.3% on Spatial Combination, and from 33.3% to 73.3% on the two-step Block Routing tasks. Failure analysis further indicates that structured demonstration alignment and causal execution memory jointly mitigate stage confusion.

関連論文

PR本紙発行元 EmplifAI