日本フィジカルAI新聞

世界のフィジカルAIを、日本語で。

週刊ニュースレター購読
マニピュレーションarXiv:2609.21753

PSR: 接触の多いマニピュレーションのための予測的感覚運動表現学習

PSR: Predictive Sensorimotor Representation Learning for Contact-Rich Manipulation

シェア:XThreadsFacebookLINEはてブBluesky

マルチモーダルな感覚運動信号から将来の接触ダイナミクスを予測する表現を事前学習し、VLAモデルの行動生成に組み込むことで、力覚を要する接触の多い操作タスクの成功率を大幅に向上させた研究。

詳しい要約

1. どんなもの?

- 接触を伴うマニピュレーション(contact-rich manipulation)向けの枠組み - Predictive Sensorimotor Representation (PSR) learning を提案 - マルチモーダルな sensorimotor 信号から予測表現の階層を学習 - それを visuomotor policy の action stream に統合 - Vision-Language-Action (VLA) モデルに実装した PSR-VLA を構築 - 6つの実世界 contact-rich タスクで評価

2. 先行研究と比べてどこがすごい?

- 既存手法は force feedback を passive に条件付けするのみ - 未来の contact dynamics を能動的に予測しないため高精度動作生成が制限 - PSR は予測表現を学習し action stream を拡張 - 結果として contact 関連の手がかりを多層的に活用可能 - PSR-VLA は π0.5 を 30.0pt、ForceVLA-π0.5 を 22.5pt、ForceVLA2-π0.5 を 19.2pt 上回る

3. 技術・手法の肝は?

- 事前学習段階で multimodal Transformer を訓練 - 未来の interaction dynamics を jointly に予測 - 予測表現の階層(hierarchy)を学習 - 学習した階層で action stream を augment - 多層の contact 関連 cue を policy が利用可能に - PSR を VLA モデルに実装し PSR-VLA を構成

4. どうやって有効だと検証した?

- 6つの実世界 contact-rich manipulation タスクで評価 - 全体成功率 91.7% を達成 - π0.5、ForceVLA-π0.5、ForceVLA2-π0.5 と比較 - それぞれ 30.0、22.5、19.2 ポイント改善 - タスク動画と stability tests を公開(https://psr-vla.pages.dev/)

5. 議論はある?

- 結果は force-aware な contact-rich manipulation への PSR の有効性を示す - 予測表現が高精度動作生成に寄与すると主張 - 限界や失敗事例、計算コスト、汎化性の議論は要旨からは不明 - 実世界6タスクのみで、sim や他ロボットへの転移は要旨からは不明

6. 次に読むべき論文は?

- π0.5 - ForceVLA-π0.5 - ForceVLA2-π0.5 - Vision-Language-Action (VLA) モデル - multimodal Transformer を用いた予測表現学習 - contact-rich manipulation における force-aware policy 学習

※ AIが要旨から生成した要約です。正確性は原文をご確認ください。

著者: Shengbao Li, Peng Xu, Chao Tang, Hao Wei, Jiaheng Wang, Hong Yin, Jiangtao Chen, Jinxuan Zhu, Zhong Zhou, Mengfan Wang, Tingguang Li

分類: cs.RO

原文アブストラクト

Contact-rich manipulation requires policies to generate precise actions by reasoning over contact forces, robot configurations, and interaction histories beyond visual observations. Existing methods passively condition on force feedback rather than actively predicting future contact dynamics, limiting their ability to generate high-precision actions. To address this problem, we introduce Predictive Sensorimotor Representation (PSR) learning, a framework that learns a hierarchy of predictive representations from multimodal sensorimotor signals and integrates them into the action stream of a visuomotor policy. Specifically, during a pretraining stage, a multimodal Transformer is trained to learn a hierarchy of predictive representations by jointly forecasting future interaction dynamics. The learned hierarchy subsequently augments the action stream, enabling the resulting policy to exploit contact-relevant cues at multiple depths. We further instantiate PSR within a Vision-Language-Action (VLA) model, resulting in PSR-VLA, and evaluate it on six real-world contact-rich manipulation tasks. Experimental results show that PSR-VLA achieves 91.7% overall success, improving over $π_{0.5}$, ForceVLA-$π_{0.5}$, and ForceVLA2-$π_{0.5}$ by 30.0, 22.5, and 19.2 percentage points, respectively. These results demonstrate the effectiveness of the proposed PSR for force-aware, contact-rich manipulation. Videos of the tasks and stability tests are available at https://psr-vla.pages.dev/.

関連論文

PR本紙発行元 EmplifAI