日本フィジカルAI新聞

世界のフィジカルAIを、日本語で。

週刊ニュースレター購読
触覚arXiv:2609.20761

Agile-WAM:接触の多いロボット制御のための俊敏な触覚ワールドアクションモデル

Agile-WAM: An Agile Tactile World Action Model for Contact-Rich Robot Control

シェア:XThreadsFacebookLINEはてブBluesky

視覚と触覚を共有潜在空間に符号化し、フローマッチングで行動と将来の視覚・触覚潜在を同時生成する触覚ワールドアクションモデルを提案。接触の時間スケールの違いを考慮したマルチホライズン予測で高精度かつ低遅延な接触リッチ操作を実現した。

詳しい要約

1. どんなもの?

- 接触の多いロボット制御のためのAgile Tactile World Action Model (Agile-WAM)を提案。 - 視覚と触覚の観測を共有潜在表現にエンコードし、直接的なvision-tactile-to-action flow-matchingプロセスで行動チャンクと将来の視覚/触覚潜在表現を同時生成。 - 視覚と触覚の時間スケールの違いに対応するため、multi-horizon multimodal predictionを導入。 - 9つのシミュレーションと5つの実世界接触操作タスクで評価。

2. 先行研究と比べてどこがすごい?

- 従来のtactile WAMは大規模事前学習生成バックボーンに依存し、推論効率と展開の柔軟性が制限されていた。 - Agile-WAMはアジャイルなアーキテクチャで、低推論レイテンシを維持しつつ高い成功率を達成。 - 実世界5実験で最強ベースラインに対し成功率29.4%の相対改善、推論レイテンシ11.9 msを実現。

3. 技術・手法の肝は?

- 視覚と触覚の観測を共有潜在表現にエンコード。 - 直接的なvision-tactile-to-action flow-matchingプロセスにより、行動チャンクと将来の視覚/触覚潜在表現を同時生成。 - 視覚と触覚の時間スケールの違いに着目し、multi-horizon multimodal predictionを導入:視覚潜在表現はより大きな時間オフセットで監督し、触覚潜在表現は次フレームで予測して細かい接触ダイナミクスを捉える。

4. どうやって有効だと検証した?

- 9つのシミュレーションと5つの実世界接触操作タスクで評価。 - 最強ベースラインを上回る成功率を示し、低推論レイテンシを維持。 - 実世界5実験で全体成功率29.4%の相対改善、推論レイテンシ11.9 msを達成。

5. 議論はある?

- 要旨からは不明。

6. 次に読むべき論文は?

- 要旨で参照/比較されている研究や関連手法は明示されていない。 - 同分野の定番として、World Action Models (WAMs)、visuomotor policies、tactile WAMs、flow-matching、multi-horizon multimodal predictionなどが関連する。

※ AIが要旨から生成した要約です。正確性は原文をご確認ください。

著者: Hanchu Zhou, Brendan Lynch, Raman Goyal, Dechen Gao, Begum Kasap, Boqi Zhao, Junshan Zhang

分類: cs.RO, cs.LG

原文アブストラクト

World Action Models (WAMs) advance beyond conventional visuomotor policies by jointly predicting future world states and robot actions, enabling the policy to learn physical dynamics that support effective control. However, recent tactile WAMs often rely on large-scale pretrained generative backbones to capture contact-rich physical dynamics, which limit their inference efficiency and flexible deployment. In this paper, we present \ABBR{}, an agile tactile World Action Model for contact-rich robot control. \ABBR{} encodes visual and tactile observations into a shared latent that serves as the source of a direct vision-tactile-to-action flow-matching process, which can jointly generate latent representations of action chunks and future visual/tactile latents. A key observation is that vision and tactile signals evolve at inherently different timescales: adjacent visual frames are often highly similar, whereas tactile signals can change abruptly upon contact. We therefore introduce multi-horizon multimodal prediction in \ABBR{}, which provides supervision for visual latent at a larger temporal offset while predicting the tactile latent in the next frame to capture fine-grained contact dynamics. Across nine simulated and five real-world contact-rich manipulation tasks, \ABBR{} demonstrates strong and robust performance, outperforming the strongest baseline in success rate while maintaining low inference latency. In particular, in five real-world experiments, \ABBR{} yields a relative gain of $\textbf{29.4\%}$ in overall success rates while achieving inference latency of $\textbf{11.9 ms}$. These results demonstrate that multimodal WAM can be achieved with an agile architecture suitable for precise and high-frequency robot control. More details are available on our project page: https://hanchuzhou.github.io/TARO_project_page/.

関連論文

PR本紙発行元 EmplifAI