日本フィジカルAI新聞

世界のフィジカルAIを、日本語で。

週刊ニュースレター購読
VLAarXiv:2610.02759

固有感覚スケッチを長期意図として用いる生成的行動ポリシー

Proprioceptive Sketches as Long-Horizon Intent for Generative Action Policies

シェア:XThreadsFacebookLINEはてブBluesky

ロボットの残りの関節空間経路を時間に依存しないコンパクトなスケッチとして生成し、それを条件に実行可能な行動チャンクを生成するPAMを提案。実機双腕タスクで成功率を47.5%から75.0%に向上させた。

詳しい要約

1. どんなもの?

生成的なロボット方策は短いaction chunkを予測するが、長期的な意図を明示的に持たない。本研究はProprioceptive Action Models (PAM)を提案し、ロボットの残りのjoint-space pathのコンパクトでtiming-freeなsketchと、密で実行可能なaction chunkを単一のtransformer denoiser内で同時生成する。sketchは時間ではなくarc lengthでpathをパラメータ化し、実行タイミングに不変な幾何学的意図を捉える。

2. 先行研究と比べてどこがすごい?

従来はlanguage plans、subgoal images、video forecastsで長期的構造を露出させていたが、生成コストが高くロボット動作への変換が必要だった。将来のロボット動作予測は変換を避けられるが、密なtime-indexed trajectoryは残りタスク全体をカバーするのに多数のパラメータを要し、短horizonではaction chunkをほぼ繰り返すだけでaction生成への指針が乏しい。PAMはtiming-free sketchによりこれを改善する。

3. 技術・手法の肝は?

sketchを時間ではなくarc lengthでパラメータ化し、実行タイミングに不変な幾何学的意図を捉える。Block-causal attentionとstaggered denoising scheduleにより、sketchからactionへの有向依存を維持し、action tokensがサンプリング全体を通じて漸進的にクリーンなsketchに条件付けられる。

4. どうやって有効だと検証した?

シミュレーションではPush-TとLIBERO-Longにおいてaction-only counterpartsより改善。実世界の4つのbimanualタスクでは成功率を47.5%から75.0%へ向上させた。

5. 議論はある?

要旨からは不明。

6. 次に読むべき論文は?

language plans、subgoal images、video forecasts、action-only counterparts、Push-T、LIBERO-Long、bimanualタスクに関連する研究。

※ AIが要旨から生成した要約です。正確性は原文をご確認ください。

著者: Fangyuan Wang, Songhao Huang, Haoxiang Sun, Shipeng Lyu, Chengyang He, Anqing Duan, Peng Zhou, David Navarro-Alarcon

分類: cs.RO

原文アブストラクト

Generative robot policies predict short action chunks but lack explicit long-horizon intent. Recent methods expose longer-horizon structure through language plans, subgoal images, or video forecasts, which are costly to generate and still need to be translated into robot motion. Predicting future robot motions avoids this translation, but a dense, time-indexed trajectory requires numerous parameters to cover the full remaining task, and over a short horizon it largely repeats the action chunk and adds little guidance for action generation. We propose Proprioceptive Action Models (PAM), which jointly generate a compact, timing-free sketch of the robot's remaining joint-space path and a dense executable action chunk within a single transformer denoiser. The sketch parameterizes the path by arc length rather than time, capturing geometric intent invariant to execution timing. Block-causal attention and a staggered denoising schedule maintain directed sketch-to-action dependence, ensuring the action tokens condition on a progressively cleaner sketch throughout sampling. In simulation, PAM improves over its action-only counterparts on Push-T and LIBERO-Long; on four real-world bimanual tasks, it raises success from 47.5% to 75.0%. Project page: https://nicehiro.github.io/pam_dp/

関連論文

PR本紙発行元 EmplifAI