日本フィジカルAI新聞

世界のフィジカルAIを、日本語で。

週刊ニュースレター購読
VLAarXiv:2609.34911

捨てた行動を再利用してロボット政策を高速化するAction Upcycling

Don't Throw Away the Tail: Action Upcycling for Policy Acceleration

シェア:XThreadsFacebookLINEはてブBluesky

行動チャンクの未使用部分を再利用し、速度が滑らかな範囲で実行時間を延ばすことで、追加学習やモデル内部情報なしに政策呼び出しを1.2〜1.7倍削減する手法を提案。

詳しい要約

1. どんなもの?

- ロボット政策は観測から未来のaction chunkを予測し、その一部(execution horizon)のみ実行して残りを破棄する。 - このhorizonの長さは反応性と効率のトレードオフを持つ。 - 本研究は、破棄されるactionを再利用するtraining-freeなアルゴリズム『Action Upcycling』を提案する。 - モデル内部にアクセスせず、追加サンプリングも行わない。 - 破棄actionと再計画actionはaction velocityが滑らかなら近いという観察に基づく。 - velocityが揺らぎ始めるまでexecution horizonを延長する。

2. 先行研究と比べてどこがすごい?

- 従来のtest-time horizon適応手法は、モデル内部を読む(アーキテクチャごとに信号選択が必要)か、追加サンプルを引く(コスト増)かのいずれか。 - Action Upcyclingはモデル内部にアクセスせず、追加サンプルも引かない。 - 破棄されるactionを再利用する点が新しい。 - 任意のchunked policyに適用でき、few-step samplingやstreaming action decodingなどの他の加速手法と直交する。 - 政策加速の新しい軸を開く。

3. 技術・手法の肝は?

- training-freeアルゴリズム。 - 政策が破棄するactionを再利用する。 - 破棄actionと再計画actionはaction velocityが滑らかな限り近いという観察を利用。 - velocityが揺らぎ始める点までexecution horizonを延長する。 - モデル内部や追加サンプリングは不要。 - 任意のchunked policyに低コストで適用可能。

4. どうやって有効だと検証した?

- シミュレーションおよび実世界のmanipulationタスクで広範な実験を実施。 - 複数のVision-Language-Action Models (VLAs)とWorld Action Model (WAM)で評価。 - 成功率を落とさずに政策呼び出しを1.2--1.7倍削減。 - 他の加速手法(few-step sampling, streaming action decoding)と直交することも示す。

5. 議論はある?

- 要旨からは不明。 - ただし、action velocityの滑らかさに依存する点や、velocityが揺らぎ始める閾値の決定方法などが議論の余地として考えられる。 - 限界や失敗ケースについての記述は要旨にはない。

6. 次に読むべき論文は?

- 要旨で参照/比較されている研究: test-time horizon適応手法、few-step sampling、streaming action decoding。 - 関連手法: Vision-Language-Action Models (VLAs)、World Action Model (WAM)。 - 同分野の定番: action chunking, receding horizon control, temporal ensembling, diffusion policy。

※ AIが要旨から生成した要約です。正確性は原文をご確認ください。

著者: Taesung Kwon, Jangho Park, Sunwoo Park, Youngmin Kim, Seonghyun Jin, Youngjun Jun, Kyumin Choi, Jong Chul Ye

分類: cs.RO, cs.AI, cs.CV, cs.LG

原文アブストラクト

Modern robot policies predict a chunk of future actions from a single observation, execute only a prefix, and discard the rest before replanning. Choosing the length of this prefix, the execution horizon, poses a trade-off between reactivity and efficiency. A short horizon keeps the policy reactive to the environment, but requires frequent policy calls. Recent test-time methods adaptively select the horizon for each chunk, but they either read model internals, where the signal must be chosen for each architecture, or draw extra samples, which adds cost. We propose *Action Upcycling*, a training-free algorithm that reuses actions the policy would otherwise discard, without accessing model internals or drawing extra samples. We find that discarded actions stay close to their replanned versions as long as the action velocity remains smooth. Action Upcycling therefore extends the execution horizon up to the point where the velocity begins to fluctuate. Extensive experiments on simulated and real-world manipulation tasks show that Action Upcycling reduces policy calls by 1.2--1.7$\times$ with no loss in success rate, across multiple Vision-Language-Action Models (VLAs) and even a World Action Model (WAM). It applies to any chunked policy at negligible cost and is orthogonal to other policy acceleration methods such as few-step sampling and streaming action decoding, opening a new axis for policy acceleration.

関連論文

PR本紙発行元 EmplifAI