日本フィジカルAI新聞

世界のフィジカルAIを、日本語で。

週刊ニュースレター購読
VLAarXiv:2610.12007

REACT: ローリングデノイジングと二重分離によるVLAモデルを用いた反応型ロボット制御

REACT: Rolling Denoising and Dual Decoupling for Reactive Robot Control with VLA Models

シェア:XThreadsFacebookLINEはてブBluesky

フロー型VLAモデルでアクションチャンクを再生成せず、アクションバッファをずらしながらデノイズし続けることで、長期的な文脈を保ちつつ反応性を高める制御フレームワークを提案。

詳しい要約

1. どんなもの?

- どんなもの? - REACTは、flow-based VLAモデルを用いたロボット制御のためのrolling-denoisingフレームワーク。 - 長いaction chunkの滑らかさと頻繁なreplanningの反応性のトレードオフを解決する。 - 持続的なaction bufferとstaggered flow timestepsを維持し、各制御ステップで最新観測に基づき全horizonをdenoiseする。 - 最もクリーンなaction blockを実行し、部分的に洗練された未来のblockを前方にシフトし、新しいノイズを末尾に追加する。 - dual decouplingによりsensing、VLM encoding、DiT denoising、action executionを分離し、実用的な計算制約下で高頻度観測更新とaction streamingを可能にする。

2. 先行研究と比べてどこがすごい?

- 先行研究と比べてどこがすごい? - 従来のflow-based VLAはaction chunkを生成するが、chunked controlでは長いchunkは滑らかさを提供する一方、頻繁なreplanningは反応性を改善するがactionの不連続性を生む。 - REACTはrolling-denoisingにより、長いhorizonの文脈を保持しつつ反応性を向上させる。 - 頻繁なreplanningや非同期ベースラインと比較して、タスク成功率を向上させ、反応遅延を低減し、より滑らかな軌道を生成する。

3. 技術・手法の肝は?

- 技術や手法の肝はどこ? - 持続的なaction bufferとstaggered flow timestepsを維持する。 - 各制御ステップで最新観測を用いて全horizonをdenoiseする。 - 最もクリーンなaction blockを実行し、部分的に洗練された未来のblockを前方にシフトし、新しいノイズを末尾に追加する。 - これにより、各実行action blockは複数の最近の観測にわたって洗練される。 - dual decouplingによりsensing、VLM encoding、DiT denoising、action executionを分離し、高頻度観測更新とaction streamingを実現する。

4. どうやって有効だと検証した?

- どうやって有効だと検証した? - RoboTwin 2.0 simulation benchmarkと、複数のロボットプラットフォームでの実世界タスク(bimanual manipulationとdynamic controlを含む)で評価。 - 頻繁なreplanningおよび非同期ベースラインと比較して、タスク成功率の向上、反応遅延の低減、より滑らかな軌道を確認。

5. 議論はある?

- 議論はある? - 要旨からは不明。

6. 次に読むべき論文は?

- 次に読むべき論文は? - flow-based VLAモデル(例:Flow-based VLA) - RoboTwin 2.0 benchmark - 頻繁なreplanningや非同期制御のベースライン手法

※ AIが要旨から生成した要約です。正確性は原文をご確認ください。

著者: Houlong Xiong, Zhenqi Qiu, Zechen Wang, Suohang Zhang, Yiyu Ren, Wanting Xu, Hongfei Niu, Chengyang He, Ge Sun, Ran Cheng, Qian Zhu

分類: cs.RO, cs.AI

原文アブストラクト

Flow-based vision-language-action (VLA) models generate action chunks for temporally coherent robot motion, but chunked control creates a fundamental closed-loop trade-off: long chunks provide smooth execution, whereas frequent replanning improves reactivity at the cost of action discontinuities. We introduce REACT, a rolling-denoising framework that makes flow-based VLAs more reactive while preserving long-horizon context. Instead of regenerating entire action chunks from scratch, REACT maintains a persistent action buffer with staggered flow timesteps. At each control step, the full horizon is denoised using the latest observation, the cleanest action block is executed, partially refined future blocks are shifted forward, and fresh noise is appended to the tail. As a result, each executed action block is refined across multiple recent observations before deployment. To support real-time control, we further introduce dual decoupling, which separates sensing, VLM encoding, DiT denoising, and action execution, enabling high-frequency observation updates and action streaming under practical compute constraints. Across the RoboTwin 2.0 simulation benchmark and real-world tasks spanning bimanual manipulation and dynamic control on multiple robot platforms, REACT improves task success and reduces reaction latency while producing smoother trajectories than frequent-replanning and asynchronous baselines.

関連論文

PR本紙発行元 EmplifAI