ReSync: 非同期ワールドアクションモデルにおける二つのクロックの再同期
ReSync: Re-Aligning the Two Clocks of Asynchronous World-Action Models
未来の映像と行動を別々のスケジュールで生成する非同期推論において、行動と映像の進捗のずれを「コミットメント・エビデンスギャップ」として定式化し、その間隔内で計算を再配置するReSyncを提案。パラメータ変更なしで成功率を4.48ポイント改善。
著者: Xi Lin, Feihong Zhang, Yulong Shi, Yanghong Mei, Zuxing Lu, Xiaofan Zhu, Zihao Liang, Zhirui Gao, Zhaowen Li
分類: cs.RO
原文アブストラクト
Jointly generating future video and actions has become a standard recipe for world-action models, and the strongest systems denoise the two streams on separate schedules: actions are decoded in few steps so control stays fast, while the video stream runs longer to keep the predicted future sharp. The design is deliberate, but it leaves the two streams on different clocks, and an action can become executable while the future that should justify it is still largely unresolved. We formalize this as a two-clock view of asynchronous inference and introduce the commitment-evidence gap, a quantity read directly from a model's own sampling schedule rather than measured by search. The gap is predictive: as it widens, candidate utility becomes harder to identify and extra candidate sampling buys less, while advancing the world stream buys more, and the two cross. Spending more world computation is therefore not simply better. The useful interval is closed at both ends, and both ends can be read off the schedule before any rollout. ReSync places the computation inside it: hold the action state, advance only the world within the supported window, then resume native denoising. No parameters change and no candidates are compared. On a frozen paired RoboCasa panel this improves success by 4.48 points, while an equal-compute control that waits without advancing the world does not move, and the same rule transfers to a second benchmark and a second backbone without retuning.