日本フィジカルAI新聞

世界のフィジカルAIを、日本語で。

週刊ニュースレター購読
VLAarXiv:2609.39526

離散フォーシング:連続デノイジングに離散ガイダンスを注入する少数ステップ行動エキスパート

Discrete Forcing: Infusing Discrete Guidance into Continuous Denoising for Few-Step Action Experts

シェア:XThreadsFacebookLINEはてブBluesky

VLAモデルにおいて、離散行動トークンで粗い行動構造を予測し、それを連続行動の精緻化のガイダンスとして用いるフローマッチング手法を提案。パラメータ数を抑えつつ高精度・高速推論を実現した。

著者: Jingbo Wang, Wenxuan Song, Wenhao Yu, Han Zhao, Xi Wang, Jiayi Chen, Donglin Wang, Yan Wang, Haoang Li

分類: cs.RO

原文アブストラクト

Efficient action generation in vision-language-action (VLA) models requires capturing both coarse action structure and fine-grained details. Discrete action tokens provide compact structural representations but sacrifice precision, while continuous action tokens offer high precision but often require multiple denoising steps. We introduce Discrete Forcing, a flow-matching framework that combines these representations through an explicit coarse-to-fine generation process. It first predicts discrete action tokens to establish a coarse action structure, then uses them to guide continuous action refinement. The discrete and continuous components share a common diffusion transformer backbone with specialized branches, maintaining a parameter count comparable to a conventional single-branch model while requiring only one forward pass per branch. Extensive evaluations across multiple benchmarks demonstrate improved performance and faster inference over a parameter-matched continuous action expert, with consistent performance gains as model capacity increases. Real-world experiments further demonstrate improvements on high-precision and dynamic manipulation tasks.

関連論文

PR本紙発行元 EmplifAI