日本フィジカルAI新聞

世界のフィジカルAIを、日本語で。

週刊ニュースレター購読
VLAarXiv:2609.38989

フローの手がかり:オープンワールド配送マニピュレーションのためのフローマッチングポリシーの操舵

Cue the Flow: Steering Flow-Matching Policies for Open-World Delivery Manipulation

シェア:XThreadsFacebookLINEはてブBluesky

事前学習済みフローマッチングVLAモデルを低レベル制御に用い、空間的手がかりでポリシーを操舵する軽量アダプタを提案。未知物体への指示追従と操作成功率を最大2倍改善。

著者: Haoxuan Wang, Griffin Galimi, Junhua Huang, Selina Song, Wayne Wu, Yan Yan, Bolei Zhou

分類: cs.RO

原文アブストラクト

Open-world goods delivery requires mobile manipulators to follow free-form user instructions and manipulate potentially novel objects. Existing dual-system approaches use high-level grounding models to convert language into grounded visual prompts, but their low-level controllers can remain brittle under noisy perception, dynamic scenes, and contact-rich interactions. We instead use a pretrained flow-matching vision-language-action model as the low-level control interface, leveraging its reactivity and robustness to environmental changes while treating the grounding output as a spatial cue for policy steering. Our key insight is that the pretrained VLA already provides a strong manipulation prior, while the spatial cue supplies the missing target information needed to guide actions under novel language--object mappings. Concretely, we introduce a lightweight cue-conditioned adapter. The adapter is first trained with contrastive objectives to produce salient and spatially discriminative cue representations, and is then supervised to predict a diagonal affine transformation over the generated action chunk, aligning policy steering with the cued target. Across tabletop and mobile-base settings, our method improves instruction following and manipulation success on both in-domain and out-of-domain objects, achieving up to near $2\times$ improvement in average task success rate with negligible inference overhead.

関連論文

PR本紙発行元 EmplifAI