日本フィジカルAI新聞

世界のフィジカルAIを、日本語で。

週刊ニュースレター購読
VLAarXiv:2610.00981

NarrativeFlow: ロボット速度場を用いたフローベース視覚-言語-行動モデル

NarrativeFlow: Flow-Based Vision-Language-Action Model Using Robot Velocity Fields

シェア:XThreadsFacebookLINEはてブBluesky

言語条件付き操作において、ロボットの速度場を連続的なフローとしてモデル化し、フローマッチングで生成する手法を提案。複数ロボットのデータを活用でき、実世界タスクで成功率が向上。

詳しい要約

1. どんなもの?

- 言語条件付きの flow-based manipulation を扱う研究。 - robot flows(robot velocity fields)を embodiment-agnostic かつ motion-centric な表現として用いる。 - 複数ロボットプラットフォームのデータを活用することを狙う。 - 言語条件付き manipulation は実用的ロボティクスに重要だが、embodiment 固有データ収集の負担が scaling を制限する。 - NarrativeFlow は言語条件付きで連続的な velocity field として robot flows をモデル化する。

2. 先行研究と比べてどこがすごい?

- 既存手法は robot flows を sparse keypoint displacements で粗く近似するか、言語条件付き manipulation を扱えない。 - NarrativeFlow は robot flows を flow-matching formulation により連続的な velocity field としてモデル化する。 - これにより実世界 manipulation と物理的に整合する robot flows を生成できる点が異なる。 - 標準データセットおよび実世界実験で代表的な baseline を上回る。

3. 技術・手法の肝は?

- robot flows を連続的な velocity field として扱う。 - flow-matching formulation を言語条件付きで用いる。 - 言語条件に基づき、物理的に整合する robot flows を生成する。 - robot flows を embodiment-agnostic な motion-centric 表現として利用する。 - 複数プラットフォームのデータ活用を可能にする枠組み。

4. どうやって有効だと検証した?

- 言語条件付き manipulation の標準データセットで実験。 - 標準評価指標において代表的な baseline 手法を上回ることを示す。 - 実世界実験を実施し、複数の manipulation タスクで baseline より高い成功率を達成。 - プロジェクトページで結果を公開。

5. 議論はある?

- 要旨からは不明。 - 限界や失敗事例、計算コスト、embodiment 間の汎化性能に関する議論は要旨では触れられていない。

6. 次に読むべき論文は?

- 要旨で参照・比較されている具体的な研究名は不明。 - 関連手法として flow-matching、vision-language-action model、language-conditioned manipulation、robot velocity fields に関する研究が挙げられる。 - 同分野の定番として RT-1、RT-2、Diffusion Policy などが考えられるが、要旨での直接参照は不明。

※ AIが要旨から生成した要約です。正確性は原文をご確認ください。

著者: Shota Kobayashi, Koki Seno, Daichi Yashima, Komei Sugiura

分類: cs.RO, cs.CV

原文アブストラクト

We focus on language-conditioned flow-based manipulation, where robot flows (robot velocity fields) serve as embodiment-agnostic, motion-centric representations for leveraging data collected from multiple robot platforms. This task is crucial because language-conditioned manipulation is essential for practical robotic systems, yet scaling robot foundation models remains limited by the labor-intensive collection of embodiment-specific data. Existing methods either coarsely approximate robot flows with sparse keypoint displacements, or cannot handle language-conditioned manipulation. To address this limitation, we propose NarrativeFlow, which models robot flows as continuous velocity fields using a flow-matching formulation conditioned on language. Accordingly, NarrativeFlow generates robot flows that are physically consistent with real-world manipulation. To validate NarrativeFlow, we have conducted experiments on standard datasets for language-conditioned manipulation. The experimental results show that NarrativeFlow outperforms representative baseline methods on standard evaluation metrics. Furthermore, through real-world experiments, we show that NarrativeFlow achieves higher success rates than baseline methods across multiple manipulation tasks. The project page is available at https://shota0520.github.io/NarrativeFlow-project-page/

関連論文

PR本紙発行元 EmplifAI