日本フィジカルAI新聞

世界のフィジカルAIを、日本語で。

週刊ニュースレター購読
VLAarXiv:2610.09718

YUBI-STAG: 自動動画言語グラウンディングによるVLAのための接触・意味豊富なアラインメント

YUBI-STAG: Contact and Semantic-Rich Alignment for VLAs via Automated Video-Language Grounding

シェア:XThreadsFacebookLINEはてブBluesky

ロボット実演に接触物体や把持動作などの詳細な意味情報を自動付与し、視覚言語行動モデルをきめ細かい言語指示と整合させるフレームワークを提案。

詳しい要約

1. どんなもの?

- 視覚言語行動(VLA)モデルを微細な操作言語に整合させるためのフレームワーク - YUBI-STAG:操作デモンストレーションに接触物体セグメンテーションと視覚言語モデルを組み合わせて時空間注釈を自動付与 - 物体の同一性、属性、状態、グリッパごとの行動、両手協調、空間的に接地された相互作用を注釈 - YUBI-VLM:YUBI-STAGを蒸留し、未分割動画から直接行動構造と注釈を少数の推論で復元 - 手首視点のみで動作可能 - YUBI-STAG-Benchで評価

2. 先行研究と比べてどこがすごい?

- 既存のロボットデモンストレーションは粗いタスク記述のみで、行動の実行方法(どのグリッパが動作するか、どの物体に接触するか、どのように把持・移動するか)を欠く - YUBI-STAGはこれらの欠落を自動的に補い、微細な操作言語との整合を可能にする - YUBI-VLMはYUBI-STAGの局所化されたシーケンスと多段階VLM推論への依存を解消し、未分割動画から直接注釈を復元 - 少数の推論呼び出しと短い実行時間で高い注釈精度を維持し、未見の操作に一般化

3. 技術・手法の肝は?

- 接触物体セグメンテーションと視覚言語モデルを組み合わせて、操作デモンストレーションに相互作用に富む意味論を自動付与 - 物体の同一性、属性、状態、グリッパごとの行動、両手協調、空間的に接地された相互作用を注釈 - YUBI-STAGを蒸留してYUBI-VLMを構築 - YUBI-VLMは未分割動画から直接行動構造と注釈を少数の推論で復元し、手首視点のみで動作

4. どうやって有効だと検証した?

- YUBI-STAG-Benchで時間的、意味的、空間的接地タスクを評価 - YUBI-VLMがYUBI-STAGの注釈精度の多くを維持しつつ、推論呼び出し回数と実行時間を削減 - 未見の操作への一般化を確認 - これらの注釈でVLAポリシーをポストトレーニングし、微細な言語と接触認識構造に整合 - 両手実験で性能と指示追従の改善を実証(物体同一性、動作グリッパ、目標位置、空間関係の制御を含む)

5. 議論はある?

- 要旨からは不明

6. 次に読むべき論文は?

- 要旨で参照/比較されている研究は明示されていない - 同分野の定番として、Vision-Language-Action (VLA) models、Vision-Language Models (VLMs)、contact-object segmentation、bimanual manipulation、robot demonstrations、post-training of VLA policies などが関連

※ AIが要旨から生成した要約です。正確性は原文をご確認ください。

著者: Masatoshi Tateno, Takehiko Ohkawa, Yueh-Hua Wu, Hanlong Li, Tatsuya Matsushima, Yoichi Sato, Kei Ota

分類: cs.RO, cs.CV

原文アブストラクト

Vision-Language-Action (VLA) models acquire broad manipulation capabilities via large-scale pretraining, yet eliciting them through language requires fine-grained alignment between instructions and physical interactions. Existing robot demonstrations typically provide only coarse task descriptions, omitting how actions are executed, including which gripper acts, which object is contacted, and how it is grasped and moved. We introduce YUBI-STAG, a framework for Spatio-Temporal Annotation and Grounding that automatically enriches manipulation demonstrations with interaction-rich semantics to align pretrained VLAs with fine-grained manipulation language. Combining contact-object segmentation with vision-language models, YUBI-STAG annotates object identities, attributes and states, per-gripper actions, bimanual coordination, and spatially grounded interactions. To address YUBI-STAG's reliance on localized sequences and multi-stage VLM inference, we distill it into YUBI-VLM. YUBI-VLM directly recovers action structure and annotations from raw, unsegmented video in few inference calls and operates from wrist views alone. We evaluate both frameworks on YUBI-STAG-Bench across temporal, semantic, and spatial grounding tasks. YUBI-VLM retains much of YUBI-STAG's annotation accuracy with fewer inference calls and shorter runtime while generalizing to unseen manipulations. Finally, post-training VLA policies on these annotations aligns them with fine-grained language and contact-aware structure. Bimanual experiments demonstrate improved performance and instruction following, including control over object identity, acting gripper, target location, and spatial relations absent from original labels.

関連論文

PR本紙発行元 EmplifAI