日本フィジカルAI新聞

世界のフィジカルAIを、日本語で。

週刊ニュースレター購読
VLAarXiv:2610.05026

GeoBridge-VLA:視覚言語行動モデルのための幾何認識残差適応

GeoBridge-VLA: Geometry-Aware Residual Adaptation for Vision-Language-Action Models

シェア:XThreadsFacebookLINEはてブBluesky

事前学習済みVLAの視覚エンコーダを凍結したまま幾何特徴を学習し、残差インターフェースで行動予測を強化する二段階手法を提案。LIBEROで70.9%、実機で74.0%の成功率を達成した。

著者: Hyun Song, Kangmin Kim, Loren Jinsoo Um, Minhui Han, Jaehyeok Park, Taewan Cho, Andrew Jaeyong Choi

分類: cs.CV, cs.RO

原文アブストラクト

Vision-language-action (VLA) models encode semantic information from vision-language pretraining, but manipulation also requires precise spatial reasoning. We present GeoBridge-VLA, a two-stage method for learning geometric features from a pretrained VLA's frozen visual encoder and using them for action prediction. Stage I trains a feature bridge and geometry decoder with depth supervision. Stage II freezes these modules and trains a gated residual interface together with the action-side projections and action expert. The residual augments the existing visual tokens without adding a second image encoder or increasing the token count. Deployment requires RGB, robot state, and language, but no depth observations. Under matched evaluation conditions, GeoBridge-VLA achieves 70.9% success on LIBERO, compared with 60.0% for SmolVLA. Disabling the residual in the same trained checkpoint reduces success from 70.90% to 69.85%, with mixed effects across suites. On a physical ROBOTIS OMY robot, GeoBridge-VLA succeeds in 148 of 200 trials (74.0%) across four tasks, compared with 108 of 200 (54.0%) for SmolVLA.

関連論文

PR本紙発行元 EmplifAI