日本フィジカルAI新聞

世界のフィジカルAIを、日本語で。

週刊ニュースレター購読
VLAarXiv:2609.15012

言語で操作可能かつ力に応答するマニピュレーションのための原子運動座標

Atomic Motion Coordinate for Language-Steerable and Force-Responsive Manipulation

シェア:XThreadsFacebookLINEはてブBluesky

言語指示と順運動学から13種の運動原子を定義し、視覚に依存せずVLA方策の動作を言語で操舵可能にするとともに、接触履歴に応じて力を適応させる手法を提案した。

詳しい要約

1. どんなもの?

- VLA policy の end effector を言語指示のみで操作可能にする手法 - Atomic Motion Coordinate (AMC) を提案 - 13 個の signed translation/rotation/hold atom を各 arm に持つ - text と forward kinematics から grounding、vision は使わない - 力応答性も備えた座標系

2. 先行研究と比べてどこがすごい?

- LA4VLA-style と比較し opposite-atom separation が大幅向上 - 単一/dual で 92.5/83.1% 対 39.1/24.0% - 視覚駆動の motion prior を抑え言語操作を実現 - OOD fruit progress を 60.5% から 87.8% に改善 - force adaptation で Plug/Vase を 59.0/71.5% から 78.5/75.2% に改善

3. 技術・手法の肝は?

- 各 arm に 13 個の signed atom を定義 - text と forward kinematics から grounding、vision は withheld - weighted codebook alignment で action-expert block に注入 - contact history が bounded spherical residual で座標を変調 - 固定 nominal latent から未実行 horizon suffix のみ再生成

4. どうやって有効だと検証した?

- 7,520 回の offline horizon intervention で評価 - opposite-atom separation を測定 - 50 回の real-robot trial をタスクごとに実施 - OOD fruit progress と force adaptation の成功率を比較

5. 議論はある?

- 言語指示のみで end effector を操作可能かという問い - 視覚駆動の motion prior が支配的かどうかの議論 - 力応答性の実現方法と限界 - 要旨からは不明な点が多い

6. 次に読むべき論文は?

- LA4VLA-style の研究 - VLA policy 関連の先行研究 - forward kinematics を用いた grounding 手法 - codebook alignment の関連研究

※ AIが要旨から生成した要約です。正確性は原文をご確認ください。

著者: Jiaqi Zhai, Jingkai Zhao, Chen Yang, Siyuan Ma, Yutian Zhang, Liwen Yang, Qinglian Wu, Weiqi Fan, Yifei Wang, Yi Zheng, Chenxi Gu, Dong Wei, Wei Zhang

分類: cs.RO

原文アブストラクト

Can changing only the language instruction redirect a VLA policy's end effector, or does the visually driven motion prior dominate? We present Atomic Motion Coordinate, a geometry-grounded coordinate for steerable and force-responsive manipulation. Each arm owns thirteen signed translation, rotation, and hold atoms grounded from text and forward kinematics with vision withheld, and the coordinate is injected into every action-expert block via weighted codebook alignment. Contact history modulates the same coordinate through a bounded spherical residual that is recomputed from a fixed nominal latent to regenerate only the unexecuted horizon suffix. Across 7,520 offline horizon interventions, opposite-atom separation reaches 92.5/83.1% (single/dual) versus 39.1/24.0% for LA4VLA-style. Across 50 real-robot trials per task, AMC raises OOD fruit progress from 60.5% to 87.8%; force adaptation raises Plug/Vase from 59.0/71.5% to 78.5/75.2%.

関連論文