日本フィジカルAI新聞

世界のフィジカルAIを、日本語で。

週刊ニュースレター購読
VLAarXiv:2610.08425

MIM-VLA: グリッパモータフィードバックから物理的相互作用表現を学習する

MIM-VLA: Learning Physical Interaction Representations from Gripper Motor Feedback

シェア:XThreadsFacebookLINEはてブBluesky

グリッパの電流・位置・速度などのモータフィードバックを128次元の相互作用トークンに符号化し、SmolVLAの把持動作経路のみを条件付けることで、触覚センサなしに物体の抵抗比較や脆い物体の把持を実現した。

詳しい要約

1. どんなもの?

- 視覚と言語と行動から把持を推論する VLA ポリシーは接触後の物理応答を明示的に表現しない。 - 本研究は gripper の current・position・velocity・signal validity を 128 次元の interaction token に符号化する motor-feedback ベースの MIM-VLA を提案。 - motor-only の Motor Interaction Module (MIM) を人間が確認した contact と interaction-phase ラベルで事前学習。 - この token は SmolVLA の gripper-action pathway のみを条件付けし、arm action と position-control interface は変更しない。 - 同じ token は候補を比較する MEM selector VLM を支え、evidence-conditioned な選択と説明を生成する。

2. 先行研究と比べてどこがすごい?

- 従来の VLA は視覚観測と robot state から把持行動を推論し、接触後の物理応答を明示的に表現しない。 - MIM-VLA は gripper から既に得られる motor feedback を用い、追加の tactile array・force-torque sensor・calibrated force estimate・direct current control を必要としない。 - 13 の物体ペアで、より抵抗の高い物体を選ぶ割合は 75.0% で、SmolVLA baseline の 48.8% を上回る。 - 視覚的に異なる物体の interaction resistance 比較、視覚的に類似した real と replica の能動的 probing による曖昧性解消、fragile 物体の gentle grasping を含む実世界設定で評価。

3. 技術・手法の肝は?

- gripper の recent current・position・velocity・signal validity を 128 次元の interaction token に符号化する。 - motor-only の Motor Interaction Module (MIM) を human-reviewed な contact と interaction-phase ラベルで pretrain する。 - この token で SmolVLA の gripper-action pathway のみを条件付けし、arm action と position-control interface は変更しない。 - 同じ token を MEM selector VLM に用い、候補 interaction を比較して evidence-conditioned な選択と説明を生成する。

4. どうやって有効だと検証した?

- 3 つの実世界設定で評価: 視覚的に異なる物体の interaction resistance 比較、視覚的に類似した real と replica の能動的 probing による曖昧性解消、fragile 物体の gentle grasping(held-out instances を含む)。 - 13 の物体ペアで、MIM-VLA はより抵抗の高い物体を 75.0% の試行で選択し、SmolVLA baseline の 48.8% を上回った。 - 評価タスクにおいて、gripper から既に得られる motor feedback を使用し、追加の tactile array・force-torque sensor・calibrated force estimate・direct current control を必要としない。

5. 議論はある?

- 要旨からは不明。 - 限界、失敗事例、計算コスト、一般化可能性、ラベル作成コスト、他ハードウェアへの転用可能性についての議論は要旨に記載がない。

6. 次に読むべき論文は?

- SmolVLA(baseline として比較) - MEM selector VLM(本研究で用いられる選択・説明モジュール) - Motor Interaction Module (MIM)(本研究の事前学習モジュール) - 関連する VLA ポリシー研究(一般名として Vision-Language-Action policies) - 触覚・力覚を用いる grasping 研究(一般名として tactile sensing / force-torque sensing を用いる手法)

※ AIが要旨から生成した要約です。正確性は原文をご確認ください。

著者: Jaeyoung Lee, Jiyeon Koo, Taehwa Kim, Yerin Cha, Andrew Jaeyong Choi

分類: cs.RO

原文アブストラクト

Vision-language-action (VLA) policies infer grasp actions primarily from visual observations and robot state, but do not explicitly represent the physical response observed after contact. We present MIM-VLA, a motor-feedback-based architecture that encodes recent gripper current, position, velocity, and signal validity as a 128-dimensional interaction token. A motor-only Motor Interaction Module (MIM) is pretrained with human-reviewed contact and interaction-phase labels and then conditions only the gripper-action pathway of SmolVLA; arm actions and the position-control interface remain unchanged. The same token supports the MEM selector VLM that compares candidate interactions and produces evidence-conditioned selections and explanations. We evaluate MIM-VLA in three real-world settings: comparing the interaction resistance of visually different objects, disambiguating visually similar real and replica objects through active probing, and gently grasping fragile objects, including held-out instances. Across 13 object pairs, MIM-VLA selects the higher-resistance object in 75.0% of trials, compared with 48.8% for the SmolVLA baseline. For the evaluated tasks, the approach uses motor feedback already available from the gripper and does not require an additional tactile array, force-torque sensor, calibrated force estimate, or direct current control.

関連論文

PR本紙発行元 EmplifAI