日本フィジカルAI新聞

世界のフィジカルAIを、日本語で。

週刊ニュースレター購読
VLAarXiv:2609.35450

Uni-VLaT:ヒューマノイドの移動操作のためのVLAポリシーへの全身触覚適応

Uni-VLaT: Whole-Body Tactile Adaptation of VLA Policies for Humanoid Loco-Manipulation

シェア:XThreadsFacebookLINEはてブBluesky

全身触覚を視覚言語行動ポリシーに統合し、未来の触覚・固有感覚・視覚表現を予測する補助タスクで学習することで、接触を伴う移動操作タスクの成功率を大幅に向上させた。

詳しい要約

1. どんなもの?

- ヒューマノイドの loco-manipulation における接触制御のための手法 Uni-VLaT を提案する研究。 - 全身の distributed tactile sensing を vision-language-action (VLA) ポリシーに統合する。 - 触覚経路の latent state を action 生成だけでなく、将来の tactile・proprioceptive・visual 表現の予測にも学習させる。 - 5つの実機タスクで評価し、平均成功率 75% を達成。

2. 先行研究と比べてどこがすごい?

- 触覚入力なしの baseline を 43 ポイント上回る。 - 予測 supervision なしの触覚入力 baseline を 7 ポイント上回る。 - 2つの pretrained VLA backbone で Table Sweeping を 30 ポイント、Back-Tap Walking を 85-90 ポイント改善。 - 従来の sparse な force/torque 測定ではなく、全身の distributed tactile sensing を活用する点が新しい。

3. 技術・手法の肝は?

- 触覚経路を追加し、その latent state を action 生成と将来予測の両方で学習。 - 予測目的により tactile-anchored multimodal context を構築し、物理世界の構造的理解を促す。 - 予測対象は将来の tactile・proprioceptive・visual 表現。 - Ablation で contextualized tactile prediction と absolute future targets が重要と示す。

4. どうやって有効だと検証した?

- 実機 5 タスクで評価:tactile-triggered locomotion、sustained physical interaction、human-robot contact、loco-manipulation。 - 平均成功率 75% を達成。 - 触覚なし baseline と予測 supervision なし触覚 baseline と比較。 - 2つの pretrained VLA backbone で Table Sweeping と Back-Tap Walking の改善を確認。 - Ablation を実施。

5. 議論はある?

- 予測的触覚学習が pretrained VLA ポリシーを全身物理インタラクションに拡張する有効な道筋であると主張。 - 限界や失敗事例、計算コスト、一般化性に関する議論は要旨からは不明。

6. 次に読むべき論文は?

- 要旨で参照/比較されている研究:触覚入力なし baseline、予測 supervision なし触覚入力 baseline、pretrained VLA backbone。 - 関連手法:vision-language-action (VLA) ポリシー、distributed tactile sensing、force/torque 測定。 - 同分野の定番:loco-manipulation、humanoid control、tactile sensing for robotics。

※ AIが要旨から生成した要約です。正確性は原文をご確認ください。

著者: Zihao Wang, Shutong Liu, Siqi Zheng, Liu Cao, Ruoqi Chen, Rundong Liu, Yanchao Yang, Mengdi Xu

分類: cs.RO

原文アブストラクト

Physical contact often determines how a humanoid should respond during loco-manipulation, yet vision and proprioception alone are often insufficient to characterize physical interaction, especially when the contact region is occluded. Unlike sparse force or torque measurements at predefined regions, distributed tactile sensing preserves spatially resolved contact patterns across the robot body. We therefore study how to integrate such whole-body tactile information into vision-language-action (VLA) policies for contact-rich control. Our approach, Uni-VLaT, introduces a tactile pathway whose latent state is trained not only for action generation, but also to predict future tactile, proprioceptive, and visual representations. This predictive objective builds a tactile-anchored multimodal context, encouraging a more structured understanding of the physical world. We evaluate Uni-VLaT on five real-robot tasks covering tactile-triggered locomotion, sustained physical interaction, human-robot contact, and loco-manipulation. Uni-VLaT achieves a 75% average success rate, outperforming a baseline without tactile input by 43 points and a tactile-input baseline without predictive supervision by 7 points. Across two pretrained VLA backbones, our method improves Table Sweeping by 30 points on both backbones and Back-Tap Walking by 85-90 points. Ablations further show that contextualized tactile prediction and absolute future targets are critical to performance. These results indicate that predictive tactile learning provides an effective route for extending pretrained VLA policies to whole-body physical interaction.

関連論文

PR本紙発行元 EmplifAI