日本フィジカルAI新聞

世界のフィジカルAIを、日本語で。

週刊ニュースレター購読
触覚arXiv:2610.10384

OpenViTac: 視覚触覚ポリシーを統一シミュレーション・実世界フレームワークで学習・評価する

OpenViTac: Learning and Benchmarking Visuo-Tactile Policies in a Unified Sim-and-Real Framework

シェア:XThreadsFacebookLINEはてブBluesky

接触を伴う操作タスクを4つの触覚能力次元で整理し、シミュレーションと実世界のペア環境でVLA・WAM・VTLAポリシーを評価するベンチマークOpenViTacを提案し、触覚表現と統合手法の比較やsim-real共訓練の分析を行った。

詳しい要約

1. どんなもの?

- 視覚触覚ポリシーをシミュレーションと実世界で統一的に評価するベンチマーク「OpenViTac」を提案。 - 接触を伴う操作を4つの触覚関連能力次元に整理し、シミュレーションと実世界のペア設定を提供。 - VLA、WAM、VTLAポリシーの一貫した評価を可能にする。 - 触覚表現と統合戦略が事前学習済みVLAモデルの性能に与える影響を調査。 - 最良の表現と統合戦略を組み合わせた触覚拡張フレームワーク「OpenVTLA」を導入。 - ペア設定を活用し、sim-real co-trainingとクロスドメインポリシー学習に影響する要因を分析。

2. 先行研究と比べてどこがすごい?

- 従来、VTLAポリシーの進歩にもかかわらず、シミュレーションと実世界にまたがる触覚を活用したロボット操作の統一ベンチマークが欠如していた。 - OpenViTacはこのギャップを埋め、シミュレーションと実世界のペア設定で一貫した評価を提供する点が新しい。 - 触覚表現と統合戦略の影響を体系的に調査し、OpenVTLAを提案。 - sim-real co-trainingの研究を可能にする点も先行研究にない特徴。

3. 技術・手法の肝は?

- 接触を伴う操作を4つの触覚関連能力次元に整理し、ベンチマークを構成。 - シミュレーションと実世界のペア設定を提供し、VLA、WAM、VTLAポリシーを評価。 - 異なる触覚表現と統合戦略を比較し、最良の組み合わせを特定。 - その結果を基にOpenVTLAフレームワークを構築。 - ペア設定を利用してsim-real co-trainingを実施し、クロスドメイン学習の要因を分析。

4. どうやって有効だと検証した?

- 要旨からは不明。具体的な実験設定や評価指標、被験者数などは記述されていない。 - ただし、ベンチマークを用いてVLA、WAM、VTLAポリシーを評価し、触覚表現と統合戦略の影響を調査したと述べられている。 - sim-real co-trainingの研究もペア設定を活用して行われた。

5. 議論はある?

- 要旨からは不明。議論や限界、今後の課題についての記述はない。 - ただし、触覚表現と統合戦略が性能に影響を与えること、クロスドメイン学習に要因があることが示唆されている。

6. 次に読むべき論文は?

- 要旨で参照/比較されている研究は明示されていない。 - 関連手法として、VLA、WAM、VTLAポリシーが挙げられる。 - 同分野の定番として、vision-tactile-language-action (VTLA) ポリシーや触覚を活用したロボット操作の研究が考えられる。

※ AIが要旨から生成した要約です。正確性は原文をご確認ください。

著者: Yifan Wu, Qin Li, Nan Min, Guojin Zhong, Haoyu Zhao, Zhiyuan Li, Houze Xu, Shengqi Xu, Xingyao Lin, Zijie Diao, Zhaoxiang Liu, Shiguo Lian, Shunlin Lu, Shihao Zhao, Ziyi Ye, Zuxuan Wu, Yu-Gang Jiang

分類: cs.RO

原文アブストラクト

Tactile feedback provides embodied agents with physical information beyond visual observations, enabling more reliable interaction with the real world. However, despite the rapid progress of vision-tactile-language-action (VTLA) policies, there remains a lack of unified benchmarks for evaluating tactile-enabled robot manipulation across simulation and the real world. To address this gap, we introduce OpenViTac, a visuo-tactile manipulation benchmark for evaluating robot policies across simulation and the real world. OpenViTac organizes contact-rich manipulation into four tactile-relevant capability dimensions and provides paired simulation-real-world settings for consistent evaluation of VLA, WAM, and VTLA policies. Building upon this benchmark, we investigate how different tactile representations and integration strategies affect the performance of pretrained VLA models. Correspondingly, we introduce OpenVTLA, a tactile augmentation framework that combines the best-performing representation and integration strategy. Furthermore, we leverage the paired benchmark setting to study sim-real co-training and analyze factors affecting cross-domain policy learning. Together, OpenViTac provides a unified platform for evaluating and advancing visuo-tactile robot manipulation.

関連論文

PR本紙発行元 EmplifAI