日本フィジカルAI新聞

世界のフィジカルAIを、日本語で。

週刊ニュースレター購読
マニピュレーションarXiv:2609.12549

STAR: 視覚・触覚・言語・行動モデルにおけるスパース触覚表現学習による巧みなマニピュレーション

STAR: Sparse Tactile Representation Learning in Vision-Tactile-Language-Action Models for Dexterous Manipulation

シェア:XThreadsFacebookLINEはてブBluesky

200時間の双腕巧みな操作データセットを構築し、触覚信号のスパース性に対処する視覚・触覚・言語・行動モデルSTARを提案。実世界4タスクで平均61%の成功率を達成した。

詳しい要約

1. どんなもの?

- 器用な多指操作のための視覚・触覚・言語・行動を統合したVTLAモデル「STAR」を提案。 - 200時間の両手器用操作データセットを構築。視覚・触覚・言語注釈付きで10,576軌道、65タスク、69.5%が多指操作。 - 触覚信号の空間的・時間的・情報的疎性に対処する統合学習レシピ。 - 実世界4タスクで平均成功率61%を達成。

2. 先行研究と比べてどこがすごい?

- 大規模実世界データ不足と疎な触覚信号からの表現抽出困難を解決。 - 従来のVTLAモデルと比べ、触覚の疎性を明示的に扱う初の統合学習レシピ。 - 200時間の両手操作データセットは既存にない規模と多様性(65タスク、多指操作69.5%)。 - タスク特化のポストトレーニングで器用な操作を実現。

3. 技術・手法の肝は?

- 視覚-触覚ジョイント事前学習:視覚と触覚の対応を学習。 - スパースグローバル触覚トークン表現:疎な触覚信号を効率的にトークン化。 - スパース未来触覚予測:将来の触覚信号を予測し、時間的疎性に対処。 - これらを統合したVTLAモデル学習レシピ。

4. どうやって有効だと検証した?

- 構築したデータセットでSTARを訓練。 - 実世界4タスクで評価、各タスク100ポストトレーニング軌道を使用。 - 平均成功率61%を達成し、器用な操作性能を実証。

5. 議論はある?

- 要旨からは不明。

6. 次に読むべき論文は?

- 要旨で参照/比較されている研究は明記されていない。 - 関連手法としてVision-Language-Action (VLA) モデル、触覚表現学習、模倣学習、大規模ロボットデータセット(例:Open X-Embodiment)が挙げられる。

※ AIが要旨から生成した要約です。正確性は原文をご確認ください。

著者: Xiangcheng Liu, Tianhao Wu, Le Zheng, Yidong Wang, Bowen Jiang, Mingjie Pan, Xinlin Ren, Yi Liu, Jianlan Luo

分類: cs.RO

原文アブストラクト

Dexterous manipulation requires coordinated multi-finger control and effective tactile feedback, yet learning these capabilities remains challenging due to the lack of large-scale real-world data and the difficulty of extracting effective representations from sparse tactile signals. We build a robot platform and teleoperation system to collect a 200-hour bimanual dexterous manipulation dataset with synchronized visual, tactile, and language annotations, comprising 10,576 trajectories across 65 tasks, 69.5% of which involve dexterous multi-finger manipulation. We further propose STAR, an integrated training recipe for vision-tactile-language-action (VTLA) models that addresses the spatial, temporal, and informational sparsity of tactile signals through visual-tactile joint pre-training, sparse-global tactile token representation, and sparse future tactile prediction. Trained on this dataset, STAR achieves a 61% average success rate across four real-world tasks with 100 post-training trajectories per task, demonstrating dexterous performance under task-specific post-training.

関連論文