Speech2Grasp: ヒューマノイドロボットにおけるテキスト条件付き把握検出の音声へのデータ効率的転移
Speech2Grasp: Data-Efficient Transfer of Text-Conditioned Grasp Detection to Speech in Humanoid Robots
テキスト条件付き把握検出モデルを音声入力に効率的に転移するフレームワークを提案し、実機ヒューマノイドでASRパイプラインより高精度かつ低遅延であることを示した。
著者: Hung Nguyen, Kim Nhat Minh Nguyen, Van Duc Vu, Van-Danh Le, Hoang Huy Le, Dinh Tuan Nguyen, Pham Tuyen Le, Van-Truong Nguyen, Quan Nguyen
分類: cs.RO, cs.CV
原文アブストラクト
Humanoid robots increasingly require multi-modal understanding for natural interaction with humans. Despite the prominence of vision-language models, they generally assume textual rather than the more natural speech inputs. In this paper, we investigate whether a well-established text-conditioned model can be transferred to speech in a data-efficient manner. Using ALBEF as a case study, we conduct diagnostic analyses showing that a lightweight MLP-based projector effectively adapts it to speech, while preserving semantic discrimination and robustness. Motivated by these findings, we introduce Speech2Grasp, a framework for data-efficient transfer of text-conditioned grasp detection to speech. Real-world humanoid robot experiments show that Speech2Grasp outperforms cascaded ASR-based pipeline, while reducing inference latency. Our findings suggest a practical paradigm for extending established text-conditioned systems to speech.
関連論文
- 否定制約付き器用把持のためのポテンシャル誘導粒子ステアリングマニピュレーション
- Facet-0: 接触を伴う精密操作のためのロボット基盤モデルマニピュレーション
- Peg-in-Bench: 高精度ロボット挿入のためのモジュール式ベンチマークマニピュレーション
- SUN: 言語に基づく制御から学習、実機への永続的プログラムマニピュレーション
- Motus2: 巧みな操作のための自己進化型汎用世界モデルマニピュレーション
- Zeva: 文脈内因果学習による汎用身体操作の実現マニピュレーション