双腕モバイルマニピュレーションのための自動ラベリング
Automatic Labelling for Bimanual Mobile Manipulation
同期した運動信号を位相に分割し、視覚言語推論で意味づけを行うことで、双腕モバイルマニピュレーションのサブタスクに自動でラベルを付与するパイプラインを提案し、実タスクで有効性を検証した。
詳しい要約
1. どんなもの?
2. 先行研究と比べてどこがすごい?
3. 技術・手法の肝は?
4. どうやって有効だと検証した?
5. 議論はある?
6. 次に読むべき論文は?
※ AIが要旨から生成した要約です。正確性は原文をご確認ください。
分類: cs.RO
原文アブストラクト
Semantically meaningful subtask labels can provide useful contexts for long-horizon policies, but automatically identifying both reliable temporal boundaries and broad semantic descriptions for annotations remains difficult. We present an automatic labelling pipeline that assigns temporal localisation to deterministic trajectory analysis and semantic interpretation to vision-language (VL) reasoning. The pipeline segments synchronised kinematic signals into phases, performs phase-localised VL reasoning to describe the contents, and aggregates the outputs for the base, left arm, and right arm actions. We evaluate this pipeline primarily on 29 real Galaxea bimanual mobile-manipulation tasks. Repeating the VL reasoning three times first produces the same output value for 87.4% on selected tasks. A review by nine participants across all 29 tasks then judgements on the labelled phases and shows positive acceptance of temporal divisions (90.5%), body labels (90.7%), and arm labels (78.7%). The results indicate that the segmentation-VL design can produce structured annotations while preserving asynchronous bimanual behaviour, providing a basis for richer semantic subtask identification and state-based verification.
関連論文
- DexTacWAM: 巧みな操作のための視触覚ワールドアクションモデルマニピュレーション
- 平面果樹園における視覚運動ロボット剪定のためのハイブリッド強化学習マニピュレーション
- CAST: 衝突を考慮した建設ロボットによる同時軌道推定と計画マニピュレーション
- ロボット構成空間における異種制約のための実行可能性距離場マニピュレーション
- InsertAnything: シミュレーションから現実への汎化可能な接触リッチ精密挿入マニピュレーション
- フレーズ単位のロボット古琴演奏:両腕動作計画と音触覚インタラクション監視マニピュレーション