UMI型ロボット教示のための高精度エンドエフェクタ位置推定
Towards Accurate End-Effector Localization for UMI-Style Robotic Manipulation Teaching
近距離マニピュレーション教示向けに実世界・シミュレーションの位置推定データセットMILDを構築し、フィッシュアイVIOとAprilTagを組み合わせたAprilVINSを提案した論文。
詳しい要約
1. どんなもの?
2. 先行研究と比べてどこがすごい?
3. 技術・手法の肝は?
4. どうやって有効だと検証した?
5. 議論はある?
6. 次に読むべき論文は?
※ AIが要旨から生成した要約です。正確性は原文をご確認ください。
著者: Junjie Zhang, Deteng Zhang, Zhisong Xu, Bo Sun, Liuyang Li, Yihong Tian, Jie Yin
分類: cs.RO
原文アブストラクト
Robot demonstration learning requires accurate and temporally complete end-effector localization during close-range manipulation and camera occlusion. Existing SLAM benchmarks emphasize navigation motions, whereas manipulation datasets prioritize policy learning over localization evaluation. We introduce MILD, a Manipulation-Interface Localization Dataset with real-world and simulation sequences. The real-world subset provides 86 sensor sequences from Insta360 X5 and Insight9 across 15 repeated tabletop tasks, calibration assets, and a per-execution robot end-effector reference trajectory. The simulation subset, MILD-Sim, extends task coverage in Isaac Sim for controlled manipulation-replay studies. Benchmarking visual-inertial and fiducial-aided systems on instrumented real-world recordings reveals large differences in both TCP-relative trajectory error and temporal coverage, even under the same nominal task. To support marker-augmented teaching workspaces without a pre-surveyed fiducial map, we present AprilVINS, which combines fisheye visual-inertial estimation with sequence-local AprilTag geometry and separates prior admission from guarded export of the jointly optimized state. On Insta360 AprilTag4 recordings, AprilVINS(full) under a unified protocol with sequence-specific profiles reaches millimeter-level SE(3)-aligned TCP-relative APE RMSE with high time completion and lower reported error than the tested routes under their respective protocols, whereas fisheye VIO without tag factors remains at centimeter scale. Ablations separate accuracy from exportability, and a MILD-Sim replay study provides task-specific tolerance references for interpreting those error magnitudes. Together, MILD and AprilVINS provide a diagnostic benchmarking framework for UMI-style demonstration collection. Code, datasets, and evaluation manifests will be released upon acceptance.
関連論文
- 巧みな把持安定性のための時間的視触覚学習マニピュレーション
- FoldBack: 長期的な衣類折り畳みのための自己修正型マスク生成ポリシーマニピュレーション
- 関節特化型ハイブリッド遠隔駆動を用いた全駆動4自由度ロボット指の設計マニピュレーション
- RoboPace: 接触を考慮した行動チャンク方策の時間最適リタイミングマニピュレーション
- 断続的な視覚喪失に頑健な実ロボットマニピュレーションのための標的モダリティドロップアウトマニピュレーション
- 腱の巻き付きとショートカットを考慮した腱駆動ロボットの設計最適化マニピュレーション