画像から直接6自由度軌道計画を行うSE(3)ニューラルポテンシャル場
SE(3) Neural Potential Fields for 6-DoF Trajectory Planning Directly from Images Without Explicit 3D Reconstruction
RGB画像からSE(3)のニューラルポテンシャル場を学習し、3D再構成なしで衝突のない6自由度把持軌道を計画する手法を提案。ナビゲーション関数による教師信号で局所解を回避し、実機で高い把持成功率を達成した。
詳しい要約
1. どんなもの?
2. 先行研究と比べてどこがすごい?
3. 技術・手法の肝は?
4. どうやって有効だと検証した?
5. 議論はある?
6. 次に読むべき論文は?
※ AIが要旨から生成した要約です。正確性は原文をご確認ください。
著者: Jeffrey Eiyike, Masoud Ataei, Elvis Gyaase, Vikas Dhiman
分類: cs.RO, cs.AI
原文アブストラクト
Reaching a 6-DoF grasp pose in clutter requires a collision-free trajectory, conventionally obtained by reconstructing the scene in 3D and planning inside that reconstruction, at the cost of its accuracy and compute. Potential fields learned directly from images remove that dependency but inherit the classical weakness of artificial potential fields: where attractive and repulsive gradients cancel, the descent grazes the obstacle instead of going around it, and can stall short of the goal. We present an SE(3) neural potential field learned from posed RGB images and supervised with a navigation function, the geodesic distance to the grasp through free space recovered from those same images during training, which removes both failures. On two tabletop scenes, from obstacle-blocked starts executed on a UR10, the field converges within 3 cm of the grasp from every start and every path it executes is collision-free against the ground-truth geometry, against 25% and 0% under image supervision alone; mean clearance rises from under a centimeter to 8.6-8.8 cm and arm-link contacts fall from 20.6-50.4% to 2.7-5.5% of executed configurations. Executed grasp success is 90.0% and 40.0% on the two scenes, the residual failures being refusals of the Cartesian executor rather than of the field. Planning takes about 2 s against 67-133 s for RRT* on a reconstruction of the same images, though under a common offline harness the two are comparable: the deployed margin is the cost of collision-checking a dense reconstruction, not planner complexity.
関連論文
- DexTacWAM: 巧みな操作のための視触覚ワールドアクションモデルマニピュレーション
- 平面果樹園における視覚運動ロボット剪定のためのハイブリッド強化学習マニピュレーション
- CAST: 衝突を考慮した建設ロボットによる同時軌道推定と計画マニピュレーション
- ロボット構成空間における異種制約のための実行可能性距離場マニピュレーション
- InsertAnything: シミュレーションから現実への汎化可能な接触リッチ精密挿入マニピュレーション
- フレーズ単位のロボット古琴演奏:両腕動作計画と音触覚インタラクション監視マニピュレーション