日本フィジカルAI新聞

世界のフィジカルAIを、日本語で。

週刊ニュースレター購読
動作計画arXiv:2608.24026

NeurRAFT: アンカーレベルのフローマッチングとクリアランス認識選好チューニングによるロボット動作計画

NeurRAFT: Robot Motion Planning via Anchor-Level Flow Matching with Clearance-Aware Preference Tuning

シェア:XThreadsFacebookLINEはてブBluesky

本論文は、アンカーウェイポイントとフローマッチングを用いた生成型ニューラル動作プランナーを提案し、選好最適化により衝突回避性能を向上させる。

詳しい要約

1. どんなもの?

NeurRAFTは、生のセンサ観測から軌道を生成するエンドツーエンドのニューラルモーションプランナーである。アンカーレベルのフローマッチングとクリアランスを考慮した選好チューニングに基づく生成的プランニングフレームワークを提案する。

2. 先行研究と比べてどこがすごい?

従来のニューラルプランナーは密なウェイポイント列をモデル化し、冗長な局所詳細や平滑性に容量を費やしていたが、NeurRAFTはコンパクトなアンカーウェイポイントに基づいて動作する。また、模倣学習では衝突回避とニアコリジョンの区別ができない問題に対し、推論時の後処理ではなく、プランナーの分布を直接安全側に整形する点が新しい。

3. 技術・手法の肝は?

手法の肝は、アンカーレベルのフローマッチングと、各アンカーのタスク空間への影響を考慮したJacobian重み付き損失を用いた学習、および推論時の2ステップ積分と3次スプライン補間による滑らかな軌道復元である。さらに、Direct Preference Optimizationを用いて、より大きな障害物クリアランスを持つ軌道に確率質量をシフトし、プランナーのパラメータに直接吸収させる。

4. どうやって有効だと検証した?

実験では、最先端のプランナーと比較して大幅な改善を示し、実世界実験では、ノイズや部分的なオクルージョンを含む深度観測下でFrankaロボットへのゼロショット転移を実証した。

5. 議論はある?

要旨からは、提案手法の限界や潜在的な欠点についての議論は不明である。また、実世界実験の詳細や、他のロボットや環境への一般化可能性についても要旨からは不明。

6. 次に読むべき論文は?

要旨で参照されている先行研究や関連手法は明示されていないが、関連する分野として、エンドツーエンドのニューラルモーションプランニング、フローマッチング、Direct Preference Optimization、および模倣学習に基づくプランニング手法が挙げられる。

※ AIが要旨から生成した要約です。正確性は原文をご確認ください。

著者: Sibo Tian, Chang Liu, Minghui Zheng, Xiao Liang

分類: cs.RO

原文アブストラクト

Recent end-to-end neural motion planners generate trajectories from raw sensor observations, avoiding the privileged geometric models required by classical planners. However, collision-free planning in cluttered environments remains challenging. We present NeurRAFT, a generative planning framework based on anchor-level flow matching and clearance-aware preference tuning. Unlike prior neural planners that model dense waypoint sequences and spend capacity on redundant local details and smoothness, NeurRAFT operates on compact anchor waypoints. We train the planner using a Jacobian-weighted loss that accounts for the task-space impact of each anchor. At inference, the anchors are generated in two integration steps, followed by cubic-spline interpolation to recover a smooth, full-resolution trajectory. Since imitation learning from positive demonstrations cannot distinguish collision-free from near-collision trajectories, collision-prone behaviors persist at test time. Rather than relying on post-hoc corrections, we directly reshape the pretrained planner's distribution toward safer solutions without augmenting inference. Specifically, Direct Preference Optimization shifts probability mass toward trajectories with larger obstacle clearance, with the resulting improvement directly absorbed into the planner parameters. Experiments show substantial improvements over state-of-the-art planners, while real-world experiments demonstrate zero-shot transfer to a Franka robot under noisy and partially occluded depth observations. Video results available at https://neurraft.github.io/.

関連論文