日本フィジカルAI新聞

世界のフィジカルAIを、日本語で。

週刊ニュースレター購読
sim2realarXiv:2609.20747

MILER: 非構造自律運転のためのsim-to-real強化学習に向けた意味的中間表現

MILER: Semantic Mid-Level Representation for Sim-to-Real Reinforcement Learning in Unstructured Autonomous Driving

シェア:XThreadsFacebookLINEはてブBluesky

意味的な中間表現シミュレータとBEVFusionによる知覚・制御のゼロショットsim-to-real転移を実現し、非構造環境での実車自律走行を達成した。

詳しい要約

1. どんなもの?

- 非構造環境での自律走行におけるsim-to-real強化学習のためのend-to-endポリシーフレームワークMILERを提案。 - オフライン訓練ではカスタムのsemantic mid-level representation (MLR)シミュレータを用い、bicycle modelに直接制御出力を適用。 - 実車展開時はcameraとLiDARをBEVFusionで処理し、MLRシミュレータと整合するsemantic bird's-eye-view表現を生成。 - ポリシーネットワークの行動は直接実車に適用せず、trajectory-alignment戦略で知覚と制御のzero-shot sim-to-real転移を実現。

2. 先行研究と比べてどこがすごい?

- 従来の強化学習は実世界自律走行、特に非構造環境への適用が稀。 - 既存のsim-to-real転移は非構造環境で困難。 - MILERはzero-shot sim-to-real転移を可能にし、知覚と制御の両方を転移。 - 具体的な先行研究との比較は要旨からは不明。

3. 技術・手法の肝は?

- オフライン訓練でcustom semantic mid-level representation (MLR)シミュレータを使用。 - 強化学習でポリシーネットワークを訓練し、制御出力をbicycle modelに直接適用。 - 実車ではcameraとLiDARをBEVFusionで処理し、MLRシミュレータと一致するsemantic bird's-eye-view表現を生成。 - ポリシー出力にtrajectory-alignment戦略を適用し、zero-shot sim-to-real転移を実現。

4. どうやって有効だと検証した?

- 多様な障害物、ヘアピンカーブ、最高33.6 km/hの速度、オフロード区間を含む3.0 kmのテストトラックで評価。 - 2台の異なる車両で合計17.3 kmを人間の介入なしで走行。 - ソフトウェアスタック全体がJetson AGX Orin上で動作。

5. 議論はある?

- 非構造環境でのsim-to-real転移の課題を克服し、zero-shot転移を実証。 - 具体的な限界や議論の詳細は要旨からは不明。

6. 次に読むべき論文は?

- 要旨で参照/比較されている研究は明示されていない。 - 関連手法としてBEVFusion、bicycle model、強化学習、sim-to-real転移が挙げられる。 - 同分野の定番としてdomain randomization、system identification、end-to-end学習などが考えられる。

※ AIが要旨から生成した要約です。正確性は原文をご確認ください。

著者: Thomas Steinecker, Denis Trescher, Alexander Bienemann, Thorsten Luettel, Mirko Maehlisch

分類: cs.RO, cs.LG

原文アブストラクト

Reinforcement learning constitutes a promising approach owing to its potential for superhuman performance and self-learned policies. However, its application to real-world autonomous driving remains scarce, particularly in unstructured environments, because of the challenges associated with sim-to-real transfer for unstructured environments. In this work, we present MILER, an end-to-end policy framework with zero-shot sim-to-real transfer. During offline training, we employ a custom semantic mid-level representation (MLR) simulator and train the policy network using reinforcement learning, with its control outputs applied directly to a bicycle model. During deployment on the real vehicle, camera and LiDAR data are processed by BEVFusion to generate a semantic bird's-eye-view representation consistent with that of the MLR simulator. The actions generated by the policy network are not applied directly to the real vehicle. Instead, we employ a trajectory-alignment strategy that enables zero-shot sim-to-real transfer of both perception and control. We extensively evaluate the proposed framework on a diverse test track comprising numerous challenges, including various obstacles, hairpin curves, velocities of up to 33.6 km/h, and off-road sections. In total, we drove 17.3 km with two different vehicles on a 3.0 km test track without human intervention, thereby demonstrating the effectiveness of our approach. Furthermore, the entire software stack runs on a Jetson AGX Orin.

関連論文

PR本紙発行元 EmplifAI