日本フィジカルAI新聞

世界のフィジカルAIを、日本語で。

週刊ニュースレター購読
sim2realarXiv:2610.06327

低コストロボットナビゲーションにおける効率的なSim-to-Real転移のためのデュアル変分オートエンコーダ

Dual Variational Autoencoders for Efficient Sim-to-Real Transfer in Low-Cost Robotic Navigation

シェア:XThreadsFacebookLINEはてブBluesky

シミュレーションと実世界の画像をデュアルVAEで共通潜在空間に整列させ、低コストロボットの屋内ナビゲーションを約91%の成功率で実現した。

詳しい要約

1. どんなもの?

- 低コストロボットの視覚ベース自律ナビゲーションにおけるsim-to-realギャップを橋渡しするハイブリッド転移学習フレームワークを提案。 - domain randomizationとfeature-level domain adaptationを組み合わせ、dual convolutional variational autoencoder(共有decoder)で両ドメインの分布を整列。 - 45225枚のシミュレーション画像と4556枚の実世界画像で学習し、実世界室内ナビゲーションの画像分類で平均成功率約91%を達成。 - 低コストロボットのreactive explorationタスクで実機展開し、Raspberry Pi 4やNVIDIA Jetson Nanoでの効率も検証。

2. 先行研究と比べてどこがすごい?

- シミュレーションのみの学習や実世界のみの学習と比較して、提案手法が大幅に優れることを実験で示した。 - 従来の直接的なpolicy transferは効果が低く、実データのみの学習は非現実的であるという課題に対し、限られた実データで効率的に適応できる点が優位。 - 具体的な先行研究名は要旨からは不明。

3. 技術・手法の肝は?

- dual convolutional variational autoencoderアーキテクチャを採用し、共有decoderでコンパクトな共通潜在表現空間を学習。 - この潜在空間でシミュレーションと実世界の分布を整列させるfeature-level domain adaptationを行う。 - domain randomizationと組み合わせ、さらに限られた実世界データを拡張する2つの相補的なdata augmentation手法を適用。

4. どうやって有効だと検証した?

- 実世界室内ナビゲーションの画像分類タスクで平均成功率約91%を達成し、シミュレーションのみ・実世界のみの学習より有意に優れることを確認。 - 実機展開により、低コストロボットのreactive explorationタスクで提案policyが成功することを検証。 - 計算コストの厳密な推定を行い、Raspberry Pi 4やNVIDIA Jetson Nanoなどのリソース制約のある組み込みプラットフォームへの適合性を確認。

5. 議論はある?

- 要旨からは、手法の限界や失敗事例、議論の詳細は不明。 - 提案手法が低コストロボット向けの実用的なソリューションであると主張しているが、一般化可能性や他のタスクへの適用性については言及されていない。

6. 次に読むべき論文は?

- 要旨で参照/比較されている研究は明示されていない。 - 同分野の関連手法として、domain randomization、feature-level domain adaptation、variational autoencoder、sim-to-real transfer、reactive explorationなどが挙げられる。

※ AIが要旨から生成した要約です。正確性は原文をご確認ください。

著者: Álvaro Díez, Fidel Aznar

分類: cs.RO, cs.CV, cs.LG

原文アブストラクト

Vision-based autonomous navigation for low-cost robots remains a fundamental challenge, primarily due to the significant gap between simulated training environments and real-world operational conditions. Direct policy transfer from simulation is often ineffective, while training exclusively on real data is impractical. We propose a hybrid transfer learning framework that effectively bridges the sim-to-real gap by combining domain randomization with feature-level domain adaptation. Our method employs a dual convolutional variational autoencoder architecture with a shared decoder, trained on an extensive set of 45225 simulated images and a minimal set of only 4556 real-world samples. This architecture learns a compact, common latent representation space that aligns the distributions of both domains. The adaptation process is further enhanced by two complementary data augmentation techniques designed to expand the limited real-world data. Experimental evaluation demonstrates that our method achieves an average success rate of almost 91% on image classification tasks for real-world indoor navigation, significantly outperforming both simulation-only and real-world-only training. We validate these findings through a direct, real-world deployment, where the proposed policy successfully guides a low-cost robot in a reactive exploration task. Furthermore, we validate the model's efficiency through a rigorous computational estimation, confirming its suitability for resource-constrained embedded platforms such as the Raspberry Pi 4 and NVIDIA Jetson Nano. This work presents a practical solution for developing effective and efficient navigation policies for low-cost robotic systems.

関連論文

PR本紙発行元 EmplifAI