日本フィジカルAI新聞

世界のフィジカルAIを、日本語で。

週刊ニュースレター購読
arXiv:2405.01107

CoViS-Net: A Cooperative Visual Spatial Foundation Model for Multi-Robot Applications

CoViS-Net: A Cooperative Visual Spatial Foundation Model for Multi-Robot Applications

シェア:XThreadsFacebookLINEはてブBluesky

著者: Jan Blumenkamp, Steven Morad, Jennifer Gielis, Amanda Prorok

分類: cs.RO, cs.MA, cs.SY, eess.SY

原文アブストラクト

Autonomous robot operation in unstructured environments is often underpinned by spatial understanding through vision. Systems composed of multiple concurrently operating robots additionally require access to frequent, accurate and reliable pose estimates. In this work, we propose CoViS-Net, a decentralized visual spatial foundation model that learns spatial priors from data, enabling pose estimation as well as spatial comprehension. Our model is fully decentralized, platform-agnostic, executable in real-time using onboard compute, and does not require existing networking infrastructure. CoViS-Net provides relative pose estimates and a local bird's-eye-view (BEV) representation, even without camera overlap between robots (in contrast to classical methods). We demonstrate its use in a multi-robot formation control task across various real-world settings. We provide code, models and supplementary material online. https://proroklab.github.io/CoViS-Net/