OVMAN: オープンボキャブラリの動作認識ナビゲーションのためのタスクとベンチマーク
OVMAN: A Task and Benchmark for Open-Vocabulary Motion-Aware Navigation
家の変化を利用して、移動した椅子や以前あった花瓶の場所など、変化に基づく目標へのナビゲーションを評価する新しいタスクとベンチマークを提案。
著者: Dibyendu Ghosh
分類: cs.RO
原文アブストラクト
Homes change between a robot's visits. Navigation benchmarks pose their goals in the world the agent currently sees, and the two-visit benchmarks that exist score recall or rearrangement rather than navigation. None of them can express go to the chair that was moved or go to where the vase used to be. OVMAN is a task in which an agent tours a scene, returns after a scripted change, and must navigate to a goal specified by the change itself. Two of its six change relations answer with a place an object has left, where nothing remains to be detected. We release 219 two-visit episodes, each certified solvable by an oracle agent that completes it three times. Two released systems fail as predicted. A zero-shot object-goal navigator reaches the change-defined target in 11.4% of episodes and essentially fails at vacated locations, 0.000 on former and 0.025 on removed. A self-maintaining open-vocabulary map answers past-tense queries at about a third of the rate of the same maps read as two visits, even when grounding is supplied. A simple two-visit reference agent reaches 45.2% when navigating, against an embodied oracle of 99.5%. An error decomposition places the remaining difficulty in open-vocabulary instance grounding rather than in geometry or in selecting the answer once positions are known.