視覚から収穫へ:マルチアームロボット果実収穫のための視覚言語モデルベンチマーク
From Vision to Harvest: Benchmarking Vision-Language Models for Multi-Arm Robotic Fruit Harvesting
実世界の果樹園画像を用いて、事前学習済み視覚言語モデルがゼロショットでマルチアーム収穫計画を生成できるかを評価する初のベンチマークを提案し、従来手法と比較した。
詳しい要約
1. どんなもの?
2. 先行研究と比べてどこがすごい?
3. 技術・手法の肝は?
4. どうやって有効だと検証した?
5. 議論はある?
6. 次に読むべき論文は?
※ AIが要旨から生成した要約です。正確性は原文をご確認ください。
著者: Vrishan Inukollu, Adyan Zaman, Anvi Kudaraya, Carlos Lazcano, Yuankai Zhu, Stavros Vougioukas, Xiaofan Yu
分類: cs.RO, cs.CV
原文アブストラクト
Multi-arm robotic harvesting offers a promising path to improve harvesting efficiency and reduce reliance on manual labor. However, practical deployment remains challenging because the system must generalize across diverse environments while efficiently coordinating multiple arms in a shared workspace. Existing methods often require substantial data collection in target environments or rely on simplifying assumptions that limit planning quality. In this work, we introduce the first comprehensive benchmark for evaluating pretrained Vision-Language Models (VLMs) on zero-shot multi-arm fruit harvesting planning. Our benchmark uses real-world apple and citrus orchard images and compares a VLM-based planning pipeline with a traditional perception-and-planning pipeline. The VLM pipeline directly generates harvesting sequences and waypoints for each arm, while a lightweight trajectory verifier checks for collisions. Our results show that frontier VLMs can generate effective multi-arm harvesting plans zero-shot, but a practical deployment remains limited by accurate 3D waypoint generation and collision-aware coordination. These results highlight both the promise and current limitations of pretrained VLMs for multi-arm robotic harvesting.