日本フィジカルAI新聞

世界のフィジカルAIを、日本語で。

週刊ニュースレター購読
全方位セグメンテーションarXiv:2610.03248

EmbPASS: 異なる身体性を超えた全方位セグメンテーションに向けて

EmbPASS: Towards Cross-Embodiment Open Panoramic Segmentation

シェア:XThreadsFacebookLINEはてブBluesky

車・ドローン・ウェアラブル・四足歩行の全方位画像を統一タクソノミで扱う異身体性ベンチマークEmbPASSと、関係認識アダプタと適応的意味転送を備えた全方位セグメンテーションネットEPONetを提案。

詳しい要約

1. どんなもの?

- 新タスク **Cross-Embodiment Open Panoramic Segmentation** を導入 - 異なるembodied platform間の観測視点・空間レイアウト差を扱う - ベンチマーク **EmbPASS** を構築 - Vehicle, Drone, Wearable, Quadruped の4プラットフォーム - 統一semantic taxonomy下のpanoramic semantic segmentation - 手法 **EPONet** を提案 - open-vocabulary panoramic semantic segmentation network - **RAMA** と **CAST** を統合

2. 先行研究と比べてどこがすごい?

- 従来はpanoramic segmentationの研究が中心で、cross-embodiment観測シフトの系統的検討は限定的 - 新タスク設定と統一taxonomyのマルチプラットフォームbenchmark **EmbPASS** を初めて提供 - **EPONet** はEmbPASSでplatform-balanced性能35.82% mIoUを達成 - 最強baselineを1.10%上回る - 既存panoramic segmentation benchmarkでも競争力維持

3. 技術・手法の肝は?

- **EPONet**: open-vocabulary panoramic semantic segmentation network - **Relation-Aware Metric Adapter (RAMA)** - 異種embodied観測下の空間モデリングを強化 - **Content-Adaptive Semantic Transfer (CAST)** - 異種観測下のsemantic transferを強化 - 詳細なアーキテクチャや学習手順は要旨からは不明

4. どうやって有効だと検証した?

- **EmbPASS** 上でextensive experimentsを実施 - 指標は **mIoU** と platform-balanced性能 - **EPONet** が35.82% mIoUで最強baselineを1.10%上回る - 既存panoramic segmentation benchmarkでも競争力を確認 - 具体的な比較手法やデータセット詳細は要旨からは不明

5. 議論はある?

- cross-embodiment観測シフトがpanoramic perceptionの一貫性・信頼性に課題 - 系統的研究が限られている点を問題提起 - **EmbPASS** が統一taxonomy下のtestbedを提供 - 限界や今後の課題についての議論は要旨からは不明

6. 次に読むべき論文は?

- 要旨で参照/比較されている研究は明示されていない - 関連手法として **panoramic semantic segmentation**, **open-vocabulary semantic segmentation**, **cross-embodiment perception** の定番研究を挙げる - 公開予定の **EmbPASS** benchmark (https://github.com/guopj1/EmbPASS) とそのbaseline群

※ AIが要旨から生成した要約です。正確性は原文をご確認ください。

著者: Pujun Guo, Yuanfan Zheng, Fei Teng, Mengfei Duan, Guoqiang Zhao, Yuheng Zhang, Kai Luo, Kailun Yang

分類: cs.CV, cs.RO

原文アブストラクト

Panoramic images provide a complete 360-degree field of view, enabling comprehensive scene understanding for embodied perception. However, heterogeneous embodied platforms exhibit substantial differences in observation viewpoints and spatial layouts, giving rise to cross-embodiment observation shifts that pose additional challenges to consistent and reliable panoramic perception, while systematic studies of this problem remain limited. To bridge this gap, we introduce a new task, termed Cross-Embodiment Open Panoramic Segmentation. Meanwhile, we establish EmbPASS, a multi-platform panoramic semantic segmentation benchmark spanning Vehicle, Drone, Wearable, and Quadruped platforms under a unified semantic taxonomy, providing a testbed for systematically studying cross-embodiment panoramic perception. We further propose EPONet, an open-vocabulary panoramic semantic segmentation network that integrates Relation-Aware Metric Adapter (RAMA) and Content-Adaptive Semantic Transfer (CAST) to enhance spatial modeling and semantic transfer under heterogeneous embodied observations. Extensive experiments show that EPONet achieves the best platform-balanced performance on EmbPASS with 35.82% mIoU, outperforming the strongest baseline by 1.10%, while remaining competitive on existing panoramic segmentation benchmarks. The source code and EmbPASS benchmark will be made publicly available at https://github.com/guopj1/EmbPASS.

PR本紙発行元 EmplifAI