AmaraSpatial-10K:空間コンピューティングと具現化AIのための空間・意味整合3Dデータセット
AmaraSpatial-10K: A Spatially and Semantically Aligned 3D Dataset for Spatial Computing and Embodied AI
実世界展開に適したメトリックスケール・アンカー付きの合成3Dアセット10,000点と評価スイートを構築し、Objaverse比でCLIP検索性能を3.4倍に向上させた。
著者: Mohammad Sadegh Salehi, Alex Perkins, Igor Maurell, Ashkan Dabbagh, Raymond Wong
分類: cs.CV, cs.AI, cs.LG
原文アブストラクト
Web-scale 3D asset collections are abundant but rarely deployment-ready, suffering from arbitrary metric scaling, incorrect pivots, brittle geometry, and incomplete textures, defects that limit their use in embodied AI, robotics, and spatial computing. We present AmaraSpatial-10K, a dataset of over 10,000 synthetic 3D assets optimised for zero-shot deployment. Each asset ships as a metric-scaled, deterministically anchored .glb with separated PBR maps, a convex collision hull, a paired reference image, and multi-sentence text metadata. Alongside the dataset we introduce a reusable evaluation suite for 3D asset banks, a continuous Scale Plausibility Score (SPS), an LLM Concept Density metric, anchor-error auditing, and a cross-modal CLIP coherence protocol, and apply it to AmaraSpatial-10K alongside matched subsets of Objaverse, HSSD, ABO, and GSO. AmaraSpatial-10K improves CLIP Recall@5 by $3.4\times$ over Objaverse ($0.612$ vs. $0.181$, median rank $267 \rightarrow 3$), achieves a $99.1\%$ physics-stability rate under Habitat-Sim with $\sim 20\times$ wall-time speed-up, and produces zero-overlap scenes when used as a drop-in asset bank for Holodeck. Controlled ablations on the same asset bank attribute the retrieval gain to description richness.
関連論文
- DARP: 多視点ロボット知覚のための校正済み双腕RGB-D-IRデータセットデータセット
- uScenes: 水中ロボット知覚のためのマルチモーダルRGB・3Dソナー画像データセットデータセット
- PRISM:マルチモーダルセンシングを備えた精密で接触豊富な実世界産業スキルデータセットデータセット
- NARRATE: 自動運転における人間中心の説明のためのマルチモーダル実世界オーストラリア運転データセットデータセット
- 衛星画像の改ざんとディープフェイク位置特定のためのベンチマークデータセット構築に向けてデータセット
- InteracVid: ライブチャット動画から構築した実インタラクティブ音声視覚応答データセットデータセット