日本フィジカルAI新聞

世界のフィジカルAIを、日本語で。

週刊ニュースレター購読
sim2realarXiv:2610.12069

LIVIN: 実在の住まいのデジタルツインで空間・身体知能を評価するベンチマーク

LIVIN: Benchmarking Spatial and Embodied Intelligence in Digital Twins of Lived-In Homes

シェア:XThreadsFacebookLINEはてブBluesky

実際に人が暮らす30世帯の家をデジタルツイン化し、物体配置や家具配置を忠実に再現したベンチマークを構築。3D検出・再構成・ナビゲーション・移動操作の4タスクで現手法を評価した。

詳しい要約

1. どんなもの?

- LIVINは、実際に人が住んでいる30の多様な家庭のデジタルツイン上に構築された、空間知能と身体性知能のベンチマーク。 - 観察された部屋のレイアウト、家具の配置、日常の持ち物を保持したレプリカ。 - 4つのタスク(3D detection、3D reconstruction、navigation、loco-manipulation)を評価する。

2. 先行研究と比べてどこがすごい?

- 既存のリソースは、規模、実世界との対応、インタラクションへの準備の間でトレードオフがあり、実際の家庭の配置を忠実かつインタラクティブに再現するものが不足していた。 - LIVINは、実在する住まいの観察に基づくデジタルツインを提供し、このギャップを埋める。

3. 技術・手法の肝は?

- 人間参加型(human-in-the-loop)ワークフローを設計。 - インスタンス認識、建築再構成、オブジェクト生成と配置の各段階で、中間結果を人間が元の観察と照合してレビュー・修正する。

4. どうやって有効だと検証した?

- LIVIN上で4つのタスク(3D detection、3D reconstruction、navigation、loco-manipulation)を評価。 - 現在の手法が、現実の住まいにある密集した物体配置、オクルージョン、限られた自由空間、制約されたインタラクション領域に依然として苦戦していることを示した。

5. 議論はある?

- 現在の手法は、現実的な住環境の密集配置、オクルージョン、限られた自由空間、制約されたインタラクション領域に課題がある。 - LIVINが、空間理解からロボットインタラクションまで、実世界の家庭における身体性AIの進歩に貢献することが期待される。

6. 次に読むべき論文は?

- 要旨からは不明(参照・比較されている研究が明記されていない)。同分野の定番として、3D detectionではScanNet、3D reconstructionではReplica、navigationではHabitat、loco-manipulationではBEHAVIORなどが考えられる。

※ AIが要旨から生成した要約です。正確性は原文をご確認ください。

著者: Peijun Xu, Chuansen Nie, Yiyang He, Yinuo Bai, Jingyang Liu, Kuixiang Shao, Yuyang Jiao, Kuanhao Xia, Jiayi Zhu, Zitian Yang, Yanqi Zhang, Tianye Tan, Shuwei Di, Junyi Xu, Jingyi Yu, Jiayuan Gu

分類: cs.CV

原文アブストラクト

Realistic household simulation must capture not only diverse environments but also the lived-in object arrangements and spatial constraints that shape robot motion and interaction. Existing resources often trade off scale, real-world correspondence, and interaction readiness, leaving a gap in faithful, interactive replicas of how real homes are actually arranged. To this end, we introduce LIVIN, a benchmark for spatial and embodied intelligence built on digital twins of 30 diverse lived-in homes. These replicas preserve observed room layouts, furniture configurations, and everyday belongings. To construct them, we design a human-in-the-loop workflow comprising instance recognition, architectural reconstruction, and object generation and placement, with intermediate results reviewed and corrected by humans against the source observations at each stage. We evaluate four tasks in LIVIN: 3D detection, 3D reconstruction, navigation, and loco-manipulation. Our evaluations show that current methods remain challenged by the dense object arrangements, occlusions, limited free space, and constrained interaction regions found in realistic lived-in homes. We hope LIVIN will help advance embodied AI in real-world homes, from spatial understanding to robotic interaction, and ultimately bring embodied intelligence into everyday home environments.

関連論文

PR本紙発行元 EmplifAI