日本フィジカルAI新聞

世界のフィジカルAIを、日本語で。

週刊ニュースレター購読
マニピュレーションarXiv:2609.40353

AssemblyWorld:汎用エージェントによる3D組立の再考

AssemblyWorld: Rethinking 3D Assembly with General-Purpose Agents

シェア:XThreadsFacebookLINEはてブBluesky

汎用エージェントが微調整なしに視覚情報から3D組立を行えるかを評価する環境とベンチマークを構築し、8システムの性能差を明らかにした。

詳しい要約

1. どんなもの?

- 3D assembly を general-purpose agents が visual interaction で実行できるか問う研究。 - AssemblyWorld: エージェントが rendered views を観察し rigid parts を操作する interactive 3D environment。 - 入力は images や assembly manuals、評価は幾何学的。 - AssemblyWorldBench: 80 objects・100 tasks(furniture, industrial assembly, fracture reassembly)。 - 8 agent systems を評価。

2. 先行研究と比べてどこがすごい?

- 従来の assembly 研究は専用 fine-tuning や mesh 直接アクセスが前提。 - 本研究は pretrained general-purpose agents を追加学習なしで用いる点が新しい。 - 2D views のみで part geometry を知覚し、mesh vertices/faces に直接触れない設定。 - 異なる agent systems を同一環境で比較可能にした。

3. 技術・手法の肝は?

- interactive 3D environment で rendered views を観察し rigid parts を操作。 - 画像または assembly manuals をガイドとして利用可能。 - 知覚は 2D views 経由、評価は resulting assemblies の幾何学的評価。 - AssemblyWorldBench で 100 tasks / 80 objects を用意。 - 詳細な手法は要旨からは不明。

4. どうやって有効だと検証した?

- 8 agent systems を AssemblyWorldBench で評価。 - 最強システムは part accuracy 80.9%、complete-assembly success 59.4%。 - open-source systems は closed-source peers より実行信頼性・組立精度で大幅に劣る。 - visual references、interaction trajectories、failures を分析。 - エージェントが組立を修正しつつ残差 positioning errors を残すことを示した。

5. 議論はある?

- 近似構造回復と精密再構成のギャップを特徴づける。 - エージェントは組立を改訂するが residual positioning errors が残る。 - open-source と closed-source の性能差が議論点。 - 具体的な限界や今後の課題は要旨からは不明。

6. 次に読むべき論文は?

- 要旨で参照/比較されている研究は明示されていない。 - 同分野の定番として 3D assembly、robot assembly、vision-language-action agents、interactive perception に関する研究が挙げられる。 - 具体的な論文名は要旨からは不明。

※ AIが要旨から生成した要約です。正確性は原文をご確認ください。

著者: Jiahao Zhang, Yeying Fan, Moitreya Chatterjee, Suhas Lohit, Bernhard Egger, Tim K. Marks, Anoop Cherian, Stephen Gould

分類: cs.CV, cs.RO

原文アブストラクト

The task of 3D assembly requires translating an understanding of parts and their relationships into precise spatial arrangements. Can pretrained general-purpose agents assemble objects through visual interaction without additional assembly-specific fine-tuning? To investigate this question, we introduce AssemblyWorld, an interactive 3D environment in which agents inspect rendered views and manipulate supplied rigid parts, guided by images or assembly manuals when available. Agents perceive part geometry through 2D views rather than direct access to mesh vertices or faces, while their resulting assemblies are evaluated geometrically. Building on this environment, we construct AssemblyWorldBench, comprising 100 assembly tasks across 80 objects spanning furniture, industrial assembly, and fracture reassembly. Evaluating eight agent systems reveals substantial differences in their capabilities. The strongest system achieves 80.9% part accuracy but 59.4% complete-assembly success. The evaluated open-source systems lag substantially behind their stronger closed-source peers in both execution reliability and assembly accuracy. Analyses of visual references, interaction trajectories, and failures show how agents revise assemblies while leaving residual positioning errors. AssemblyWorld provides a common setting for both assessing the capabilities of interactive assembly agents and characterizing the gap between approximate structure recovery and precise reconstruction.

関連論文

PR本紙発行元 EmplifAI