日本フィジカルAI新聞

世界のフィジカルAIを、日本語で。

週刊ニュースレター購読
sim2realarXiv:2610.06643

AffordCraft: 単一画像からタスク対応シミュレーション資産をスケーラブルに構築

AffordCraft: Scalable Construction of Task-Ready Simulation Assets from Single Images

シェア:XThreadsFacebookLINEはてブBluesky

単一のRGB画像とタスク指示から、関節や可動部を保った物理的に妥当なシミュレーション用オブジェクトを検索ベースで生成する手法を提案。生成手法より高成功率・低計算コストで、ライブラリ拡張により性能が向上する。

詳しい要約

1. どんなもの?

- 単一の RGB 画像と task instruction から、simulation で使える articulated asset を構築する手法 AffordCraft を提案。 - 生成ではなく retrieval により、object と操作対象 part を特定し、library から一致する articulated asset を選択して画像に fit する。 - part と joint を保持したまま物理的に有効な asset を生成し、manipulation task の構築にも用いる。

2. 先行研究と比べてどこがすごい?

- 既存の generative model は part や joint を予測しても simulation で settle/move しにくく、general-purpose agent は画像ごとに多数の model call を要する。 - AffordCraft は retrieval により、box や mask なしで 2,000 枚中 1,703 枚の物理的に有効な asset を生成。 - 5 つの generative 手法は同じ画像で最大 45% しか通過せず、中央値で有効 asset あたり 10〜78 倍の GPU time を要する。

3. 技術・手法の肝は?

- 単一 RGB 画像と task instruction から object と操作すべき part を locate する。 - 事前構築した articulated asset の library から一致する entry を選択し、part と joint を壊さずに画像へ fit する。 - library を 141 から 11,372 entry に拡張しても手法の変更は不要で、category coverage と requested label の選択率が向上する。

4. どうやって有効だと検証した?

- 2,000 枚の写真(31 カテゴリ)で、box/mask なしに 1,703 枚の物理的に有効な asset を生成。 - 5 つの generative 手法と比較し、通過率と GPU time で優位性を確認。 - 50 枚の cluttered 画像では、自動検出後の 237 annotated object 中 162 が同じ物理テストを通過。 - 構築した asset から単一物体および composed scene の manipulation task を構築し、scripted demonstration で訓練した policy が未見の初期状態から両種のタスクを完了。

5. 議論はある?

- library の拡張が手法変更なしに category coverage と label 一致率を改善する点を議論。 - 一方、生成手法との比較における物理妥当性や計算コストの詳細、失敗事例の分析は要旨からは不明。 - cluttered 画像での検出・選択の限界や、より多様なタスクへの一般化可能性は要旨からは不明。

6. 次に読むべき論文は?

- 要旨で比較されている 5 つの generative 手法(具体的名称は要旨からは不明)。 - general-purpose agent を用いた asset 構築手法。 - articulated object の生成・再構成に関する研究(例: generative model による part/joint 予測)。 - simulation における articulated asset library や retrieval ベースの手法。

※ AIが要旨から生成した要約です。正確性は原文をご確認ください。

著者: Haoyun Yang, Xueyang Zhou, Ziyi Xie, Yongchao Chen

分類: cs.RO, cs.AI, cs.CV

原文アブストラクト

Robot learning in simulation depends on the objects the simulator offers. Many tasks need objects with separate parts, joints that allow the required motion, and physical properties that remain valid under contact. Existing methods recover this structure anew for every image: generative models predict parts and joints that mostly fail to settle or move in simulation, and general-purpose agents need a long session of model calls for each photograph. AffordCraft builds such an asset from a single RGB image and a task instruction by retrieval instead of generation: it locates the object and the part to operate, selects a matching entry from a library of articulated assets, and fits it to the image while keeping its parts and joints intact. Without any box or mask marking the object, AffordCraft produces a physically valid asset for 1,703 of 2,000 photographs from 31 categories. Five generative methods pass on at most 45% of the same photographs and, at the median, need 10 to 78 times our GPU time per valid asset. On 50 cluttered images, 162 of 237 annotated objects pass the same physical test after automatic detection. Growing the library from 141 to 11,372 entries needs no change to the method and raises category coverage from 46% to 100% and the share of selections with the requested label from 18% to 51%. We also build manipulation tasks from the constructed assets, both with single objects and in composed scenes; policies trained on scripted demonstrations complete both kinds of tasks from initial states unseen in training.

関連論文

PR本紙発行元 EmplifAI