日本フィジカルAI新聞

世界のフィジカルAIを、日本語で。

週刊ニュースレター購読
組み立て計画/マルチモーダル生成arXiv:2610.00487

ScaffoldM3C: 生成的な安定構築計画のためのマルチモーダル逐次モンテカルロフレームワーク

ScaffoldM3C: A Multimodal Sequential Monte Carlo Framework for Generative Stable Construction Planning

シェア:XThreadsFacebookLINEはてブBluesky

テキスト・画像・スケッチを条件に、足場ブロックを活用しながら安定したブロック構造の組み立て手順を生成する軽量マルチモーダルモデルを提案。逐次モンテカルロで複数の組み立て候補を並行探索し、従来比4倍小型で5〜20倍高速化を実現した。

詳しい要約

1. どんなもの?

- 3D構造物を自律構築するための生成的な安定構築計画フレームワーク。 - ScaffoldM3Cは、マルチモーダル(テキスト、画像、スケッチ)条件付けが可能な軽量自己回帰モデル。 - 次のブロック候補を提案し、Sequential Monte Carlo (SMC) で複数の組み立てシーケンスを維持。 - 足場(scaffold)の役割を明示的に考慮し、補助的なscaffold block tokenを導入。 - シミュレーションと実世界のロボット組み立てデモで有効性を実証。

2. 先行研究と比べてどこがすごい?

- 先行研究は大規模言語モデルをテキストベースの生成構築に微調整していたが、マルチモーダル条件付けができず、足場の役割を無視し、推論が遅かった。 - ScaffoldM3Cは、競合ベースラインより4倍小さく、推論速度が5倍から20倍高速。 - 構築品質は最先端手法と同等で、全体的な安定性はより高い。 - マルチモーダル条件付けと足場の明示的考慮を実現。

3. 技術・手法の肝は?

- 構築を確率的な次ブロック生成タスクとして定式化。複数の可能な組み立てアクションと複数のタスク条件付けモダリティを考慮。 - 補助的なscaffold block tokenを導入し、足場の有用性を明示的に考慮。 - 軽量な自己回帰モデルが次ステップの候補ブロックを提案。 - Sequential Monte Carlo (SMC) を用いて、可能な組み立てシーケンスの母集団を維持し、複数の組み立て方向を同時に考慮。 - StableText2Brickデータセットを拡張し、画像条件付けプロンプトと足場で安定化された構築シーケンスを含める。

4. どうやって有効だと検証した?

- シミュレーションと実世界のロボット組み立てデモを通じて有効性を実証。 - 競合ベースラインと比較して、4倍小さいモデルサイズ、5倍から20倍の推論速度向上を達成。 - 構築品質は最先端手法と同等で、全体的な安定性はより高いことを示した。

5. 議論はある?

- 要旨からは不明。

6. 次に読むべき論文は?

- 要旨で参照/比較されている研究: 大規模言語モデルをテキストベースの生成構築に微調整した最先端手法、StableText2Brickデータセット。 - 関連手法: Sequential Monte Carlo (SMC)。 - 同分野の定番: 生成的な構築計画、マルチモーダル条件付け、ロボット組み立て。

※ AIが要旨から生成した要約です。正確性は原文をご確認ください。

著者: Gadiel Sznaier Camps, Chengyang He, Guillaume Sartoretti, Eduardo Montijano, Mac Schwager

分類: cs.RO, cs.LG

原文アブストラクト

Autonomously constructing physically realizable 3D structures remains a significant challenge due to combinatorial action spaces, interchangeable components, equifinal assembly sequences, and strict stability requirements during construction. State-of-the-art methods fine-tune large language models for text-based generative construction. However, these approaches do not allow for Multimodal (text, image, sketch) conditioning, overlook the practical role of scaffolding for stabilizing intermediate structures, and suffer from slow inference speeds. Therefore, we formulate construction as a probabilistic next-block generation task with multiple potential assembly actions and multiple potential task conditioning modalities. Concurrently, we explicitly consider the utility of scaffolding by introducing an auxiliary scaffold block token. We present Scaffold Multimodal Monte Carlo (ScaffoldM3C), a multimodal, lightweight, auto-regressive model for stable block-based construction, that proposes a set of next-step candidate blocks. Leveraging these candidates, we utilize Sequential Monte Carlo (SMC) to maintain a population of possible assembly sequences, allowing us to consider multiple, potentially different, assembly directions simultaneously. We train our multimodal architecture by extending the StableText2Brick dataset to contain image conditioning prompts and scaffold-stabilized build sequences. ScaffoldM3C is 4x smaller than competing baselines, yielding a 5x to 20x speedup during inference, while achieving comparable construction quality to state-of-the-art methods and higher overall stability. We demonstrate the effectiveness of our approach through simulations and real-world robot assembly demonstrations.

PR本紙発行元 EmplifAI