日本フィジカルAI新聞

世界のフィジカルAIを、日本語で。

週刊ニュースレター購読
sim2realarXiv:2609.21045

DEXTERA: 単一画像から実機展開可能な巧みなマニピュレーションへ向けたReal-to-Sim-to-Real

DEXTERA: From a Single Image to Deployable Dexterous Manipulation via Real-to-Sim-to-Real

シェア:XThreadsFacebookLINEはてブBluesky

1枚のRGB画像からシーンを分解し、物理パラメータ推定・ロボットキャリブレーション・軌道生成を自動化して、シミュレーションのみで学習した巧みなマニピュレーション方策を実機にゼロショット展開できるようにしたフレームワーク。

詳しい要約

1. どんなもの?

- 単一のRGB画像から、dexterous manipulationのためのdeployable policyを自動生成するreal-to-sim-to-realフレームワークDEXTERAを提案。 - 4段階の統合プロセス:(1) 単一画像を静的Gaussian背景とinteractive rigid/articulated assetsに分解し、VLMで物理パラメータを推定、(2) metric scene global alignment、object canonicalization、morphology-balanced robot calibration、(3) simulator task primitive構築、VR teleoperation、object-centric trajectory synthesis、(4) imitation learningとreinforcement learningをサポートする共有multimodal policy interface。 - 13のtask-embodimentペア、2つのdexterous robotプラット…

2. 先行研究と比べてどこがすごい?

- 手動でのdigital twin構築が不要で、単一画像から自動的にreal-to-sim-to-realを実現。 - 生成ベースラインと比較して、優れたvisual fidelityと3D geometric reconstructionを達成。 - シミュレーションのみで訓練したpolicyがzero-shot real-robot deploymentを可能にし、simulation-real co-trainingで平均物理policy成功率を29.2%から61.9%に大幅改善。 - 従来のsim-to-realにおける視覚・幾何・動力学ギャップを軽減。

3. 技術・手法の肝は?

- 単一画像からのscene factorization:静的Gaussian背景とinteractive assets(rigid/articulated)に分離し、VLMで物理パラメータを推定。 - metric scene global alignment、object canonicalization、morphology-balanced robot calibrationによる幾何・動力学の整合。 - scalable simulator task primitive構築、VR teleoperation、object-centric trajectory synthesis。 - 共有multimodal policy interfaceでimitation learningとreinforcement learningを統合。

4. どうやって有効だと検証した?

- 13のtask-embodimentペア、2つのdexterous robotプラットフォーム、6つのpolicyアーキテクチャで評価。 - 生成ベースラインと比較してvisual fidelityと3D geometric reconstructionが優れることを確認。 - cross-domain trajectory replaysで物理的相互作用の一貫性を検証。 - simulation-only trained policyによるzero-shot real-robot deploymentを実証。 - simulation-real co-trainingで平均物理policy成功率が29.2%から61.9%に向上。

5. 議論はある?

- 要旨からは、限界や議論の詳細は不明。 - 残存するvisual, geometric, dynamics gapsの完全な解消については言及されていない。 - 実世界展開の成功率61.9%は改善を示すが、完全な成功ではない。

6. 次に読むべき論文は?

- 要旨で参照/比較されている研究:generative baselines(具体的名称は要旨からは不明)。 - 関連手法:real-to-sim-to-real、sim-to-real transfer、dexterous manipulation、Gaussian splatting、VLM、imitation learning、reinforcement learning。 - 同分野の定番:domain randomization、digital twin、teleoperation、multimodal policy learning。

※ AIが要旨から生成した要約です。正確性は原文をご確認ください。

著者: Jin Wu, Lianjie Yuan, Zeyan Sun, Yuanyuan Lei, Disi A, Bicheng Han, Fangzhou Xia

分類: cs.RO

原文アブストラクト

Collecting real-world robot data for dexterous manipulation is costly and time-consuming. While high-fidelity physics simulators enable scalable data synthesis and policy learning, constructing deployment-ready digital twins manually remains labor-intensive, and residual visual, geometric, and dynamics gaps hinder reliable sim-to-real transfer. We present DEXTERA, an automated real-to-sim-to-real framework that transforms a single RGB image into deployable policies for dexterous manipulation across four unified stages: (1) single-image scene factorization into a static Gaussian background and interactive rigid or articulated assets with VLM-inferred physical parameters; (2) metric scene global alignment, object canonicalization, and morphology-balanced robot calibration; (3) scalable simulator task primitive construction, VR teleoperation, and object-centric trajectory synthesis; and (4) a shared multimodal policy interface supporting both imitation learning and reinforcement learning. We evaluate DEXTERA across 13 task-embodiment pairs, 2 dexterous robot platforms, and 6 policy architectures. Experimental results demonstrate that DEXTERA achieves superior visual fidelity and 3D geometric reconstruction compared to generative baselines, while cross-domain trajectory replays validate strong physical interaction consistency. Furthermore, simulation-only trained policies enable viable zero-shot real-robot deployment, while simulation-real co-training substantially improves mean physical policy success from 29.2% to 61.9% across diverse policy architectures.

関連論文

PR本紙発行元 EmplifAI