日本フィジカルAI新聞

世界のフィジカルAIを、日本語で。

週刊ニュースレター購読
sim2realarXiv:2610.05407

TUCO: シミュレーション実演のキュレーションによるSim-to-Realロボット方策の共訓練

TUCO: Curating Simulation Demonstrations for Sim-to-Real Robot Policy Co-Training

シェア:XThreadsFacebookLINEはてブBluesky

シミュレーション実演を能動的に選別するデータキュレーション手法TUCOを提案し、影響関数を用いて軌道の有用性と集合の網羅性を評価することで、Sim-to-Real方策共訓練の性能を向上させた。

詳しい要約

1. どんなもの?

- シミュレーションのデモンストレーションを実世界データと共に使う sim-to-real ロボット方策の co-training のためのデータキュレーション手法 TUCO を提案する研究。 - シミュレーションデモを能動的に選択するデータキュレーションの価値を体系的に検討した初の研究と主張。 - Trajectory-level Utility and set-level Coverage Optimization の略で、軌跡レベルの有用性と集合レベルの被覆を最適化する。

2. 先行研究と比べてどこがすごい?

- 既存のキュレーション手法は closed-loop の target behavior から軌跡レベルの有用性と集合レベルの被覆を測る統一基準を欠いていた。 - TUCO は influence functions を用いて各 source demonstration が target-domain scoring rollouts に与える影響を追跡し、共通の closed-loop 基準を提供する。 - 単一シミュレータ、sim-to-sim、sim-to-real の設定で state-of-the-art 性能を達成。

3. 技術・手法の肝は?

- influence functions で各 source demonstration が target-domain scoring rollouts に与える影響を追跡。 - その影響を target return への全体的寄与と rollouts 間の変動に分解し、軌跡有用性と集合被覆の共通 closed-loop 基準とする。 - これらの指標を統合した curation objective で冗長性を減らし相補的なデモを選ぶ performance-aligned subset optimizer を提案。

4. どうやって有効だと検証した?

- RoboMimic と OmniReset で広範な実験を実施。 - 単一シミュレータ、sim-to-sim、sim-to-real の各設定で TUCO が state-of-the-art 性能を達成することを示した。 - 能動的なシミュレーションデータキュレーションが sim-to-real 方策 co-training に有効であることを確立。

5. 議論はある?

- 要旨からは不明。

6. 次に読むべき論文は?

- RoboMimic - OmniReset - influence functions を用いたデータ選択・キュレーションの関連研究

※ AIが要旨から生成した要約です。正確性は原文をご確認ください。

著者: Ning Zhu, Mengfei Zhao, Yikai Tang, Zhangyujie Sun, Peihao Li, Dongyue Ni, Jindou Jia, Jianfei Yang

分類: cs.RO

原文アブストラクト

Simulation demonstrations can supplement scarce real-world data for robot policy co-training. However, the value of using data curation to actively select these demonstrations for sim-to-real co-training remains underexplored. Existing curation methods also lack a unified criterion for measuring trajectory-level utility and set-level coverage from closed-loop target behavior. To address these gaps, we present the first systematic study of data curation for sim-to-real robot policy co-training and propose Trajectory-level Utility and set-level Coverage Optimization (TUCO). TUCO uses influence functions to trace how each source demonstration affects target-domain scoring rollouts. Our key insight is that these effects can be decomposed into an overall contribution to target return and variation across rollouts, providing a common closed-loop basis for measuring trajectory utility and set coverage. We further propose a performance-aligned subset optimizer that combines these measures in a unified curation objective to reduce redundancy and select complementary demonstrations. Extensive experiments on RoboMimic and OmniReset establish the value of active simulation data curation for sim-to-real policy co-training and show that TUCO achieves state-of-the-art performance across single-simulator, sim-to-sim, and sim-to-real settings.

関連論文

PR本紙発行元 EmplifAI