日本フィジカルAI新聞

世界のフィジカルAIを、日本語で。

週刊ニュースレター購読
データ拡張/模倣学習arXiv:2610.02339

NEEDLEWORK: 検証済み局所ステッチによるロボットデータのオフライン書き換え

NEEDLEWORK: Offline Rewriting of Robot Data with Verified Local Stitches

シェア:XThreadsFacebookLINEはてブBluesky

RGB画像と固有感覚のみを用いて、ロボットのデモンストレーション間に短く検証済みの行動ブリッジをオフラインで追加し、失敗軌道も活用してデータ拡張する手法を提案。実機タスクで成功率を平均21ポイント改善した。

詳しい要約

1. どんなもの?

- ロボットのデモンストレーションから、非効率・失敗を含むエピソードでも有用な行動を抽出し、軌道を接続(stitching)して訓練データを拡張するオフライン手法 NEEDLE を提案。 - 高次元ロボットデータ(RGB画像、proprioception、エピソード単位の結果)のみを用い、新たな環境相互作用や特権的な物体状態なしで、記録済み観測間に短く検証済みの action bridge を追加する。 - これにより、suboptimal な回り道を回避し、action coverage を広げ、失敗軌道も活用したデータセット拡張を実現する。

2. 先行研究と比べてどこがすごい?

- 従来の trajectory stitching は低次元状態表現に依存することが多く、高次元ロボットデータでは有用な接続の発見と feasibility の検証が困難だった。 - NEEDLE は RGB 画像と proprioception のみを用い、新たな環境相互作用や privileged object state を必要とせずに、高次元データでの接続生成と検証を可能にした。 - 実ロボットタスクで最強ベースラインに対し平均 21 パーセントポイントの成功率向上を達成。

3. 技術・手法の肝は?

- 記録済み観測間に短い action bridge を追加するオフライン・データ拡張アルゴリズム。 - まず、suboptimal な回り道を迂回し、action coverage を広げ、失敗軌道も活用する接続を特定・生成。 - 次に、受け入れられた bridge をポリシー訓練に組み込むサンプリング手法を提案。中間画像を合成せず、元のデモンストレーションを破棄しないため、元データの coverage を保持しつつ代替行動を学習可能。

4. どうやって有効だと検証した?

- 実ロボットタスクで評価を実施。 - 各タスクで最強ベースラインと比較し、成功率が平均 21 パーセントポイント向上。 - 詳細な実験設定やタスクの種類は要旨からは不明。動画と補足資料が https://needle-work.github.io/ で公開。

5. 議論はある?

- 要旨からは、手法の限界や失敗事例、計算コスト、スケーラビリティに関する議論は不明。 - 実ロボットタスクでの成功率向上は示されているが、他のタスクや環境への一般化可能性については言及されていない。

6. 次に読むべき論文は?

- 要旨で参照・比較されている具体的な先行研究は明示されていない。 - 関連手法として trajectory stitching、offline dataset augmentation、imitation learning、offline reinforcement learning などが挙げられる。 - 同分野の定番として、RoboMimic、D4RL、Decision Transformer、Diffusion Policy などの論文を読むことが推奨される。

※ AIが要旨から生成した要約です。正確性は原文をご確認ください。

著者: Juntao Ren, Yifan Hou, Shuran Song

分類: cs.RO, cs.LG

原文アブストラクト

Robot demonstrations may contain useful behavior even when individual episodes are inefficient or unsuccessful. Trajectory stitching offers a way to compose these behaviors into improved training data, but identifying useful connections and verifying their feasibility is difficult in high-dimensional robot data, where many prior methods rely on low-dimensional state representations. We introduce NEEDLE, an offline dataset-augmentation algorithm that addresses these challenges by adding short, verified action bridges between recorded observations in high-dimensional robot demonstrations. First, NEEDLE identifies and creates connections that bypass suboptimal detours, broaden action coverage, and augment the original dataset with failed trajectories, using only RGB images, proprioception, and episode-level outcomes, without new environment interaction or privileged object state. Next, we present a sampling technique that incorporates accepted bridges into policy training without synthesizing intermediate images or discarding the original demonstrations, allowing policies to learn alternative actions while retaining the original dataset's coverage. On real-robot tasks, NEEDLE improves success rate over the strongest baseline on each task by an average of 21 percentage points. Videos and supplementary materials are on https://needle-work.github.io/.

PR本紙発行元 EmplifAI