日本フィジカルAI新聞

世界のフィジカルAIを、日本語で。

週刊ニュースレター購読
VLAarXiv:2610.04540

DreamFormer: 言語条件付きロボット操作のためのTransformer世界モデルによる夢模倣

DreamFormer: Dream Imitation with a Transformer World Model for Language-Conditioned Robotic Manipulation

シェア:XThreadsFacebookLINEはてブBluesky

学習した世界モデルの潜在空間内で専門家のデモを模倣し、言語条件付き多タスクスキルを獲得するモデルベースエージェントを提案。CALVINベンチマークで既存手法を上回る。

詳しい要約

1. どんなもの?

- 言語条件付き多タスクロボット操作のための model-based agent - 学習した world model の latent imagination 内で expert demonstrations を模倣 - まず task-agnostic Transformer world model を unstructured play data から学習 - 次に latent space で agent rollouts を expert demonstrations に整合させる intrinsic reward を最適化 - 長期 horizon の imagination を可能にするため、高解像度 multi-view 観測を単一 input token に符号化

2. 先行研究と比べてどこがすごい?

- 先行 Transformer world models は multi-token 表現や spatial downsampling を使用 - DreamFormer は高解像度 multi-view 観測を単一 token に符号化し、長期 imagination を実用的に - offline behavioral cloning の covariate shift を、imagination 内 on-policy 訓練で緩和 - CALVIN benchmark で LUMOS を上回る (2.52 vs 2.34) - HULC と比較し zero-shot transfer でほぼ倍 (1.30 vs 0.67)

3. 技術・手法の肝は?

- task-agnostic Transformer world model を unstructured play data から学習 - latent imagination 内で on-policy に policy を訓練 - intrinsic reward で agent rollouts を expert demonstrations の latent space に整合 - 高解像度 multi-view 観測を単一 input token に符号化 - spatial downsampling と multi-token 表現を回避

4. どうやって有効だと検証した?

- 長期 horizon の CALVIN benchmark で評価 - single-environment evaluation で LUMOS と比較 (2.52 vs 2.34 average tasks completed per chain of five) - zero-shot transfer to unseen environment で HULC と比較 (1.30 vs 0.67) - imagination rollouts の tractability も確認

5. 議論はある?

- world model が学習した dynamics は直接 clone した policy より転移しやすい可能性 - 生物学的 agent における internal models の役割と整合 - 未見状況での behavior を支援 - 詳細な議論や限界は要旨からは不明

6. 次に読むべき論文は?

- LUMOS - HULC - CALVIN benchmark - Transformer world models - behavioral cloning

※ AIが要旨から生成した要約です。正確性は原文をご確認ください。

著者: Mostafa Kotb, Cornelius Weber, Muhammad Burhan Hafez, Stefan Wermter

分類: cs.RO

原文アブストラクト

We introduce DreamFormer, a model-based agent that acquires language-conditioned, multi-task skills by imitating expert demonstrations within the latent imagination of a learned world model. DreamFormer first learns a task-agnostic Transformer world model from unstructured play data, then acquires task-specific behaviors by optimizing an intrinsic reward that aligns agent-generated rollouts with expert demonstrations in latent space. Since the policy is trained on-policy inside imagination, it is exposed to its own errors during training, mitigating the covariate shift inherent to offline behavioral cloning. To make long-horizon imagination affordable, DreamFormer encodes a high-resolution multi-view robotic observation into a single input token, avoiding both spatial downsampling and the multi-token representations used by prior Transformer world models. On the long-horizon CALVIN benchmark, DreamFormer outperforms LUMOS, the comparable model-based agent, on single-environment evaluation (2.52 vs 2.34 average tasks completed per chain of five) while keeping imagination rollouts tractable. Against HULC, the behavior cloning baseline, it nearly doubles performance on zero-shot transfer to an unseen environment (1.30 vs 0.67), indicating that dynamics learned by the world model transfer more readily than a directly cloned policy. This is consistent with the role attributed to internal models in biological agents, where a model of environment dynamics supports behavior in situations not previously encountered.

関連論文

PR本紙発行元 EmplifAI