日本フィジカルAI新聞

世界のフィジカルAIを、日本語で。

週刊ニュースレター購読
世界モデル/sim2realarXiv:2610.02691

DeltaWorld: 行動条件付き潜在差分学習による物理的に整合な対話型世界シミュレータ

DeltaWorld: Physically Consistent Interactive World Simulators via Action-Conditioned Latent Increment Learning

シェア:XThreadsFacebookLINEはてブBluesky

ロボット行動による潜在特徴の変化量を予測して次状態に加算する手法を提案し、物体の貫通や過剰変形を抑えた物理的に整合な操作シミュレータを実現した。

詳しい要約

1. どんなもの?

- ロボット操作のための物理的に整合する対話型world simulator「DeltaWorld」を提案。 - 行動条件付きの将来画像列を生成し、ロボット計画・policy学習・評価をスケール可能にする。 - 長期的なrobot-object相互作用のダイナミクスを保つことを目指す。 - IWS manipulation benchmarkと複数embodimentの自己収集cross-robot datasetで評価。

2. 先行研究と比べてどこがすごい?

- 既存world modelは次のlatent state全体を予測し、ロボット行動による微妙な変化を捉え損ねる。 - その結果、object interpenetrationや過剰変形など物理的に不自然な結果を生む。 - DeltaWorldはaction-induced latent feature changesを予測して加算するDelta-LTMでこの問題に対処。 - cross-robot datasetでIWS baseline比FVD 46.6%減、LPIPS 31.1%減を達成。

3. 技術・手法の肝は?

- Delta Latent Transition Model (Delta-LTM):次のlatent stateを直接予測せず、行動起因のlatent feature変化を予測し現在状態に加算。 - Interaction-aware Latent Alignment:counterfactual interaction regionsを構築し、相互作用関連のlatent変化を監督。 - これによりobject interpenetrationと過剰変形を軽減。 - 長期的なaction-conditioned video predictionを可能にする。

4. どうやって有効だと検証した?

- IWS manipulation benchmarkで評価。 - 複数のrobot embodimentsとmanipulation tasksを含む自己収集cross-robot datasetで評価。 - cross-robot datasetにおいてIWS baseline比でFVD 46.6%減、LPIPS 31.1%減。 - 長期的なaction-conditioned video predictionへの可能性を示す。

5. 議論はある?

- 既存world modelのlatent state全体予測が微妙な行動変化を捉えず、物理的不整合を生む点を指摘。 - Delta-LTMとInteraction-aware Latent Alignmentで対処。 - 評価結果は有効性を示すが、限界や失敗ケース、計算コスト、一般化性の議論は要旨からは不明。

6. 次に読むべき論文は?

- IWS baseline(Interactive World Simulator) - Delta Latent Transition Model (Delta-LTM) 関連のlatent dynamicsモデル - Interaction-aware Latent Alignment 関連のcounterfactual supervision研究 - 同分野の定番としてDreamer、PlaNetなどのworld model

※ AIが要旨から生成した要約です。正確性は原文をご確認ください。

著者: Boyuan Hou, Xiaoge Cao, Chaofan Zhang, Shuo Wang, Shaowei Cui

分類: cs.RO

原文アブストラクト

Interactive world simulators can provide scalable environments for robot planning, policy training, and evaluation by predicting action consequences while reducing reliance on repeated physical rollouts. To serve these applications, they must generate future image sequences that respond faithfully to robot actions and preserve the dynamics of robot-object interactions over long horizons. However, existing world models typically predict the entire next latent state and often fail to capture subtle changes induced by robot actions. Such omissions can produce physically implausible outcomes, including object interpenetration and excessive deformation. To address this limitation, we propose DeltaWorld, a physically consistent interactive world simulator for robotic manipulation. Our method introduces the Delta Latent Transition Model (Delta-LTM), which predicts action-induced latent feature changes and adds them to the current latent state to obtain the next state, rather than predicting the next latent state directly. To mitigate object interpenetration and excessive deformation in predicted future frames, Interaction-aware Latent Alignment is introduced to construct counterfactual interaction regions and supervise interaction-related latent changes. DeltaWorld is evaluated on the IWS manipulation benchmark and a self-collected cross-robot dataset covering multiple robot embodiments and manipulation tasks. On the cross-robot dataset, DeltaWorld reduces FVD by 46.6% and LPIPS by 31.1% relative to the IWS baseline. These results highlight the potential of DeltaWorld for long-horizon action-conditioned video prediction in robotic manipulation.

PR本紙発行元 EmplifAI