日本フィジカルAI新聞

世界のフィジカルAIを、日本語で。

週刊ニュースレター購読
世界モデル/計画arXiv:2609.35138

FlexiWorld: 複数時間スケールにわたる柔軟なアクションチャンクによる学習と計画

FlexiWorld: Learning and Planning via Flexible Action Chunks Across Multiple Time Scales

シェア:XThreadsFacebookLINEはてブBluesky

可変長アクションチャンクと混合スパン目標監督を組み合わせたJEPAベースの世界モデルを提案し、長期的制御の成功率を向上させた。

詳しい要約

1. どんなもの?

- FlexiWorldは、JEPAベースのworld modelで、可変長action chunksと混合スパン目標監督を組み合わせ、長期的制御を改善する。 - 訓練中に目標スパンを変化させ、行動をランダムに可変長チャンクに分割する。 - 因果的action encoderと自己回帰actorを共同訓練し、計画にはARCEMを用いる。 - 4つのベンチマークで平均成功率89.29%を達成。

2. 先行研究と比べてどこがすごい?

- 既存手法は固定長チャンクを用い、目標条件付き行動生成を省略するか、短い目標スパンに監督を限定していた。 - FlexiWorldは混合スパン目標監督と可変長action chunksを組み合わせ、長期的制御を改善。 - 最強ベースラインの83.98%に対し、89.29%の平均成功率を達成。 - 再訓練なしで異なる計画チャンク長をサポートし、長いチャンクでARCEMを約1.3倍高速化。

3. 技術・手法の肝は?

- JEPAベースのworld modelで、混合スパン目標監督と可変長action chunksを統合。 - 訓練時に目標スパンを変化させ、行動をランダムに可変長チャンクに分割。 - 因果的action encoderが可変長チャンクを埋め込み、自己回帰actorがプリミティブ行動を逐次生成。 - Student Forcingで生成行動プレフィックスを訓練し、exposure biasを低減。 - 計画にはARCEMを使用し、action-residual searchとチャンク内自己回帰フィードバック、チャンク境界潜在予測を組み合わせる。

4. どうやって有効だと検証した?

- 4つのベンチマークと目標距離で評価し、平均成功率89.29%を達成(最強ベースラインは83.98%)。 - PushTアブレーションで、混合スパン監督、可変長チャンク、Student Forcingが直接制御を改善することを示す。 - 再訓練なしで異なる計画チャンク長をサポートし、長いチャンクでARCEMを約1.3倍高速化しつつ平均成功率を維持。

5. 議論はある?

- 要旨からは不明。

6. 次に読むべき論文は?

- 要旨で参照/比較されている研究や関連手法は明示されていない。同分野の定番として、JEPA、world models、action chunks、Cross-Entropy Method (CEM) などが挙げられる。

※ AIが要旨から生成した要約です。正確性は原文をご確認ください。

著者: Shidu Ren, Qilin Gu, Zhenghao Ni, Junhan Sun, Jiaqi Wang, Damien Scieur, Yunze Liu

分類: cs.LG

原文アブストラクト

Latent world models predict future states for goal-directed planning using action chunks spanning multiple primitive steps. Existing methods typically use fixed-length chunks and either omit goal-conditioned action generation or limit their supervision to short goal spans. We introduce FlexiWorld, a JEPA-based world model that combines mixed-span goal supervision with variable-length action chunks to improve long-horizon control. During training, we sample varying goal spans and randomly partition the actions into variable-length chunks. We jointly train the world model with a causal action encoder that embeds variable-length chunks and an autoregressive actor that generates primitive actions sequentially. Student Forcing reduces exposure bias by training on generated action prefixes. For planning, Actor-Residual Cross-Entropy Method (ARCEM) combines action-residual search with within-chunk autoregressive feedback and chunk-boundary latent prediction. Across four benchmarks and goal distances, FlexiWorld with ARCEM achieves 89.29% mean success, compared with 83.98% for the strongest baseline. PushT ablations show improved direct control from mixed-span supervision, variable-length chunks, and Student Forcing. Without retraining, FlexiWorld supports different planning chunk lengths: longer chunks accelerate ARCEM by approximately $1.3\times$ on average while maintaining comparable average success.

関連論文

PR本紙発行元 EmplifAI