日本フィジカルAI新聞

世界のフィジカルAIを、日本語で。

週刊ニュースレター購読
ワールドモデルarXiv:2609.17909

Zing-0.5: リアルタイムな操作とテキスト制御で遊べる世界生成へ

Zing-0.5: Toward Playable Worlds with Real-Time Joint Action and Text Control

シェア:XThreadsFacebookLINEはてブBluesky

キーボード操作とテキスト指示を同時に受け付け、リアルタイムで生成される世界を探索・操作できる5Bの自己回帰型ワールドモデルを提案。

詳しい要約

1. どんなもの?

- Zing-0.5は、5Bのautoregressive world modelであり、ユーザーが生成された世界を探索し、出来事に影響を与え、キーボードとオンラインテキスト制御を通じてフィードバックに応答できるplayabilityを目指す。 - 832 x 480の解像度で24 FPSのリアルタイム推論を、ストリーム分あたり約0.009 USDのサーバーレンタルコストで実現する。 - モデル重み、推論コード、Zing-SGLang serving実装を公開し、playable generated worldsの研究を支援する。

2. 先行研究と比べてどこがすごい?

- 先行研究と比べて、unified action and text conditioning、event-scale supervision、low-cost real-time interactionの3つの技術的貢献を統合している点がすごい。 - 特に、joint keyboard and online text controlを可能にし、ナビゲーションとイベント制御を同一シーケンスで学習する点が新しい。 - リアルタイム性能と低コストを両立し、158 WBench Navigationケースでoverall score 81.0、consistency score 88.5を達成。

3. 技術・手法の肝は?

- 技術の肝は3点:(1) Unified action and text conditioning:magnitude-aware keyboard inputsとtemporally aligned text instructions、jointly annotated videosを組み合わせ、ナビゲーションとイベント制御を同一シーケンスで学習。 - (2) Event-scale supervision for incremental generation:connected multi-prompt videosで訓練されたsegment-level teacherが、block-level causal studentをdistribution-matching distillationで監督。 - (3) Low-cost real-time interaction:four-step generationとcontext-preserving streamingを組み合わせ、832 x 480で24 FPSを実現。

4. どうやって有効だと検証した?

- 158 WBench Navigationケースで評価し、overall score 81.0、consistency score 88.5を達成。 - joint-control demonstrationで、生成を再起動せずにナビゲーション継続中にテキスト指示によるイベント変更を示した。 - リアルタイム推論性能(24 FPS)とコスト(約0.009 USD/stream-minute)を推定。

5. 議論はある?

- 要旨からは、限界や議論点についての明示的な記述はない。 - ただし、playable generated worldsの実現に向けた技術的貢献と評価結果が示されており、今後の課題は要旨からは不明。

6. 次に読むべき論文は?

- 要旨で参照/比較されている研究は明示されていない。 - 関連手法として、autoregressive world models、distribution-matching distillation、WBench Navigationが挙げられる。 - 同分野の定番として、World Models、Dreamer、Genieなどの一般名が考えられるが、要旨に直接の参照はない。

※ AIが要旨から生成した要約です。正確性は原文をご確認ください。

著者: Mingyang Chen, Shengdong Chen, Xiaoxiao Fu, Bosheng Gong, Haoyuan Guo, Bowen Li, Jiawen Li, Kejun Li, Tianpeng Li, Yin Liu, Haoze Sun, Zeyang Tian, Meng Wang, Xinmiao Wu, Jiangqiao Yan, Zining Zhao

分類: cs.CV, cs.LG

原文アブストラクト

We introduce Zing-0.5, a 5B autoregressive world model designed for playability: users can explore generated worlds, influence unfolding events, and respond to the resulting feedback through joint keyboard and online text control. Our approach brings together three technical contributions: (1) Unified action and text conditioning, combining magnitude-aware keyboard inputs with temporally aligned text instructions and jointly annotated videos to learn navigation and event control within the same sequence; (2) Event-scale supervision for incremental generation, using a segment-level teacher trained on connected multi-prompt videos to supervise a block-level causal student through distribution-matching distillation; and (3) Low-cost real-time interaction, combining four-step generation with context-preserving streaming to support 832 x 480 inference at 24 FPS at an estimated server rental cost of approximately USD 0.009 per stream-minute. Zing-0.5 achieves an overall score of 81.0 and a consistency score of 88.5 across 158 WBench Navigation cases. A joint-control demonstration shows a text-directed event change during continued navigation without restarting generation. We release the model weights, inference code, and Zing-SGLang serving implementation to support further work on playable generated worlds.

関連論文

PR本紙発行元 EmplifAI