日本フィジカルAI新聞

世界のフィジカルAIを、日本語で。

週刊ニュースレター購読
ワールドモデルarXiv:2609.20034

Astronex-World 1.0: リアルタイム対話型ワールドモデル基盤

Astronex-World 1.0: Real-Time Interactive World Model Foundation

シェア:XThreadsFacebookLINEはてブBluesky

テキストや画像からカメラ軌道・連続行動・身体IDに沿って未来映像を予測する制御可能な動画ワールドモデルを提案し、双方向モデルと因果モデルを構築して24fpsでリアルタイム生成を実現した。

詳しい要約

1. どんなもの?

- Astronex-World 1.0は、オープンで制御可能なビデオ世界モデル基盤。 - テキストまたは初期画像から未来の視覚状態を予測。 - フレーム整合したカメラ軌道、連続アクション、embodiment識別子を入力。 - ロールアウトの指定位置にテキストイベントを挿入可能。 - 双方向モデルと因果モデルを提供。 - 双方向モデルは完全コンテキスト生成、因果モデルはblock-causal attentionとcross-block KV cachingで永続生成。 - 両方ともWan2.2-TI2V-5B prior上に構築。

2. 先行研究と比べてどこがすごい?

- 5BモデルながらWBench Fullで13.6B LongCat-Videoや14B Heliosを上回る。 - 22B LTX-2.3とは1ポイント差。 - 同じ5B priorからNVIDIA A100でpost-trainされたYUME 1.5を上回る。 - リアルタイムストリーミングを単一GPUで実現。 - 全5段階の学習を2台のNVIDIA L20 48GB GPUで実行可能。

3. 技術・手法の肝は?

- Wan2.2-TI2V-5B priorをベース。 - PRoPEでカメラintrinsicsとextrinsicsを注入。 - 64次元action streamが各Transformer層を変調。 - 5段階学習:双方向カメラ・アクション制御、block-causal生成への変換、few-step studentの蒸留、混合ドメイン動力学の復元、非対称DMD/DMD2分布マッチング。 - 因果モデルはblock-causal attentionとcross-block KV cachingを採用。

4. どうやって有効だと検証した?

- WBench Naviで73.5、WBench Fullで70.0を記録。 - WBench Fullで13.6B LongCat-Video、14B Helios、YUME 1.5を上回り、22B LTX-2.3に迫る。 - 因果モデルは832x480解像度、24fpsでリアルタイムストリーミング。 - 全5段階の学習が2台のNVIDIA L20 48GB GPUで実行可能。

5. 議論はある?

- 予約されたaction入力・出力インターフェースにより、embodied intelligenceや自動運転へのpost-trainingが可能。 - その他の議論や限界は要旨からは不明。

6. 次に読むべき論文は?

- Wan2.2-TI2V-5B - LongCat-Video - Helios - LTX-2.3 - YUME 1.5 - DMD/DMD2 - PRoPE

※ AIが要旨から生成した要約です。正確性は原文をご確認ください。

著者: Xin Zhou, Cong Miao

分類: cs.CV, cs.AI, cs.RO

原文アブストラクト

We present Astronex-World 1.0, an open controllable video world-model foundation. Given a text prompt (text-to-video) or an initial observation (image-to-video), the model predicts future visual states under frame-aligned camera trajectories, continuous actions, and an embodiment identifier, and accepts text events inserted at a specified position of a rollout. The family provides a bidirectional model for full-context generation and a causal model with block-causal attention and cross-block KV caching for persistent generation, both built on the Wan2.2-TI2V-5B prior. PRoPE injects camera intrinsics and extrinsics, while a 64-dimensional action stream modulates every Transformer layer. A five-stage training path develops bidirectional camera and action control, converts the backbone to block-causal generation, distills a few-step student, restores mixed-domain dynamics, and applies asymmetric DMD/DMD2 distribution matching. The causal model generates 832x480 video at 24 fps. All five training stages run on two NVIDIA L20 48 GB GPUs, and the causal model streams in real time on one. It scores 73.5 on WBench Navi and 70.0 on WBench Full. On Full, this 5B model is above the 13.6B LongCat-Video and the 14B Helios, within one point of the 22B LTX-2.3, and above YUME 1.5, which is post-trained from the same 5B prior on NVIDIA A100 GPUs. The reserved action input and output interfaces allow post-training for embodied intelligence and autonomous driving.

関連論文

PR本紙発行元 EmplifAI