日本フィジカルAI新聞

世界のフィジカルAIを、日本語で。

週刊ニュースレター購読
世界モデルarXiv:2609.27656

InternW0: 効率的な実世界インタラクションのための基盤的物理世界モデル

InternW0: A Foundational Physical World Model for Efficient Real-World Interactions

シェア:XThreadsFacebookLINEはてブBluesky

上海AI Labが、視覚予測とロボット制御を非対称なvideo-actionアーキテクチャで統合した物理世界モデルInternW0を提案。約7,200時間のデータで学習し、化学合成やピペッティングなどの実世界タスクで評価した。

詳しい要約

1. どんなもの?

- 上海AI LaboratoryのInternW物理世界モデルシリーズの最初の実装 - omnimodal interfaces、非同期多周波処理、部分観測と外部影響下の局所物理モデリングを中心に構築 - 未来の視覚ダイナミクスと連続ロボット制御を非対称video-actionアーキテクチャとflow matchingで共同学習 - 大容量video expertが長期予測コンテキストを提供し、軽量action expertが高速タイムスケールで動作 - ドメイン固有インターフェースとsoft promptsで異種embodimentsをサポート - contact-aware post-trainingで力と触覚信号を組み込み、接触リッチなマニピュレーションに対応

2. 先行研究と比べてどこがすごい?

- 従来の物理世界モデルは予測が行動可能であり続けることを保証しにくい - InternW0は非対称video-actionアーキテクチャにより、予測と制御を異なるタイムスケールで統合 - 行動更新ごとに未来を再生成せず、layerwise K/Vを再利用し観測条件付きcontext routingで新状態に適応 - 異種embodiments向けのドメイン固有インターフェースとsoft promptsを導入 - 力・触覚信号を組み込むcontact-aware post-trainingを実施 - 約7,200時間の異種ロボット・一人称視点データで学習し、EgoLab(275時間の実ラボ一人称データセット)を含む

3. 技術・手法の肝は?

- 非対称video-actionアーキテクチャとflow matchingを採用 - 大容量video expertが長期予測コンテキストを提供 - 軽量action expertが高速タイムスケールで動作 - 行動更新ごとに未来を再生成せず、layerwise K/Vを再利用 - 観測条件付きcontext routingで新たに観測された状態に適応 - ドメイン固有インターフェースとsoft promptsで異種embodimentsをサポート - contact-aware post-trainingで力と触覚信号を組み込む

4. どうやって有効だと検証した?

- シミュレーションベンチマークと実世界科学タスクで評価 - 15段階のmetal-organic framework合成ワークフロー - 5段階の接触・力対応の器用なマニピュレーション(汎用量的ピペッティング) - 約7,200時間の異種ロボット・一人称データ(EgoLab含む)で学習

5. 議論はある?

- 要旨からは不明

6. 次に読むべき論文は?

- InternWシリーズの他の論文 - flow matchingを用いたvideo-actionモデル - 物理世界モデルに関する先行研究 - EgoLabデータセット - metal-organic framework合成 - 接触・力対応マニピュレーション

※ AIが要旨から生成した要約です。正確性は原文をご確認ください。

著者: Jisong Cai, Yao Mu, Ganlin Yang, Zhe Cao, Zhangzheng Tu, Xing Gao, Kailin Li, Xinyu Zhan, Lixin Yang, Yangkun Zhu, Haoxiang Ma, Ming Zhou, Qiaojun Yu, Yufei Xue, Liqun He, Yifei Yao, Yifan Zhu, Long Ling, Bingqi Jiang, Haoyu Guo, Xueyue Zhu, Bowen Zhou, Bin Zhao, Tianfan Xue, Chunhua Shen, Weinan Zhang

分類: cs.RO, cs.AI

原文アブストラクト

Physical intelligence requires more than predicting how the world may evolve: predictions must remain actionable as the world continues to change. We introduce InternW0, the first instantiation of the InternW physical world model series from Shanghai AI Laboratory, built around omnimodal interfaces, asynchronous multi-frequency processing, and local physical modeling under partial observations and external influences. InternW0 jointly learns future visual dynamics and continuous robot control through an asymmetric video--action architecture with flow matching. A high-capacity video expert provides longer-horizon predictive context, while a lightweight action expert operates at a faster timescale. Instead of regenerating the future for every action update, InternW0 reuses layerwise K/V and adapts it to newly observed states through observation-conditioned context routing. Domain-specific interfaces and soft prompts support heterogeneous embodiments, while contact-aware post-training incorporates force and tactile signals for contact-rich manipulation. We train InternW0 on approximately 7,200 hours of heterogeneous robot and egocentric data, including EgoLab, a 275-hour real-laboratory egocentric dataset. Evaluation spans simulation benchmarks and real-world scientific tasks, including a 15-stage metal--organic framework synthesis workflow and 5-stage contact- and force-aware dexterous manipulation for general-purpose quantitative pipetting. These results advance scalable, asynchronous, and science-native physical world models for universal and efficient real-world interactions.

関連論文

PR本紙発行元 EmplifAI