日本フィジカルAI新聞

世界のフィジカルAIを、日本語で。

週刊ニュースレター購読
VLAarXiv:2609.14973

PhysBrain 1.5:視覚言語モデルから物理基盤モデルへ

PhysBrain 1.5: From Vision-Language Models to Physical Foundation Models

シェア:XThreadsFacebookLINEはてブBluesky

人間のインタラクション動画から身体動作を学習し、行動生成と未来状態予測を統合した8Bの物理基盤モデルを構築、28の身体理解ベンチマークでオープンソース最高性能を達成した。

詳しい要約

1. どんなもの?

- 物理環境の理解・行動生成・将来状態予測を統合したモデル - 観察・相互作用・環境変化の物理ループに基づく - 視覚言語モデルから出発し、言語応答・end-effector運動・dense visual targetsを離散系列として自己回帰的に学習 - 8Bモデルで28のembodied understandingベンチマーク平均72.5を達成

2. 先行研究と比べてどこがすごい?

- 従来の視覚言語モデルを物理基盤モデルへ拡張 - オープンソースで新たなSOTAを達成し、GPT-6-AstraやGemini 3.6 Flashなどの商用モデルと同等の性能 - 14のベンチマークで最高のオープンソース結果 - 一般的なマルチモーダル能力も維持

3. 技術・手法の肝は?

- 一般的なvision-language modelを基盤とする - 言語応答、end-effector motion、dense visual targetsを離散系列にエンコード - 自己回帰的なnext-token predictionで統合最適化 - 事前学習は人間の相互作用動画からembodied supervisionを取得 - タスク中心のエピソードで意味的・空間的文脈と復元された運動・後続観察をペアリング - 人間のデモ、ロボット軌道、シミュレーション経験の混合でsupervised fine-tuning

4. どうやって有効だと検証した?

- 28のembodied understandingベンチマークで評価 - 8Bモデルが平均スコア72.5を達成 - 14のベンチマークで最高のオープンソース結果 - 定性的にend-effector軌道生成と将来シーン予測(空間的に整合したRGB、depth、robot-mask出力)を実証

5. 議論はある?

- 要旨からは不明

6. 次に読むべき論文は?

- GPT-6-Astra - Gemini 3.6 Flash - 一般的なvision-language model - embodied understandingベンチマーク

※ AIが要旨から生成した要約です。正確性は原文をご確認ください。

著者: DeepCybo Team, Yu Bin, Haipeng Cao, Zheng Chang, Kai Chen, Youning Chen, Kailin Deng, Yichao Du, Xiaotong Fu, Haoyang Ge, Yunlong Guo, Chenliu Hao, Jiyan He, Xuguo He, Yakun Hou, Kai Hu, Cong Huang, Tuopusen Huang, Yu Huang, Hong Li, Peize Li, Shijie Lian, Xiaopeng Lin, Yun Lin, Haibao Liu, Haochen Liu, Qiuzhi Liu, Shengcai Liu, Zhiqiang Liu, Tao Luo, Peng Ren, Shuo Ren, Chaoyi Ruan, Zhaolong Shen, Yukun Shi, Qiyuan Su, Yuxuan Tian, Yining Wang, Changti Wu, Hao Wu, Xueyin Xu, Ruoqi Yang, Zhaoyang Yang, Hang Yuan, Zhaoyang Zeng, Hanwen Zhang, Ruimeng Zhang, Yao Zhang, Yibo Zhang, Yuxiang Zhang, Zhirui Zhang, Ziyi Zhang, Zubin Zheng, Zishen Zhuang

分類: cs.CV, cs.RO

原文アブストラクト

We present PhysBrain 1.5, a unified model for understanding physical environments, generating actions, and predicting future states. Motivated by the physical loop of observation, interaction, and environmental change, we bring these capabilities into a common learning framework. Starting from a general vision--language model, we encode language responses, end-effector motion, and dense visual targets as discrete sequences and jointly optimize them with autoregressive next-token prediction. Pre-training draws its embodied supervision entirely from human interaction videos, using task-centered episodes to pair semantic and spatial context with recovered motion and subsequent observations. We then adapt the model through supervised fine-tuning on a mixture of human demonstrations, robot trajectories, and simulated experience. Across 28 embodied understanding benchmarks, our 8B model achieves an average score of 72.5, setting a new open-source state of the art and performing on par with leading proprietary models such as GPT-6-Astra and Gemini 3.6 Flash. It achieves the best open-source results on 14 benchmarks while retaining general multimodal capabilities. Beyond these understanding evaluations, qualitative examples show the model's ability to produce end-effector trajectories and predict future scenes through spatially aligned RGB, depth, and robot-mask outputs.

関連論文