日本フィジカルAI新聞

世界のフィジカルAIを、日本語で。

週刊ニュースレター購読
VLAarXiv:2609.24526

ME-VLM:身体性認知とエージェント連携を統合した視覚言語モデル

ME-VLM:A Unified VLM for Embodied Cognition and Agent Coordination

シェア:XThreadsFacebookLINEはてブBluesky

身体性認知とマルチモーダルエージェント能力を統合したVLMを提案し、身体性・エージェント・自動運転・ナビゲーションのベンチマークで競争力のある性能を達成。エッジ展開向けの最適化も行った。

詳しい要約

1. どんなもの?

- 提案: MachEmbodied-VLM (ME-VLM) - 4B と 35B-A3B の2 variant を持つ unified VLM - 目的: embodied cognition と multimodal agent 能力の統合 - 対象: 物理環境とデジタル環境 - physical perception と spatiotemporal reasoning - planning, interaction, outcome assessment - 学習データ: embodied と multimodal agent タスク - execution observations と feedback を含む - 学習 pipeline: 3段階 - embodied capability injection - embodied と multimodal-agent experts の separate RL - multi-teacher on-policy distillation で統合 - 評価: embodied, agent,…

2. 先行研究と比べてどこがすごい?

- 従来の VLM は visual-linguistic understanding 中心 - 実世界の environmental constraints や execution feedback を扱いにくい - 従来の embodied AI は特定タスクや環境に限定されがち - ME-VLM の差分 - embodied cognition と multimodal agent 能力を単一モデルに統合 - 4B と 35B-A3B の2規模で提供 - execution observations と feedback を学習データに含める - outcome assessment と decision refinement を支援 - digital と physical の両環境を対象 - edge 展開向け最適化で 4B を M100 上で動作 - 具体的な先行研究名や数値比較は要旨からは不明

3. 技術・手法の肝は?

- 学習データ構築 - embodied と multimodal agent タスクを横断 - execution observations と feedback を含める - 学習 pipeline - embodied capability injection - embodied expert と multimodal-agent expert を separate RL で学習 - multi-teacher on-policy distillation で両 expert の補完的能力を単一モデルへ統合 - モデル variant - 4B と 35B-A3B - edge 展開技術 - visual token compression - W4A8 quantization - hardware-software co-optimization - 対象能力 - physical perception, spatiotemporal reasoning - planning, interaction, outcome asses…

4. どうやって有効だと検証した?

- 評価 benchmark - embodied benchmarks - agent benchmarks - autonomous-driving tasks - embodied-navigation tasks - 結果 - 両 benchmark で competitive performance - edge 検証 - 4B variant を M100 上で on-device inference - prefill latency を 400 ms から 188 ms に削減 - 具体的な数値や比較対象の詳細は要旨からは不明

5. 議論はある?

- 要旨で明示的な limitation や議論は記載されていない - 想定される論点 - embodied と agent の能力統合の一般性 - execution feedback の質と量への依存 - multi-teacher on-policy distillation の安定性 - edge 展開時の精度と latency の trade-off - これらは要旨からは不明

6. 次に読むべき論文は?

- 要旨で参照/比較されている研究は明示されていない - 関連手法として以下が挙げられる - vision-language model (VLM) - embodied AI - multimodal agent - reinforcement learning (RL) - on-policy distillation - W4A8 quantization - visual token compression - 同分野の定番として embodied navigation や autonomous driving の benchmark 論文を次に読むべき

※ AIが要旨から生成した要約です。正確性は原文をご確認ください。

著者: Foundation Model, Li Auto Inc

分類: cs.CV

原文アブストラクト

Physical AI requires models to ground visual and linguistic understanding in real-world environments while accounting for environmental constraints and execution feedback. We introduce MachEmbodied-VLM (ME-VLM), a unified vision-language model with two variants, 4B and 35B-A3B, that brings together embodied cognition and multimodal agent capabilities. Our work emphasizes physical perception and spatiotemporal reasoning, together with planning, interaction, and outcome assessment in both digital and physical environments. We construct training data spanning embodied and multimodal agent tasks, including execution observations and feedback to support outcome assessment and decision refinement. The training pipeline comprises embodied capability injection, separate reinforcement learning of embodied and multimodal-agent experts, and multi-teacher on-policy distillation that consolidates their complementary capabilities into a single model. Experiments show competitive performance on both embodied and agent benchmarks, as well as on autonomous-driving and embodied-navigation tasks. For edge deployment, visual token compression, W4A8 quantization, and hardware--software co-optimization enable on-device inference of the 4B variant on the M100, reducing prefill latency from 400 ms to 188 ms. Project Page: https://machembodied.com/ME-Brain/ME-VLM.html Code Repository: https://github.com/MachEmbodied/ME-VLM

関連論文

PR本紙発行元 EmplifAI