日本フィジカルAI新聞

世界のフィジカルAIを、日本語で。

週刊ニュースレター購読
VLAarXiv:2609.32253

樹状突起に着想を得た頑健な行動制御のための視覚-言語-行動モデルDS-VLA

DS-VLA: A Dendritic-inspired Vision-Language-Action Model for Robust Action Control

シェア:XThreadsFacebookLINEはてブBluesky

脳の樹状突起の動態を模したアクションニューロン構造をVLAモデルに導入し、行動が一時的に乱される状況でも高い成功率を維持できる頑健な制御を実現した。

著者: Yaxing Lyu, Jingyi Li, Mingkun Xu, Yujie Wu

分類: cs.RO, cs.AI, cs.NE

原文アブストラクト

Vision-language-action (VLA) models have achieved strong performance in language-conditioned manipulation, yet success under nominal evaluation does not necessarily translate into robust closed-loop behavior when executed actions are transiently corrupted. We introduce DS-VLA, a dendritic-inspired action architecture that incorporates dendritic spiking dynamics into VLA control to address this limitation. Specifically, to enable modularized feature processing and temporal information integration, DS-VLA equips action neurons with multiple sparsely connected dendritic branches, each featuring heterogeneous, learned decay factors. Furthermore, to suppress unreliable state updates while preserving task-relevant historical information, we introduce a neuron-wise inhibitory gate that adaptively regulates the admission of new multimodal evidence into dendritic states prior to somatic dynamics. We evaluate DS-VLA on all four LIBERO suites under both nominal rollouts and a unified closed-loop action-perturbation protocol. DS-VLA achieves a 91.6\% average nominal success rate and an 87.35\% average perturbed success rate, retaining 95.4\% of its nominal performance. Under the same reported perturbation setting, OpenVLA-OFT, FAST, $π_0$, and GR00T achieve 39.45\%, 23.90\%, 28.55\%, and 30.75\%, respectively. A controlled ablation isolates the contribution of neuron-wise shared inhibition, while analyses of neural dynamics and post-perturbation trajectories associate robust performance with selective evidence suppression and effective behavioral recovery. Together, these results demonstrate that integrating brain-inspired computational mechanisms offers a promising architectural prior for robust embodied intelligence beyond merely scaling vision-language backbones or generative action decoders.

関連論文

PR本紙発行元 EmplifAI