Zhide Zhong
The Hong Kong University of Science and Technology (Guangzhou)
収録論文 19本 ・ フィジカルAI/ロボット学習
VLAワールドモデルVLA/クロス身体操作
※arXiv著者名で収集。同姓同名の別人の論文が含まれる場合があります。
論文
- SG-WAM: テキスト接地と空間認識を備えた意味的ガイダンスによる世界行動モデルVLA2026/8/9
世界行動モデル(WAM)の将来ビデオ生成と行動予測を言語指示に整合させるため、VLMベースのプランナーで意味的先見を生成し、それを高レベルな意味的ガイダンスとして注入する手法を提案。シミュレーションと実世界で精度と指示追従性を実証。
- 前方予測だけで十分か?JEPAワールドモデルのための物理状態接地ワールドモデル2026/8/7
JEPAベースのワールドモデルに、ロボットの自己受容状態と関節角変化を接地する2つの目的を追加し、潜在表現の識別性と下流タスク性能を向上させる手法を提案した。
- DyPES-VLA: 共有ダイナミクス事前分布と身体特化制御の学習によるクロス身体操作VLA/クロス身体操作2026/8/6
異なるロボット身体間で共有できるダイナミクス事前分布を学習し、身体ごとの専門家ネットワークで直接制御するVLAモデルを提案。手動の行動形式変換を不要にし、シミュレーションと実機で最高性能を達成した。
- Robust-WAM: 生成事前学習と意味的予見を橋渡しするワールド・アクションモデルVLA2026/8/6
ロボット制御用のワールド・アクションモデルにおいて、VAE潜在空間の生成事前学習を保ちつつ、意味的潜在空間の頑健性を組み込む後処理手法を提案した。外観変化に頑健なアクション予測を実現する。
- DyPES-VLA: Learning Shared Dynamics Priors and Embodiment-Specific Control for Cross-Embodiment Manipulation2026/8/1
- SG-WAM: Text-Grounded and Spatial-aware Semantic Guidance for World-Action Models2026/8/1
- Robust-WAM: Bridging Generative Pretraining and Semantic Foresight in World-Action Models2026/8/1
- Is Forward Prediction Enough? Physical State Grounding for JEPA World Models2026/8/1
- S-VAM: Shortcut Video-Action Model by Self-Distilling Geometric and Semantic Foresight2026/3/1
- DFM-VLA: Iterative Action Refinement for Robot Manipulation via Discrete Flow Matching2026/3/1
- VLA-OPD: Bridging Offline SFT and Online RL for Vision-Language-Action Models via On-Policy Distillation2026/3/1
- DyGeoVLN: Infusing Dynamic Geometry Foundation Model into Vision-Language Navigation2026/3/1
- DualCoT-VLA: Visual-Linguistic Chain of Thought via Parallel Reasoning for Vision-Language-Action Models2026/3/1
- Rethinking the Practicality of Vision-language-action Model: A Comprehensive Benchmark and An Improved Baseline2026/2/1
- Embodied Robot Manipulation in the Era of Foundation Models: Planning and Learning Perspectives2025/12/1
- Towards a Unified Understanding of Robot Manipulation: A Comprehensive Survey2025/10/1
- FlowVLA: Visual Chain of Thought-based Motion Reasoning for Vision-Language-Action Models2025/8/1
- PD-VLA: Accelerating Vision-Language-Action Model Integrated with Action Chunking via Parallel Decoding2025/3/1
- ASSIST: Interactive Scene Nodes for Scalable and Realistic Indoor Simulation2023/11/1