Songen Gu
Fudan University
収録論文 15本 ・ フィジカルAI/ロボット学習
VLA
※arXiv著者名で収集。同姓同名の別人の論文が含まれる場合があります。
論文
- Vid2WAM: ビデオ拡散事前知識を世界行動モデルへ蒸留するVLA2026/8/9
大規模ビデオ基盤モデルの拡散事前知識をコンパクトな世界行動モデルに蒸留し、専門家デモの少ない環境でもロボットの汎化性能とデータ効率を向上させる手法を提案した。
- Vid2WAM: Distilling Video Diffusion Priors into World Action Models2026/8/1
- VT-WAM: Visual-Tactile World Action Model for Contact-Rich Manipulation2026/7/1
- $\tau_0$-WM: A Unified Video-Action World Model for Robotic Manipulation2026/6/1
- TacForeSight: Force-Guided Tactile World Model for Contact-Rich Manipulation2026/6/1
- VistaBot: View-Robust Robot Manipulation via Spatiotemporal-Aware View Synthesis2026/4/1
- PokeVLA: Empowering Pocket-Sized Vision-Language-Action Model with Comprehensive World Knowledge Guidance2026/4/1
- OmniVTA: Visuo-Tactile World Modeling for Contact-Rich Robotic Manipulation2026/3/1
- Say, Dream, and Act: Learning Video World Models for Instruction-Driven Robot Manipulation2026/2/1
- World In Your Hands: A Large-Scale and Open-Source Ecosystem for Learning Human-Centric Manipulation in the Wild2025/12/1
- Mimir: Hierarchical Goal-Driven Diffusion with Uncertainty Propagation for End-to-End Autonomous Driving2025/12/1
- Data Scaling Laws for Imitation Learning-Based End-to-End Autonomous Driving2024/12/1
- ComDrive: Comfort-Oriented End-to-End Autonomous Driving2024/10/1
- GaussianGrasper: 3D Language Gaussian Splatting for Open-vocabulary Robotic Grasping2024/3/1
- ASSIST: Interactive Scene Nodes for Scalable and Realistic Indoor Simulation2023/11/1