Hang Zhao
収録論文 68本 ・ フィジカルAI/ロボット学習
VLA
※arXiv著者名で収集。同姓同名の別人の論文が含まれる場合があります。
論文
- G0.5: ロボットの推論と行動のための単一自己回帰ストリームVLA2026/8/12
事前学習済みのVLMと別の行動エキスパートを組み合わせる従来のVLAモデルに対し、単一のトランスフォーマーデコーダが推論トークンと行動トークンを単一の目的で生成する自己回帰VLAモデルG0.5を提案。クロスエンボディメント行動トークナイザー、ネイティブな思考連鎖ストリーム、視覚メモリモジュールにより、基礎モデル規模での学習を可能にし、複数のベンチマークで最先端を達成。
- SLAMFormer-$\infty$: Infinite SLAM Transformer for Unbounded Frontend and Backend Processing2026/8/1
- G0.5: One Autoregressive Stream for Robot Reasoning and Action2026/8/1
- Agent-driven Long-tail Simulation for Autonomous Driving2026/7/1
- Automated Synthesis of Facial Mechanisms for Conversational Animatronic Robots2026/7/1
- Is Your Trajectory Displacement Safe in Long-tail?2026/6/1
- OMG: Omni-Modal Motion Generation for Generalist Humanoid Control2026/6/1
- PAIWorld: A 3D-Consistent World Foundation Model for Robotic Manipulation2026/6/1
- TopoRetarget: Interaction-Preserving Retargeting for Dexterous Manipulation2026/6/1
- Coarse-to-Control: Action-Token Planning for Vision-Language-Action Models2026/6/1
- Dexora: Open-source VLA for High-DoF Bimanual Dexterity2026/5/1
- UniDex: A Robot Foundation Suite for Universal Dexterous Hand Control from Egocentric Human Videos2026/3/1
- VolumeDP: Modeling Volumetric Representation for Manipulation Policy Learning2026/3/1
- Generative Control as Optimization: Time Unconditional Flow Matching for Adaptive and Robust Robotic Control2026/3/1
- TTT-Parkour: Rapid Test-Time Training for Perceptive Robot Parkour2026/2/1
- Embodied Intelligence for Flexible Manufacturing: A Survey2026/2/1
- ActionCodec: What Makes for Good Action Tokenizers2026/2/1
- Hiking in the Wild: A Scalable Perceptive Parkour Framework for Humanoids2026/1/1
- Deep Whole-body Parkour2026/1/1
- FASTer: Toward Efficient Autoregressive Vision Language Action Modeling via Neural Action Tokenization2025/12/1
- RoboBPP: Benchmarking Robotic Online Bin Packing with Physics-based Simulation2025/12/1
- RoboCOIN: An Open-Sourced Bimanual Robotic Data Collection for Integrated Manipulation2025/11/1
- Galaxea Open-World Dataset and G0 Dual-System VLA Model2025/9/1
- SLAM-Former: Putting SLAM into One Transformer2025/9/1
- Discrete Diffusion for Reflective Vision-Language-Action Models in Autonomous Driving2025/9/1
- SAMP: Spatial Anchor-based Motion Policy for Collision-Aware Robotic Manipulators2025/9/1
- OccVLA: Vision-Language-Action Model with Implicit 3D Occupancy Supervision2025/9/1
- Morpheus: A Neural-driven Animatronic Face with Hybrid Actuation and Diverse Emotion Control2025/7/1
- Conditioning Matters: Training Diffusion Policies is Faster Than You Think2025/5/1
- Designing Pin-pression Gripper and Learning its Dexterous Grasping with Online In-hand Adjustment2025/5/1
- Deliberate Planning of 3D Bin Packing on Packing Configuration Trees2025/4/1
- PIN-WM: Learning Physics-INformed World Models for Non-Prehensile Manipulation2025/4/1
- Learning Personalized Driving Styles via Reinforcement Learning from Human Feedback2025/3/1
- RoboEngine: Plug-and-Play Robot Data Augmentation with Semantic Robot Segmentation and Background Generation2025/3/1
- MoE-Loco: Mixture of Experts for Multitask Locomotion2025/3/1
- Physics-informed Neural Network Predictive Control for Quadruped Locomotion2025/3/1
- Embrace Collisions: Humanoid Shadowing for Deployable Contact-Agnostics Motions2025/2/1
- VR-Robo: A Real-to-Sim-to-Real Framework for Visual Robot Navigation and Locomotion2025/2/1
- Generalizing Motion Planners with Mixture of Experts for Autonomous Driving2024/10/1
- Robust Robot Walker: Learning Agile Locomotion over Tiny Traps2024/9/1
- Playful DoggyBot: Learning Agile and Precise Quadrupedal Locomotion2024/9/1
- SARO: Space-Aware Robot System for Terrain Crossing via Vision-Language Model2024/7/1
- Humanoid Parkour Learning2024/6/1
- LiDAR-based 4D Occupancy Completion and Forecasting2023/10/1
- Large Trajectory Models are Scalable Motion Predictors and Planners2023/10/1
- GPT-Driver: Learning to Drive with GPT2023/10/1
- Boosting Offline Reinforcement Learning for Autonomous Driving with Hierarchical Latent Skills2023/9/1
- Robot Parkour Learning2023/9/1
- A Universal Semantic-Geometric Representation for Robotic Manipulation2023/6/1
- Programmatically Grounded, Compositionally Generalizable Robotic Manipulation2023/4/1
- Learning Physically Realizable Skills for Online Packing of General 3D Shapes2022/12/1
- P4P: Conflict-Aware Motion Prediction for Planning in Autonomous Driving2022/11/1
- InterSim: Interactive Traffic Simulation via Explicit Relation Modeling2022/10/1
- ViP3D: End-to-end Visual Trajectory Prediction via 3D Agent Queries2022/8/1
- VectorFlow: Combining Images and Vectors for Traffic Occupancy and Flow Prediction2022/8/1
- VectorMapNet: End-to-end Vectorized HD Map Learning2022/6/1
- M2I: From Factored Marginal Trajectory Prediction to Interactive Prediction2022/2/1
- IFR-Explore: Learning Inter-object Functional Relationships in 3D Indoor Scenes2021/12/1
- DETR3D: 3D Object Detection from Multi-view Images via 3D-to-2D Queries2021/10/1
- Learning Practically Feasible Policies for Online 3D Bin Packing2021/8/1
- DenseTNT: End-to-end Trajectory Prediction from Dense Goal Sets2021/8/1
- DenseTNT: Waymo Open Dataset Motion Prediction Challenge 1st Place Solution2021/6/1
- Large Scale Interactive Motion Forecasting for Autonomous Driving : The Waymo Open Motion Dataset2021/4/1
- Predictive Visual Tracking: A New Benchmark and Baseline Approach2021/3/1
- Unsupervised Monocular Depth Learning in Dynamic Scenes2020/10/1
- CLOUD: Contrastive Learning of Unsupervised Dynamics2020/10/1
- SEMI: Self-supervised Exploration via Multisensory Incongruity2020/9/1
- TNT: Target-driveN Trajectory Prediction2020/8/1