Tai Wang
収録論文 39本 ・ フィジカルAI/ロボット学習
※arXiv著者名で収集。同姓同名の別人の論文が含まれる場合があります。
論文
- InternVLA-A1.5: Unifying Understanding, Latent Foresight, and Action for Compositional Generalization2026/7/1
- X-NavDP: Generalizing Navigation Diffusion Policy to Novel Behavior and Embodiments with Group Q-score Reweighted Matching2026/7/1
- Exploratory, Communicative, and Deployable: Vision-Driven Embodied Agents for Open-World Mobile Manipulation2026/7/1
- Cortex: A Bidirectionally Aligned Embodied Agent Framework for Long-horizon Manipulation2026/7/1
- EBench: Elemental Diagnosis of Generalist Mobile Manipulation Policies2026/6/1
- ReactiveBFM: Reactive Closed-Loop Motion Planning Towards Universal Humanoid Whole-Body Control2026/6/1
- RoboInter: A Holistic Intermediate Representation Suite Towards Robotic Manipulation2026/2/1
- Demystifying Action Space Design for Robotic Manipulation Policies2026/2/1
- Nimbus: A Unified Embodied Synthetic Data Generation Framework2026/1/1
- InternVLA-A1: Unifying Understanding, Generation and Action for Robotic Manipulation2026/1/1
- LoGoPlanner: Localization Grounded Navigation Policy with Metric-aware Visual Geometry2025/12/1
- VL-LN Bench: Towards Long-horizon Goal-oriented Navigation with Active Dialogs2025/12/1
- Ground Slow, Move Fast: A Dual-System Foundation Model for Generalizable Vision-and-Language Navigation2025/12/1
- InternVLA-M1: A Spatially Guided Vision-Language-Action Framework for Generalist Robot Policy2025/10/1
- X-VLA: Soft-Prompted Transformer as Scalable Cross-Embodiment Vision-Language-Action Model2025/10/1
- InternScenes: A Large-scale Simulatable Indoor Scene Dataset with Realistic Layouts2025/9/1
- InstructVLA: Vision-Language-Action Instruction Tuning from Understanding to Manipulation2025/7/1
- StreamVLN: Streaming Vision-and-Language Navigation via SlowFast Context Modeling2025/7/1
- Rethinking the Embodied Gap in Vision-and-Language Navigation: A Holistic Study of Physical and Visual Disparities2025/7/1
- GENMANIP: LLM-driven Simulation for Generalizable Instruction-Following Manipulation2025/6/1
- CronusVLA: Towards Efficient and Robust Manipulation via Multi-Frame Vision-Language-Action Modeling2025/6/1
- GLEAM: Learning Generalizable Exploration Policy for Active Mapping in Complex 3D Indoor Scenes2025/5/1
- NavDP: Learning Sim-to-Real Navigation Diffusion Policy with Privileged Information Guidance2025/5/1
- LabUtopia: High-Fidelity Simulation and Hierarchical Benchmark for Scientific Embodied Agents2025/5/1
- Towards Latency-Aware 3D Streaming Perception for Autonomous Driving2025/4/1
- RoboGround: Robotic Manipulation with Grounded Vision-Language Priors2025/4/1
- VLM-Grounder: A VLM Agent for Zero-Shot 3D Visual Grounding2024/10/1
- Open-Vocabulary Object-Goal Navigation by Generalizing Semantic Mapping with Dense CLIP2024/7/1
- GRUtopia: Dream General Robots in a City at Scale2024/7/1
- MMScan: A Multi-Modal 3D Scene Dataset with Hierarchical Grounded Language Annotations2024/6/1
- CooHOI: Learning Cooperative Human-Object Interaction with Manipulated Object Dynamics2024/6/1
- An Empirical Study of Training State-of-the-Art LiDAR Segmentation Models2024/5/1
- GenNBV: Generalizable Next-Best-View Policy for Active 3D Reconstruction2024/2/1
- EmbodiedScan: A Holistic Multi-Modal 3D Perception Suite Towards Embodied AI2023/12/1
- Scene as Occupancy2023/6/1
- Monocular 3D Object Detection with Depth from Motion2022/7/1
- MV-FCOS3D++: Multi-View Camera-Only 4D Object Detection with Pretrained Monocular Backbones2022/7/1
- FCOS3D: Fully Convolutional One-Stage Monocular 3D Object Detection2021/4/1
- FLAVA: Find, Localize, Adjust and Verify to Annotate LiDAR-Based Point Clouds2020/11/1