Liang Wang
Institute of Automation, Chinese Academy of Sciences
収録論文 33本 ・ フィジカルAI/ロボット学習
ワールドモデルVLA/操作
※arXiv著者名で収集。同姓同名の別人の論文が含まれる場合があります。
論文
- XEWorld:行動条件付きワールドモデルは未知のロボット形態に汎化できるか?ワールドモデル2026/8/6
ロボット操作のための行動条件付きワールドモデルが、未見のロボット形態に対して物理ダイナミクスを正しく予測できるかを検証するため、クロスエンボディメントテストベッドXEWorldを導入し、既存モデルの限界を分析した。
- BridgeVLA++: データ効率的で汎化性が高く、メモリ拡張された3D操作のための視覚-言語-行動フレームワークVLA/操作2026/8/5
事前学習済み視覚言語モデルを活用した3Dロボット操作フレームワークBridgeVLAを拡張し、空間的・時間的メモリを統合することで、データ効率と汎化性を保ちつつ、記憶依存の操作タスクで最先端の性能を達成した。
- SyncPlan: Long-Horizon LLM Coordination with Explicit Synchronization and Adaptive Correction2026/8/1
- BridgeVLA++: A Data-Efficient, Generalizable, and Memory-Augmented Vision-Language-Action Framework for 3D Manipulation2026/8/1
- XEWorld: Can Action-Conditioned World Models Generalize to Unseen Robot Embodiments?2026/8/1
- Stop to Decide: Latency-Aware Proprioceptive Navigation Primitives for Mapping-Free Quadruped Inspection2026/7/1
- FlowWAM: Optical Flow as a Unified Action Representation for World Action Models2026/7/1
- E-TTS: A New Embodied Test-Time Scaling Framework for Robotic Manipulation2026/6/1
- DIM-WAM: World-Action Modeling with Diverse Historical Event Memory2026/6/1
- Improving Vision-Language-Action Model Fine-Tuning with Structured Stage and Keyframe Supervision2026/6/1
- GUIDE: Goal-Initialized Directional Understanding for End-to-End Visual Navigation2026/6/1
- SpatialVAM:Spatial-Aware Multi-View Video Diffusion as a Data-Efficient Robot Policy2026/4/1
- FloorPlan-VLN: A New Paradigm for Floor Plan Guided Vision-Language Navigation2026/3/1
- UAOR: Uncertainty-aware Observation Reinjection for Vision-Language-Action Models2026/2/1
- BridgeV2W: Bridging Video Generation Models to Embodied World Models via Embodiment Masks2026/2/1
- PUMA: Perception-driven Unified Foothold Prior for Mobility Augmented Quadruped Parkour2026/1/1
- Embodied Co-Design for Rapidly Evolving Agents: Taxonomy, Frontiers, and Challenges2025/12/1
- RflyUT-Sim: A Simulation Platform for Development and Testing of Complex Low-Altitude Traffic Control2025/12/1
- VERM: Leveraging Foundation Models to Create a Virtual Eye for Efficient 3D Robotic Manipulation2025/12/1
- SimScale: Learning to Drive via Real-World Simulation at Scale2025/11/1
- ZJUNlict Extended Team Description Paper 20252025/11/1
- EgoDemoGen: Egocentric Demonstration Generation for Viewpoint Generalization in Robotic Manipulation2025/9/1
- EC-Flow: Enabling Versatile Robotic Manipulation from Action-Unlabeled Videos via Embodiment-Centric Flow2025/7/1
- BridgeVLA: Input-Output Alignment for Efficient 3D Manipulation Learning with Vision-Language Models2025/6/1
- Orchestrating Joint Offloading and Scheduling for Low-Latency Edge SLAM2025/2/1
- MCRL4OR: Multimodal Contrastive Representation Learning for Off-Road Environmental Perception2025/1/1
- Constraint-Aware Zero-Shot Vision-Language Navigation in Continuous Environments2024/12/1
- TopoSD: Topology-Enhanced Lane Segment Perception with SDMap Prior2024/11/1
- GR-MG: Leveraging Partially Annotated Data via Multi-Modal Goal-Conditioned Policy2024/8/1
- Pose-Graph Attentional Graph Neural Network for Lidar Place Recognition2023/9/1
- ETPNav: Evolving Topological Planning for Vision-Language Navigation in Continuous Environments2023/4/1
- BEVBert: Multimodal Map Pre-training for Language-guided Navigation2022/12/1
- 1st Place Solutions for RxR-Habitat Vision-and-Language Navigation Competition (CVPR 2022)2022/6/1