Yan Huang
Institute of Automation, Chinese Academy of Sciences
収録論文 35本 ・ フィジカルAI/ロボット学習
ワールドモデルVLA/操作
※arXiv著者名で収集。同姓同名の別人の論文が含まれる場合があります。
論文
- XEWorld:行動条件付きワールドモデルは未知のロボット形態に汎化できるか?ワールドモデル2026/8/6
ロボット操作のための行動条件付きワールドモデルが、未見のロボット形態に対して物理ダイナミクスを正しく予測できるかを検証するため、クロスエンボディメントテストベッドXEWorldを導入し、既存モデルの限界を分析した。
- BridgeVLA++: データ効率的で汎化性が高く、メモリ拡張された3D操作のための視覚-言語-行動フレームワークVLA/操作2026/8/5
事前学習済み視覚言語モデルを活用した3Dロボット操作フレームワークBridgeVLAを拡張し、空間的・時間的メモリを統合することで、データ効率と汎化性を保ちつつ、記憶依存の操作タスクで最先端の性能を達成した。
- XEWorld: Can Action-Conditioned World Models Generalize to Unseen Robot Embodiments?2026/8/1
- BridgeVLA++: A Data-Efficient, Generalizable, and Memory-Augmented Vision-Language-Action Framework for 3D Manipulation2026/8/1
- DA-Nav: Direction-Aware City-Scale Vision-Language Navigation2026/7/1
- Towards Trustworthy Embodied Intelligence: A Systems Framework and Graded Trustworthiness Levels2026/7/1
- FlowWAM: Optical Flow as a Unified Action Representation for World Action Models2026/7/1
- SKIP: Sparse Keyframe Interpolation Paradigm for Efficient Embodied World Models2026/6/1
- DIM-WAM: World-Action Modeling with Diverse Historical Event Memory2026/6/1
- Improving Vision-Language-Action Model Fine-Tuning with Structured Stage and Keyframe Supervision2026/6/1
- E-TTS: A New Embodied Test-Time Scaling Framework for Robotic Manipulation2026/6/1
- Decentralized Pose Graph Riemannian Optimization for Object-based Multi-Robot SLAM2026/6/1
- WAM-Nav: Asymmetric Latent World-Action Modeling for Unified Visual Navigation2026/6/1
- Efficient-WAM: A 1B-Parameter World-Action Model with Low-Cost Future Imagination2026/6/1
- SpatialVAM:Spatial-Aware Multi-View Video Diffusion as a Data-Efficient Robot Policy2026/4/1
- FloorPlan-VLN: A New Paradigm for Floor Plan Guided Vision-Language Navigation2026/3/1
- BridgeV2W: Bridging Video Generation Models to Embodied World Models via Embodiment Masks2026/2/1
- UAOR: Uncertainty-aware Observation Reinjection for Vision-Language-Action Models2026/2/1
- UniBYD: A Unified Framework for Learning Robotic Manipulation Across Embodiments Beyond Imitation of Human Demonstrations2025/12/1
- VERM: Leveraging Foundation Models to Create a Virtual Eye for Efficient 3D Robotic Manipulation2025/12/1
- VP-AutoTest: A Virtual-Physical Fusion Autonomous Driving Testing Platform2025/12/1
- UMIGen: A Unified Framework for Egocentric Point Cloud Generation and Cross-Embodiment Robotic Imitation Learning2025/11/1
- EgoDemoGen: Egocentric Demonstration Generation for Viewpoint Generalization in Robotic Manipulation2025/9/1
- CoCoL: A Communication Efficient Decentralized Collaborative Method for Multi-Robot Systems2025/8/1
- UltraTac: Integrated Ultrasound-Augmented Visuotactile Sensor for Enhanced Robotic Perception2025/8/1
- EC-Flow: Enabling Versatile Robotic Manipulation from Action-Unlabeled Videos via Embodiment-Centric Flow2025/7/1
- BridgeVLA: Input-Output Alignment for Efficient 3D Manipulation Learning with Vision-Language Models2025/6/1
- Constraint-Aware Zero-Shot Vision-Language Navigation in Continuous Environments2024/12/1
- HGSFusion: Radar-Camera Fusion with Hybrid Generation and Synchronization for 3D Object Detection2024/12/1
- GR-MG: Leveraging Partially Annotated Data via Multi-Modal Goal-Conditioned Policy2024/8/1
- Chemistry3D: Robotic Interaction Benchmark for Chemistry Experiments2024/6/1
- Dual-modal Tactile E-skin: Enabling Bidirectional Human-Robot Interaction via Integrated Tactile Perception and Feedback2024/2/1
- ETPNav: Evolving Topological Planning for Vision-Language Navigation in Continuous Environments2023/4/1
- BEVBert: Multimodal Map Pre-training for Language-guided Navigation2022/12/1
- 1st Place Solutions for RxR-Habitat Vision-and-Language Navigation Competition (CVPR 2022)2022/6/1