Peiyan Li
Institute of Automation, Chinese Academy of Sciences
収録論文 24本 ・ フィジカルAI/ロボット学習
ワールドモデルVLA/操作
※arXiv著者名で収集。同姓同名の別人の論文が含まれる場合があります。
論文
- XEWorld:行動条件付きワールドモデルは未知のロボット形態に汎化できるか?ワールドモデル2026/8/6
ロボット操作のための行動条件付きワールドモデルが、未見のロボット形態に対して物理ダイナミクスを正しく予測できるかを検証するため、クロスエンボディメントテストベッドXEWorldを導入し、既存モデルの限界を分析した。
- BridgeVLA++: データ効率的で汎化性が高く、メモリ拡張された3D操作のための視覚-言語-行動フレームワークVLA/操作2026/8/5
事前学習済み視覚言語モデルを活用した3Dロボット操作フレームワークBridgeVLAを拡張し、空間的・時間的メモリを統合することで、データ効率と汎化性を保ちつつ、記憶依存の操作タスクで最先端の性能を達成した。
- XEWorld: Can Action-Conditioned World Models Generalize to Unseen Robot Embodiments?2026/8/1
- BridgeVLA++: A Data-Efficient, Generalizable, and Memory-Augmented Vision-Language-Action Framework for 3D Manipulation2026/8/1
- FlowWAM: Optical Flow as a Unified Action Representation for World Action Models2026/7/1
- Xiaomi-Robotics-1: Scaling Vision-Language-Action Models with over 100K Hours of Real-World Trajectories2026/7/1
- Revisiting Embodied Chain-of-Thought for Generalizable Robot Manipulation2026/6/1
- Improving Vision-Language-Action Model Fine-Tuning with Structured Stage and Keyframe Supervision2026/6/1
- E-TTS: A New Embodied Test-Time Scaling Framework for Robotic Manipulation2026/6/1
- SKIP: Sparse Keyframe Interpolation Paradigm for Efficient Embodied World Models2026/6/1
- RotVLA: Rotational Latent Action for Vision-Language-Action Model2026/5/1
- SpatialVAM:Spatial-Aware Multi-View Video Diffusion as a Data-Efficient Robot Policy2026/4/1
- Unified 4D World Action Modeling from Video Priors with Asynchronous Denoising2026/4/1
- Scaling World Model for Hierarchical Manipulation Policies2026/2/1
- UAOR: Uncertainty-aware Observation Reinjection for Vision-Language-Action Models2026/2/1
- BridgeV2W: Bridging Video Generation Models to Embodied World Models via Embodiment Masks2026/2/1
- VERM: Leveraging Foundation Models to Create a Virtual Eye for Efficient 3D Robotic Manipulation2025/12/1
- EgoDemoGen: Egocentric Demonstration Generation for Viewpoint Generalization in Robotic Manipulation2025/9/1
- EC-Flow: Enabling Versatile Robotic Manipulation from Action-Unlabeled Videos via Embodiment-Centric Flow2025/7/1
- BridgeVLA: Input-Output Alignment for Efficient 3D Manipulation Learning with Vision-Language Models2025/6/1
- What Matters in Building Vision-Language-Action Models for Generalist Robots2024/12/1
- GR-MG: Leveraging Partially Annotated Data via Multi-Modal Goal-Conditioned Policy2024/8/1
- Leveraging Large Language Model for Heterogeneous Ad Hoc Teamwork Collaboration2024/6/1
- Demonstrating HumanTHOR: A Simulation Platform and Benchmark for Human-Robot Collaboration in a Shared Workspace2024/6/1