Jiangmiao Pang
Shanghai Artificial Intelligence Laboratory
収録論文 100本 ・ フィジカルAI/ロボット学習
強化学習/ヒューマノイドシミュレーション評価
※arXiv著者名で収集。同姓同名の別人の論文が含まれる場合があります。
論文
- RoboStriker: 自律型ヒューマノイドボクシングのための潜在空間戦略ゲーム強化学習/ヒューマノイド2026/8/17
ヒューマノイドボクシングを2プレイヤーの潜在空間ゼロサムマルコフゲームとして定式化し、戦略的探索と物理的実現性の矛盾を解決する階層的フレームワークを提案した。
- GAUGE: 物理的忠実性を測定するための実世界基盤ベンチマークシミュレーション評価2026/8/6
シミュレーションエンジンと生成ビデオワールドモデルの物理的忠実性を、実世界の軌跡に基づいて診断するベンチマークを提案した。
- GAUGE: A Measurement-Grounded Benchmark for Physical Fidelity in Simulation Engines and Video World Models2026/8/1
- KAI: A Kinematic-Aware Interface for Data-Efficient Articulated Object Manipulation2026/7/1
- RoboInter1.5: A Holistic Intermediate Representation Suite for Embodied World Modeling and Robotic Manipulation2026/7/1
- InternVLA-A1.5: Unifying Understanding, Latent Foresight, and Action for Compositional Generalization2026/7/1
- X-NavDP: Generalizing Navigation Diffusion Policy to Novel Behavior and Embodiments with Group Q-score Reweighted Matching2026/7/1
- Exploratory, Communicative, and Deployable: Vision-Driven Embodied Agents for Open-World Mobile Manipulation2026/7/1
- Scaling Behavior Foundation Model for Humanoid Robots2026/7/1
- Cortex: A Bidirectionally Aligned Embodied Agent Framework for Long-horizon Manipulation2026/7/1
- EBench: Elemental Diagnosis of Generalist Mobile Manipulation Policies2026/6/1
- MemoryWAM: Efficient World Action Modeling with Persistent Memory2026/6/1
- ReactiveBFM: Reactive Closed-Loop Motion Planning Towards Universal Humanoid Whole-Body Control2026/6/1
- Imagine2Real: Towards Zero-shot Humanoid-Object Interaction via Video Generative Priors2026/5/1
- STABLE: Simulation-Ready Tabletop Layout Generation via a Semantics-Physics Dual System2026/5/1
- HiVLA: A Visual-Grounded-Centric Hierarchical Embodied Manipulation System2026/4/1
- SIM1: Physics-Aligned Simulator as Zero-Shot Data Scaler in Deformable Worlds2026/4/1
- ForceVLA2: Unleashing Hybrid Force-Position Control with Force Awareness for Contact-Rich Manipulation2026/3/1
- FutureVLA: Joint Visuomotor Prediction for Vision-Language-Action Model2026/3/1
- Towards Human-Like Manipulation through RL-Augmented Teleoperation and Mixture-of-Dexterous-Experts VLA2026/3/1
- UltraDexGrasp: Learning Universal Dexterous Grasping for Bimanual Robots with Synthetic Data2026/3/1
- Tac2Real: Reliable and GPU Visuotactile Simulation for Online Reinforcement Learning and Zero-Shot Real-World Deployment2026/3/1
- Feel Robot Feels: Tactile Feedback Array Glove for Dexterous Manipulation2026/3/1
- One-Policy-Fits-All: Geometry-Aware Action Latents for Cross-Embodiment Manipulation2026/3/1
- Scalable and General Whole-Body Control for Cross-Humanoid Locomotion2026/2/1
- Demystifying Action Space Design for Robotic Manipulation Policies2026/2/1
- RoboInter: A Holistic Intermediate Representation Suite Towards Robotic Manipulation2026/2/1
- DexImit: Learning Bimanual Dexterous Manipulation from Monocular Human Videos2026/2/1
- SoMA: A Real-to-Sim Neural Simulator for Robotic Soft-body Manipulation2026/2/1
- ST4VLA: Spatially Guided Training for Vision-Language-Action Models2026/2/1
- Embodiment-Aware Generalist Specialist Distillation for Unified Humanoid Whole-Body Control2026/2/1
- HumanX: Toward Agile and Generalizable Humanoid Interaction Skills from Human Videos2026/2/1
- Robo3R: Enhancing Robotic Manipulation with Accurate Feed-Forward 3D Reconstruction2026/2/1
- RoboStriker: Hierarchical Decision-Making for Autonomous Humanoid Boxing2026/1/1
- InternVLA-A1: Unifying Understanding, Generation and Action for Robotic Manipulation2026/1/1
- Nimbus: A Unified Embodied Synthetic Data Generation Framework2026/1/1
- UniCon: A Unified System for Efficient Robot Learning Transfers2026/1/1
- RoboVIP: Multi-View Video Generation with Visual Identity Prompting Augments Robot Manipulation2026/1/1
- VLNVerse: A Benchmark for Vision-Language Navigation with Versatile, Embodied, Realistic Simulation and Evaluation2025/12/1
- LoGoPlanner: Localization Grounded Navigation Policy with Metric-aware Visual Geometry2025/12/1
- MM-ACT: Learn from Multimodal Parallel Generation to Act2025/12/1
- VL-LN Bench: Towards Long-horizon Goal-oriented Navigation with Active Dialogs2025/12/1
- H-Zero: Cross-Humanoid Locomotion Pretraining Enables Few-shot Novel Embodiment Transfer2025/12/1
- Ground Slow, Move Fast: A Dual-System Foundation Model for Generalizable Vision-and-Language Navigation2025/12/1
- MM-ACT: Learn from Multimodal Parallel Generation to Act2025/11/30
- Gallant: Voxel Grid-based Humanoid Locomotion and Local-navigation across 3D Constrained Terrains2025/11/1
- InternData-A1: Pioneering High-Fidelity Synthetic Data for Pre-training Generalist Policy2025/11/1
- Towards Adaptable Humanoid Control via Adaptive Motion Tracking2025/10/1
- X-VLA: Soft-Prompted Transformer as Scalable Cross-Embodiment Vision-Language-Action Model2025/10/1
- InternVLA-M1: A Spatially Guided Vision-Language-Action Framework for Generalist Robot Policy2025/10/1
- Expertise need not monopolize: Action-Specialized Mixture of Experts for Vision-Language-Action Learning2025/10/1
- PhysHSI: Towards a Real-World Generalizable and Natural Humanoid-Scene Interaction System2025/10/1
- Humanoid Goalkeeper: Learning from Position Conditioned Task-Motion Constraints2025/10/1
- SimpleVLA-RL: Scaling VLA Training via Reinforcement Learning2025/9/11
- InternScenes: A Large-scale Simulatable Indoor Scene Dataset with Realistic Layouts2025/9/1
- MesaTask: Towards Task-Driven Tabletop Scene Generation via 3D Spatial Reasoning2025/9/1
- A Vision-Language-Action-Critic Model for Robotic Real-World Reinforcement Learning2025/9/1
- Behavior Foundation Model for Humanoid Robots2025/9/1
- SimpleVLA-RL: Scaling VLA Training via Reinforcement Learning2025/9/1
- F1: A Vision-Language-Action Model Bridging Understanding and Generation to Actions2025/9/1
- Discrete Diffusion VLA: Bringing Discrete Diffusion to Action Decoding in Vision-Language-Action Policies2025/8/27
- Discrete Diffusion VLA: Bringing Discrete Diffusion to Action Decoding in Vision-Language-Action Policies2025/8/1
- Rethinking the Embodied Gap in Vision-and-Language Navigation: A Holistic Study of Physical and Visual Disparities2025/7/1
- UniTracker: Learning Universal Whole-Body Motion Tracker for Humanoid Robots2025/7/1
- InstructVLA: Vision-Language-Action Instruction Tuning from Understanding to Manipulation2025/7/1
- StreamVLN: Streaming Vision-and-Language Navigation via SlowFast Context Modeling2025/7/1
- GENMANIP: LLM-driven Simulation for Generalizable Instruction-Following Manipulation2025/6/1
- CronusVLA: Towards Efficient and Robust Manipulation via Multi-Frame Vision-Language-Action Modeling2025/6/1
- LabUtopia: High-Fidelity Simulation and Hierarchical Benchmark for Scientific Embodied Agents2025/5/1
- GLEAM: Learning Generalizable Exploration Policy for Active Mapping in Complex 3D Indoor Scenes2025/5/1
- TeleOpBench: A Simulator-Centric Benchmark for Dual-Arm Dexterous Teleoperation2025/5/1
- NavDP: Learning Sim-to-Real Navigation Diffusion Policy with Privileged Information Guidance2025/5/1
- Towards Latency-Aware 3D Streaming Perception for Autonomous Driving2025/4/1
- Novel Demonstration Generation with Gaussian Splatting Enables Robust One-Shot Manipulation2025/4/1
- RoboGround: Robotic Manipulation with Grounded Vision-Language Priors2025/4/1
- Gripper Keypose and Object Pointflow as Interfaces for Bimanual Robotic Manipulation2025/4/1
- Aether: Geometric-Aware Unified World Modeling2025/3/1
- A Data-Centric Revisit of Pre-Trained Vision Models for Robot Learning2025/3/1
- AgiBot World Colosseo: A Large-scale Manipulation Platform for Scalable and Intelligent Embodied Systems2025/3/1
- HOMIE: Humanoid Loco-Manipulation with Isomorphic Exoskeleton Cockpit2025/2/1
- Re$^3$Sim: Generating High-Fidelity Simulation Data via 3D-Photorealistic Real-to-Sim for Robotic Manipulation2025/2/1
- VB-Com: Learning Vision-Blind Composite Humanoid Locomotion Against Deficient Perception2025/2/1
- BeamDojo: Learning Agile Humanoid Locomotion on Sparse Footholds2025/2/1
- Learning Humanoid Standing-up Control across Diverse Postures2025/2/1
- A Unified and General Humanoid Whole-Body Controller for Versatile Locomotion2025/2/1
- Predictive Inverse Dynamics Models are Scalable Learners for Robotic Manipulation2024/12/1
- Learning Humanoid Locomotion with Perceptive Internal Model2024/11/1
- VLM-Grounder: A VLM Agent for Zero-Shot 3D Visual Grounding2024/10/1
- Open-Vocabulary Object-Goal Navigation by Generalizing Semantic Mapping with Dense CLIP2024/7/1
- GRUtopia: Dream General Robots in a City at Scale2024/7/1
- CooHOI: Learning Cooperative Human-Object Interaction with Manipulated Object Dynamics2024/6/1
- MMScan: A Multi-Modal 3D Scene Dataset with Hierarchical Grounded Language Annotations2024/6/1
- Learning H-Infinity Locomotion Control2024/4/1
- RoboKeyGen: Robot Pose and Joint Angles Estimation via Diffusion-based 3D Keypoint Generation2024/3/1
- RoboDuet: Learning a Cooperative Policy for Whole-body Legged Loco-Manipulation2024/3/1
- GenNBV: Generalizable Next-Best-View Policy for Active 3D Reconstruction2024/2/1
- EmbodiedScan: A Holistic Multi-Modal 3D Perception Suite Towards Embodied AI2023/12/1
- Hybrid Internal Model: Learning Agile Legged Locomotion with Simulated Robot Response2023/12/1
- Monocular 3D Object Detection with Depth from Motion2022/7/1
- FCOS3D: Fully Convolutional One-Stage Monocular 3D Object Detection2021/4/1