Yang Liu
Harbin Institute of Technology
収録論文 85本 ・ フィジカルAI/ロボット学習
VLA/ヒューマノイドタスク計画
※arXiv著者名で収集。同姓同名の別人の論文が含まれる場合があります。
論文
- EATR-Stereo: 身体性を考慮したステレオ証拠のルーティングによるヒューマノイド視覚言語行動制御VLA/ヒューマノイド2026/8/18
頭部搭載ステレオカメラを持つヒューマノイドのVLA制御において、主視点のトークンを保持しつつ補助視点の情報を身体状態に応じて選択的に統合するフレームワークを提案し、実機で高い成功率を達成した。
- GraphThink: グラフ強化LLM思考による長期的身体性タスク計画タスク計画2026/8/8
LLMベースの身体性エージェントの計画における幻覚や長期タスクへの汎化不足を解決するため、タスクグラフとシーングラフを統合したフレームワークを提案。ALFREDベンチマークで最先端性能を達成した。
- GraphThink: Graph-Enhanced LLM Thinking for Long-Horizon Embodied Task Planning2026/8/1
- Bridge-WA: Predicting Where and How the World Changes for Robotic Action2026/7/1
- PhyAgentOS: A Self-Evolving Operating System for Embodied Agents with Decoupled Cognitive Planning and Physical Execution2026/7/1
- HCPG-Flow:Hierarchical Contact-Progress Guidance for Flow-Policy Robot Manipulation2026/7/1
- Understanding and Mitigating the Video-Action Generalization Gap via Temporal Ratio2026/7/1
- Event-Conditioned Diagnostics of Kinematic, Contact, and Object-Permanence Fields in Passive Object-State World Models2026/6/1
- RoboProcessBench: Benchmarking Process-Aware Understanding in Vision-Language Robotic Manipulation2026/6/1
- VLAMotor: Test-Guided Enhancement of Vision-Language-Action Models via Agent-BasedData Synthesis2026/6/1
- ARP: Enhancing Quantized Skill Abstractions via Visual Alignment and Iterative Refinement for Robotic Manipulation2026/6/1
- In-Context Model Predictive Generation: Open-Vocabulary Motion Synthesis from Language Models to Physics2026/6/1
- Seeing Touch from Motion: A Unified Modality-Aware Visuo-Tactile Policy with Tactile Motion Correlation2026/6/1
- STELLAR: Scaling 3D Perception Large Models for Autonomous Driving2026/5/1
- RoVLA: Multi-Consistency Constraints for Robust Vision-Language-Action Models2026/5/1
- Uncovering Linguistic Fragility in Vision-Language-Action Models via Diversity-Aware Red Teaming2026/4/1
- BrainMem: Brain-Inspired Evolving Memory for Embodied Agent Task Planning2026/4/1
- AcceRL: A Distributed Asynchronous Reinforcement Learning and World Model Framework for Vision-Language-Action Models2026/3/19
- ACCURATE: Arbitrary-shaped Continuum Reconstruction Under Robust Adaptive Two-view Estimation2026/3/1
- ReMemNav: A Rethinking and Memory-Augmented Framework for Zero-Shot Object Navigation2026/3/1
- TacMamba: A Tactile History Compression Adapter Bridging Fast Reflexes and Slow VLA Reasoning2026/3/1
- MMaDA-VLA: Large Diffusion Vision-Language-Action Model with Unified Multi-Modal Instruction and Generation2026/3/1
- OccTrack360: 4D Panoptic Occupancy Tracking from Surround-View Fisheye Cameras2026/3/1
- Now You See That: Learning End-to-End Humanoid Locomotion from Raw Pixels2026/2/1
- DDP-WM: Disentangled Dynamics Prediction for Efficient World Models2026/2/1
- Eva-Tracker: ESDF-update-free, Visibility-aware Planning with Target Reacquisition for Robust Aerial Tracking2026/2/1
- FRAPPE: Infusing World Modeling into Generalist Policies via Multiple Future Representation Alignment2026/2/1
- PUMA: Perception-driven Unified Foothold Prior for Mobility Augmented Quadruped Parkour2026/1/1
- Preparation and Motion Study of Magnetically Driven Micro Soft Robot Mimicking the Cownose Ray2026/1/1
- Detecting Non-Optimal Decisions of Embodied Agents via Diversity-Guided Metamorphic Testing2025/12/1
- HiF-VLA: Hindsight, Insight and Foresight through Motion Representation for Vision-Language-Action Models2025/12/1
- An Anatomy of Vision-Language-Action Models: From Modules to Milestones and Challenges2025/12/1
- Phase-Based Multi-Gait Learning for a Salamander-Like Robot2025/11/1
- HDCNet: A Hybrid Depth Completion Network for Grasping Transparent and Reflective Objects2025/11/1
- Beyond Description: Cognitively Benchmarking Fine-Grained Action for Embodied Agents2025/11/1
- RoVer: Robot Reward Model as Test-Time Verifier for Vision-Language-Action Model2025/10/1
- Learning to Throw-Flip2025/10/1
- Bridging Perception and Planning: Towards End-to-End Planning for Signal Temporal Logic Tasks2025/9/1
- Mars Traversability Prediction: A Multi-modal Self-supervised Approach for Costmap Generation2025/9/1
- Learning to See and Act: Task-Aware Virtual View Exploration for Robotic Manipulation2025/8/1
- Long-VLA: Unleashing Long-Horizon Capability of Vision Language Action Model for Robot Manipulation2025/8/1
- AR-VRM: Imitating Human Motions for Visual Robot Manipulation with Analogical Reasoning2025/8/1
- ULC: A Unified and Fine-Grained Controller for Humanoid Loco-Manipulation2025/7/1
- Vibration-aware Lidar-Inertial Odometry based on Point-wise Post-Undistortion Uncertainty2025/7/1
- GFM-Planner: Perception-Aware Trajectory Planning with Geometric Feature Metric2025/7/1
- AutoLayout: Closed-Loop Layout Synthesis via Slow-Fast Collaborative Reasoning2025/7/1
- Learning Accurate Whole-body Throwing with High-frequency Residual Policy and Pullback Tube Acceleration2025/6/1
- DCIRNet: Depth Completion with Iterative Refinement for Dexterous Grasping of Transparent and Reflective Objects2025/6/1
- Overcoming Environmental Meta-Stationarity in MARL via Adaptive Curriculum and Counterfactual Group Advantage2025/6/1
- OpenHelix: A Short Survey, Empirical Analysis, and Open-Source Dual-System VLA Model for Robotic Manipulation2025/5/1
- Towards Robust Multi-UAV Collaboration: MARL with Noise-Resilient Communication and Attention Mechanisms2025/3/1
- Learning Perceptive Humanoid Locomotion over Challenging Terrain2025/3/1
- Learning Humanoid Locomotion with World Model Reconstruction2025/2/1
- Neural-Network-Driven Reward Prediction as a Heuristic: Advancing Q-Learning for Mobile Robot Path Planning2024/12/1
- InfiniteWorld: A Unified Scalable Simulation Framework for General Visual-Language Robot Interaction2024/12/1
- Depth-PC: A Visual Servo Framework Integrated with Cross-Modality Fusion for Sim2Real Transfer2024/11/1
- Aligning Cyber Space with Physical World: A Comprehensive Survey on Embodied AI2024/7/1
- SuperVINS: A Real-Time Visual-Inertial SLAM Framework for Challenging Imaging Conditions2024/7/1
- Human-centered In-building Embodied Delivery Benchmark2024/6/1
- MAS-SAM: Segment Any Marine Animal with Aggregated Features2024/4/1
- NEDS-SLAM: A Neural Explicit Dense Semantic SLAM Framework using 3D Gaussian Splatting2024/3/1
- Bringing Robots Home: The Rise of AI Robots in Consumer Electronics2024/3/1
- R$\times$R: Rapid eXploration for Reinforcement Learning via Sampling-based Reset Distributions and Imitation Pre-training2024/1/1
- GarchingSim: An Autonomous Driving Simulator with Photorealistic Scenes and Minimalist Workflow2024/1/1
- Mobile-Seed: Joint Semantic Segmentation and Boundary Detection for Mobile Robots2023/11/1
- Alexa, play with robot: Introducing the First Alexa Prize SimBot Challenge on Embodied AI2023/8/1
- NNPP: A Learning-Based Heuristic Model for Accelerating Optimal Path Planning on Uneven Terrain2023/8/1
- A reproducible approach to merging behavior analysis based on High Definition Map2023/3/1
- Sampling-based Exploration for Reinforcement Learning of Dexterous Manipulation2023/3/1
- FLYOVER: A Model-Driven Method to Generate Diverse Highway Interchanges for Autonomous Vehicle Testing2023/1/1
- The Design and Realization of Multi-agent Obstacle Avoidance based on Reinforcement Learning2022/10/1
- OA-Bug: An Olfactory-Auditory Augmented Bug Algorithm for Swarm Robots in a Denied Environment2022/9/1
- Research on Stable Obstacle Avoidance Control Strategy for Tracked Intelligent Transportation Vehicles in Non-structural Environment Based on Deep Learning2022/8/1
- A Solution to Adaptive Mobile Manipulator Throwing2022/7/1
- Design and experimental investigation of a vibro-impact self-propelled capsule robot with orientation control2022/2/1
- A Survey on Scenario-Based Testing for Automated Driving Systems in High-Fidelity Simulation2021/12/1
- The Unified Mathematical Framework for IMU Preintegration in Inertial-Aided Navigation System2021/11/1
- VISITRON: Visual Semantics-Aligned Interactively Trained Object-Navigator2021/5/1
- The Role of the Hercules Autonomous Vehicle During the COVID-19 Pandemic: An Autonomous Logistic Vehicle for Contactless Goods Transportation2020/4/1
- Predicting Unobserved Space For Planning via Depth Map Augmentation2019/11/1
- Federated Transfer Reinforcement Learning for Autonomous Driving2019/10/1
- DF-SLAM: A Deep-Learning Enhanced Visual SLAM System based on Deep Local Features2019/1/1
- Characterization of a RS-LiDAR for 3D Perception2017/9/1
- A Control Performance Index for Multicopters Under Off-nominal Conditions2017/5/1
- Motion Imitation Based on Sparsely Sampled Correspondence2016/7/1