Yan Wang
Tsinghua University
収録論文 54本 ・ フィジカルAI/ロボット学習
操作マニピュレーションVLA/世界モデルVLA
※arXiv著者名で収集。同姓同名の別人の論文が含まれる場合があります。
論文
- StageWAM: ロボット操作におけるワールドアクションモデルのためのジョイント埋め込みステージ予測操作2026/8/11
ロボット操作タスクにおいて、短期的な物理的未来に加えて、タスクの進行段階を表すセマンティックな未来を予測するStageWAMを提案し、成功率と実行効率を向上させた。
- JEPA-WAM: ロボット操作のためのワールドアクションモデルにおける段階レベル結合埋め込み予測マニピュレーション2026/8/11
ロボット操作タスクにおいて、短期的な物理的未来と段階的な意味的未来を区別し、段階レベルの潜在目標を予測するJEPA-WAMを提案。50のタスクで成功率90.25%を達成し、実行ステップ数を削減した。
- JEPA-WAM: 共同埋め込み世界モデリングによる視覚・言語・行動ポリシーの学習VLA/世界モデル2026/8/10
事前学習済みV-JEPA空間に基づく潜在世界モデルを導入し、遷移予測と行動生成を共有予測器で結合することで、ロボット制御の性能と汎化を向上させた。
- 4D-WAM: 軌跡フィールドによる世界行動モデルへの時空間認識の注入VLA2026/8/8
ロボットの行動生成と動画予測を統合する世界行動モデルに、3次元軌跡フィールドの時空間知識を表現整合で注入する訓練戦略を提案。局所的な動き整合と長期的な目的地整合の2つの目的関数により、軌跡レベルの時空間表現を学習し、空間理解や実行精度、汎化性を向上させる。
- 4D-WAM: 軌跡フィールドによる世界行動モデルへの時空間認識の注入VLA2026/8/8
ロボットの行動生成と動画予測を統合する世界行動モデルに、3次元軌跡フィールドの時空間知識を表現整合で注入する訓練戦略を提案し、空間理解と実行精度を向上させた。
- JEPA-WAM: Learning Vision-Language-Action Policies with Joint-Embedding World Modeling2026/8/1
- 4D-WAM: Infusing Spatiotemporal Awareness into World Action Models through Trajectory Fields2026/8/1
- JEPA-WAM: Stage-Level Joint-Embedding Prediction for World-Action Models in Robot Manipulation2026/8/1
- Adaptive-WAM: Quality-Guided Early-Exit Planning from Intermediate Video-Diffusion Features2026/8/1
- DINS-IO: Learned Inertial Odometry via Differentiable INS Consistency2026/7/1
- ST-WAM: Semantic-Temporal World Action Model for Robust Manipulation under Visual Distribution Shifts2026/7/1
- ASSCG: Just-Right Gating over Chattering for Fast-Slow LLM Planning in Autonomous Driving2026/6/1
- MV-Actor: Aligning Multi-View Semantics and Spatial Awareness for Bimanual Manipulation2026/6/1
- Planning-aligned Token Compression for Long-Context Autonomous Driving2026/6/1
- A Physics-Grounded Benchmark for Multi-Agent Dynamics in World Models2026/6/1
- Cosmos 3: Omnimodal World Models for Physical AI2026/6/1
- AR Forcing: Towards Long-Horizon Robot Navigation World Model2026/5/1
- CLOVER: Closed-Loop Value Estimation and Ranking for End-to-End Autonomous Driving Planning2026/5/1
- Reactive Planning based Control for Mobile Robots in Obstacle-Cluttered Environments2026/5/1
- RoboMemArena: A Comprehensive and Challenging Robotic Memory Benchmark2026/5/1
- CapVector: Learning Transferable Capability Vectors in Parametric Space for Vision-Language-Action Models2026/5/1
- Fast-dVLA: Accelerating Discrete Diffusion VLA to Real-Time Performance2026/3/1
- Devil is in Narrow Policy: Unleashing Exploration in Driving VLA Models2026/3/1
- DFM-VLA: Iterative Action Refinement for Robot Manipulation via Discrete Flow Matching2026/3/1
- Accelerating Structured Chain-of-Thought in Autonomous Vehicles2026/2/1
- From Representational Complementarity to Dual Systems: Synergizing VLM and Vision-Only Backbones for End-to-End Driving2026/2/1
- FRAPPE: Infusing World Modeling into Generalist Policies via Multiple Future Representation Alignment2026/2/1
- Counterfactual VLA: Self-Reflective Vision-Language-Action Model with Adaptive Reasoning2025/12/1
- Latent Chain-of-Thought World Modeling for End-to-End Driving2025/12/1
- Alpamayo-R1: Bridging Reasoning and Action Prediction for Generalizable Autonomous Driving in the Long Tail2025/11/1
- MoE-Based Learned Inertial Odometry for Bicycle Localization2025/10/1
- Efficient Multi-Camera Tokenization with Triplanes for End-to-End Driving2025/6/1
- SELC: Self-Supervised Efficient Local Correspondence Learning for Low Quality Images2025/4/1
- Low-Complexity Cooperative Payload Transportation for Nonholonomic Mobile Robots Under Scalable Constraints2025/2/1
- IROAM: Improving Roadside Monocular 3D Object Detection Learning from Autonomous Vehicle Data Domain2025/1/1
- Extrapolated Urban View Synthesis Benchmark2024/12/1
- MR-ULINS: A Tightly-Coupled UWB-LiDAR-Inertial Estimator with Multi-Epoch Outlier Rejection2024/8/1
- Revisiting Sparse Rewards for Goal-Reaching Reinforcement Learning2024/7/1
- Asynchronous Large Language Model Enhanced Planner for Autonomous Driving2024/6/1
- A New Method in Facial Registration in Clinics Based on Structure Light Images2024/5/1
- SceneVerse: Scaling 3D Vision-Language Learning for Grounded Scene Understanding2024/1/1
- Ordering-Flexible Multi-Robot Coordination for MovingTarget Convoying Using Long-TermTask Execution2024/1/1
- Interpretable Reinforcement Learning for Robotics and Continuous Control2023/11/1
- Waymax: An Accelerated, Data-Driven Simulator for Large-Scale Autonomous Driving Research2023/10/1
- Predictive Control for Autonomous Driving with Uncertain, Multi-modal Predictions2023/10/1
- DynaVIG: Monocular Vision/INS/GNSS Integrated Navigation and Object Tracking for AGV in Dynamic Scenes2022/11/1
- Real-Time Reinforcement Learning for Vision-Based Robotics Utilizing Local and Remote Computers2022/10/1
- Vehicle Trajectory Tracking Through Magnetic Sensors: A Case Study of Two-lane Road2022/9/1
- Robotic Imitation of Human Assembly Skills Using Hybrid Trajectory and Force Learning2021/3/1
- A Reliable Gravity Compensation Control Strategy for dVRK Robotic Arms With Nonlinear Disturbance Forces2020/1/1
- LDLS: 3-D Object Segmentation Through Label Diffusion From 2-D Images2019/10/1
- Motion Planning through Demonstration to Deal with Complex Motions in Assembly Process2019/10/1
- A Convex Optimization-based Dynamic Model Identification Package for the da Vinci Research Kit2019/2/1
- Model-Driven Feed-Forward Prediction for Manipulation of Deformable Objects2016/7/1