Wenxuan Song
収録論文 39本 ・ フィジカルAI/ロボット学習
評価基盤ロボットポリシー評価VLAワールドモデル
※arXiv著者名で収集。同姓同名の別人の論文が含まれる場合があります。
論文
- XPolicyLab: ロボットポリシー評価と展開のための統一標準・オープンエコシステム評価基盤2026/8/10
ロボットポリシーの評価と展開を統一する標準規格とオープンエコシステムを提案し、N個のポリシーとM個の環境の接続コストをO(NM)からO(N+M)に削減する。
- XPolicyLab: ロボットポリシー評価と展開のための統一標準・オープンエコシステムロボットポリシー評価2026/8/10
ロボットポリシーの評価と展開を統一する標準規格とオープンエコシステムを提案し、N個のポリシーとM個の環境の接続コストをO(NM)からO(N+M)に削減した。
- 4D-WAM: 軌跡フィールドによる世界行動モデルへの時空間認識の注入VLA2026/8/8
ロボットの行動生成と動画予測を統合する世界行動モデルに、3次元軌跡フィールドの時空間知識を表現整合で注入する訓練戦略を提案。局所的な動き整合と長期的な目的地整合の2つの目的関数により、軌跡レベルの時空間表現を学習し、空間理解や実行精度、汎化性を向上させる。
- 4D-WAM: 軌跡フィールドによる世界行動モデルへの時空間認識の注入VLA2026/8/8
ロボットの行動生成と動画予測を統合する世界行動モデルに、3次元軌跡フィールドの時空間知識を表現整合で注入する訓練戦略を提案し、空間理解と実行精度を向上させた。
- 前方予測だけで十分か?JEPAワールドモデルのための物理状態接地ワールドモデル2026/8/7
JEPAベースのワールドモデルに、ロボットの自己受容状態と関節角変化を接地する2つの目的を追加し、潜在表現の識別性と下流タスク性能を向上させる手法を提案した。
- Robust-WAM: 生成事前学習と意味的予見を橋渡しするワールド・アクションモデルVLA2026/8/6
ロボット制御用のワールド・アクションモデルにおいて、VAE潜在空間の生成事前学習を保ちつつ、意味的潜在空間の頑健性を組み込む後処理手法を提案した。外観変化に頑健なアクション予測を実現する。
- Is Forward Prediction Enough? Physical State Grounding for JEPA World Models2026/8/1
- XPolicyLab: A Unified Standard and Open Ecosystem for Robot Policy Evaluation and Deployment2026/8/1
- 4D-WAM: Infusing Spatiotemporal Awareness into World Action Models through Trajectory Fields2026/8/1
- Robust-WAM: Bridging Generative Pretraining and Semantic Foresight in World-Action Models2026/8/1
- DreamTrajectory: Trajectory-Guided Action Generation with World Model Alignment for Mobile Manipulation2026/8/1
- Xiaomi-Robotics-1: Scaling Vision-Language-Action Models with over 100K Hours of Real-World Trajectories2026/7/1
- Revisiting Embodied Chain-of-Thought for Generalizable Robot Manipulation2026/6/1
- SEDualVLN: A Spatially-Enhanced Dual-System for Vision-Language Navigation2026/5/1
- CapVector: Learning Transferable Capability Vectors in Parametric Space for Vision-Language-Action Models2026/5/1
- RoboMemArena: A Comprehensive and Challenging Robotic Memory Benchmark2026/5/1
- DFM-VLA: Iterative Action Refinement for Robot Manipulation via Discrete Flow Matching2026/3/1
- Fast-dVLA: Accelerating Discrete Diffusion VLA to Real-Time Performance2026/3/1
- VAMPO: Policy Optimization for Improving Visual Dynamics in Video Action Models2026/3/1
- S-VAM: Shortcut Video-Action Model by Self-Distilling Geometric and Semantic Foresight2026/3/1
- MMaDA-VLA: Large Diffusion Vision-Language-Action Model with Unified Multi-Modal Instruction and Generation2026/3/1
- Rethinking the Practicality of Vision-language-action Model: A Comprehensive Benchmark and An Improved Baseline2026/2/1
- FRAPPE: Infusing World Modeling into Generalist Policies via Multiple Future Representation Alignment2026/2/1
- HiF-VLA: Hindsight, Insight and Foresight through Motion Representation for Vision-Language-Action Models2025/12/1
- Embodied Robot Manipulation in the Era of Foundation Models: Planning and Learning Perspectives2025/12/1
- Unified Diffusion VLA: Vision-Language-Action Model via Joint Discrete Denoising Diffusion Process2025/11/1
- Spatial Forcing: Implicit Spatial Representation Alignment for Vision-language-action Model2025/10/1
- Towards a Unified Understanding of Robot Manipulation: A Comprehensive Survey2025/10/1
- VLA^2: Empowering Vision-Language-Action Models with an Agentic Framework for Unseen Concept Manipulation2025/10/1
- VLA-Adapter: An Effective Paradigm for Tiny-Scale Vision-Language-Action Model2025/9/1
- FlowVLA: Visual Chain of Thought-based Motion Reasoning for Vision-Language-Action Models2025/8/1
- ReconVLA: Reconstructive Vision-Language-Action Model as Effective Robot Perceiver2025/8/1
- RationalVLA: A Rational Vision-Language-Action Model with Dual System2025/6/1
- CEED-VLA: Consistency Vision-Language-Action Model with Early-Exit Decoding2025/6/1
- OpenHelix: A Short Survey, Empirical Analysis, and Open-Source Dual-System VLA Model for Robotic Manipulation2025/5/1
- PD-VLA: Accelerating Vision-Language-Action Model Integrated with Action Chunking via Parallel Decoding2025/3/1
- MoRE: Unlocking Scalability in Reinforcement Learning for Quadruped Vision-Language-Action Models2025/3/1
- GeRM: A Generalist Robotic Model with Mixture-of-experts for Quadruped Robot2024/3/1
- QUAR-VLA: Vision-Language-Action Model for Quadruped Robots2023/12/1