Yu-Gang Jiang
Fudan University
収録論文 41本 ・ フィジカルAI/ロボット学習
VLA/強化学習画像編集/ロボティクスエージェントアーキテクチャ
※arXiv著者名で収集。同姓同名の別人の論文が含まれる場合があります。
論文
- StructRL: フローベースVLAのための構造化アクション空間探索VLA/強化学習2026/8/15
フローベースの視覚言語行動モデル(VLA)のオンライン強化学習において、ノイズをアクション空間に直接注入する構造化探索手法StructRLを提案し、シミュレーションと実世界タスクで性能を向上させた。
- HandEdit: 身体性を考慮した人間からロボットへの器用な手の画像編集のための統一ベンチマーク画像編集/ロボティクス2026/8/12
人間の手の画像を様々なロボットハンドに変換する大規模な画像編集データセットとベンチマークを構築し、既存の編集モデルの評価と身体性を考慮した編集モデルの発展を促進する。
- ETA: 身体性タスクのための新しいエージェントパラダイムエージェントアーキテクチャ2026/8/4
ロボットのChatGPT的瞬間を目指し、プランナー・インターフェース・ワールドのループで構成される新しい身体性タスクエージェント(ETA)を提案し、オープンソース実装OpenETAを公開した。
- HandEdit: A Unified Benchmark for Egocentric Human-to-Robot Dexterous Hand Image Editing2026/8/1
- ETA: A New Agentic Paradigm for Embodied Tasks2026/8/1
- EgoRecovery: Acquiring Failure Recovery Ability Through Human Recovery Demonstration2026/7/1
- SegDiff: Segmented Trajectory Diffusion for Consistent and Adaptive Robot Manipulation2026/7/1
- Coarse-to-Control: Action-Token Planning for Vision-Language-Action Models2026/6/1
- ActiveMimic: Egocentric Video Pretraining with Active Perception2026/6/1
- Two Bridges, One Pathway: From VLMs to Generalizable VLAs with Embodied Trajectory-Coupled Data2026/6/1
- UniDexTok: A Unified Dexterous Hand Tokenizer from Real Data2026/6/1
- Seeing Touch from Motion: A Unified Modality-Aware Visuo-Tactile Policy with Tactile Motion Correlation2026/6/1
- Advancing Omnimodal Embodied Agents from Isolated Skills to Everyday Physical Autonomy2026/6/1
- ThinkingVLA: Interleaved Vision and Language Reasoning for Robotic Manipulation2026/6/1
- Bench2Drive-Robust: Benchmarking Closed-Loop Autonomous Driving under Deployment Perturbations2026/5/1
- GuidedVLA: Specifying Task-Relevant Factors via Plug-and-Play Action Attention Specialization2026/5/1
- VLA-Pro: Cross-Task Procedural Memory Transfer for Vision-Language-Action Models2026/5/1
- Safety in Embodied AI: A Survey of Risks, Attacks, and Defenses2026/5/1
- Afford-VLA: Action-Aligned Visual Planning via Internalized Affordance2026/5/1
- World Action Models: The Next Frontier in Embodied AI2026/5/1
- HazardArena: Evaluating Semantic Safety in Vision-Language-Action Models2026/4/1
- AssemLM: A Spatial Reasoning Multimodal Large Language Model for Robotic Assembly2026/4/1
- Robotic Grasping and Placement Controlled by EEG-Based Hybrid Visual and Motor Imagery2026/3/1
- OCRA: Object-Centric Learning with 3D and Tactile Priors for Human-to-Robot Action Transfer2026/3/1
- RC-NF: Robot-Conditioned Normalizing Flow for Real-Time Anomaly Detection in Robotic Manipulation2026/3/1
- FRoM-W1: Towards General Humanoid Whole-Body Control with Language Instructions2026/1/1
- Schr\"odinger's Navigator: Imagining an Ensemble of Futures for Zero-Shot Object Navigation2025/12/1
- HiMoE-VLA: Hierarchical Mixture-of-Experts for Generalist Vision-Language-Action Policies2025/12/1
- Unify Robot Actions in Camera Frame2025/11/1
- TP-MDDN: Task-Preferenced Multi-Demand-Driven Navigation with Autonomous Decision-Making2025/11/1
- RoboOmni: Proactive Robot Manipulation in Omni-modal Context2025/10/1
- DropVLA: An Action-Level Backdoor Attack on Vision-Language-Action Models2025/10/1
- Ask-to-Clarify: Resolving Instruction Ambiguity through Multi-turn Dialogue2025/9/1
- Embodied AI: From LLMs to World Models2025/9/1
- Actor-Critic for Continuous Action Chunks: A Reinforcement Learning Framework for Long-Horizon Robotic Manipulation with Sparse Reward2025/8/1
- HumanoidGen: Data Generation for Bimanual Dexterous Manipulation via LLM Reasoning2025/7/1
- TriVLA: A Triple-System-Based Unified Vision-Language-Action Model with Episodic World Modeling for General Robot Control2025/7/1
- You Only Estimate Once: Unified, One-stage, Real-Time Category-level Articulated Object 6D Pose Estimation for Robotic Grasping2025/6/1
- Human2Robot: Learning Robot Actions from Paired Human-Robot Videos2025/2/1
- SparseGrasp: Robotic Grasping via 3D Semantic Gaussian Splatting from Sparse Multi-View RGB Images2024/12/1
- VLABench: A Large-Scale Benchmark for Language-Conditioned Robotics Manipulation with Long-Horizon Reasoning Tasks2024/12/1