Hao Shi
収録論文 53本 ・ フィジカルAI/ロボット学習
※arXiv著者名で収集。同姓同名の別人の論文が含まれる場合があります。
論文
- Label-Free Target-Domain Adaptation for Unconstrained Event-Image Feature Matching via Dual-Stage Distillation2026/7/1
- PS-MOT: Cultivating Instance Awareness from Point Seeds for Multi-Object Tracking2026/6/1
- VL2Spike: Spike-driven Distillation from VLMs for Low-Power Visual Perception in Embodied AI2026/6/1
- MemoryVLA++: Temporal Modeling via Memory and Imagination in Vision-Language-Action Models2026/6/1
- EgoEV-HandPose: Egocentric 3D Hand Pose Estimation and Gesture Recognition with Stereo Event Cameras2026/5/1
- E-VLA: Event-Augmented Vision-Language-Action Model for Dark and Blurred Scenes2026/4/1
- OccTrack360: 4D Panoptic Occupancy Tracking from Surround-View Fisheye Cameras2026/3/1
- O3N: Omnidirectional Open-Vocabulary Occupancy Prediction for Urban Autonomous Agents2026/3/1
- RMBench: Memory-Dependent Robotic Manipulation Benchmark with Insights into Policy Design2026/3/1
- ExoGS: A 4D Real-to-Sim-to-Real Framework for Scalable Manipulation Data Collection2026/1/1
- Forging Spatial Intelligence: A Roadmap of Multi-Modal Data Pre-Training for Autonomous Systems2025/12/1
- OneOcc: Semantic Occupancy Prediction for Legged Robots with a Single Panoramic Camera2025/11/1
- OmniTrack++: Omnidirectional Multi-Object Tracking by Learning Large-FoV Trajectory Feedback2025/11/1
- SpatialActor: Exploring Disentangled Spatial Representations for Robust Robotic Manipulation2025/11/1
- Dexbotic: Open-Source Vision-Language-Action Toolbox2025/10/1
- Event-guided 3D Gaussian Splatting for Dynamic Human and Scene Reconstruction2025/9/1
- MemoryVLA: Perceptual-Cognitive Memory in Vision-Language-Action Models for Robotic Manipulation2025/8/1
- QuaDreamer: Controllable Panoramic Video Generation for Quadruped Robots2025/8/1
- GeoVLA: Empowering 3D Representations in Vision-Language-Action Models2025/8/1
- Unlocking Constraints: Source-Free Occlusion-Aware Seamless Segmentation2025/6/1
- Resource-Efficient Affordance Grounding with Complementary Depth and Semantic Prompts2025/3/1
- EgoEvGesture: Gesture Recognition Based on Egocentric Event Camera2025/3/1
- Omnidirectional Multi-Object Tracking2025/3/1
- HierDAMap: Towards Universal Domain Adaptive BEV Mapping via Hierarchical Perspective Priors2025/3/1
- TS-CGNet: Temporal-Spatial Fusion Meets Centerline-Guided Diffusion for BEV Mapping2025/3/1
- Event-aided Semantic Scene Completion2025/2/1
- Benchmarking the Robustness of Optical Flow Estimation to Corruptions2024/11/1
- E-3DGS: Gaussian Splatting with Exposure and Motion Events2024/10/1
- EI-Nexus: Towards Unmediated and Flexible Inter-Modality Local Feature Extraction and Matching for Event-Image Data2024/10/1
- GenMapping: Unleashing the Potential of Inverse Perspective Mapping for Robust Online HD Map Construction2024/9/1
- SF-TIM: A Simple Framework for Enhancing Quadrupedal Robot Jumping Agility by Combining Terrain Imagination and Measurement2024/8/1
- Occlusion-Aware Seamless Segmentation2024/7/1
- Label-efficient Semantic Scene Completion with Scribble Annotations2024/5/1
- DTCLMapper: Dual Temporal Consistent Learning for Vectorized HD Map Construction2024/5/1
- MambaMOS: LiDAR-based 3D Moving Object Segmentation with Motion-aware State Space Model2024/4/1
- Representing Domain-Mixing Optical Degradation for Real-World Computational Aberration Correction via Vector Quantization2024/3/1
- Offboard Occupancy Refinement with Hybrid Propagation for Autonomous Driving2024/3/1
- Towards Precise 3D Human Pose Estimation with Multi-Perspective Spatial-Temporal Relational Transformers2024/1/1
- Exploring Event-based Human Pose Estimation with 3D Event Representations2023/11/1
- CoBEV: Elevating Roadside 3D Object Detection with Depth and Height Complementarity2023/10/1
- FocusFlow: Boosting Key-Points Optical Flow Estimation for Autonomous Driving2023/8/1
- Towards Anytime Optical Flow Estimation with Event Cameras2023/7/1
- LF-PGVIO: A Visual-Inertial-Odometry Framework for Large Field-of-View Cameras using Points and Geodesic Segments2023/6/1
- Bi-Mapper: Holistic BEV Semantic Mapping for Autonomous Driving2023/5/1
- FishDreamer: Towards Fisheye Semantic Completion via Unified Image Outpainting and Segmentation2023/3/1
- PanoVPR: Towards Unified Perspective-to-Equirectangular Visual Place Recognition via Sliding Windows across the Panoramic View2023/3/1
- LF-VISLAM: A SLAM Framework for Large Field-of-View Cameras with Negative Imaging Plane on Mobile Agents2022/9/1
- Behind Every Domain There is a Shift: Adapting Distortion-aware Vision Transformers for Panoramic Semantic Segmentation2022/7/1
- Efficient Human Pose Estimation via 3D Event Point Cloud2022/6/1
- Review on Panoramic Imaging and Its Applications in Scene Understanding2022/5/1
- CSFlow: Learning Optical Flow via Cross Strip Correlation for Autonomous Driving2022/2/1
- LF-VIO: A Visual-Inertial-Odometry Framework for Large Field-of-View Cameras with Negative Plane2022/2/1
- PanoFlow: Learning 360{\deg} Optical Flow for Surrounding Temporal Understanding2022/2/1