日本フィジカルAI新聞

世界のフィジカルAIを、日本語で。

週刊ニュースレター購読
VLAarXiv:2609.39323

HiWE: 視覚キーポイント強化による階層的ワールド知識モデルを用いたゼロショット3D経路計画

HiWE: Hierarchical World Knowledge Model with Visual Keypoint Enhancement for Zero-Shot 3D Path Planning

シェア:XThreadsFacebookLINEはてブBluesky

視覚的グラウンディングと言語ベースの計画を点ベースのインターフェースで結び、ゼロショットで3D操作経路を計画する階層的フレームワークを提案。

著者: Guoqing Ma, Mingqi Yuan, Chen Gao, Jiayu Chen, Shan Yu

分類: cs.RO, cs.AI

原文アブストラクト

Robot demonstration generation requires a system to identify where an interaction should occur, plan a feasible motion, and execute the required contact. HiWE connects these decisions through a point-based interface between visual grounding and language-based planning. PointVLM is instruction-tuned to associate task-relevant objects with image coordinates using a mixture of point annotations, segmentation-derived samples, robot observations, and visual question answering data. Depth measurements lift these predictions into a semantic 3D representation. A language planner, 3DLLM, uses this representation to specify end-effector waypoints and gripper commands, while a hybrid grasping module resolves local grasp poses. The evaluation covers 14 simulated manipulation tasks and four physical-robot tasks, together with ablations of the visual training data, spatial inputs, and grasp selection. Here, zero-shot execution refers to deployment without task-specific demonstration training; the visual model uses existing robot data during fine-tuning. This paper describes the original point-based formulation of the framework; its relationship to the subsequent GeneralVLA extension is detailed in the introduction.

関連論文

PR本紙発行元 EmplifAI