日本フィジカルAI新聞

世界のフィジカルAIを、日本語で。

週刊ニュースレター購読
内視鏡ロボティクスarXiv:2608.22093

EndoNav:言語誘導ロボット内視鏡検査のための意味論から幾何学への接地

EndoNav: Semantic-to-Geometric Grounding for Language-Guided Robotic Endoscopic Examination

シェア:XThreadsFacebookLINEはてブBluesky

外科医の音声コマンドを患者固有の解剖学的シーン表現に基づく視点目標に変換し、内視鏡の自律的な可視化動作を実現するフレームワークを提案した。

詳しい要約

1. どんなもの?

EndoNavは、言語指示によるロボット内視鏡検査のためのフレームワークである。外科医の音声コマンドを転写・解釈し、患者固有の解剖学的シーン表現に基づいて、内視鏡の自律的可視化行動(視点目標、検査軌道)を生成する。生成された目標は、幾何学的制約付き動作計画と関節空間制御により実行される。

2. 先行研究と比べてどこがすごい?

従来のロボット内視鏡システムは、内視鏡の安定化や位置変更は可能だが、効果的な可視化支援のための関連文脈(解剖学的知識)を持たない。EndoNavは、高レベルの外科医コマンドを患者固有の解剖学的目標に接地(semantic-to-geometric grounding)し、自律的な可視化行動に変換する点で新しい。

3. 技術・手法の肝は?

手法の肝は、エンドスコピック視点エージェントがロボット動作を直接生成するのではなく、構造化された可視化目標(視点と検査軌道)を生成し、それを幾何学的制約付き動作計画と関節空間制御で実行する点。また、患者固有の解剖学的シーン表現に基づいてコマンドを解釈する点。

4. どうやって有効だと検証した?

3つのCT由来解剖モデルを用いた構造化3パス副鼻腔検査で評価。1つの死体標本では、2人のレジデント外科医による検査と比較し、平均可視化IoUは87.04%と84.37%(外科医間IoUは87.44%)、外科医が観察した解剖学的表面の92.91%と93.20%を回復した。

5. 議論はある?

要旨からは、限界や議論は不明。ただし、死体標本1体のみの比較であり、一般化にはさらなる検証が必要と考えられる。

6. 次に読むべき論文は?

要旨で参照されている関連研究は明示されていないが、同分野の定番として、言語指示によるロボットナビゲーション(例:Vision-and-Language Navigation)や、内視鏡ロボットの自律制御に関する論文が挙げられる。

※ AIが要旨から生成した要約です。正確性は原文をご確認ください。

著者: Jecia Z. Y. Mao, Hisashi Ishida, Kathryn Jung, Masaru Ishii, Russell H. Taylor, Manish Sahu

分類: cs.RO

原文アブストラクト

Minimally invasive procedures performed within confined anatomical spaces depend on continuous endoscopic visualization. Current robotic endoscope systems can stabilize or reposition an endoscope, but they do not possess relevant context to provide effective visualization assistance. We present EndoNav, an anatomy-grounded natural-language framework that translates high-level surgeon commands into autonomous endoscopic visualization behaviors within patient-specific sinonasal anatomy. Spoken surgeon commands are transcribed and interpreted by an endoscopic viewpoint agent conditioned on a patient-specific anatomical scene representation. Rather than generating robot motion directly, the viewpoint agent generates structured visualization objectives that are converted into target viewpoints and inspection trajectories, which are then executed through geometry-constrained endoscope motion planning and joint-space control. We evaluate EndoNav using a structured three-pass sinus examination across three CT-derived anatomical models. For one cadaveric specimen, autonomous visualization is compared with sinus examinations performed by two resident surgeons. EndoNav achieved mean visualization IoUs of 87.04% and 84.37% relative to the two surgeon examinations, compared with an inter-surgeon IoU of 87.44%, while recovering 92.91% and 93.20% of surgeon-observed anatomical surfaces, respectively. These results demonstrate the feasibility of grounding high-level anatomical commands into patient-specific geometric objectives and translating them into anatomically constrained robotic visualization behaviors.