EndoLIFT: 言語で曖昧性を解消する潜在条件付き整流フローによる双方向内視鏡制御
EndoLIFT: Language-Disambiguated Latent-Conditioned Rectified Flow for Bidirectional Endoscopic Control
内視鏡検査の双方向制御において、同じ視覚情報でも逆の動作が必要となる「意図の曖昧性」を解決するため、言語指示と軌跡潜在変数を組み合わせた視覚言語行動ポリシーを提案した。
著者: Chi Kit Ng, Yidong Zhang, Lui Siu Hing, Jinsong Lin, Tianchun Wu, Ho Yin Chim, Zhiqing Tang, Tao Yang, Huxin Gao, Trevor Yeung, Raymond Shing-Yan Tang, Hongliang Ren
分類: cs.RO
原文アブストラクト
Routine gastrointestinal endoscopy is intrinsically bidirectional: the instrument is advanced to reach target anatomy and later withdrawn or retroflexed for inspection, while an external cue may require earlier reversal. When the requested phase changes before the visual scene does, nearly identical observations can require opposite axial actions. We identify and formalize this ambiguity in bidirectional endoscopic control as intent aliasing. We propose EndoLIFT (Endoscopic Language-Instruction Flow with Trajectory Latents), a vision-language-action policy that combines explicit language-based intent conditioning with a latent-conditioned rectified-flow action expert. The policy receives RGB, a language instruction, and the previous-action state; a 32-D variational trajectory latent stochastically conditions continuous action-chunk generation. Controlled same-observation instruction swaps establish that language selects the axial mode, independently of whether the trajectory latent is present. Relative to the matched model without latent conditioning, EndoLIFT improves navigation-direction accuracy by 11.1 percentage points and reduces wrong-direction advance by 83\%. An architecture-controlled 1-bit mode-flag reference exhibits weaker canonical-anchor switching, while EndoLIFT retains 82.8\% intent-following accuracy across 44 held-out linguistic variants. In closed-loop evaluation, EndoLIFT improves overall success by 30 percentage points over EndoLIFT w/o VTL on both the seen colon phantom and the unseen lung and stomach phantoms, and completes 10/10 ex-vivo porcine-trachea trials. These results separate language-based intent selection from the trajectory latent's contribution to directional correctness and robust retraction.
関連論文
- 密集制約環境下でのロボットによる歯牙形成のカバレッジ計画医療ロボティクス
- US-VLA: 腹部超音波検査のための視覚・言語・動作モデル医療ロボティクス
- RoSE: 内視鏡ステント試験用のロボット軟性食道医療ロボティクス
- MRI室内での針ベース遠隔操作のためのマスタースレーブロボットマニピュレータ医療ロボティクス