日本フィジカルAI新聞

世界のフィジカルAIを、日本語で。

週刊ニュースレター購読
医療ロボティクスarXiv:2608.20478

EndoLIFT: 言語で曖昧性を解消する潜在条件付き整流フローによる双方向内視鏡制御

EndoLIFT: Language-Disambiguated Latent-Conditioned Rectified Flow for Bidirectional Endoscopic Control

シェア:XThreadsFacebookLINEはてブBluesky

内視鏡検査の双方向制御において、同じ視覚情報でも逆の動作が必要となる「意図の曖昧性」を解決するため、言語指示と軌跡潜在変数を組み合わせた視覚言語行動ポリシーを提案した。

詳しい要約

1. どんなもの?

EndoLIFTは、内視鏡操作における双方向制御(前進・後退)の曖昧性(intent aliasing)を解決するためのvision-language-actionポリシーである。RGB画像、言語指示、前回のアクション状態を入力とし、32次元の変分軌道潜在変数(trajectory latent)で連続アクションチャンク生成を条件付けつつ、明示的な言語ベースの意図条件付けを組み合わせる。

2. 先行研究と比べてどこがすごい?

先行研究では、視覚シーンが変化する前に指示が変わると、ほぼ同一の観察に対して逆の軸方向アクションが必要となる曖昧性が未解決だった。EndoLIFTは、言語指示による意図選択と軌道潜在変数の寄与を分離し、言語が軸方向モードを選択することを示した点が新しい。また、潜在条件付けなしのモデルと比較して、ナビゲーション方向精度を11.1ポイント向上させ、誤方向前進を83%削減した。

3. 技術・手法の肝は?

手法の肝は、言語指示による意図条件付けと、latent-conditioned rectified-flowアクション生成の組み合わせである。具体的には、RGB、言語指示、前回のアクション状態を入力とし、32次元の変分軌道潜在変数(VTL)が連続アクションチャンク生成を確率的に条件付ける。また、同一観察での指示入れ替え実験により、言語が軸方向モードを選択することを検証した。

4. どうやって有効だと検証した?

有効性は、同一観察での指示入れ替え実験、潜在条件付けなしモデルとの比較、1ビットモードフラグ参照との比較、44の未見言語バリエーションでの意図追従精度評価、閉ループ評価(seen colon phantom、unseen lung/stomach phantoms、ex-vivo porcine trachea)で検証した。その結果、EndoLIFTは全体成功率を30ポイント向上させ、ex-vivo試験で10/10成功した。

5. 議論はある?

議論として、言語ベースの意図選択と軌道潜在変数の寄与が分離され、言語が軸方向モードを選択することが示されたが、軌道潜在変数が方向正確性と堅牢な後退に寄与することも示唆された。また、1ビットモードフラグ参照は弱いcanonical-anchor切り替えしか示さず、言語指示の重要性が強調された。ただし、実臨床での有効性や他の解剖学的部位への一般化については要旨からは不明。

6. 次に読むべき論文は?

要旨で参照されている関連研究として、vision-language-action policy、latent-conditioned rectified flow、variational trajectory latent、bidirectional endoscopic controlに関する論文が挙げられる。具体的には、同分野の定番としてRT-2やDiffusion Policyなどが関連するが、要旨に明示されていないため、次に読むべき論文は要旨からは不明。

※ AIが要旨から生成した要約です。正確性は原文をご確認ください。

著者: Chi Kit Ng, Yidong Zhang, Lui Siu Hing, Jinsong Lin, Tianchun Wu, Ho Yin Chim, Zhiqing Tang, Tao Yang, Huxin Gao, Trevor Yeung, Raymond Shing-Yan Tang, Hongliang Ren

分類: cs.RO

原文アブストラクト

Routine gastrointestinal endoscopy is intrinsically bidirectional: the instrument is advanced to reach target anatomy and later withdrawn or retroflexed for inspection, while an external cue may require earlier reversal. When the requested phase changes before the visual scene does, nearly identical observations can require opposite axial actions. We identify and formalize this ambiguity in bidirectional endoscopic control as intent aliasing. We propose EndoLIFT (Endoscopic Language-Instruction Flow with Trajectory Latents), a vision-language-action policy that combines explicit language-based intent conditioning with a latent-conditioned rectified-flow action expert. The policy receives RGB, a language instruction, and the previous-action state; a 32-D variational trajectory latent stochastically conditions continuous action-chunk generation. Controlled same-observation instruction swaps establish that language selects the axial mode, independently of whether the trajectory latent is present. Relative to the matched model without latent conditioning, EndoLIFT improves navigation-direction accuracy by 11.1 percentage points and reduces wrong-direction advance by 83\%. An architecture-controlled 1-bit mode-flag reference exhibits weaker canonical-anchor switching, while EndoLIFT retains 82.8\% intent-following accuracy across 44 held-out linguistic variants. In closed-loop evaluation, EndoLIFT improves overall success by 30 percentage points over EndoLIFT w/o VTL on both the seen colon phantom and the unseen lung and stomach phantoms, and completes 10/10 ex-vivo porcine-trachea trials. These results separate language-based intent selection from the trajectory latent's contribution to directional correctness and robust retraction.

関連論文