日本フィジカルAI新聞

世界のフィジカルAIを、日本語で。

週刊ニュースレター購読
医療ロボティクス/操作ガイダンスarXiv:2610.06008

世界モデルと検索ベース行動計画による超音波操作ガイダンス

Ultrasound Operator Guidance Using World Modeling and Retrieval Based Action Planning

シェア:XThreadsFacebookLINEはてブBluesky

超音波画像からプローブ操作を案内するため、V-JEPA 2.1で潜在空間に符号化し、参照データベースから類似状態を検索して非パラメトリックな遷移モデルを構築、再cedingホライズン計画で目標断面へ導く手法を提案した。

詳しい要約

1. どんなもの?

- 超音波検査の操作者ガイダンスシステム - 訓練が不十分なユーザーに対し、プローブを目標ビューへ動かす指示を提供 - 超音波取得のダイナミクスをretrieval-induced latent transition modelとして定式化 - マルチステップ計画とworld model内でのretrievalを組み合わせ - V-JEPA 2.1をバックボーンとして使用 - 展開時はライブ超音波画像のみでガイダンスを生成し、プローブ追跡ハードウェア不要

2. 先行研究と比べてどこがすごい?

- 従来の操作者ガイダンスシステムと比較して - パラメトリックな遷移関数を学習する代わりに、データベースからの物理的に実行された遷移を直接利用 - 非パラメトリックなretrieval-induced transition modelを構築 - 再ceding-horizon planningをサポート - 頸動脈超音波において、目標ビュー到達率86%を達成 - 代表的なベースラインの52%と43%を上回る - 全ての目標ビューで両ベースラインを上回り、困難な縦方向内頸動脈・外頸動脈ビューも含む

3. 技術・手法の肝は?

- V-JEPA 2.1バックボーンで観測を潜在空間にエンコード - 解剖学的に関連するビューが近くに配置される - 参照データベースから類似ビューを取得 - データベースにはエンコードされた潜在状態と対応するプローブ位置・姿勢が含まれる - パラメトリック遷移関数を学習せず、データベースからの物理的に実行された遷移を直接使用 - 非パラメトリックなretrieval-induced transition modelを構築 - 再ceding-horizon planningを実行 - 展開時はライブ超音波画像フィードのみからガイダンスを生成

4. どうやって有効だと検証した?

- 頸動脈超音波への適用 - 回顧的閉ループエピソードで目標ビュー到達率86%を達成 - ベースラインは52%と43% - 全ての目標ビューでベースラインを上回る - 未見のボランティアに対する前向き実現可能性研究 - リアルタイムでCPU上で蒸留を用いて実行 - 目標ビュー到達率83%を達成

5. 議論はある?

- 計画が任意のエンコード可能な目標潜在状態への近接性によって駆動されるため、同じworld modelで以前に取得した患者固有のフレームに戻ることが可能 - 周術期やフォローアップモニタリングなどの再現可能な縦断的イメージングをサポート - その他の議論や限界については要旨からは不明

6. 次に読むべき論文は?

- V-JEPA 2.1 - retrieval-induced latent transition model - 再ceding-horizon planning - 頸動脈超音波におけるベースライン手法 - その他の関連研究は要旨からは不明

※ AIが要旨から生成した要約です。正確性は原文をご確認ください。

著者: Noortje I. P. Schueler, Hans van Gorp, Ruud J. G. van Sloun

分類: cs.LG, cs.CV

原文アブストラクト

Ultrasound is widely used, but acquisition quality is heavily dependent on the operator's knowledge and expertise. With demand for examinations outpacing the supply of trained sonographers, operator-guidance systems aim to close this gap by instructing a less trained user how to move the probe toward a target view. In this paper, we propose a retrieval-induced latent transition model for ultrasound acquisition dynamics, formulating ultrasound operator guidance as multi-step planning and retrieval in a world model. Using a V-JEPA 2.1 backbone, observations are first encoded into a latent space where anatomically related views lie close together. We then retrieve similar views from a reference database containing encoded latent states and corresponding probe positions and orientations. Rather than learning a parametric transition function, we directly use physically executed transitions from the database to establish our nonparametric, retrieval-induced transition model that supports receding-horizon planning. At deployment, guidance is generated from the live ultrasound image feed alone, without any probe tracking hardware. Applied to carotid ultrasound, the proposed planner reaches the target view in 86% of retrospective closed-loop episodes, versus 52% and 43% for representative baselines, outperforming both on every target view, including the challenging longitudinal internal and external carotid artery views. A prospective feasibility study on unseen volunteers, run in real time on a CPU using distillation, reaches 83% target-view reachability. Because planning is driven by proximity to any encodable goal latent, the same world model can navigate back to any previously acquired, patient-specific frame, supporting reproducible longitudinal imaging for e.g. perioperative or follow-up monitoring.

PR本紙発行元 EmplifAI