日本フィジカルAI新聞

世界のフィジカルAIを、日本語で。

週刊ニュースレター購読
VLAarXiv:2608.15009v1

ForceU-VLA: 力覚を考慮した身体化超音波スキャンのための視覚・言語・行動モデル

ForceU-VLA: A Force-Aware Vision-Language-Action Model for Embodied Ultrasound Scanning

シェア:XThreadsFacebookLINEはてブBluesky

超音波スキャン中に力信号と超音波画像を融合し、スキャン段階に応じて適応的に特徴を調整する新しいVLAモデルを提案し、実世界の力覚対応データセットも構築した。

詳しい要約

1. どんなもの?

ForceU-VLAは、力覚信号と超音波画像フィードバックを活用した、自律的な身体化超音波スキャンのためのForce-Aware Vision-Language-Actionモデルである。Force-Ultrasound Synergistic Fusion Module (FUSFM)とStage-Adaptive Modulation Mechanism (SAMM)を導入し、プローブの動きを安定・信頼性高く導く。また、実世界の力覚対応データセットForceU-VLA-Dataを構築し、2臓器・5つの臨床スキャンビュー・450軌道・約10万フレームを含む。

2. 先行研究と比べてどこがすごい?

既存手法は力覚と超音波モダリティの結合が疎であり、スキャン段階の認識が不足していた。ForceU-VLAは、力覚と超音波画像を相乗的に融合し、段階適応変調により動的なプローブ-組織相互作用を捉える点で優れている。

3. 技術・手法の肝は?

手法の肝は、FUSFMによる力覚と超音波視覚情報の相乗的融合と、SAMMによるスキャン段階に応じたマルチモーダル特徴の適応的変調である。これにより、プローブの接触安定性と圧力調整を向上させる。

4. どうやって有効だと検証した?

実世界のForceU-VLA-Dataデータセットを用いた広範な実験により、接触安定性とプローブ圧力調整の改善、タスク実行品質とシステム信頼性の向上を実証した。

5. 議論はある?

要旨からは、提案手法の限界や他の手法との比較に関する議論は不明。ただし、データセットが2臓器に限定されていることや、実環境での検証が中心である点が議論の余地となる可能性がある。

6. 次に読むべき論文は?

要旨で参照されている関連研究は明示されていないが、同分野の定番として、Vision-Language-Actionモデル、超音波スキャンの自動化、力覚制御を用いたロボット手術などの論文が挙げられる。

※ AIが要旨から生成した要約です。正確性は原文をご確認ください。

著者: Xingzheng Wu, Cheng Zhang, Guihao Yan, Xifeng Hu, Zhi Liu, Qing Cai

分類: cs.RO, cs.CV

原文アブストラクト

Embodied intelligent ultrasound scanning enables the automation and standardization of the ultrasound examination process by integrating perception, decision-making, and execution capabilities. However, existing methods suffer from loosely coupled modeling between force and ultrasound modalities and lack awareness of scanning stages, which limits their ability to capture dynamic probe-tissue interactions. To address these issues, we propose ForceU-VLA, a force-aware Vision-Language-Action model for autonomous embodied ultrasound scanning, which leverages force signals and ultrasound image feedback throughout the scanning process to enable accurate and high-quality ultrasound acquisition. Firstly, we propose a Force-Ultrasound Synergistic Fusion Module (FUSFM) that synergistically fuses ultrasound visual and force-feedback information to provide stable, reliable guidance for probe motion. Secondly, a Stage-Adaptive Modulation Mechanism (SAMM) is proposed to accommodate the task requirements across different scanning stages by adaptively modulating multimodal features to enhance their representation quality. Additionally, we introduce ForceU-VLA-Data, a real-world, force-aware embodied ultrasound dataset that integrates visual, force, and action signals, including data from two organs across five representative clinical scanning views, and comprising 450 expert-collected trajectories with approximately 100,000 synchronized multimodal frames. Extensive experimental results demonstrate that ForceU-VLA significantly improves contact stability and probe pressure regulation in embodied ultrasound scanning, thereby effectively enhancing task execution quality and overall system reliability. The source code is available at https://github.com/VMVLab/ForceU-VLA.