日本フィジカルAI新聞

世界のフィジカルAIを、日本語で。

週刊ニュースレター購読
手術ロボティクスarXiv:2609.27227

単眼動画からの手術キネマティクス:学習された関節運動制約による再構成

Surgical Kinematics from Monocular Video with Learned Articulated Motion Constraints

シェア:XThreadsFacebookLINEはてブBluesky

単眼動画から手術器具の位置・姿勢・ジョー角度を推定するキネマティック再構成ネットワークを提案し、Open-Hデータセットで経路長誤差と動作セグメンテーション精度を改善した。

詳しい要約

1. どんなもの?

- 単眼動画から手術器具のkinematicsを再構成する研究 - 推定対象は器具のposition、orientation、jaw angle - ロボット手術の客観評価に用いるinstrument kinematicsを動画のみから得ることを目的とする - kinematic reconstruction networkを提案 - 評価はOpen-Hの2,802エピソードで実施

2. 先行研究と比べてどこがすごい?

- 先行研究としてLiveMAEと比較 - 主要なOpen-Hベンチマークでpath-length MAEを0.45cmから0.34cmへ低減 - motion segmentationのtemporal mean average precisionを44.54%から54.44%へ向上 - 単眼動画からの再構成で先行手法を上回る性能を示す - 具体的な優位性は上記2指標で示される

3. 技術・手法の肝は?

- frozen DINOv3 featuresのglobal attention poolingを利用 - fine-tuned SAM 3.1 masksによる器具landmarkのlocal poolingを組み合わせる - 共有Transformer encoderとtemporal convolutional headsで表現を統合 - mask geometry、monocular depth、arm-specific multilayer regression networksによるvisual state estimatesを入力に含む - position branchはdisplacement magnitudeとdirectionを分離予測し移動距離を保持 - 予測状態観測とmotion incrementsをdifferentiable weighted least squaresでtrajectory fitting - quaternion観測を累積予測回転に対して表現しquadratic orientation objectiveを構成

4. どうやって有効だと検証した?

- Open-Hの2,802エピソードでreconstructionを評価 - 主要Open-HベンチマークでLiveMAEと比較 - path-length mean absolute errorを0.45cmから0.34cmへ改善 - motion segmentationのtemporal mean average precisionを44.54%から54.44%へ改善 - これらの定量比較により有効性を検証

5. 議論はある?

- 要旨からは不明 - 限界や失敗事例、計算コスト、一般化性に関する議論は要旨に記載なし - 比較対象や評価指標の詳細な考察も要旨からは不明

6. 次に読むべき論文は?

- LiveMAE(要旨で比較されている先行研究) - DINOv3(利用されているvisual feature) - SAM 3.1(利用されているsegmentation model) - Open-H(評価データセット) - monocular depth estimationやsurgical video-based kinematic reconstructionの関連研究

※ AIが要旨から生成した要約です。正確性は原文をご確認ください。

著者: Mehmet Kerem Turkcan, Soham Samal, Zoran Kostic

分類: cs.CV, cs.RO

原文アブストラクト

Objective assessment of robotic surgery uses instrument kinematics, which must be reconstructed when only video is available. We introduce a kinematic reconstruction network for estimating instrument position, orientation and jaw angle from monocular video. Our visual representation combines global attention pooling of frozen DINOv3 features with local pooling at instrument landmarks from fine-tuned SAM 3.1 masks. Our shared Transformer encoder and temporal convolutional heads integrate this representation with mask geometry, monocular depth and visual state estimates from arm-specific multilayer regression networks. Our position branch predicts displacement magnitude and direction separately to preserve traveled distance. We fit trajectories to predicted state observations and motion increments by differentiable weighted least squares, expressing quaternion observations relative to cumulative predicted rotations to obtain a quadratic orientation objective. We evaluate reconstruction across 2,802 Open-H episodes. Compared with LiveMAE on the main Open-H benchmark, our method reduces path-length mean absolute error from 0.45 to 0.34\,cm and increases temporal mean average precision for motion segmentation from 44.54\% to 54.44\%.

関連論文

PR本紙発行元 EmplifAI