日本フィジカルAI新聞

世界のフィジカルAIを、日本語で。

週刊ニュースレター購読
宇宙ロボティクス/相対航法arXiv:2610.07231

未知宇宙機に対する単眼画像を用いたTransformer支援カルマンフィルタによる相対航法

Monocular Navigation Relative to Unknown Spacecraft Using a Transformer-Aided Kalman Filter

シェア:XThreadsFacebookLINEはてブBluesky

単眼カメラ画像のみから未知の宇宙機の姿勢と位置を推定するため、Transformerによるオドメトリ推定とMSCKFを組み合わせた手法を提案し、SPE3Rデータセットで評価した。

詳しい要約

1. どんなもの?

単眼画像のみを用いて未知の宇宙機の相対姿勢(位置・姿勢)を推定する学習ベースのパイプライン。servicerの単一カメラから得られる画像列を入力とし、TransformerネットワークとMulti-State Constraint Kalman Filter (MSCKF)を組み合わせる。ランデブー・近接運用中に、目標宇宙機の形状や慣性特性の事前知識、depth/lidar/stereoなどの追加センサを必要とせず、未知のターゲットに汎化する。

2. 先行研究と比べてどこがすごい?

既存の視覚ベース手法は、ターゲット形状や慣性特性の事前知識を必要としたり、depth・lidar・stereoなどの追加センサに依存したり、スケール不定の並進しか復元できないことが多い。提案手法は単眼カメラのみで、未知の宇宙機に対しても汎化し、フィルタにより相対軌道要素・姿勢・角速度を直接推定する点が優れている。

3. 技術・手法の肝は?

TransformerネットワークがSuperPoint特徴をLightGlueでマッチングし、画像間のオドメトリ(スケール不定の姿勢変化)を推定。これを疑似観測としてMSCKFに与え、軌道・姿勢運動学モデルと融合する。フィルタは相対軌道要素、カメラに対するターゲット姿勢、関連する角速度を直接推定する。単眼かつ近距離のため、servicerの姿勢マヌーバによりターゲットまでの距離の完全可観測性を回復する。

4. どうやって有効だと検証した?

SPE3Rデータセットを再レンダリングした高解像度版(103機の宇宙機の合成画像を含む)で学習・評価。うち11機を学習時にホールドアウトし、未知ターゲットへの汎化を評価。ホールドアウト宇宙機のレンダリング軌道に対してMonte Carloシミュレーションを実施。結果、未知ターゲット周辺のナビゲーションで姿勢誤差中央値3.7°、ROEのレンジ誤差中央値2.2%を達成。

5. 議論はある?

単眼かつ近距離という条件下で距離の完全可観測性を得るためにservicerの姿勢マヌーバが不可欠である点が議論の対象。また、学習ベースのフロントエンドとKalmanフィルタの組み合わせの有効性が示されたが、実機環境や多様な照明条件への適用可能性については要旨からは不明。

6. 次に読むべき論文は?

SuperPoint、LightGlue、Multi-State Constraint Kalman Filter (MSCKF)、SPE3R dataset。関連手法として、depth/lidar/stereoを用いる視覚ベースの宇宙機姿勢推定や、スケール不定の単眼オドメトリ手法が挙げられる。

※ AIが要旨から生成した要約です。正確性は原文をご確認ください。

著者: Pol Francesch Huc, Simone D'Amico

分類: cs.RO, cs.CV

原文アブストラクト

This work presents a novel learning-based pipeline for pose estimation of unknown spacecraft using only monocular images from a single servicer. The approach combines a transformer-based neural network with a Multi-State Constraint Kalman Filter (MSCKF) to estimate the pose (i.e., position and orientation of the target spacecraft relative to the camera) throughout rendezvous and proximity operations. Unlike existing vision-based methods that require prior knowledge of the target shape or inertia properties, rely on additional sensing modalities such as depth, lidar, or stereo, or only recover translation up to scale, the proposed pipeline generalizes to previously unseen spacecraft using a single monocular camera. The transformer network estimates the odometry, the change in pose between images up to scale, from SuperPoint features matched by LightGlue. The MSCKF uses these pseudo-measurements along with an orbit and attitude kinematics model to estimate the pose of the target. In particular, the relative orbit elements, the target's attitude with respect to the servicer's camera, and the associated angular velocity are estimated directly by the filter. Given the monocular approach and short distance to the target, the full observability of the range to the target is recovered via attitude maneuvers by the servicer. The method is trained and evaluated on a re-rendered high-resolution version of the SPE3R dataset, which includes synthetic images of 103 spacecraft. Eleven of these spacecraft are held out during training to evaluate the generalization to unseen targets. Monte Carlo simulations are then used to evaluate the navigation pipeline on rendered trajectories of the held out spacecraft. The results demonstrate that learned vision pipelines as a front-end for Kalman filters provide median errors of 3.7° in attitude and 2.2% of range in ROE when navigating about unknown targets.

PR本紙発行元 EmplifAI