日本フィジカルAI新聞

世界のフィジカルAIを、日本語で。

週刊ニュースレター購読
故障診断/LLMarXiv:2609.20620

AUV故障復旧のためのシミュレーションプラットフォーム:LLMベース診断戦略の探求

A Simulation Platform for AUV Fault Recovery: Exploring LLM-Based Diagnostic Strategies

シェア:XThreadsFacebookLINEはてブBluesky

AUVの異常時にLLMを診断・復旧プランナとして呼び出すアーキテクチャを検討し、物理ベースの故障注入とLLM評価を統合した閉ループシミュレータSPARを構築して、質量移動故障に対する複数LLMの診断性能を比較した。

詳しい要約

1. どんなもの?

AUVの故障回復を目的に、通常運用は従来のdeterministic layered control autonomyが担い、onboard anomaly detectionが想定外を検知した際にLLMをdiagnostic/recovery plannerとして呼び出すアーキテクチャを検討。これを評価するclosed-loop simulation architecture『SPAR (Simulation Platform for AUV Recovery)』を提案。real-time C vehicle softwareと上位orchestration layerを結合し、physics-based fault injection、structured prompting、language-model interaction、mission file generation、validation、execution、LLM-judge scoringを行う。

2. 先行研究と比べてどこがすごい?

従来のAUV自律制御はdeterministic layered controlが中心で、想定外故障の回復は検知止まりか人手依存だった。本研究はLLMをdiagnostic/recovery plannerとして組み込み、検知からmitigationまで拡張するarchitectureを提示。さらにLLMのstochastic性を踏まえ、個別デモではなくensemble testingで厳密評価する方法論を提案した点が新しい。

3. 技術・手法の肝は?

SPARはreal-time C vehicle softwareとhigher-level orchestration layerを結合したclosed-loop simulation。physics-based fault injection、structured prompting、language-model interaction、mission file generation、validation、execution、LLM-judge scoringを統合。fault realizations、prompt structures、reasoning models、mission conditionsを変化させ、mass-shift faultで480 trialsを実施。frontier modelとoff-the-shelf locally deployable LLM 3種を比較。

4. どうやって有効だと検証した?

mass-shift faultに対し480 SPAR trialsを実施。fault realizations、prompt structures、reasoning models、mission conditionsを変化させ、frontier modelとlocal LLM 3種を評価。Model choiceがdiagnosisを支配し、frontier modelはCG-shift mechanismをtop three hypothesesに85-90%で含めるが、best local modelは60-78%。local-model成功はcomplete diagnostic procedureの遵守と関連し、弱いモデルはactuatorがcommandに追従しているのにelevator failureへ早合点する傾向。diagnosisとoperational decision performanceは本datasetではcoupleしていない。

5. 議論はある?

LLMのstochastic性のためrigorous evaluationにはensemble testingが必要と主張。Model choiceがdiagnosis性能を支配し、local modelでもcomplete diagnostic procedure遵守で成功が関連。一方、diagnosisとoperational decision performanceはcoupleしない可能性が示唆され、弱いモデルはpremature commitmentする傾向。ただし要旨からは具体的な議論の詳細や限界は不明。

6. 次に読むべき論文は?

要旨で参照/比較されているのはfrontier modelとoff-the-shelf locally deployable LLM 3種、およびconventional deterministic layered control autonomy。関連手法としてLLM-based diagnostic/recovery planning、ensemble testing、LLM-judge scoring、physics-based fault injectionが挙げられる。同分野の定番としてAUV fault detection and recovery、layered control autonomy、LLM-assisted mission managementが次に読むべき候補。

※ AIが要旨から生成した要約です。正確性は原文をご確認ください。

著者: Khalid Halba, Kylie Cooper, James G. Bellingham

分類: cs.RO, cs.AI

原文アブストラクト

Autonomous underwater vehicles (AUVs) operating beyond reliable communications must recover from failures without human intervention. We investigate an architecture in which conventional deterministic layered control autonomy manages normal operations, while an invokable large language model (LLM) serves as a diagnostic and recovery planner when onboard anomaly detection identifies performance outside expected limits. Because language models are stochastic, rigorous evaluation requires ensemble testing rather than individual demonstrations. We present a closed-loop simulation architecture that couples real-time C vehicle software with a higher-level orchestration layer for physics-based fault injection, structured prompting, language-model interaction, mission file generation, validation, execution, and LLM-judge scoring. The framework, which we call SPAR (Simulation Platform for AUV Recovery), supports evaluation across fault realizations, prompt structures, reasoning models, and mission conditions. We vary these for a mass-shift fault over 480 SPAR trials, evaluating a frontier model and three off-the-shelf locally deployable LLMs. Model choice dominates diagnosis: the frontier model places the CG-shift mechanism in its top three hypotheses in 85-90% of trials, versus 60-78% for the best local model. Reasoning analysis indicates that local-model success is associated with following the complete diagnostic procedure, whereas weaker models often commit prematurely to elevator failure even though the actuator tracks its command. Diagnosis and operational decision performance do not appear to be coupled in this dataset. The contributions are an architecture extending unanticipated-fault recovery from detection to mitigation and an ensemble methodology for evaluating LLM-assisted mission management on low-power AUVs.

PR本紙発行元 EmplifAI