日本フィジカルAI新聞

世界のフィジカルAIを、日本語で。

週刊ニュースレター購読
音声アシスタントarXiv:2609.21109

話しかけて、ジャービス:自律走行レースカー向けのオープンソース・エッジ展開可能な音声アシスタントフレームワーク

Talk to Me, Jarvis: An Open-Source Edge-Deployable Voice Assistant Framework for Autonomous Racecars

シェア:XThreadsFacebookLINEはてブBluesky

自律走行車の高レベルな行動指示を音声で行うためのオフライン音声アシスタント「Jarvis」を開発し、Mistral 7Bのドメイン特化ファインチューニングにより97.63%の意図認識精度と平均1.39秒の低遅延を実現した。

詳しい要約

1. どんなもの?

- 自律走行レースカー向けのオフライン音声アシスタント「Jarvis」を開発 - 音声認識・合成と自然言語コマンド分類を統合した軽量ローカルフレームワーク - エッジデプロイ可能で、高レベル行動コマンドを扱う - 中核はMistral 7Bをドメイン特化ファインチューニングしたテキスト・トゥ・コマンド分類器 - オープンソース実装を提供

2. 先行研究と比べてどこがすごい?

- 従来のオンラインホスト型LLMはネットワーク依存と推論遅延が課題 - 本手法はオフラインで動作し、より大規模なオンラインモデルを上回る性能 - 意図認識精度97.63%、平均処理遅延1.39秒を達成 - 時間-criticalな自律走行アプリケーションに適する

3. 技術・手法の肝は?

- Mistral 7Bモデルをドメイン特化でファインチューニング - 音声認識・合成と自然言語コマンド分類を統合 - 軽量ローカルフレームワークとしてエッジデプロイ可能 - 低遅延推論を実現

4. どうやって有効だと検証した?

- 実験的評価により、オンラインホスト型のより大規模なモデルを上回ることを示す - 意図認識精度97.63%、平均処理遅延1.39秒を達成 - 迅速な応答が求められる操作に適することを確認

5. 議論はある?

- オンラインモデルのネットワーク依存と遅延問題を解決 - オフラインで動作するため、時間-criticalなアプリケーションに適する - オープンソース実装を提供し、さらなる研究とファインチューニングを支援 - 具体的な議論や限界については要旨からは不明

6. 次に読むべき論文は?

- 要旨で参照/比較されている研究は明示されていない - 関連手法として、Mistral 7B、大規模言語モデル(LLM)を用いた音声アシスタント、オンラインホスト型モデルが挙げられる - 同分野の定番として、音声認識(ASR)、テキスト・トゥ・スピーチ(TTS)、意図分類(intent classification)に関する研究が考えられる

※ AIが要旨から生成した要約です。正確性は原文をご確認ください。

著者: Daniel Henel, Frederik Werner, Alexander Langmann, Johannes Betz

分類: cs.LG, cs.RO

原文アブストラクト

Recent advances in large language models have improved their effectiveness as back-end components for voice assistants, particularly in intent understanding and context-aware input classification. However, online-hosted models introduce network dependency and variable inference latency, limiting their suitability for time-critical autonomous driving applications. In this work, we address these issues by developing Jarvis, an offline voice assistant for high-level behavioral commands of autonomous vehicles. Its architecture integrates speech recognition and synthesis with natural language command classification into a lightweight, local framework. Jarvis core component is a text-to-command classifier, built using a domain-specific fine-tuning of the Mistral 7B model, demonstrating low-latency inference. Our experimental evaluation demonstrates that our solution outperforms larger online-hosted models, achieving 97.63 % intent recognition accuracy with an average processing latency of 1.39 s, making it well-suited for operations requiring quick response times. To support further research and fine-tuning, we provide an open-source implementation.

PR本紙発行元 EmplifAI