日本フィジカルAI新聞

世界のフィジカルAIを、日本語で。

週刊ニュースレター購読
音声対話arXiv:2609.36903

MultiTalk: 長時間・多人数・バイリンガル会話への全二重音声モデルの拡張

MultiTalk: Scaling Full-Duplex Speech Models to Long, Multi-Party, Bilingual Conversation

シェア:XThreadsFacebookLINEはてブBluesky

長時間・多人数・英中バイリンガルの全二重音声対話モデルを実現するため、57.6k時間の合成データセットと実録音ベースの評価ベンチマークを構築した研究。

詳しい要約

1. どんなもの?

- 長文・多人数・バイリンガル会話に拡張したfull-duplex音声モデル - Moshiパラダイムをlong-horizonとmulti-partyの2軸で拡張 - 英語と中国語に対応 - 合成学習データ57.6k時間(MultiTalkPT, MultiTalkFT)を公開 - 評価用MultiTalkBenchも導入

2. 先行研究と比べてどこがすごい?

- 既存のopen-source full-duplex音声モデルはlong-context robustnessとmulti-party interactionが限定的 - 既存のmulti-party音声コーパスは小規模でcodec-frame-level full-duplex modeling向けでない - 既存のlong-audioベンチマークはpassive listening中心、speech-to-speechベンチマークは短くdyadic - 本研究は長文・多人数・バイリンガルを同時に扱う点で先行研究を超える

3. 技術・手法の肝は?

- Moshi-styleのbilingualモデルを訓練 - 合成データ生成で長さ、参加者数、turn-taking、overlap、backchannels、interruptions、addressee shifts、long-range coreferenceを制御 - MultiTalkBenchは実人間録音から構築し、long-range entity tracking、topic coherence、addressee selectionのprobeを含む - 会話平均32.6分の長文多人数バイリンガルfull-duplex対話を評価

4. どうやって有効だと検証した?

- MultiTalkBench上で評価 - 提案モデルはMoshi、MiniCPM-o-4.5、Qwen3-Omni-30B-A3B-Instructなどのopen-sourceベースラインを大幅に上回る - 長文多人数英中会話で一貫性を維持できることを確認

5. 議論はある?

- データと評価の両面が制約だったと指摘 - 合成データの質や実世界への一般化可能性は要旨からは不明 - 長文多人数会話における課題(例: long-range coreference)への対処は示唆されるが、限界は要旨からは不明

6. 次に読むべき論文は?

- Moshi - MiniCPM-o-4.5 - Qwen3-Omni-30B-A3B-Instruct - MultiTalkPT, MultiTalkFT, MultiTalkBench

※ AIが要旨から生成した要約です。正確性は原文をご確認ください。

著者: Ke Wang, Houxing Ren, Zimu Lu, Yunqiao Yang, Zhuofan Zong, Mingjie Zhan, Hongsheng Li

分類: cs.CL, cs.AI, cs.SD

原文アブストラクト

End-to-end full-duplex speech models have brought open-source machine conversation closer to human-like interaction, yet existing systems remain limited in two intertwined dimensions: long-context robustness and multi-party interaction. Real-world scenarios such as meetings, group lessons, and social-robot reception require a single model to track, contextualize, and respond to multiple speakers over extended durations. Progress is constrained by both data and evaluation: open multi-party speech corpora remain small and are not designed for codec-frame-level full-duplex modeling, while existing long-audio benchmarks focus on passive listening and speech-to-speech benchmarks are mostly short and dyadic. We extend the Moshi paradigm jointly along the long-horizon and multi-party axes in English and Chinese. First, we release 57.6k hours of synthetic training data ($\href{https://huggingface.co/datasets/MultiTalk/MultiTalkPT}{MultiTalkPT}$ and $\href{https://huggingface.co/datasets/MultiTalk/MultiTalkFT}{MultiTalkFT}$) for long-form, multi-party, English-Chinese full-duplex dialogue, with controllable length, participant count, turn-taking, overlap, backchannels, interruptions, addressee shifts, and long-range coreference. Second, we introduce $\href{https://huggingface.co/datasets/MultiTalk/MultiTalkBench}{MultiTalkBench}$, built from real human recordings, for evaluating long-form, multi-party, bilingual full-duplex dialogue. Conversations average 32.6 minutes and include probes for long-range entity tracking, topic coherence, and addressee selection. Third, we train a bilingual Moshi-style model that sustains coherent multi-party English-Chinese conversations over extended durations and substantially outperforms open-source baselines including Moshi, MiniCPM-o-4.5, and Qwen3-Omni-30B-A3B-Instruct on MultiTalkBench.

関連論文

PR本紙発行元 EmplifAI