汎用セマンティックおよび目標指向通信のための基盤モデル
Foundation Models for Generalizable Semantic and Goal-Oriented Communication
視覚言語基盤モデルと拡散モデルを活用し、低ビットレートでも汎化性と高忠実度を両立するセマンティック通信フレームワークFMSGOCを提案した。
著者: Boliang Liu, Wint Yi Poe, Riccardo Trivisonno, Giuseppe Caire
分類: cs.LG, cs.AI, cs.RO, eess.IV
原文アブストラクト
Semantic and goal-oriented communication is increasingly studied for 6G, but generalization beyond seen data remains a key weakness under tight rate budgets. Many existing systems overfit their training data and degrade sharply at very low bit rates because they attempt to compress the entire signal. We introduce Foundation Model-Guided Semantic and Goal-Oriented Communication (FMSGOC), a framework that uses broad visual-linguistic Foundation Model priors to mitigate overfitting. It further improves rate efficiency by concentrating bits on sparse, goal-aligned anchors and relying on generative foundation-model priors to reconstruct the masked regions. By decoupling what to send from how to reconstruct, a vision-language foundation model selects and transmits a sparse set of semantic anchors, while a pretrained diffusion model, fine-tuned for masked completion, reconstructs the image at the receiver. In our experiments, FMSGOC reaches 0.039 bits per pixel (BPP), maintains high semantic fidelity (cosine similarity 0.87-0.90 on CIFAR-10), remains robust on previously unseen inputs (0.83-0.86 on ImageNet), and shows good perceptual similarity (0.1278/0.1558, CIFAR-10/ImageNet), outperforming strong end-to-end baselines at lower bit rates.