日本フィジカルAI新聞

世界のフィジカルAIを、日本語で。

週刊ニュースレター購読
口唇同期arXiv:2608.18832

EfficientSync: 変形ベースの参照テクスチャ混合によるリアルタイム口唇同期

EfficientSync: Real-Time Lip Synchronization via Deformation-Based Reference Texture Mixing

シェア:XThreadsFacebookLINEはてブBluesky

音声駆動の口唇同期を、参照フレームのテクスチャを保持したまま変形ベースでリアルタイムに実現する新しいフレームワークを提案した論文。

詳しい要約

1. どんなもの?

EfficientSyncは、音声駆動の口唇同期(lip synchronization)をリアルタイムで行うフレームワーク。入力音声に合わせて話者動画の口元領域を編集し、頭部姿勢・同一性・背景を保持する。従来手法がGANや拡散モデルで下顔全体を再構築するのに対し、EfficientSyncは参照フレームの実テクスチャを変形ベースで保持・混合することで、歯や唇のしわなどの口腔内詳細の幻覚を防ぎ、高速処理(単一GPUで166 FPS)を実現する。

2. 先行研究と比べてどこがすごい?

従来手法は下顔全体をGANや拡散モデルで再構築するため、レイテンシが大きく、歯や唇のしわなどの口腔内詳細を幻覚し、本来のテクスチャを保存できない。EfficientSyncは、同一性保存のボトルネックは参照フレームの不足ではなく、参照フレームに含まれる本物のテクスチャを忠実に転送するメカニズムの欠如にあると主張し、再合成ではなく参照テクスチャを保持する変形ベースのアプローチを提案。これにより、テクスチャの整合性を低コストで維持しつつ、リアルタイム処理を達成している点が優れている。

3. 技術・手法の肝は?

手法の核は3つの要素からなる。1) Dynamic Texture Mixer: 複数参照フレームの融合をチャネル単位の選択問題として再定式化し、各参照を空間的に整列させてグローバル文脈で評価し、チャネル単位の重み付き和で集約することで、テクスチャの整合性を低コストで保持する。2) Spatio-Temporal Shifted Adaptive Masking: ソースフレームを口唇生成条件と独立した背景事前情報に分解し、下顔領域の漏れを抑制しながら、合成された口を背景にシームレスにブレンドする。3) STAR Sampling: ゼロオーバーヘッドの前処理ステップで、最も鮮明でトポロジー的に多様な参照フレームを取得する。

4. どうやって有効だと検証した?

HDTFおよびVFHQデータセットを用いて実験を実施。視覚品質と同一性保存の指標で最先端(state-of-the-art)の結果を示し、単一GPU上で166 FPSのリアルタイム処理を達成した。定量的評価に加え、動画デモを公開している。

5. 議論はある?

要旨からは、議論の詳細は不明。ただし、従来手法の幻覚問題を指摘し、テクスチャ保存の重要性を主張している点が議論の中心と考えられる。また、リアルタイム性と品質のトレードオフや、参照フレーム選択の影響などが今後の課題として考えられるが、要旨には明記されていない。

6. 次に読むべき論文は?

要旨で参照・比較されている研究は明示されていないが、音声駆動の口唇同期分野の定番として、Wav2Lip、PC-AVS、Diff2Lip、SadTalkerなどが関連する。また、変形ベースの手法としては、First Order Motion ModelやThin-Plate Spline Motion Modelなどが関連する。

※ AIが要旨から生成した要約です。正確性は原文をご確認ください。

著者: Fa-Ting Hong, Runzhen Liu, Luchuan Song, Hongmin Cai, Chuhua Xian

分類: cs.CV

原文アブストラクト

Audio-driven lip synchronization manipulates the mouth region of a talking-face video to match the driving audio while preserving head pose, identity, and background. Although the task is inherently local editing, prevailing approaches reconstruct the entire lower face with heavy GAN- or diffusion-based decoders, incurring substantial latency and, more critically, hallucinating intra-oral details such as teeth and lip wrinkles instead of preserving authentic textures. We contend that the bottleneck in identity preservation is not the scarcity of reference frames, but the lack of a mechanism that faithfully transfers the genuine textures they already contain. We therefore present EfficientSync, a real-time deformation-based framework that retains reference textures rather than resynthesizing them. First, the Dynamic Texture Mixer reformulates multi-reference fusion as channel-wise selection, evaluating each spatially aligned reference in a global context and aggregating them by channel-wise weighted summation, preserving textural integrity at low cost. Second, Spatio-Temporal Shifted Adaptive Masking decomposes the source frame into lip-generation conditions and an independent background prior, suppressing lower-face leakage while blending the synthesized mouth seamlessly into the background. Third, STAR Sampling, a zero-overhead pre-processing step, retrieves the sharpest and most topologically diverse reference frames. Experiments on HDTF and VFHQ show state-of-the-art visual quality and identity preservation at 166 FPS on a single GPU. Video demos: https://alunaticat.github.io/EfficientSync/index.html.