Multilingual Voice-Video Cloning with Multimodal Diffusion Model

What can Diff-Dolly do?

1. Multilingual Voice Cloning

Input Ref Speech

English

German

Spanish

French

Portuguese

Russain

Input Ref Speech

English

German

Spanish

French

Portuguese

Russain


2. Video Lip Cloning

Instruction: Please unmute manually 🔇.

English

German

Spanish

French

Portuguese

Russain

English

German

Spanish

French

Portuguese

Russain


3. Video Face Cloning

Instruction: Please unmute manually 🔇.

Text

aber auch eine ganz andere, eine sehr persönliche bedeutung.

Speech

Facial Image

Talking Video

Result


Main Results

Qualitative Results

1. Voice Cloning Results

Instruction: Please compare the test speech with the Reference TTS Speech for their speech naturaness (semantic naturaness); and compare the test speech and the Reference Voice for their voice similarity.

Demo 1: but it's not enough (English)

GT Speeech:

Reference Voice:

Ours

VITS

FastSpeech 2

HPM

V2C

Demo 2: she wanted to change policy at the government level (English)

GT Speeech:

Reference Voice:

Ours

VITS

FastSpeech 2

HPM

V2C

Demo 3: und leichte flugzeuge fliegen weiter(German)

GT Speeech:

Reference Voice:

Ours

VITS

FastSpeech 2

HPM

V2C

Demo 4: ese pedido es un mandato(Spanish)

GT Speeech:

Reference Voice:

Ours

VITS

FastSpeech 2

HPM

V2C

Demo 5: elas faziam testes de aptidão musical (Portuguese)

GT Speeech:

Reference Voice:

Ours

VITS

FastSpeech 2

HPM

V2C

See Mel-Spectrogram Comparison

As shown in Figure, in the first row, VITS exhibits duration asynchronization with the ground truth, resulting in significant differences. In the third row, we observe severe over-smoothing in FastSpeech 2, HPM, and V2C, causing a degradation in the reconstruction of details. In contrast, our results are closer to the ground truth, benefiting from enhanced detail reconstruction and duration synchronization capabilities of Diff-Dolly.


2. Lip Motion Cloning Results

Instruction: Please compare the test video with the Reference video for their naturalness and lip sync (Please unmute manually 🔇).

Demo 1: Show abuse the light of day by talking about it with your children your coworkers your friends and (English)

Demo 2: I see now i never was one and not the other (English)

Demo 3: Donc, si j'essaye de faire un poème. (French)

Demo 4: denn ich weiß, ihr wollt das zeugdoch jetzt schon. gebt es zu. (German)

Demo 5: ou por não ser utilizada , porque não passou para a próxima geração (Portuguese)

See Image Comparison

As shown in Figure on speaker B, other methods produce blurry mouth while we shows more clear results. On speaker C, our results show fewer artifacts and blend more naturally with the surrounding skin. In contrast, IP-LAP displays blurry artifacts, and TalkLip generates boundary artifacts. Furthermore, in terms of lip-sync, Diff-Dolly shows improvements of 0.05 in LSE-D and 0.59 (28% relative improvement) in LMD compared to the second-best method, which indicates that our method achieves better synchronization. On speaker A, Diff-Dolly produces videos with realistic lip shapes and natural motions. In contrast, TalkLip shows significant differences in the pronunciation of $/w/$.


3. Face Cloning Results

See Image Comparison

Specifically, we observe that the baselines tend to align the attributes of the reference video, resulting in reduced diversity of the generated content due to one-to-one mapping. Our Diff-Dolly incorporates only the landmark and background information from the reference video during inference, reducing dependency on the reference video and enhancing consistency with the source speaker's style. We also observe that DiffSwap, despite using the diffusion model, shows poor temporal consistency in videos focusing on image generation. In contrast, Diff-Dolly achieves strong temporal consistency in videos and remains compatible with image generation.


4. Unified Voice-Video Cloning Results

Instruction: Please unmute manually 🔇.

Text

aber auch eine ganz andere, eine sehr persönliche bedeutung.

Speech

Facial Image

Talking Video

Result

Text

weil es einfach eine problematik ist,die mich selbst betroffen hat.

Speech

Facial Image

Talking Video

Result

This website is licensed under a Creative Commons Attribution-ShareAlike 4.0 International License.

Thanks to StreamV2V for demo inspiration, and DreamBooth and Nerfies for website template.