Multilingual Voice-Video Cloning with Multimodal Diffusion Model
Instruction: Please unmute manually 🔇.
Instruction: Please unmute manually 🔇.
Result
Instruction: Please compare the test speech with the Reference TTS Speech for their speech naturaness (semantic naturaness); and compare the test speech and the Reference Voice for their voice similarity.
Demo 1: but it's not enough (English)
Demo 2: she wanted to change policy at the government level (English)
Demo 3: und leichte flugzeuge fliegen weiter(German)
Demo 4: ese pedido es un mandato(Spanish)
Demo 5: elas faziam testes de aptidão musical (Portuguese)
As shown in Figure, in the first row, VITS exhibits duration asynchronization with the ground truth, resulting in significant differences. In the third row, we observe severe over-smoothing in FastSpeech 2, HPM, and V2C, causing a degradation in the reconstruction of details. In contrast, our results are closer to the ground truth, benefiting from enhanced detail reconstruction and duration synchronization capabilities of Diff-Dolly.
Instruction: Please compare the test video with the Reference video for their naturalness and lip sync (Please unmute manually 🔇).
Demo 1: Show abuse the light of day by talking about it with your children your coworkers your friends and (English)
Demo 2: I see now i never was one and not the other (English)
Demo 3: Donc, si j'essaye de faire un poème. (French)
Demo 4: denn ich weiß, ihr wollt das zeugdoch jetzt schon. gebt es zu. (German)
Demo 5: ou por não ser utilizada , porque não passou para a próxima geração (Portuguese)
As shown in Figure on speaker B, other methods produce blurry mouth while we shows more clear results. On speaker C, our results show fewer artifacts and blend more naturally with the surrounding skin. In contrast, IP-LAP displays blurry artifacts, and TalkLip generates boundary artifacts. Furthermore, in terms of lip-sync, Diff-Dolly shows improvements of 0.05 in LSE-D and 0.59 (28% relative improvement) in LMD compared to the second-best method, which indicates that our method achieves better synchronization. On speaker A, Diff-Dolly produces videos with realistic lip shapes and natural motions. In contrast, TalkLip shows significant differences in the pronunciation of $/w/$.
Specifically, we observe that the baselines tend to align the attributes of the reference video, resulting in reduced diversity of the generated content due to one-to-one mapping. Our Diff-Dolly incorporates only the landmark and background information from the reference video during inference, reducing dependency on the reference video and enhancing consistency with the source speaker's style. We also observe that DiffSwap, despite using the diffusion model, shows poor temporal consistency in videos focusing on image generation. In contrast, Diff-Dolly achieves strong temporal consistency in videos and remains compatible with image generation.
Instruction: Please unmute manually 🔇.
Result
Result
This website is licensed under a Creative Commons Attribution-ShareAlike 4.0 International License.
Thanks to StreamV2V for demo inspiration, and DreamBooth and Nerfies for website template.