https://speechresearch.github.io/naturalspeech2/ NaturalSpeech 2: Latent Diffusion Models are Natural and Zero-Shot Speech and Singing Synthesizers [Paper] [Reddit Discussion] [Hack News] Kai Shen*, Zeqian Ju*, Xu Tan*, Yanqing Liu, Yichong Leng, Lei He, Tao Qin, Sheng Zhao, Jiang Bian Microsoft Research Asia & Microsoft Azure Speech Abstract. While text-to-speech (TTS) systems (e.g., NaturalSpeech) have achieved high speech quality on single-speaker recording-studio datasets, these datasets are not enough to capture the diversity in human speech such as speaker identities, prosodies, styles (e.g., singing). When scaling to large-scale, multi-speaker, and in- the-wild datasets, current TTS systems usually quantize speech into discrete tokens and use language models to generate these tokens one by one, which suffer from unstable prosody, word skipping/repeating issue, and poor voice quality. In this paper, we develop NaturalSpeech 2, a TTS system that uses a latent diffusion model to synthesize natural voices with high expressiveness/robustness/ fidelity and strong zero-shot ability. Specifically, we leverage a neural audio codec with residual vector quantizers to reconstruct speech waveform and get the quantized latent vectors, and then use a diffusion model to generate these latent vectors conditioned on text input. To enhance the zero-shot capability, we design a speech prompting mechanism to facilitate in-context learning in the duration /pitch predictor and diffusion model. We scale NaturalSpeech 2 to large-scale datasets with 44K hours of speech and singing data and evaluate its voice quality on unseen (zero-shot) speakers. NaturalSpeech 2 outperforms previous TTS systems by a large margin in terms of prosody/timbre similarity, robustness, and voice quality, and can perform novel zero-shot singing synthesis with only a speech prompt. This research is done in alignment with Microsoft's responsible AI principles. This page is for research demonstration purposes only. Overview [overview] NaturalSpeech 2 consists of an audio codec encoder/decoder and a latent diffusion model conditioned on a prior (a phoneme encoder and a duration/pitch predictor). [table1] The comparison between NaturalSpeech 2 and previous large-scale TTS systems. LibriSpeech Samples Text Prompt Ground Baseline NaturalSpeech Truth 2 Your Your Your Indeed, there were only one browser browser browser Your browser or two strangers who could does not does not does not does not be admitted among the support support support support the sisters without producing the the the audio the same result. audio audio audio element. element. element. element. Your Your Your browser browser browser Your browser For if he's anywhere on the does not does not does not does not farm, we can send for him in support support support support the a minute. the the the audio audio audio audio element. element. element. element. Their piety would be like their names, like their faces, like their clothes, Your Your Your and it was idle for him to browser browser browser Your browser tell himself that their does not does not does not does not humble and contrite hearts support support support support the it might be paid a the the the audio far-richer tribute of audio audio audio element. devotion than his had ever element. element. element. been. A gift tenfold more acceptable than his elaborate adoration. Come, come returned Hawkeye, uncasing his honest countenance, the better to Your Your Your assure the wavering browser browser browser Your browser confidence of his companion. does not does not does not does not You may see a skin which, if support support support support the it be not as white as one of the the the audio the gentle ones, has no audio audio audio element. tinge of red to it that the element. element. element. winds of the heaven and the sun have not bestowed. Now, let us to business. Your Your Your The air and the earth are browser browser browser Your browser curiously mated and does not does not does not does not intermingled as if the one support support support support the were the breath of the the the the audio other. audio audio audio element. element. element. element. I had always known him to be Your Your Your restless in his manner, but browser browser browser Your browser on this particular occasion does not does not does not does not he was in such a state of support support support support the uncontrollable agitation the the the audio that it was clear something audio audio audio element. very unusual had occurred. element. element. element. Your Your Your browser browser browser Your browser His death in this does not does not does not does not conjuncture was a public support support support support the misfortune. the the the audio audio audio audio element. element. element. element. Your Your Your browser browser browser Your browser It is this that is of does not does not does not does not interest to theory of support support support support the knowledge. the the the audio audio audio audio element. element. element. element. Your Your Your For a few miles, she browser browser browser Your browser followed the line hitherto does not does not does not does not presumably occupied by the support support support support the coast of Algeria, but no the the the audio land appeared to the south. audio audio audio element. element. element. element. VCTK Samples Text Prompt Ground Truth Baseline NaturalSpeech 2 Your browser Your browser Your browser Your browser It is an does not does not does not does not absolute support the support the support the support the nonsense. audio audio audio audio element. element. element. element. Your browser Your browser Your browser Your browser Truth is the does not does not does not does not child of time. support the support the support the support the audio audio audio audio element. element. element. element. Your browser Your browser Your browser Your browser We have a long does not does not does not does not way to go this support the support the support the support the week. audio audio audio audio element. element. element. element. Maybe we Your browser Your browser Your browser Your browser expected too does not does not does not does not much from the support the support the support the support the fixture. audio audio audio audio element. element. element. element. Your browser Your browser Your browser Your browser We will turn does not does not does not does not the corner. support the support the support the support the audio audio audio audio element. element. element. element. It will also Your browser Your browser Your browser Your browser require a does not does not does not does not lengthy series support the support the support the support the of clinical audio audio audio audio trials. element. element. element. element. Your browser Your browser Your browser Your browser Subs not used, does not does not does not does not McKenzie, support the support the support the support the Ritchie. audio audio audio audio element. element. element. element. Your browser Your browser Your browser Your browser That's the only does not does not does not does not thing I will support the support the support the support the say. audio audio audio audio element. element. element. element. Your browser Your browser Your browser Your browser Every aspect of does not does not does not does not our play was support the support the support the support the first class. audio audio audio audio element. element. element. element. Samples compared with VALL-E Text Prompt Ground VALL-E NaturalSpeech Truth 2 Your Your Your Your browser And lay me down in my browser browser browser does not cold bed and leave my does not does not does not support the shining lot. support support support audio the audio the audio the audio element. element. element. element. Yea, his honourable Your Your Your Your browser worship is within, but browser browser browser does not he hath a godly minister does not does not does not support the or two with him, and support support support audio likewise a leech. the audio the audio the audio element. element. element. element. Your Your Your Your browser The army found the browser browser browser does not people in poverty and does not does not does not support the left them in comparative support support support audio wealth. the audio the audio the audio element. element. element. element. Thus did this humane and Your Your Your right minded father browser browser browser Your browser comfort his unhappy does not does not does not does not daughter, and her mother support support support support the embracing her again, did the audio the audio the audio audio all she could to soothe element. element. element. element. her feelings. Singing Samples Text Prompt Prompt NaturalSpeech 2 Type So listen Your browser does not Your browser does not very Speech support the audio support the audio carefully. element. element. So listen Your browser does not Your browser does not very Singing support the audio support the audio carefully. element. element. Text Prompt Prompt NaturalSpeech 2 Type To keep it all Your browser does not Your browser does not in like you Speech support the audio support the audio do. element. element. To keep it all Your browser does not Your browser does not in like you Singing support the audio support the audio do. element. element. Text Prompt Prompt NaturalSpeech 2 Type In the middle Your browser does not Your browser does not of the night. Speech support the audio support the audio element. element. In the middle Your browser does not Your browser does not of the night. Singing support the audio support the audio element. element. Ethics Statement NaturalSpeech 2 can synthesize speech with good expressiveness/ fidelity and good similarity with a speech prompt, which could be potentially misused, such as speaker mimicking and voice spoofing. To avoid potential issues, we appeal to our practitioners to not abuse this technology and to develop defending tools to detect AI-synthesized voices. We will always take Microsoft AI Principles as guidelines to develop such AI models. Other Related Works NaturalSpeech: End-to-End Text to Speech Synthesis with Human-Level Quality Speech research conducted at Microsoft Research Asia