[HN Gopher] PlayHT2.0: State-of-the-Art Generative Voice AI Mode...
       ___________________________________________________________________
        
       PlayHT2.0: State-of-the-Art Generative Voice AI Model for
       Conversational Speech
        
       Author : smusamashah
       Score  : 23 points
       Date   : 2023-08-11 17:21 UTC (5 hours ago)
        
 (HTM) web link (news.play.ht)
 (TXT) w3m dump (news.play.ht)
        
       | tasubotadas wrote:
       | What models/architecture are they using?
        
         | kastnerkyle wrote:
         | Previously TortoiseTTS was associated with PlayHT in some way,
         | although the exact connection is a bit vague [0].
         | 
         | From the descriptions here it sounds a lot like AudioLM / SPEAR
         | TTS / some of Meta's recent multilingual TTS approaches,
         | although those models are not open source, sounds like PlayHT's
         | approach is in a similar spirit. The discussion of "mel tokens"
         | is closer to what I would call the classic TTS pipeline in many
         | ways... PlayHT has generally been kind of closed about what
         | they used, would be interesting to know more.
         | 
         | If you are interested in some recent open to sample-from work
         | pushing on this kind of random expressiveness (sometimes at the
         | expense of typical "quality" in terms of TTS), Bark is pretty
         | interesting [1]. Though the audio quality suffers a bit from
         | how they realize sequences -> waveforms, the prosody and timing
         | is really interesting.
         | 
         | I assume the key factor here is high quality, emotive audio
         | with good data cleaning processes. Probably not even a lot of
         | data, at least in the scale of "a lot" in speech, e.g. ASR
         | (millions of hours) or TTS (hundreds to thousands). As opposed
         | to some radically new architectural piece never before seen in
         | the literature, there are lots of really nice tools for emotive
         | and expressive TTS buried in recent years of publications.
         | 
         | Tacotron 2 is perfectly capable of this type of stuff as well,
         | as shown by Dessa [2] a few years ago (this writeup is a nice
         | intro to TTS concepts). With the limit largely being, at some
         | point you haven't heard certain phonetic sounds before in a
         | voice, and need to do _something_ to get plausible outcomes for
         | new voices.
         | 
         | [0] Discussion here https://github.com/neonbjb/tortoise-
         | tts/issues/182#issuecomm...
         | 
         | [1]
         | https://www.tiktok.com/@jonathanflyfly/video/722513498370947...
         | 
         | [1a] Bark github https://github.com/suno-ai/bark
         | 
         | [2] https://medium.com/dessa-news/realtalk-how-it-
         | works-94c1afda...
        
           | narrationbox wrote:
           | Mel + multispeaker vocoder is very much a classic (tacotron
           | era) TTS approach
        
       | jasonjmcghee wrote:
       | Self-proclaimed state of the art. A year ago, i would have been
       | blown away, today, this is dramatically worse than Eleven Labs.
       | Lower quality audio, strange cadence, pretty monotonic. It's not
       | what people sound like.
       | 
       | I think it's impressive, but i wouldn't call it state of the art.
        
       | ilaksh wrote:
       | It says closed alpha, but also says available through the API. Is
       | it closed or open now?
        
         | sattoshi wrote:
         | I'm guessing it's closed because you can't download it. It's
         | only available through the API.
        
       ___________________________________________________________________
       (page generated 2023-08-11 23:01 UTC)