[HN Gopher] PlayHT2.0: State-of-the-Art Generative Voice AI Mode...
___________________________________________________________________
PlayHT2.0: State-of-the-Art Generative Voice AI Model for
Conversational Speech
Author : smusamashah
Score : 23 points
Date : 2023-08-11 17:21 UTC (5 hours ago)
(HTM) web link (news.play.ht)
(TXT) w3m dump (news.play.ht)
| tasubotadas wrote:
| What models/architecture are they using?
| kastnerkyle wrote:
| Previously TortoiseTTS was associated with PlayHT in some way,
| although the exact connection is a bit vague [0].
|
| From the descriptions here it sounds a lot like AudioLM / SPEAR
| TTS / some of Meta's recent multilingual TTS approaches,
| although those models are not open source, sounds like PlayHT's
| approach is in a similar spirit. The discussion of "mel tokens"
| is closer to what I would call the classic TTS pipeline in many
| ways... PlayHT has generally been kind of closed about what
| they used, would be interesting to know more.
|
| If you are interested in some recent open to sample-from work
| pushing on this kind of random expressiveness (sometimes at the
| expense of typical "quality" in terms of TTS), Bark is pretty
| interesting [1]. Though the audio quality suffers a bit from
| how they realize sequences -> waveforms, the prosody and timing
| is really interesting.
|
| I assume the key factor here is high quality, emotive audio
| with good data cleaning processes. Probably not even a lot of
| data, at least in the scale of "a lot" in speech, e.g. ASR
| (millions of hours) or TTS (hundreds to thousands). As opposed
| to some radically new architectural piece never before seen in
| the literature, there are lots of really nice tools for emotive
| and expressive TTS buried in recent years of publications.
|
| Tacotron 2 is perfectly capable of this type of stuff as well,
| as shown by Dessa [2] a few years ago (this writeup is a nice
| intro to TTS concepts). With the limit largely being, at some
| point you haven't heard certain phonetic sounds before in a
| voice, and need to do _something_ to get plausible outcomes for
| new voices.
|
| [0] Discussion here https://github.com/neonbjb/tortoise-
| tts/issues/182#issuecomm...
|
| [1]
| https://www.tiktok.com/@jonathanflyfly/video/722513498370947...
|
| [1a] Bark github https://github.com/suno-ai/bark
|
| [2] https://medium.com/dessa-news/realtalk-how-it-
| works-94c1afda...
| narrationbox wrote:
| Mel + multispeaker vocoder is very much a classic (tacotron
| era) TTS approach
| jasonjmcghee wrote:
| Self-proclaimed state of the art. A year ago, i would have been
| blown away, today, this is dramatically worse than Eleven Labs.
| Lower quality audio, strange cadence, pretty monotonic. It's not
| what people sound like.
|
| I think it's impressive, but i wouldn't call it state of the art.
| ilaksh wrote:
| It says closed alpha, but also says available through the API. Is
| it closed or open now?
| sattoshi wrote:
| I'm guessing it's closed because you can't download it. It's
| only available through the API.
___________________________________________________________________
(page generated 2023-08-11 23:01 UTC)