[HN Gopher] Eleven v3
       ___________________________________________________________________
        
       Eleven v3
        
       Author : robertvc
       Score  : 112 points
       Date   : 2025-06-05 18:41 UTC (4 hours ago)
        
 (HTM) web link (elevenlabs.io)
 (TXT) w3m dump (elevenlabs.io)
        
       | artninja1988 wrote:
       | Sounds absolutely amazing, like 99% indistinguishable from real
       | professional voice actors to me. I couldn't find any pricing
       | though. Anyone know what they charge for it?
        
         | minimaxir wrote:
         | > Public API for Eleven v3 (alpha) is coming soon. For early
         | access, please contact sales.
         | 
         | I suspect they themselves don't know the exact pricing yet and
         | want to assess demand first.
        
         | delgaudm wrote:
         | Ouch. Professional Voice Actor here.
        
           | razemio wrote:
           | Just here to say the oposite. It is astonshing how far away
           | it still is from a professional voice actor while being
           | really good. Emotion is completely missing. Instead it seems
           | to try to hard to express exactly that. I cant really put my
           | finger on it. It feels predictable, flat and the timing is
           | strange.
        
             | mrkstu wrote:
             | Better by a mile than most anime voice work, but lacks the
             | detail that a good voice narrator has on an audio book.
        
       | minimaxir wrote:
       | > Eleven v3 is 80% off until the end of June 2025 for self-serve
       | users using it through the UI.
       | 
       | That's definitely one way to loss-lead.
        
         | lostmsu wrote:
         | Open source stuff like Kokoro and the recent Chatterbox are hot
         | on their heels.
         | 
         | https://www.reddit.com/r/MachineLearning/comments/1kxv01f/p_...
        
           | minimaxir wrote:
           | It's definitely a response to Chatterbox, which is very
           | funny.
        
       | lostmsu wrote:
       | Hm, is it good in all languages? Russian sounds very robotic.
        
         | lharries wrote:
         | It's a research preview for now but it should work well in 70+
         | languages. Voices make a big difference, can you try with a few
         | Russian IVCs?
        
         | GrayShade wrote:
         | Romanian sounds awful too, like the TTSes from 15 years ago.
        
           | lharries wrote:
           | can you try with a Romanian voice?
        
             | GrayShade wrote:
             | I'm not sure what you mean. I chose Romanian from the
             | language selector and tried Matilda, Alice and Laura. Laura
             | actually sounds like an English TTS trying to pronounce
             | Romanian.
        
               | gozzoo wrote:
               | Exactly the same thing with Bulgarian voices.
        
               | lharries wrote:
               | It should work well for the logged in voices here:
               | https://elevenlabs.io/app/voice-library?language=ro
               | 
               | We are in the process of updated the homepage voices for
               | the new languages
        
         | NewMountain wrote:
         | There's something very wrong with the Russian one. The first
         | example "Jessica | Tell History", is British woman speaking
         | British English transliterated from Russian. It's absolute
         | murder of the Russian language and painful to listen to.
         | 
         | The second example "Jessica | Record a commercial" is perfect.
         | Confidence restored.
         | 
         | The third example "Laura | Help a client" is back to glass in
         | your ears. This time an American is speaking American English
         | transliterated from Russian.
         | 
         | Yikes. The English sounded fine, but the Russian has serious
         | issues. Either there's a bug in your configuration (I hope) or
         | your evals for Russian are unsound.
         | 
         | Edit: dial back the editorializing.
        
       | wewewedxfgdf wrote:
       | I did not see an British accent example.
       | 
       | Generally it appears the TTS systems all do US accents and the
       | British accent tends to sound like Frasier - an American faking
       | an British accent.
        
         | lharries wrote:
         | We have lots of great British voices in our voice library! Or
         | if you want to hear an american trying to do a british accent
         | add "[British accent]" at the start of the generation
        
           | not_your_mentat wrote:
           | I kept an English prompt, selected a French voice, and was
           | delighted to hear an British English woman. :shrug:
        
             | lharries wrote:
             | If you'd like it to sound like a french person speaking
             | french this voice works great:
             | https://elevenlabs.io/app/voice-
             | library?voiceId=xTZlmU8dKXdy...
             | 
             | Or if you want a french person speaking english with a
             | french accent use that voice with "[French accent]" before
             | it
        
           | wewewedxfgdf wrote:
           | It would be good if your demos made it more obvious. There's
           | a vast arrays of AI developments wanting me to check them out
           | - you have seconds to get my attention.
        
         | fakedang wrote:
         | ElevenLabs v2's accented voices are still much stronger than
         | any of its competition. And I've tried it with Arabic, French,
         | Hindi and English.
        
       | drag0s wrote:
       | English sounds really great, congrats! other languages I've tried
       | doesn't sound that good, you can hear a strong english accent
        
         | dustincoates wrote:
         | The French one sounded like an Alabaman who took a semester of
         | college French.
         | 
         | But the English sounds really good.
        
           | lharries wrote:
           | If you're trying to make an audiobook about an Alabaman
           | visiting Paris this might be quite useful... But in
           | seriousness try it with this voice:
           | https://elevenlabs.io/app/voice-
           | library?voiceId=rbFGGoDXFHtV...
        
             | dustincoates wrote:
             | I'll give it a check. I was playing the sample on the v3
             | page.
        
         | lharries wrote:
         | Can you try with a voice that was trained on that language?
         | This research preview is more variable based on the voice
         | chosen
        
         | k__ wrote:
         | German sounds okay.
        
           | lharries wrote:
           | There's lots of great german voices here which should be
           | better: https://elevenlabs.io/app/voice-
           | library/collections/SHEPnUB9...
           | 
           | The voice selection matters a lot for this research preview
        
           | shafyy wrote:
           | I tried German in the preview box there, and it had a very
           | strong English accent.
        
         | 8f2ab37a-ed6c wrote:
         | With Italian, it starts reading the text with an absolutely
         | comical American accent, but then about 10-20 words in it
         | gradually snaps into a natural Italian pronunciation and it
         | sounds fantastic from that point on. Not sure what's going on
         | behind the scenes, but it sounds like it starts with an en-us
         | baseline and then somehow zones in on the one you specified.
         | Using Alice.
        
       | ianbicking wrote:
       | I've been using OpenAI's new models a lot lately
       | (https://www.openai.fm/)... separating instructions from the
       | spoken word is an interesting choice, and I'm assuming also has a
       | lot to do with OpenAI/GPT using "instructions" across their
       | products, and maybe they are just more comfortable and familiar
       | generating the data and do the training for that style.
       | 
       | Separate instructions is a bit awkward, but does allow mixing
       | general instructions with specific instructions. Like I can
       | concatenate output-specific instructions like "voice lowers to a
       | whisper after 'but actually', and a touch of fear" with a general
       | instruction like "a deep voice with a hint of an English accent"
       | and it mostly figures it out.
       | 
       | The result with OpenAI feels much less predictable and of lower
       | production quality than Eleven Labs. But the range of prosidy is
       | much larger, almost overengaged. The range of _voices_ is much
       | smaller with OpenAI... you can instruct the voices to sound
       | different, but it feels a little like the same person doing
       | different voices.
       | 
       | But in the end OpenAI's biggest feature is that it's 10x cheaper
       | and completely pay-as-you-go. (Why are all these TTS services
       | doing subscriptions on top of limits and credits? Blech!)
        
         | lharries wrote:
         | > The result with OpenAI feels much less predictable and of
         | lower production quality than ElevenLabs
         | 
         | Thank you Ian! Credit to our research team for making this
         | possible
         | 
         | For the prosidy, if you choose an expressive voice the prosidy
         | should be larger
        
           | Velorivox wrote:
           | The word is "prosody", right?
        
       | zamadatix wrote:
       | The (American English) voices are absolutely amazing but the tags
       | for laughs still feel more like an "inserted dedicated laugh
       | section" than a "laugh at this point in speaking" type thing.
       | I.e. it can't seem to reliably know when to giggle while saying a
       | word, "just" giggle leading up to a word.
        
         | echelon wrote:
         | They're also still too expensive, and that's creating a lot of
         | opportunity for other players.
         | 
         | Even though ElevenLabs remains the quality leader, the others
         | aren't that far behind.
         | 
         | There are even a bunch of good TTS models being released as
         | fully open source, especially by cutting-edge Chinese labs and
         | companies. Perhaps in a bid to cut off the legs of American AI
         | companies or to commoditize their compliment. Whatever the
         | case, it's great for consumers.
         | 
         | YCombinator-backed PlayHT has been releasing some of their good
         | stuff too.
        
           | taf2 wrote:
           | What would say are some of the best open source TTS -
           | chatterbox maybe?
        
             | jsemrau wrote:
             | I had good results with Nemo + xTTS_v2
             | 
             | https://docs.nvidia.com/nemo-framework/user-
             | guide/latest/nem...
             | 
             | https://huggingface.co/coqui/XTTS-v2
        
           | monkeywork wrote:
           | could you list 2 or 3 of the ones you think are best quality
           | to $?
        
         | lharries wrote:
         | If you edit the text so that laugh makes sense in the context
         | it should be much more natural like this one:
         | https://x.com/elevenlabsio/status/1930689782331412811
        
           | zamadatix wrote:
           | The first laugh in that "<LAUGHS> Hey, Dr. Von Fusion" is a
           | dedicated laugh section, which the model does extremely well,
           | but it works because that's a natural place to laugh before
           | actually speaking the following words. Skip ahead to
           | "...robot chuckle. Jessica: <LAUGHS> I know right!" and you
           | get an awkwardly time/toned light chuckle completely
           | separated from the "I know" you'd naturally continue saying
           | while making that chuckle.
           | 
           | You can always rewrite the text to avoid times where one
           | would naturally laugh through the next couple of following
           | words but that's just attempting to avoid the problem and do
           | a different kind of laugh instead.
        
             | Davidzheng wrote:
             | have to say that this human can't tell the difference
             | between this and other real humans so...
        
       | carlosjobim wrote:
       | Their non-English (automated?) localization of the front page is
       | ridiculously badly translated.
        
         | lharries wrote:
         | Which language isn't good and I'll get that fixed asap?
        
           | carlosjobim wrote:
           | You need native or at least fluent speakers to help you, to
           | get the expressions right. For example Swedish is written
           | like a word-for-word translation from English.
        
       | sojuz151 wrote:
       | Polish is quite good, expected based on the founders' background
        
       | ricketycricket wrote:
       | From the example: "Oh no, I'm really sorry to hear you're having
       | trouble with your new device. That sounds frustrating."
       | 
       | Being patronized by a machine when you just want help is going to
       | feel absolutely terrible. Not looking forward to this future.
        
         | mjamesaustin wrote:
         | "I can help you get a replacement. Here let me pull up a
         | totally hallucinated order number and a link that goes nowhere.
         | Did that solve your problem?"
        
           | rhet0rica wrote:
           | Look at it this way--if someone were trying to _sabotage_ the
           | entire tech support industry, convincing companies to ditch
           | all their existing staff and infrastructure and replace them
           | with our cheerfully unhelpful and fault-prone AI friends
           | would be a great start!
        
         | SoftTalker wrote:
         | Yeah it's irritating enough when humans do it, it's so
         | transparently insincere. Just help me with my problem.
         | 
         | I guess I am just old now but I hate talking to computers, I
         | never use Siri or any other voice interfaces, and I don't want
         | computers talking to me as if they are human. Maybe if it were
         | like Star Trek and the computer just said "Working..." and then
         | gave me the answer it would be tolerable. Just please cut out
         | all the conversation.
        
       | hek2sch wrote:
       | The actual title of the release: Eleven v3 -- The most expensive
       | Text to Speech model
        
       | riebschlager wrote:
       | I didn't see anything about this in the documentation or
       | prompting guide, but... is it supposed to be able to sing?
       | 
       | Since I am a fundamentally unserious person, I copied in the
       | Friends theme song lyrics into the demo and what came out was a
       | singing voice with guitar. In another test, I added [verse] and
       | [chorus] labels and it's singing acappella.
       | 
       | [1] and [2] were prompted with just the lyrics. [3] was with the
       | verse/chorus tags. I tried other popular songs, but for whatever
       | reason, those didn't flip the switch to have it sing.
       | 
       | [1] http://the816.com/x/friends-1.mp3 [2]
       | http://the816.com/x/friends-2.mp3 [3]
       | http://the816.com/x/friends-3.mp3
        
         | yawnxyz wrote:
         | They have some singing in their demo! So I'm guessing that's
         | baked into the model
        
           | louisjoejordan wrote:
           | Might take a few tries, but it will.
        
         | londons_explore wrote:
         | interestingly not very similar to the actual friends intro -
         | suggesting it isn't a matter of overfitting onto something
         | rather common in the training data.
        
       | jurgenaut23 wrote:
       | French is atrocious. It sounds like beginner-level english
       | speakers trying to decipher a text without understanding it.
        
         | lharries wrote:
         | Can you try with this voice? https://elevenlabs.io/app/voice-
         | library?voiceId=xTZlmU8dKXdy...
         | 
         | Voice selection matters more for this model
        
       | louisjoejordan wrote:
       | quick note that that voice selection matters a lot with our new
       | v3 model, especially voice language!
       | 
       | We have a curated list of v3 voices in the library, but feel free
       | to try others to find what works. Make sure language <> voice
       | language match.
        
         | politelemon wrote:
         | Unfortunately many of the foreign language generation sounds
         | unnatural, with a strong American accent. I've tried the
         | Spanish, Galician, Tagalog, German. I did try the curated
         | samples.
        
           | lharries wrote:
           | Can you choose a voice that's native in that language in the
           | voice library: https://elevenlabs.io/app/voice-
           | library?language=es
        
       | code51 wrote:
       | High probability your v2 voice will break with this.
        
       | brian_herman wrote:
       | Unfortunately voice actors will be replaced by someThing like
       | this hopefully they will find someThing else To do
        
         | geuis wrote:
         | I dunno. It's definitely a concern in the community. But real
         | people are still getting work.
         | 
         | Audible has ruined their catalog listings with their "Virtual
         | voice" thing and no option to filter them out. They're mostly
         | low quality books narrated by subpar AI voice that don't sell
         | at all, while making it extremely difficult to find quality new
         | books to listen to.
        
       | maxglute wrote:
       | What's the state of open source tts? I'm a heavy TTS user,
       | anything that can run at 3x-4x speed off enthusiast hardware?
        
         | tomr75 wrote:
         | expressive: https://github.com/resemble-ai/chatterbox
         | 
         | dialogue like notebooklm: https://github.com/nari-labs/dia
        
           | omnimus wrote:
           | Are there any good ones that do languages other than US
           | english?
        
       | christophilus wrote:
       | We're using elevenlabs in a new prototype, and it gets confused
       | by its own voice which my mic picks up. Unless I wear headphones,
       | it thinks I'm talking, and it gets into a loop.
       | 
       | I hope this release fixes that bug!
        
         | thomasfromcdnjs wrote:
         | That doesn't sound like a problem they need to solve.
         | 
         | On your client you need to implement some form of echo
         | cancellation.
        
         | jhgg wrote:
         | This is not a model issue - you just have not properly
         | implemented acoustic echo cancellation on your end.
        
       | palisade wrote:
       | For reference in case anyone is wondering, it is based on:
       | 
       | https://github.com/152334H/tortoise-tts-fast
       | 
       | The developer of tortoise tts fast was hired by Eleven labs.
        
       | moralestapia wrote:
       | >Is this available over API?
       | 
       | >Public API for Eleven v3 (alpha) is coming soon.
       | 
       | There is zero use for this without an API endpoint. At least is
       | coming.
        
       | hadrien01 wrote:
       | The French language examples on that page are atrocious. One of
       | them starts reading French like a native English speaker, then
       | mid-sentence switches to a proper accent. Another one does some
       | words with a Canadian-French accent, but not all of them. And the
       | only one with a proper and constant accent from start to end
       | sounds worse than the default Windows TTS...
        
       | flakiness wrote:
       | Japanese: Better than v2, but still far from "natural". Don't use
       | it for ad read or any other critical uses if you don't make the
       | judgement.
        
       ___________________________________________________________________
       (page generated 2025-06-05 23:00 UTC)