[HN Gopher] Eleven v3
___________________________________________________________________
Eleven v3
Author : robertvc
Score : 112 points
Date : 2025-06-05 18:41 UTC (4 hours ago)
(HTM) web link (elevenlabs.io)
(TXT) w3m dump (elevenlabs.io)
| artninja1988 wrote:
| Sounds absolutely amazing, like 99% indistinguishable from real
| professional voice actors to me. I couldn't find any pricing
| though. Anyone know what they charge for it?
| minimaxir wrote:
| > Public API for Eleven v3 (alpha) is coming soon. For early
| access, please contact sales.
|
| I suspect they themselves don't know the exact pricing yet and
| want to assess demand first.
| delgaudm wrote:
| Ouch. Professional Voice Actor here.
| razemio wrote:
| Just here to say the oposite. It is astonshing how far away
| it still is from a professional voice actor while being
| really good. Emotion is completely missing. Instead it seems
| to try to hard to express exactly that. I cant really put my
| finger on it. It feels predictable, flat and the timing is
| strange.
| mrkstu wrote:
| Better by a mile than most anime voice work, but lacks the
| detail that a good voice narrator has on an audio book.
| minimaxir wrote:
| > Eleven v3 is 80% off until the end of June 2025 for self-serve
| users using it through the UI.
|
| That's definitely one way to loss-lead.
| lostmsu wrote:
| Open source stuff like Kokoro and the recent Chatterbox are hot
| on their heels.
|
| https://www.reddit.com/r/MachineLearning/comments/1kxv01f/p_...
| minimaxir wrote:
| It's definitely a response to Chatterbox, which is very
| funny.
| lostmsu wrote:
| Hm, is it good in all languages? Russian sounds very robotic.
| lharries wrote:
| It's a research preview for now but it should work well in 70+
| languages. Voices make a big difference, can you try with a few
| Russian IVCs?
| GrayShade wrote:
| Romanian sounds awful too, like the TTSes from 15 years ago.
| lharries wrote:
| can you try with a Romanian voice?
| GrayShade wrote:
| I'm not sure what you mean. I chose Romanian from the
| language selector and tried Matilda, Alice and Laura. Laura
| actually sounds like an English TTS trying to pronounce
| Romanian.
| gozzoo wrote:
| Exactly the same thing with Bulgarian voices.
| lharries wrote:
| It should work well for the logged in voices here:
| https://elevenlabs.io/app/voice-library?language=ro
|
| We are in the process of updated the homepage voices for
| the new languages
| NewMountain wrote:
| There's something very wrong with the Russian one. The first
| example "Jessica | Tell History", is British woman speaking
| British English transliterated from Russian. It's absolute
| murder of the Russian language and painful to listen to.
|
| The second example "Jessica | Record a commercial" is perfect.
| Confidence restored.
|
| The third example "Laura | Help a client" is back to glass in
| your ears. This time an American is speaking American English
| transliterated from Russian.
|
| Yikes. The English sounded fine, but the Russian has serious
| issues. Either there's a bug in your configuration (I hope) or
| your evals for Russian are unsound.
|
| Edit: dial back the editorializing.
| wewewedxfgdf wrote:
| I did not see an British accent example.
|
| Generally it appears the TTS systems all do US accents and the
| British accent tends to sound like Frasier - an American faking
| an British accent.
| lharries wrote:
| We have lots of great British voices in our voice library! Or
| if you want to hear an american trying to do a british accent
| add "[British accent]" at the start of the generation
| not_your_mentat wrote:
| I kept an English prompt, selected a French voice, and was
| delighted to hear an British English woman. :shrug:
| lharries wrote:
| If you'd like it to sound like a french person speaking
| french this voice works great:
| https://elevenlabs.io/app/voice-
| library?voiceId=xTZlmU8dKXdy...
|
| Or if you want a french person speaking english with a
| french accent use that voice with "[French accent]" before
| it
| wewewedxfgdf wrote:
| It would be good if your demos made it more obvious. There's
| a vast arrays of AI developments wanting me to check them out
| - you have seconds to get my attention.
| fakedang wrote:
| ElevenLabs v2's accented voices are still much stronger than
| any of its competition. And I've tried it with Arabic, French,
| Hindi and English.
| drag0s wrote:
| English sounds really great, congrats! other languages I've tried
| doesn't sound that good, you can hear a strong english accent
| dustincoates wrote:
| The French one sounded like an Alabaman who took a semester of
| college French.
|
| But the English sounds really good.
| lharries wrote:
| If you're trying to make an audiobook about an Alabaman
| visiting Paris this might be quite useful... But in
| seriousness try it with this voice:
| https://elevenlabs.io/app/voice-
| library?voiceId=rbFGGoDXFHtV...
| dustincoates wrote:
| I'll give it a check. I was playing the sample on the v3
| page.
| lharries wrote:
| Can you try with a voice that was trained on that language?
| This research preview is more variable based on the voice
| chosen
| k__ wrote:
| German sounds okay.
| lharries wrote:
| There's lots of great german voices here which should be
| better: https://elevenlabs.io/app/voice-
| library/collections/SHEPnUB9...
|
| The voice selection matters a lot for this research preview
| shafyy wrote:
| I tried German in the preview box there, and it had a very
| strong English accent.
| 8f2ab37a-ed6c wrote:
| With Italian, it starts reading the text with an absolutely
| comical American accent, but then about 10-20 words in it
| gradually snaps into a natural Italian pronunciation and it
| sounds fantastic from that point on. Not sure what's going on
| behind the scenes, but it sounds like it starts with an en-us
| baseline and then somehow zones in on the one you specified.
| Using Alice.
| ianbicking wrote:
| I've been using OpenAI's new models a lot lately
| (https://www.openai.fm/)... separating instructions from the
| spoken word is an interesting choice, and I'm assuming also has a
| lot to do with OpenAI/GPT using "instructions" across their
| products, and maybe they are just more comfortable and familiar
| generating the data and do the training for that style.
|
| Separate instructions is a bit awkward, but does allow mixing
| general instructions with specific instructions. Like I can
| concatenate output-specific instructions like "voice lowers to a
| whisper after 'but actually', and a touch of fear" with a general
| instruction like "a deep voice with a hint of an English accent"
| and it mostly figures it out.
|
| The result with OpenAI feels much less predictable and of lower
| production quality than Eleven Labs. But the range of prosidy is
| much larger, almost overengaged. The range of _voices_ is much
| smaller with OpenAI... you can instruct the voices to sound
| different, but it feels a little like the same person doing
| different voices.
|
| But in the end OpenAI's biggest feature is that it's 10x cheaper
| and completely pay-as-you-go. (Why are all these TTS services
| doing subscriptions on top of limits and credits? Blech!)
| lharries wrote:
| > The result with OpenAI feels much less predictable and of
| lower production quality than ElevenLabs
|
| Thank you Ian! Credit to our research team for making this
| possible
|
| For the prosidy, if you choose an expressive voice the prosidy
| should be larger
| Velorivox wrote:
| The word is "prosody", right?
| zamadatix wrote:
| The (American English) voices are absolutely amazing but the tags
| for laughs still feel more like an "inserted dedicated laugh
| section" than a "laugh at this point in speaking" type thing.
| I.e. it can't seem to reliably know when to giggle while saying a
| word, "just" giggle leading up to a word.
| echelon wrote:
| They're also still too expensive, and that's creating a lot of
| opportunity for other players.
|
| Even though ElevenLabs remains the quality leader, the others
| aren't that far behind.
|
| There are even a bunch of good TTS models being released as
| fully open source, especially by cutting-edge Chinese labs and
| companies. Perhaps in a bid to cut off the legs of American AI
| companies or to commoditize their compliment. Whatever the
| case, it's great for consumers.
|
| YCombinator-backed PlayHT has been releasing some of their good
| stuff too.
| taf2 wrote:
| What would say are some of the best open source TTS -
| chatterbox maybe?
| jsemrau wrote:
| I had good results with Nemo + xTTS_v2
|
| https://docs.nvidia.com/nemo-framework/user-
| guide/latest/nem...
|
| https://huggingface.co/coqui/XTTS-v2
| monkeywork wrote:
| could you list 2 or 3 of the ones you think are best quality
| to $?
| lharries wrote:
| If you edit the text so that laugh makes sense in the context
| it should be much more natural like this one:
| https://x.com/elevenlabsio/status/1930689782331412811
| zamadatix wrote:
| The first laugh in that "<LAUGHS> Hey, Dr. Von Fusion" is a
| dedicated laugh section, which the model does extremely well,
| but it works because that's a natural place to laugh before
| actually speaking the following words. Skip ahead to
| "...robot chuckle. Jessica: <LAUGHS> I know right!" and you
| get an awkwardly time/toned light chuckle completely
| separated from the "I know" you'd naturally continue saying
| while making that chuckle.
|
| You can always rewrite the text to avoid times where one
| would naturally laugh through the next couple of following
| words but that's just attempting to avoid the problem and do
| a different kind of laugh instead.
| Davidzheng wrote:
| have to say that this human can't tell the difference
| between this and other real humans so...
| carlosjobim wrote:
| Their non-English (automated?) localization of the front page is
| ridiculously badly translated.
| lharries wrote:
| Which language isn't good and I'll get that fixed asap?
| carlosjobim wrote:
| You need native or at least fluent speakers to help you, to
| get the expressions right. For example Swedish is written
| like a word-for-word translation from English.
| sojuz151 wrote:
| Polish is quite good, expected based on the founders' background
| ricketycricket wrote:
| From the example: "Oh no, I'm really sorry to hear you're having
| trouble with your new device. That sounds frustrating."
|
| Being patronized by a machine when you just want help is going to
| feel absolutely terrible. Not looking forward to this future.
| mjamesaustin wrote:
| "I can help you get a replacement. Here let me pull up a
| totally hallucinated order number and a link that goes nowhere.
| Did that solve your problem?"
| rhet0rica wrote:
| Look at it this way--if someone were trying to _sabotage_ the
| entire tech support industry, convincing companies to ditch
| all their existing staff and infrastructure and replace them
| with our cheerfully unhelpful and fault-prone AI friends
| would be a great start!
| SoftTalker wrote:
| Yeah it's irritating enough when humans do it, it's so
| transparently insincere. Just help me with my problem.
|
| I guess I am just old now but I hate talking to computers, I
| never use Siri or any other voice interfaces, and I don't want
| computers talking to me as if they are human. Maybe if it were
| like Star Trek and the computer just said "Working..." and then
| gave me the answer it would be tolerable. Just please cut out
| all the conversation.
| hek2sch wrote:
| The actual title of the release: Eleven v3 -- The most expensive
| Text to Speech model
| riebschlager wrote:
| I didn't see anything about this in the documentation or
| prompting guide, but... is it supposed to be able to sing?
|
| Since I am a fundamentally unserious person, I copied in the
| Friends theme song lyrics into the demo and what came out was a
| singing voice with guitar. In another test, I added [verse] and
| [chorus] labels and it's singing acappella.
|
| [1] and [2] were prompted with just the lyrics. [3] was with the
| verse/chorus tags. I tried other popular songs, but for whatever
| reason, those didn't flip the switch to have it sing.
|
| [1] http://the816.com/x/friends-1.mp3 [2]
| http://the816.com/x/friends-2.mp3 [3]
| http://the816.com/x/friends-3.mp3
| yawnxyz wrote:
| They have some singing in their demo! So I'm guessing that's
| baked into the model
| louisjoejordan wrote:
| Might take a few tries, but it will.
| londons_explore wrote:
| interestingly not very similar to the actual friends intro -
| suggesting it isn't a matter of overfitting onto something
| rather common in the training data.
| jurgenaut23 wrote:
| French is atrocious. It sounds like beginner-level english
| speakers trying to decipher a text without understanding it.
| lharries wrote:
| Can you try with this voice? https://elevenlabs.io/app/voice-
| library?voiceId=xTZlmU8dKXdy...
|
| Voice selection matters more for this model
| louisjoejordan wrote:
| quick note that that voice selection matters a lot with our new
| v3 model, especially voice language!
|
| We have a curated list of v3 voices in the library, but feel free
| to try others to find what works. Make sure language <> voice
| language match.
| politelemon wrote:
| Unfortunately many of the foreign language generation sounds
| unnatural, with a strong American accent. I've tried the
| Spanish, Galician, Tagalog, German. I did try the curated
| samples.
| lharries wrote:
| Can you choose a voice that's native in that language in the
| voice library: https://elevenlabs.io/app/voice-
| library?language=es
| code51 wrote:
| High probability your v2 voice will break with this.
| brian_herman wrote:
| Unfortunately voice actors will be replaced by someThing like
| this hopefully they will find someThing else To do
| geuis wrote:
| I dunno. It's definitely a concern in the community. But real
| people are still getting work.
|
| Audible has ruined their catalog listings with their "Virtual
| voice" thing and no option to filter them out. They're mostly
| low quality books narrated by subpar AI voice that don't sell
| at all, while making it extremely difficult to find quality new
| books to listen to.
| maxglute wrote:
| What's the state of open source tts? I'm a heavy TTS user,
| anything that can run at 3x-4x speed off enthusiast hardware?
| tomr75 wrote:
| expressive: https://github.com/resemble-ai/chatterbox
|
| dialogue like notebooklm: https://github.com/nari-labs/dia
| omnimus wrote:
| Are there any good ones that do languages other than US
| english?
| christophilus wrote:
| We're using elevenlabs in a new prototype, and it gets confused
| by its own voice which my mic picks up. Unless I wear headphones,
| it thinks I'm talking, and it gets into a loop.
|
| I hope this release fixes that bug!
| thomasfromcdnjs wrote:
| That doesn't sound like a problem they need to solve.
|
| On your client you need to implement some form of echo
| cancellation.
| jhgg wrote:
| This is not a model issue - you just have not properly
| implemented acoustic echo cancellation on your end.
| palisade wrote:
| For reference in case anyone is wondering, it is based on:
|
| https://github.com/152334H/tortoise-tts-fast
|
| The developer of tortoise tts fast was hired by Eleven labs.
| moralestapia wrote:
| >Is this available over API?
|
| >Public API for Eleven v3 (alpha) is coming soon.
|
| There is zero use for this without an API endpoint. At least is
| coming.
| hadrien01 wrote:
| The French language examples on that page are atrocious. One of
| them starts reading French like a native English speaker, then
| mid-sentence switches to a proper accent. Another one does some
| words with a Canadian-French accent, but not all of them. And the
| only one with a proper and constant accent from start to end
| sounds worse than the default Windows TTS...
| flakiness wrote:
| Japanese: Better than v2, but still far from "natural". Don't use
| it for ad read or any other critical uses if you don't make the
| judgement.
___________________________________________________________________
(page generated 2025-06-05 23:00 UTC)