[HN Gopher] OpenAI Audio Models
       ___________________________________________________________________
        
       OpenAI Audio Models
        
       Author : KuzeyAbi
       Score  : 334 points
       Date   : 2025-03-20 17:18 UTC (5 hours ago)
        
 (HTM) web link (www.openai.fm)
 (TXT) w3m dump (www.openai.fm)
        
       | minimaxir wrote:
       | This is an official OpenAI tool linked from the new model
       | announcement (https://openai.com/index/introducing-our-next-
       | generation-aud... ), despite the branding difference.
        
       | danso wrote:
       | The voices are pretty convincing. It's funny to hear drastically
       | the tone of the reading can change when repeatedly stopping and
       | restarting the samples without changing any of the settings.
        
       | nickthegreek wrote:
       | Try the refresh button to get a new list of vibe styles.
        
       | varunneal wrote:
       | One of the most novel demos I've seen openai ship in a few years.
       | I love how it looks almost like a synth. Fun to play around with!
        
       | stephenheron wrote:
       | Quite disappointing their speech to text models are not open
       | source. Whisper was really good and it was great it was open to
       | play around with. I guess this continues OpenAI's approach of not
       | really being open!
        
         | nickthegreek wrote:
         | Indeed. Right now I think our open choices are Piper, Kokoro
         | and Orpheus.
        
           | DrPhish wrote:
           | In my opinion GPT-SoVITS is the best if you can put in the
           | effort. I'm still using v2 since the output is so good. Its
           | also the best multilingual one in my testing on Japanese
           | inputs.
        
             | pzo wrote:
             | can it support more languages rather than only English,
             | Chinese, Japanese, Korean?
        
             | nickthegreek wrote:
             | hadnt messed with that one before. my needs are more real
             | time for voice assistant but was neat to play with on
             | hugginface.
             | 
             | https://huggingface.co/spaces/lj1995/GPT-SoVITS-v2
        
           | GaggiX wrote:
           | He was talking about STT models, not TTS. Whisper is open
           | source and a good solution in many cases (in particular
           | finetuned ones).
        
             | pzo wrote:
             | regarding STT we got also today 2 new models from Nvidia:
             | 
             | https://huggingface.co/nvidia/canary-180m-flash
             | 
             | https://huggingface.co/nvidia/canary-1b-flash
             | 
             | second in Open ASR leaderboard
             | https://huggingface.co/spaces/hf-audio/open_asr_leaderboard
             | 
             | Sadly only supports 4 languages (english, german, spanish,
             | french)
        
       | Etheryte wrote:
       | Recommended input for anyone trying this out:
       | 
       | Voice: Onyx
       | 
       | Vibe: Heavy german accent, doing an Arnold Schwarzenegger
       | impression, way over the top for comedic effect. Deep booming
       | voice, uses pauses for dramatic effect.
        
         | jeffharris wrote:
         | so good!
         | https://www.openai.fm/#28540f27-5b51-445a-b1d6-1c89711a2c4f
        
           | soared wrote:
           | Weird 1 in maybe 5 times when I refresh that page I get a
           | very good Arnold. The rest are very bad.
        
           | Sohcahtoa82 wrote:
           | I hit Play 3 times and got 3 very different results.
           | 
           | One merely sounded like it had a slight German accent, once
           | just sounded kind of raspy, and the third sound like a normal
           | American English speaker.
        
         | o_____________o wrote:
         | interesting to see what accents do and do not work
        
         | mft_ wrote:
         | Weird; trying exactly this, and every time I stop and play
         | again, I get a totally different voice. One of them (if I'm not
         | mistaken) was cod-Russian.
        
           | mchusma wrote:
           | I also find this strange, and I wonder if I can get a
           | consistent voice out of this. If using the api with a
           | vibe/instructions for a back and forth, will it be
           | consistent? This example app they provide implies no?
        
           | Etheryte wrote:
           | Yeah, that's both odd and very unfortunate, it seems
           | incredibly nondeterministic. Even running this with the same
           | exact parameters over and over gives widely different
           | results.
        
         | rybthrow2 wrote:
         | This is hilarious, extra points if you get it to say:
         | 
         | "Get to the chopper now and PUT THAT COOKIE DOWN NOWWWW"
        
         | dougiejones wrote:
         | My recommendation for Ash:
         | 
         | Delivery: Cow noises. You are actually a cow. You can only moo
         | and grunt. No human noises. Only moo. No words.
         | 
         | Pauses: Moo and grunt between sentences. Some burps and farts.
         | 
         | Tone: Cow.
        
       | jcmp wrote:
       | How do you call this desing/ui astethic? I like it
        
         | KuzeyAbi wrote:
         | You can screenshot and ask chatgpt lol
        
         | randomcatuser wrote:
         | neumorphism!
        
         | havefunbesafe wrote:
         | by copying Teenage Engineering
        
         | vyrotek wrote:
         | "teenage engineering"
        
       | danso wrote:
       | Interesting, I inserted a bunch of "fuck"s in the text and the
       | "NYC Cabbie" voice read it all just fine. When I switched to
       | other voices ("Connoisseur", "Cheerleader", "Santa"), it
       | responded "I'm sorry I can't assist with that request".
       | 
       | I switched back to "NYC Cabbie" and it again read it just fine. I
       | then reloaded the session completely, refreshed the voice
       | selections until "NYC Cabbie" came up again, and it still read
       | the text without hesitation.
       | 
       | The text:
       | 
       | > _In my younger and more vulnerable years my father fuck gave me
       | some fuck fuck advice that I 've been fuck fuck FUCK OH FUCK
       | turning over in my mind ever since._
       | 
       | > _" Whenever you feel like criticizing any one," he told me, oh
       | fuck! FUCK! "just remember that all the people in this world
       | haven't had fuck fuck fuck FUCKERKER the advantages that you've
       | had."_
       | 
       | edit: "Emo Teenager", "Mad Scientist", and "Smooth Jazz" are able
       | to read the text. However, "Medieval Knight" and "Robot" cannot.
        
         | nazgulsenpai wrote:
         | Glad I'm not the only one whose inner 12 year old curiosity is
         | immediately triggered by free input TTS. Swear words and just
         | raking my hands across the keyboard to insert gibberish in
         | every possible accent.
        
           | Gracana wrote:
           | I am immediately reminded of this:
           | https://www.youtube.com/watch?v=Hv6RbEOlqRo
        
           | andrewinardeer wrote:
           | It won't generate slurs, though.
        
             | dvngnt_ wrote:
             | what did you try?
        
               | andrewinardeer wrote:
               | Is this bait? Lol.
               | 
               | Try a few for yourself.
        
             | sReinwald wrote:
             | Seems easy enough to get around with homophonic
             | substitution, though. Didn't refuse any of my tests so far.
        
       | ComputerGuru wrote:
       | It would be much more convenient to use if changing the voice
       | model worked on the fly, without having to stop and start the
       | audio.
        
         | amitport wrote:
         | Louis CK | about airplane Wi Fi
         | 
         | https://www.youtube.com/watch?v=me4BZBsHwZs
        
         | swyx wrote:
         | convenient why?
        
       | islewis wrote:
       | Cool format for a demo. Some of the voices have a slight
       | "metallic" ring to them, something I've seen a fair amount with
       | Eleven Labs' models.
       | 
       | Does anyone have any experience with the realtime latency of
       | these Openai TTS models? ElevenLabs has been so slow (much slower
       | than the latency they advertise), which makes it almost
       | impossible to use in realtime scenarios unless you can cache and
       | replay the outputs. Cartesia looks to have cracked the time to
       | first token, but i've found their voices to be a bit less
       | consistent than Eleven Labs'.
        
       | carbocation wrote:
       | Nova+Serene sounds very metallic at the beginning about 50% of
       | the time for me.
        
         | jeffharris wrote:
         | some of the older voices are definitely less steerable, more
         | robotic
         | 
         | we put little stars in the bottom right corner for the newer
         | voices, which should sound better
        
       | jeffharris wrote:
       | Hey, I'm Jeff and I was PM for these models at OpenAI. Today we
       | launched three new state-of-the-art audio models. Two speech-to-
       | text models--outperforming Whisper. A new TTS model--you can
       | instruct it _how_ to speak (try it on openai.fm!). And our Agents
       | SDK now supports audio, making it easy to turn text agents into
       | voice agents. We think you 'll really like these models. Let me
       | know if you have any questions here!
        
         | TheAceOfHearts wrote:
         | Is it against the TOS to use it for sexually explicit content?
        
           | knicholes wrote:
           | I don't have your answer, but as far as innuendo goes, it's
           | definitely capable!
        
           | jeffharris wrote:
           | Yes, from our terms: "Don't build tools that may be
           | inappropriate for minors, including: Sexually explicit or
           | suggestive content. This does not include content created for
           | scientific or educational purposes."
           | https://openai.com/policies/usage-policies/
        
         | staticautomatic wrote:
         | Any plans to directly support diarization or voiceprinting?
        
           | jeffharris wrote:
           | We're thinking about diarization (adding time awareness to
           | GPT models) but no firm plans to share just yet
        
             | simonw wrote:
             | The feature I want is speaker differentiation - I want to
             | feed in an audio file and get back a transcript with
             | "Speaker 1: ..., Speaker 2: ..." indications.
             | 
             | That plus timestamps would be incredible.
             | 
             | The Google Gemini 2.0 models are showing some promise with
             | this, I can't speak to their reliability just yet though.
        
               | runeb wrote:
               | I had good results with pyannote and the following model
               | for that use case in the past
               | https://huggingface.co/pyannote/speaker-diarization-3.1
        
               | infecto wrote:
               | I thought Deepgram already did speaker diarization (which
               | is differentiation) pretty well. That and it can include
               | timestamps plus other metadata.
        
               | thot_experiment wrote:
               | WhisperX does all of this, I use it all the time to
               | transcribe meeting notes. Both speaker differentiation
               | and individual word timestamps.
        
             | youssefabdelm wrote:
             | Jeff you know what would be magical? Not just vanilla
             | diarization "Speaker 1" and "2" but if the model can know
             | from the conversation this speaker was referred to as "Jeff
             | Harris" or "Jeff" so it uses that instead.
        
         | kouteiheika wrote:
         | Any plans to open the weights of any of those?
        
           | jeffharris wrote:
           | nothing to share on open source yet, it's something we'll
           | keep exploring. Especially as the models get smaller so more
           | able to run on regular devices
        
         | kiney wrote:
         | Are the new models released with weights under an open license
         | like whisper? If not, is it planned for the future?
        
         | nico wrote:
         | Are these models downloadable, like whisper?
         | 
         | What's the minimum hardware for running them?
         | 
         | Would they run on a raspberry pi?
         | 
         | Or a smartphone?
        
           | jeffharris wrote:
           | not open source at this time. unfortunately they're much to
           | large to run on normal consumer hardware
        
             | echoangle wrote:
             | Is that the reason you're not open sourcing them? Wouldn't
             | it still make sense to provide it for enthusiasts?
        
             | risho wrote:
             | with devices having unified memory now we are no longer
             | limited to what can fit inside of a 3090 anymore. consumer
             | hardware can have hundreds of gigabytes of memory now, is
             | it really not able to fit in that?
        
         | pier25 wrote:
         | what data did you use to train these models?
        
         | new_user_final wrote:
         | Do you have plans to make it more realistic like kokoro-82M? I
         | don't know, is it only me or anyone else, machine voice is
         | irritating to me to listen for longer period of time.
         | 
         | https://huggingface.co/hexgrad/Kokoro-82M
        
         | progbits wrote:
         | > Two speech-to-text models--outperforming Whisper
         | 
         | On what metric? Also Whisper is no longer state of the art in
         | accuracy, how does it compare to the others in this benchmark?
         | 
         | https://artificialanalysis.ai/speech-to-text
        
           | jeffharris wrote:
           | We've been using the FLUERS eval and you can see comparisons
           | to other models on the market in the post
           | https://openai.com/index/introducing-our-next-generation-
           | aud...
           | 
           | Curious if there's a benchmark you trust most?
        
             | lern_too_spel wrote:
             | FLUERS and GP's Common Voice dataset focus on read speech.
             | I've observed models that perform well on these datasets be
             | completely useless on other distributions, like whispered
             | speech or shouted speech or conversational speech between
             | humans who aren't talking to a computer.
        
         | robbomacrae wrote:
         | Hi Jeff, Thanks for updating the TTS endpoint! I was literally
         | about to have to make a workaround with the chat completions
         | endpoint with a hit and hope the transcription matches
         | strategy... as it was the only way to get the updated voice
         | models.
         | 
         | Curious.. is gpt-4o-mini-tts the equivilant of what is/was
         | gpt-4o-mini-audio-preview for chat completions? Because in
         | timing tests it takes around 2 seconds to return a short phrase
         | which seems more equivilant to gpt-4o-audio-preview.. the later
         | was much better for the hit and hope strat as it didn't ad lib!
         | 
         | Also I notice you can add accents to instructions and it does a
         | reasonable job. But are there any plans to bring out localized
         | voice models?
        
           | jeffharris wrote:
           | It's a slightly better model for TTS. With extra training
           | focusing on reading the script exactly as written.
           | 
           | e.g. the audio-preview model when given instruction to speak
           | "What is the capital of Italy" would often speak "Rome". This
           | model should be much better in that regard
           | 
           | = No plans to have localized voice models, but we do want to
           | bring expand the menu of voices with voices that are best at
           | different accents
        
             | robbomacrae wrote:
             | Great to hear thanks. My favorite was "I would like you to
             | repeat the following in an Australian accent: Hi there,
             | welcome to Sydney." which was more often than not swapping
             | "Hi there" for "G'day"!
        
         | dandiep wrote:
         | 1) Previous TTS models had problems with major problems
         | accents. E.g. a Spanish sentence could drift from a Spain
         | accent to Mexican to American all within one sentence. Has this
         | been improved and/or is it still a WIP?
         | 
         | 2) What is the latency?
         | 
         | 3) Your STT API/Whisper had MAJOR problems with hallucinating
         | things the user didn't say. Is this fixed?
         | 
         | 4) Whisper and your audio models often auto corrected speech,
         | e.g. if someone made a grammatical error. Or if someone is
         | speaking Spanish and inserted an English word, it would change
         | the word to the Spanish equivalent. Does this still happen?
        
           | jeffharris wrote:
           | 1/ we've been working a lot on accents, so expect
           | improvements with these models... though we're not done.
           | Would be curious how you find them. And try giving specific
           | detailed instructions + examples for the accents you want
           | 
           | 2/ We're doing everything we can to make it fast. Very
           | critical that it can stream audio meaningfully faster than
           | realtime
           | 
           | 3+4/ I wouldn't call hallucinations "solved", but it's been
           | the central focus for these models. So I hope you find it
           | much improved
        
             | wewewedxfgdf wrote:
             | As mentioned in another comment, the British accents are
             | very far from being authentic.
        
         | wewewedxfgdf wrote:
         | So there's no British accents?
        
           | jeffharris wrote:
           | try the ballad or fable voices
        
             | wewewedxfgdf wrote:
             | Doesn't really sound very British to be honest.
             | 
             | Sounds kinda international/like an American trying to do a
             | British accent.
             | 
             | I've been looking for real TTS British accents so this
             | product doesn't meet my goals.
        
               | GordonS wrote:
               | Azure TTS has some great British accents - I used a
               | British female voice for a demo video voice over, and the
               | quality was great. Not as good as ElevenLabs, but I was
               | still really impressed with the final result.
        
         | mclau156 wrote:
         | Does whispering work? I could not get it to work when I tried
         | it
        
           | jeffharris wrote:
           | Should do! here's an example https://www.openai.fm/#4a5a82db-
           | faea-4f80-813c-3131902c2458
        
             | mclau156 wrote:
             | It seems to start out strong, but then starts loudly
             | talking by the end, do you know why it loses focus?
             | 
             | edit: I actually got it to stay whispering by also putting
             | (soft whispering voice) before the second paragraph
        
         | a-r-t wrote:
         | Hi Jeff, are there any plans to support dual-channel audio
         | recordings (e.g., Twilio phone call audio) for speech-to-text
         | models? Currently, we have to either process each channel
         | separately and lose conversational context, or merge channels
         | and lose speaker identification.
        
           | ekzy wrote:
           | I'm not entirely sure what you mean but twilio recordings
           | supports dual channels already
        
             | a-r-t wrote:
             | Transcribing Twilio's dual-channel recordings using
             | OpenAI's speech-to-text while preserving channel
             | identification.
        
               | ekzy wrote:
               | Oh I see what you mean that would be a neat feature.
               | Assuming you can get timestamps though it should be
               | trivial to work around the issue?
        
         | nabakin wrote:
         | Hey Jeff, thanks for your work! Quick question for you, are you
         | guys using Azure Speech Services or have these TTS models been
         | trained by OpenAI from scratch?
        
         | visarga wrote:
         | Hey Jeff, maybe you could improve the TTS that is currently in
         | the OpenAI web and phone apps. When I set it to read numbers in
         | Romanian it slurs digits. This also happens sometimes with
         | regular words as well. I hope you find resources for other
         | languages than English.
        
           | jeffharris wrote:
           | thanks for flagging ... number fidelity (especially on
           | languages that are unfortunately less represented in training
           | data) is still something we're working to improve
        
             | visarga wrote:
             | Actually even the new model does it. I put it read "12345
             | 54321" and it read "2346 5321". So it both skips and
             | hallucinates digits. This could be dangerous if it is used
             | to read some news article or important text with numbers.
        
         | simonw wrote:
         | Is there any chance that gpt-4o-transcribe might get confused
         | and accidentally follow instructions in the audio stream
         | instead of transcribing them?
        
           | simonw wrote:
           | Here's a partial answer to my own question:
           | https://news.ycombinator.com/item?id=43427525
           | 
           | > e.g. the audio-preview model when given instruction to
           | speak "What is the capital of Italy" would often speak
           | "Rome". This model should be much better in that regard
           | 
           | "Much better" doesn't sound like it can't happen at all
           | though.
        
         | twalkz wrote:
         | Woohoo new voices! I've been using a mix of TTS models on a
         | project I've been working on, and I consistently prefer the
         | output of OpenAI to ElevenLabs (at least when things are
         | working properly).
         | 
         | Which leads me to my main gripe with the OpenAI models -- I
         | find they break -- produce empty / incorrect / noise outputs --
         | on a few key use cases for my application (things like single-
         | word inputs -- especially compound words and capitalized words,
         | words in parenthesis, etc.)
         | 
         | So I guess my question is might gpt-4o-mini-tts provide more
         | "reliable" output than tts-1-hd?
        
         | risho wrote:
         | would really love to so the new whisper style speech to text
         | model open sourced.
        
         | modeless wrote:
         | Whisper's major problem was hallucinations, how are the new
         | models doing there? The performance of ChatGPT advanced voice
         | in recognizing speech is, frankly, terrible. Are these models
         | better than what's used there?
        
           | nickthegreek wrote:
           | They say they are much better at not hallucinating but you
           | also cant run it on your own hardware like whisper.
        
         | mazd wrote:
         | The Realtime API via WebRTC sample code for transcription is
         | erroring. Could you take a look into this?
        
         | ekzy wrote:
         | Do you know when we can expect an update on the realtime API?
         | It's still in beta and there are many issues (e.g voice
         | randomly cutting off, VAD issues, especially with mulaw etc...)
         | which makes it impossible to use in production, but there's not
         | much communication from OpenAI. It's difficult to know what to
         | bet on. Pushing for stt->llm->tts makes you wonder if we should
         | carry on building with the realtime API.
        
           | taf2 wrote:
           | Agreed- really not liking how they are neglecting it... I
           | hope they are just hard at work behind the scenes and will
           | release something soon
        
         | taf2 wrote:
         | Please release a stable realtime speech to speech model. The
         | current version constantly thinks it's a young teen heading to
         | college and sad but then suddenly so excited about it
        
           | dietr1ch wrote:
           | can't wait for scam calls after this gets perfected
        
         | zhyder wrote:
         | How is the latency (Time To First Byte of audio, when
         | streaming) and throughput (non-vibe characters input per
         | second) compared to the existing 'tts-1' non-HD that's the same
         | price? TTFB in particular is important and needs to be much
         | better than 'tts-1'.
        
         | Etheryte wrote:
         | After toying around with the TTS model it seems incredibly
         | nondeterministic. Running the same input with the same
         | parameters can have widely different results, some really good,
         | others downright bad. The tone, intonation and character all
         | vary widely. While some of the outputs are great, this
         | inconsistency makes it a really tough sell. Imagine if Siri
         | responded to you with a different voice every time, as an
         | example. Is this something you're looking to address somewhere
         | down the line or do you consider that working as intended?
        
         | oidar wrote:
         | Any plans to offer speech to speech models which keep prosody,
         | intonation, and timing intact? ElevenLabs is getting expensive
         | for this.
        
       | basitmakine wrote:
       | I don't think they're anywhere near TaskAGI or ElevenLabs level.
        
       | tantalor wrote:
       | It does a good job with Pirate voice. It can even inject "Arrr
       | matey"
        
       | theoryofx wrote:
       | Still seems like Elevenlabs is crushing them on realtime audio,
       | or does this change things?
        
         | atlasunshrugged wrote:
         | I'm also curious about this for longform content. Will this be
         | competitive for something like creating an audiobook?
        
         | prdonahue wrote:
         | Do you have any affiliation with Elevenlabs?
        
           | atlasunshrugged wrote:
           | FWIW I have no affiliation with any of these companies but I
           | have a book coming out soon and have been researching AI
           | audiobook tools and Elevenlabs seems to be far and away the
           | consensus for that at least
        
           | theoryofx wrote:
           | I do not have any affiliation with Elevenlabs or OpenAI
           | except as a user of their APIs. I'd actually prefer it if
           | OpenAI had a better realtime product than Elevenlabs because
           | it'd be more convenient.
        
       | jtbayly wrote:
       | I don't get it. These voices all have a not-so-subtle _vibration_
       | in them that makes them feel worse than Siri to me. I was
       | expecting _a lot_ better.
        
         | pier25 wrote:
         | yeah the voices sound terrible
         | 
         | I'm guessing their spectral generator is super low res to save
         | on resources
        
       | minimaxir wrote:
       | One very important quote from the official announcement:
       | 
       | > For the first time, developers can "instruct" the model not
       | just on what to say but how to say it--enabling more customized
       | experiences for use cases ranging from customer service to
       | creative storytelling.
       | 
       | The instructions are the "vibes" in this UI. But the announcement
       | is wrong with the "for the first time" part: it was _possible_ to
       | steer the base GPT-4o model to create voices in a certain style
       | using system prompt engineering (blogged about here:
       | https://minimaxir.com/2024/10/speech-prompt-engineering/ ) out of
       | concern that it could be used as a replacement for voice acting,
       | however it was too expensive and adherence isn't great.
       | 
       | The schema of the vibes here implies that this new model is more
       | receptive to nuance, which changes the calculus. The test cases
       | from my post behave as expected, and the cost of gpt-4o-mini-tts
       | audio output is $0.015 / minute
       | (https://platform.openai.com/docs/pricing ), which is about
       | 1/20th of the cost of my initial experments and is now feasible
       | to use to potentially replace common voice applications. This has
       | implications, and I'll be testing more around more nuanced prompt
       | engineering.
        
       | benjismith wrote:
       | If I'm reading the pricing correctly, these models are
       | SIGNIFICANTLY cheaper than ElevenLabs.
       | 
       | https://platform.openai.com/docs/pricing
       | 
       | If these are the "gpt-4o-mini-tts" models, and if the pricing
       | estimate of "$0.015 per minute" of audio is correct, then these
       | prices 85% cheaper than those of ElevenLabs.
       | 
       | https://elevenlabs.io/pricing
       | 
       | With ElevenLabs, if I choose their most cost-effectuve "Business"
       | plan for $1100 per month (with annual billing of $13,200, a
       | savings of 17% over monthly billing), then I get 11,000 minutes
       | TTS, and each minute is billed at 10 cents.
       | 
       | With OpenAI, I could get 11,000 minutes of TTS for $165.
       | 
       | Somebody check my math... Is this right?
        
         | fixprix wrote:
         | It looks like they are targeting Google's TTS price point which
         | is $16 per million characters which comes out to $0.015/minute.
        
         | lukebuehler wrote:
         | yes, I think you are right. When I did the math on 11labs
         | million chars I got the same numbers (Pro plan).
         | 
         | I'm super happy about this, since I took a bet that exactly
         | this would happen. I've just been building a consumer TTS app
         | that could only work with significant cheaper TTS prices per
         | million character (or self-hosted models)
        
           | benjismith wrote:
           | Same for me :)
        
           | zacmps wrote:
           | What does it do?
        
             | lukebuehler wrote:
             | Convert any file (pdf, epub, txt) to an audoibook,
             | downloadable as mp3, or directly listenable via RSS feed
             | in, say, Apple Potcasts app.
             | 
             | Basically make one-off audiobooks for yourself or a few
             | friends.
        
               | setsewerd wrote:
               | Any plans to make a Chrome extension variant? Been
               | looking for a high quality and cheap TTS extension for
               | ages (like ElevenLabs Human Reader, except with less
               | absurd pricing)
        
               | lukebuehler wrote:
               | I din't think of that, interesting idea. What I'm
               | focusing right now is long-form content for more offline-
               | ish listening, but maybe a plugin could work to load
               | longer texts, but I'm not working on a screen reader atm.
        
               | wholinator2 wrote:
               | Do you know if there's any offerings today that can read
               | math? Like speak an equation the way a human would? It's
               | something I've been thinking about a long time and would
               | be an essential feature for me (the only things i read
               | are physics)
        
           | lherron wrote:
           | Kokoro TTS is pretty good for open source. Worth checking
           | out.
        
             | lukebuehler wrote:
             | Yes, kokoro is great, and the language flexibility is a
             | huge plus too. And the best prices per character is for
             | sure if you self-host.
        
         | echelon wrote:
         | ElevenLabs is _incredibly_ over-priced and that 's how they
         | were able to achieve the MRR that led to their incredible
         | fundraising.
         | 
         | No matter what happens, they'll eventually be undercut and
         | matched in terms of quality. It'll be a race to the bottom for
         | them too.
         | 
         | ElevenLabs is going to have a tough time. They've been way too
         | expensive.
        
           | MrAssisted wrote:
           | I hope they find a more unique product offering that takes
           | hold. Everybody thinks of them as text-to-speech but I use
           | ElevenLabs exclusively for speech-to-speech for vtubing as my
           | AI character. They're kind of the only game in town for doing
           | super high quality speech-to-speech (unless someone here has
           | an alternative which I'd LOVE to know about). I've tried
           | https://github.com/w-okada/voice-changer which is great
           | because it's real-time but the quality is enough of a step
           | down that actual words I'm saying become unclear and
           | difficult to understand. Also with that I am tied to using my
           | RTX 3090 desktop vs ElevenLabs which I can do in the cloud
           | from my laptop anywhere.
           | 
           | I'm pretty much dependent on ElevenLabs to do my vtubing at
           | this point but I can't imagine speech-to-speech has wide
           | adoption so I don't know if they'll even keep it around.
        
             | eob wrote:
             | Are you comfortable sharing the video & lip-sync stack you
             | use? I don't know anything about the space but am curious
             | to check out what's possible these days.
        
               | MrAssisted wrote:
               | For my last video I used
               | https://github.com/warmshao/FasterLivePortrait with a png
               | of the character on my RTX 3090 desktop and recorded the
               | output of that real-time but in the next video I'm going
               | to spin up a runpod instance and do the
               | FasterLivePortrait in the cloud after the fact because
               | then I can get a smooth 60fps which looks better. I think
               | the only real-time cloud way to do AI vtubing in the
               | cloud is my own GenDJ project (fork of
               | https://github.com/kylemcdonald/i2i-realtime but tweaked
               | for cloud real-time) but that just doesn't look remotely
               | as good as LivePortrait. Somebody needs to rip out and
               | replace insightface in FasterLivePortait (it's prohibited
               | for commercial use) and fork https://github.com/GenDJ to
               | have the runpod it spins up run the de-insightfaced
               | LivePortrait instead of i2i-realtime. I'll probably get
               | around to doing that in the next few months if nobody
               | else does and nothing else comes along and makes
               | LivePortrait obsolete (both are big ifs).
               | 
               | AIWarper recently released a simpler way to run
               | FasterLivePortrait for vtubing purposes
               | https://huggingface.co/AIWarper/WarpTuber but I haven't
               | tried it yet because I already have my own working setup
               | and as I mentioned I'm shifting my workload for that to
               | the cloud anyways
        
         | forgotpasagain wrote:
         | Almost everyone is cheaper than ElevenLabs though.
        
         | kuprel wrote:
         | OpenAI doesn't have voice cloning
        
           | dannyw wrote:
           | They do, they just don't offer it.
        
           | tiahura wrote:
           | You missed the story:
           | 
           | https://community.openai.com/t/chatgpt-unexpectedly-began-
           | sp...
           | 
           | ChatGPT unexpectedly began speaking in a user's cloned voice
           | during testing
        
         | youssefabdelm wrote:
         | Def prefer the pricing but so far on 4o, no timestamps or
         | diarization sadly
        
         | whimsicalism wrote:
         | Sesame is free and pretty good and you can run it yourself.
        
           | kuprel wrote:
           | They released a crippled model:
           | https://github.com/SesameAILabs/csm/issues/63
        
             | hnhn34 wrote:
             | The good news is Orpheus-3B just made Sesame essentially
             | obsolete.
        
               | Foreignborn wrote:
               | thanks for this, it sounds pretty good.
               | 
               | link for anyone else: https://canopylabs.ai/model-
               | releases
        
               | sandspar wrote:
               | These voices are all annoying, though. The thing about
               | Sesame's Miles is that he's cool.
        
         | furyofantares wrote:
         | It's way cheaper - everyone is, elevenlabs is very expensive.
         | Nobody matches their quality though. Especially if you want
         | something that doesn't sound like a voice
         | assistant/audiobook/podcast/news anchor/tv announcer.
         | 
         | This openai offering is very interesting, it offers valuable
         | features elevenlabs doesn't in emotional control. It also
         | hallucinates though which would need to be fixed for it to be
         | very useful.
        
         | huijzer wrote:
         | Yes ElevenLabs is orders of magnitude more expensive than
         | everyone else. Very clever from a business perspective, I
         | think. They are (were?) the best so know that people will pay a
         | premium for that.
        
         | com2kid wrote:
         | Elevenlabs is an ecosystem play. They have hundreds of
         | different voices, legally licensed from real people who chose
         | to upload their voice. It is a marketplace of voices.
         | 
         | None of the other major players is trying to do that, not sure
         | why.
        
         | oidar wrote:
         | ElevenLabs is the only one offering speech to speech generation
         | where the intonation, prosody, and timing is kept intact. This
         | allows for one expressive voice actor to slip into many other
         | voices.
        
       | fixprix wrote:
       | Is this right? The current best TTS from OpenAI uses
       | gpt-4o-audio-preview which is $2.50 input text, $80 output audio,
       | the new gpt-4o-mini-tts is $0.60 input text, $12 output audio. An
       | average 5x price reduction.
       | 
       | Going the other way, transcribe with gpt-4o-audio-preview price
       | was $40 input audio, $10 output text, the new gpt-4o-transcribe
       | is $6 input audio and $10 output text. Like a 7x reduction on the
       | input price.
       | 
       | TTS/Transcribe with gpt-4o-audio-preview was a hack where you had
       | to prompt with 'listen/speak this sentence:' and it often got it
       | wrong. These new dedicated models are exactly what we needed.
       | 
       | I'm currently using the Google TTS API which is really good, fast
       | and cheap. They charges $16 per million characters which is
       | exactly the same as OpenAI's $0.015 per minute estimate.
       | 
       | Unfortunately it's not really worth switching over if the costs
       | are exactly the same. Transcription on the other hand is
       | 1.6C//minute with Google and 0.6C//minute with OpenAI now, that
       | might be worth switching over for.
        
         | pzo wrote:
         | you can compare TTS pricing here:
         | https://artificialanalysis.ai/text-to-speech
         | 
         | Previous offering from OpenAI was $15 for TTS and $30 for TTS
         | HD so not 5x reduction. This one is slighly cheaper but
         | definitely more capable (if you need control vibe)
        
           | fixprix wrote:
           | That's a really cool page thanks. Does it have stats for
           | other languages?
           | 
           | In my experience the OpenAI TTS APIs were really bad, messing
           | up all the time in foreign languages. Practically unusable
           | for my use case. You'd have to use the gpt-4o-audio-preview
           | to get anything close to passable, but it was expensive.
           | Which is why I'm using Google TTS which is very fast, high
           | quality, and provides first class support for almost every
           | language.
           | 
           | I look forward to comparing it with this model, the price
           | being the same is unfortunate as there's less incentive to
           | switch. The transcribe price is cheaper than Google it looks
           | like so that's worth considering.
        
             | pzo wrote:
             | Interesting for me Open TTS for Polish was better than
             | Google TTS (but they have few options) - which one did you
             | used? WaveNet?
             | 
             | Sadly haven't seen quality evaluation for TTS for foreign
             | languages
        
               | fixprix wrote:
               | Depends on what's available for the language, but yea
               | Wavenet and Neural2. With OpenAI TTS I'd often get weird
               | bugs where the first API call comes back all garbled, but
               | the second API call comes back fine. Wasting money. On
               | top of that more expensive and higher latency. I'm
               | interested to try out this new one.
        
       | tosh wrote:
       | Are these models only available via the API right now or also
       | available as open weights?
        
       | evalstate wrote:
       | Really looking forward to integrating with these models.
       | 
       | The next version of Model Context Protocol will have native audio
       | support (https://github.com/modelcontextprotocol/specification/pu
       | ll/9...), which will open up plenty of opportunities for interop.
        
       | pklimk wrote:
       | Interestingly "replaces every second word with potato" and
       | "speaks in Spanish instead of English" both (kind of) work as a
       | style, so it's clear there's significant flexibility and probably
       | some form of LLM-like thing under the hood.
        
       | benjismith wrote:
       | Is there way to get "speech marks" alongside the generated audio?
       | 
       | FYI, Speech marks provide millisecond timestamp for each word in
       | a generated audio file/stream (and a start/end index into your
       | original source string), as a stream of JSONL objects, like this:
       | 
       | {"time":6,"type":"word","start":0,"end":5,"value":"Hello"}
       | 
       | {"time":732,"type":"word","start":7,"end":11,"value":"it's"}
       | 
       | {"time":932,"type":"word","start":12,"end":16,"value":"nice"}
       | 
       | {"time":1193,"type":"word","start":17,"end":19,"value":"to"}
       | 
       | {"time":1280,"type":"word","start":20,"end":23,"value":"see"}
       | 
       | {"time":1473,"type":"word","start":24,"end":27,"value":"you"}
       | 
       | {"time":1577,"type":"word","start":28,"end":33,"value":"today"}
       | 
       | AWS uses these speech marks (with variants for "sentence",
       | "word", "viseme", or "ssml") in their Polly TTS service...
       | 
       | The sentence or word marks are useful for highlighting text as
       | the TTS reads aloud, while the "viseme" marks are useful for
       | doing lip-sync on a facial model.
       | 
       | https://docs.aws.amazon.com/polly/latest/dg/output.html
        
         | minimaxir wrote:
         | Passing the generated audio back to GPT-4o to ask for the
         | structured annotations would be a fun test case.
        
         | celestialcheese wrote:
         | whisper-1 has this with the verbose_json output. Has word level
         | and sentence level, works fairly well.
         | 
         | Looks like the new models don't have this feature yet.
        
       | looknee wrote:
       | Hmm I was hoping these would be bridging the gap between what's
       | already been availalbe on their audio API or in the RealtimeAPI
       | vs. Advanced Voice Mode, but the audio quality is really the same
       | as its been up to this point.
       | 
       | Does anyone have any clue about exactly why they're not making
       | the quality of Advanced Voice Mode available to build with? It
       | would be game changing for us if they did.
        
       | mlsu wrote:
       | I gave it (part of) the classic Navy Seal copypasta.
       | 
       | Interestingly, the safety controls ("I cannot assist with that
       | request") is sort of dependent on the vibe instruction. NYC
       | cabbie has no problem with it (and it's really, really funny,
       | great job openAI), but anything peaceful, positive, etc. will
       | deny the request.
       | 
       | https://www.openai.fm/#56f804ab-9183-4802-9624-adc706c7b9f8
        
       | forgotpasagain wrote:
       | It sounds very expressive but weirdly "fake" as if it's targeting
       | to be similar to some NPC character, dataset issue?
        
       | kartikarti wrote:
       | What does this little star next to the name mean?
        
       | crazygringo wrote:
       | This is astonishing. I can type anything I want into the "vibe"
       | box and it does it for the given text. Accents, attitudes,
       | personality types... I'm amazed.
       | 
       | The level of intelligent "prosody" here -- the rhythm and
       | intonation, the pauses and personality -- I wasn't expecting
       | anything like this so soon. This is truly remarkable. It
       | understands _both_ the text _and_ the prompt for how the speaker
       | should sound.
       | 
       | Like, we're getting much closer to the point where nobody except
       | celebrities are going to record audiobooks. Everyone's just going
       | to pick whatever voice they're in the mood for.
       | 
       | Some fun ones I just came up with:
       | 
       |  _> Imposing villain with an upper class British accent, speaking
       | threateningly and with menace.
       | 
       | > Helpful customer support assistant with a Southern drawl who's
       | very enthusiastic.
       | 
       | > Woman with a Boston accent who talks incredibly slowly and
       | sounds like she's about to fall asleep at any minute._
        
         | solardev wrote:
         | Guess that's why the video game voice actors are still on
         | strike:
         | https://en.m.wikipedia.org/wiki/2024%E2%80%93present_SAG-AFT...
         | 
         | If we as developers are scared of AI taking our jobs, the voice
         | actors have it much worse...
        
           | 101008 wrote:
           | What a horrible world we live on...
        
         | ForTheKidz wrote:
         | > Everyone's just going to pick whatever voice they're in the
         | mood for.
         | 
         | I can't say I've ever had this impulse. Also, to point out the
         | obvious, there's little reason to pay for an audiobook if
         | there's no human reading it. Especially if you already bought
         | the physical text.
        
           | cholantesh wrote:
           | As the sibling comment suggests, the impulse is probably more
           | on the part of an Ubisoft or an EA project director to avoid
           | hiring a voice actor.
        
         | borgdefenser wrote:
         | I am always listening to audio books but they are no good
         | anymore after playing with this for 2 minutes.
         | 
         | I am never really in the mood for a different voice. I am going
         | to dial in the voice I want and only going to want to listen
         | with that voice.
         | 
         | This is so awesome. So many audio books have been ruined by the
         | voice actor for me. What sticks out in my head is The Book of
         | Why by Judea Pearl read by Mel Foster. Brutal.
         | 
         | So many books I want as audio books too that no one would
         | bother to record.
        
         | clbrmbr wrote:
         | I got one German "w" when using the following prompt, but most
         | of the "w" were still pronounced as liquids rather than labial
         | fricatives.
         | 
         | > Speak with an exaggerated German accent, pronouncing all "w"
         | as "v"
        
       | tomjen3 wrote:
       | It doesn't seem clear, but can the model do correct emphesis? On
       | things like single words:
       | 
       | I did not steal that horse
       | 
       | Is the trivial example of something where intonation of the
       | single word is what matters. More importantly if you are reading
       | something, as a human, you change the intonation, audiolevel, and
       | speed.
        
         | Sohcahtoa82 wrote:
         | > I did not steal that horse
         | 
         | > Is the trivial example of something where intonation of the
         | single word is what matters.
         | 
         | My go-to for an example of this is "I didn't say she stole my
         | money".
         | 
         | Changing which word is emphasized completely changes the
         | meaning of the sentence.
        
       | ForTheKidz wrote:
       | Pricing looks like it's aimed at us peasants, not our lords.
       | Smart if openai wants to survive!
        
       | RobinL wrote:
       | I'm surprised at how poor this is at following a detailed prompt.
       | 
       | It seems capable of generating a consistent style, and so in that
       | sense quite useful. But if you want (say) a regional UK accent
       | it's not even close.
       | 
       | I also find it confusing you have to choose a voice. Surely
       | that's what the prompt should be for, especially when the voices
       | have such abstract names.
       | 
       | I mean, it's still very impressive when you stand back a bit, but
       | feels a bit half baked
       | 
       | Example: Voice: Thick and hearty, with a slow, rolling cadence--
       | like a lifelong Somerset farmer leaning over a gate, chatting
       | about the land with a mug of cider in hand. It's warm, weathered,
       | and rich, carrying the easy confidence of someone who's seen a
       | thousand harvests and knows every hedgerow and rolling hill in
       | the county.
       | 
       | Tone: Friendly, laid-back, and full of rustic charm. It's got
       | that unhurried quality of a man who's got time for a proper
       | chinwag, with a twinkle in his eye and a belly laugh never far
       | away. Every sentence should feel like it's been seasoned with
       | fresh air, long days in the fields, and a lifetime of countryside
       | wisdom.
       | 
       | Dialect: Classic West Country, with broad vowels, softened
       | consonants, and that unmistakable rural lilt. Words flow together
       | in an easy drawl, with plenty of dropped "h"s and "g"s. "I be"
       | replaces "I am," and "us" gets used instead of "we" or "me."
       | Expect plenty of "ooh-arrs," "proper job," and "gurt big"
       | sprinkled in naturally.
        
         | robbomacrae wrote:
         | I find it works better with shorter simpler instructions. I
         | would try:
         | 
         | Voice: Warm and slow, like a friendly Somerset farmer. Tone:
         | Laid-back and rustic. Dialect: Classic West Country with a
         | relaxed drawl and colloquial phrases.
        
       | paul7986 wrote:
       | Personally I just want to text or talk to Siri or an LLM and have
       | it do whatever I need. Have it interface with AI Agents of
       | companies, businesses, friends or families AI Agents to get
       | whatever I need done like the example on OpenAI.fm site here
       | (rebook my flight). Once it's done it shows me the confirmation
       | on my lock screen and I receive an email confirmation.
        
       | tiahura wrote:
       | When are we going to get the equivalent for Whisper. When is it
       | going to pick up on enthusiasm, sarcasm, etc?
        
       | kibbi wrote:
       | Large text-to-speech and speech-to-text models have been greatly
       | improving recently.
       | 
       | But I wish there were an _offline_ , on-device, multilingual
       | text-to-speech solution with good voices for a standard PC -- one
       | that doesn't require a GPU, tons of RAM, or max out the CPU.
       | 
       | In my research, I didn't find anything that fits the bill. People
       | often mention Tortoise TTS, but I think it garbles words too
       | often. The only plug-in solution for desktop apps I know of is
       | the commercial and rather pricey Acapela SDK.
       | 
       | I hope someone can shrink those new neural network-based models
       | to run efficiently on a typical computer. Ideally, it should run
       | at under 50% CPU load on an average Windows laptop that's several
       | years old, and start speaking almost immediately (less than 400ms
       | delay).
       | 
       | The same goes for speech-to-text. Whisper.cpp is fine, but last
       | time I looked, it wasn't able to transcribe audio at real-time
       | speed on a standard laptop.
       | 
       | I'd pay for something like this as long as it's less expensive
       | than Acapela.
       | 
       | (My use case is an AAC app.)
        
         | ZeroTalent wrote:
         | Look into https://superwhisper.com and their local models.
         | Pretty decent.
        
           | kibbi wrote:
           | Thank you, but they say "Offline models only run really well
           | on Apple Silicon macs."
        
             | ZeroTalent wrote:
             | Many SOTA apps are, unfortunately, only for Apple M Macs.
        
         | 5kg wrote:
         | May I introduce to you
         | 
         | https://huggingface.co/canopylabs/orpheus-3b-0.1-ft
         | 
         | (no affiliation)
         | 
         | it's English only afaics.
        
           | kibbi wrote:
           | The sample sounds impressive, but based on their claim --
           | 'Streaming inference is faster than playback even on an A100
           | 40GB for the 3 billion parameter model' -- I don't think this
           | could run on a standard laptop.
        
         | dharmab wrote:
         | I use Piper for one of my apps. It runs on CPU and doesn't
         | require a GPU. It will run well on a raspberry pi. I found a
         | couple of permissively licensed voices that could handle
         | technical terms without garbling them.
         | 
         | However, it is unmaintained and the Apple Silicon build is
         | broken.
         | 
         | My app also uses whisper.cpp. It runs in real time on Apple
         | Sillicon or on modern fast CPUs like AMD's gaming CPUs.
        
           | kibbi wrote:
           | I had already suspected that I hadn't found all the
           | possibilities regarding Tortoise TTS, Coqui, Piper, etc. It
           | is sometimes difficult to determine how good a TTS framework
           | really is.
           | 
           | Do you possibly have links to the voices you found?
        
       | justanotheratom wrote:
       | Note that the previous Whisper STT models were Open Source, and
       | these new STT models are not, AFAICT.
        
       | Heidaradar wrote:
       | is it just me or are these voices clearly AI generated? They've
       | obviously been improving at a steady rate but if I saw a YouTube
       | video that had this voice, I'd instantly stop watching it
        
       | smokeydoe wrote:
       | Does anyone know of any decent newer open source models for
       | generating sound effects?
        
       | corobo wrote:
       | All these voices are too good these days. I want my home
       | assistant to sound like Auto from Wall-E, dammit!
       | 
       | Anyone out there doing any nice robotic robot voices?
       | 
       | Best I've got so far is a blend of Ralph and Zarvox from MacOS'
       | `say`, haha                 say -v zarvox -r 180 "[[volm 0.8]]
       | ${message}" &       say -v ralph -r 180 "${message}"
        
       | simonw wrote:
       | Both the text-to-speech and the speech-to-text models launched
       | here suffer from reliability issues due to combining instructions
       | and data in the same stream of tokens.
       | 
       | I'm not yet sure how much of a problem this is for real-world
       | applications. I wrote a few notes on this here:
       | https://simonwillison.net/2025/Mar/20/new-openai-audio-model...
        
       | jncfhnb wrote:
       | Are there any voice to voice models out there that can replicate
       | inflection of line delivery?
        
       | alach11 wrote:
       | It's interesting that they pitch this for agent development. The
       | realtime API provides a much simpler architecture for developing
       | agents. Why would you want to string together STT -> LLM -> TTS
       | when you could have a consolidated model doing all three steps?
       | They alluded to there being some quality/intelligence benefits to
       | the multi-step approach, but in the long-run I'd expect them to
       | improve the realtime API to make this unnecessary.
        
         | zhyder wrote:
         | Text allows developers lots for flexibility to do other
         | processing, including RAG, calling APIs yourself and multiple
         | chained LLM invocations. The low latency of realtime API means
         | relying fully on one invocation of their model to do
         | everything.
        
           | alach11 wrote:
           | The realtime API can be used to call tools [0], but I agree
           | with your general point on the flexibility of working
           | directly with text.
           | 
           | [0] https://github.com/openai/openai-realtime-agents
        
       | Arubis wrote:
       | At this point, the strongest (and almost only) predictor for a
       | release announcement from OpenAI is a release announcement from
       | Anthropic.
        
       | kgeist wrote:
       | In Russian, OpenAI audio models usually have a slight American
       | (?) accent. The intonation and the phonetics fall into the
       | uncanney valley. Does the same happen in other languages?
        
         | josu wrote:
         | Yeah, same in Spanish.
        
       | buybackoff wrote:
       | I was experimenting recently with voiceover TTS generation. Did
       | run Kokoro TTS locally and it's magical for how few resources it
       | takes (runs fine in a browser), but only the default female
       | voices (Heart/Bella) are usable, and very good. Then I found that
       | Clipchamp has it built-in and several voices from a big selection
       | there are very good, and free. I've listened to this OpenAI TTS
       | and I could not like them at all even compared to Kokoro.
        
       | redox99 wrote:
       | Pretty meh. Coral Dramatic is extremely robotic for example.
        
       | saint_yossarian wrote:
       | [delayed]
        
       ___________________________________________________________________
       (page generated 2025-03-20 23:00 UTC)