[HN Gopher] OpenAI Audio Models
___________________________________________________________________
OpenAI Audio Models
Author : KuzeyAbi
Score : 334 points
Date : 2025-03-20 17:18 UTC (5 hours ago)
(HTM) web link (www.openai.fm)
(TXT) w3m dump (www.openai.fm)
| minimaxir wrote:
| This is an official OpenAI tool linked from the new model
| announcement (https://openai.com/index/introducing-our-next-
| generation-aud... ), despite the branding difference.
| danso wrote:
| The voices are pretty convincing. It's funny to hear drastically
| the tone of the reading can change when repeatedly stopping and
| restarting the samples without changing any of the settings.
| nickthegreek wrote:
| Try the refresh button to get a new list of vibe styles.
| varunneal wrote:
| One of the most novel demos I've seen openai ship in a few years.
| I love how it looks almost like a synth. Fun to play around with!
| stephenheron wrote:
| Quite disappointing their speech to text models are not open
| source. Whisper was really good and it was great it was open to
| play around with. I guess this continues OpenAI's approach of not
| really being open!
| nickthegreek wrote:
| Indeed. Right now I think our open choices are Piper, Kokoro
| and Orpheus.
| DrPhish wrote:
| In my opinion GPT-SoVITS is the best if you can put in the
| effort. I'm still using v2 since the output is so good. Its
| also the best multilingual one in my testing on Japanese
| inputs.
| pzo wrote:
| can it support more languages rather than only English,
| Chinese, Japanese, Korean?
| nickthegreek wrote:
| hadnt messed with that one before. my needs are more real
| time for voice assistant but was neat to play with on
| hugginface.
|
| https://huggingface.co/spaces/lj1995/GPT-SoVITS-v2
| GaggiX wrote:
| He was talking about STT models, not TTS. Whisper is open
| source and a good solution in many cases (in particular
| finetuned ones).
| pzo wrote:
| regarding STT we got also today 2 new models from Nvidia:
|
| https://huggingface.co/nvidia/canary-180m-flash
|
| https://huggingface.co/nvidia/canary-1b-flash
|
| second in Open ASR leaderboard
| https://huggingface.co/spaces/hf-audio/open_asr_leaderboard
|
| Sadly only supports 4 languages (english, german, spanish,
| french)
| Etheryte wrote:
| Recommended input for anyone trying this out:
|
| Voice: Onyx
|
| Vibe: Heavy german accent, doing an Arnold Schwarzenegger
| impression, way over the top for comedic effect. Deep booming
| voice, uses pauses for dramatic effect.
| jeffharris wrote:
| so good!
| https://www.openai.fm/#28540f27-5b51-445a-b1d6-1c89711a2c4f
| soared wrote:
| Weird 1 in maybe 5 times when I refresh that page I get a
| very good Arnold. The rest are very bad.
| Sohcahtoa82 wrote:
| I hit Play 3 times and got 3 very different results.
|
| One merely sounded like it had a slight German accent, once
| just sounded kind of raspy, and the third sound like a normal
| American English speaker.
| o_____________o wrote:
| interesting to see what accents do and do not work
| mft_ wrote:
| Weird; trying exactly this, and every time I stop and play
| again, I get a totally different voice. One of them (if I'm not
| mistaken) was cod-Russian.
| mchusma wrote:
| I also find this strange, and I wonder if I can get a
| consistent voice out of this. If using the api with a
| vibe/instructions for a back and forth, will it be
| consistent? This example app they provide implies no?
| Etheryte wrote:
| Yeah, that's both odd and very unfortunate, it seems
| incredibly nondeterministic. Even running this with the same
| exact parameters over and over gives widely different
| results.
| rybthrow2 wrote:
| This is hilarious, extra points if you get it to say:
|
| "Get to the chopper now and PUT THAT COOKIE DOWN NOWWWW"
| dougiejones wrote:
| My recommendation for Ash:
|
| Delivery: Cow noises. You are actually a cow. You can only moo
| and grunt. No human noises. Only moo. No words.
|
| Pauses: Moo and grunt between sentences. Some burps and farts.
|
| Tone: Cow.
| jcmp wrote:
| How do you call this desing/ui astethic? I like it
| KuzeyAbi wrote:
| You can screenshot and ask chatgpt lol
| randomcatuser wrote:
| neumorphism!
| havefunbesafe wrote:
| by copying Teenage Engineering
| vyrotek wrote:
| "teenage engineering"
| danso wrote:
| Interesting, I inserted a bunch of "fuck"s in the text and the
| "NYC Cabbie" voice read it all just fine. When I switched to
| other voices ("Connoisseur", "Cheerleader", "Santa"), it
| responded "I'm sorry I can't assist with that request".
|
| I switched back to "NYC Cabbie" and it again read it just fine. I
| then reloaded the session completely, refreshed the voice
| selections until "NYC Cabbie" came up again, and it still read
| the text without hesitation.
|
| The text:
|
| > _In my younger and more vulnerable years my father fuck gave me
| some fuck fuck advice that I 've been fuck fuck FUCK OH FUCK
| turning over in my mind ever since._
|
| > _" Whenever you feel like criticizing any one," he told me, oh
| fuck! FUCK! "just remember that all the people in this world
| haven't had fuck fuck fuck FUCKERKER the advantages that you've
| had."_
|
| edit: "Emo Teenager", "Mad Scientist", and "Smooth Jazz" are able
| to read the text. However, "Medieval Knight" and "Robot" cannot.
| nazgulsenpai wrote:
| Glad I'm not the only one whose inner 12 year old curiosity is
| immediately triggered by free input TTS. Swear words and just
| raking my hands across the keyboard to insert gibberish in
| every possible accent.
| Gracana wrote:
| I am immediately reminded of this:
| https://www.youtube.com/watch?v=Hv6RbEOlqRo
| andrewinardeer wrote:
| It won't generate slurs, though.
| dvngnt_ wrote:
| what did you try?
| andrewinardeer wrote:
| Is this bait? Lol.
|
| Try a few for yourself.
| sReinwald wrote:
| Seems easy enough to get around with homophonic
| substitution, though. Didn't refuse any of my tests so far.
| ComputerGuru wrote:
| It would be much more convenient to use if changing the voice
| model worked on the fly, without having to stop and start the
| audio.
| amitport wrote:
| Louis CK | about airplane Wi Fi
|
| https://www.youtube.com/watch?v=me4BZBsHwZs
| swyx wrote:
| convenient why?
| islewis wrote:
| Cool format for a demo. Some of the voices have a slight
| "metallic" ring to them, something I've seen a fair amount with
| Eleven Labs' models.
|
| Does anyone have any experience with the realtime latency of
| these Openai TTS models? ElevenLabs has been so slow (much slower
| than the latency they advertise), which makes it almost
| impossible to use in realtime scenarios unless you can cache and
| replay the outputs. Cartesia looks to have cracked the time to
| first token, but i've found their voices to be a bit less
| consistent than Eleven Labs'.
| carbocation wrote:
| Nova+Serene sounds very metallic at the beginning about 50% of
| the time for me.
| jeffharris wrote:
| some of the older voices are definitely less steerable, more
| robotic
|
| we put little stars in the bottom right corner for the newer
| voices, which should sound better
| jeffharris wrote:
| Hey, I'm Jeff and I was PM for these models at OpenAI. Today we
| launched three new state-of-the-art audio models. Two speech-to-
| text models--outperforming Whisper. A new TTS model--you can
| instruct it _how_ to speak (try it on openai.fm!). And our Agents
| SDK now supports audio, making it easy to turn text agents into
| voice agents. We think you 'll really like these models. Let me
| know if you have any questions here!
| TheAceOfHearts wrote:
| Is it against the TOS to use it for sexually explicit content?
| knicholes wrote:
| I don't have your answer, but as far as innuendo goes, it's
| definitely capable!
| jeffharris wrote:
| Yes, from our terms: "Don't build tools that may be
| inappropriate for minors, including: Sexually explicit or
| suggestive content. This does not include content created for
| scientific or educational purposes."
| https://openai.com/policies/usage-policies/
| staticautomatic wrote:
| Any plans to directly support diarization or voiceprinting?
| jeffharris wrote:
| We're thinking about diarization (adding time awareness to
| GPT models) but no firm plans to share just yet
| simonw wrote:
| The feature I want is speaker differentiation - I want to
| feed in an audio file and get back a transcript with
| "Speaker 1: ..., Speaker 2: ..." indications.
|
| That plus timestamps would be incredible.
|
| The Google Gemini 2.0 models are showing some promise with
| this, I can't speak to their reliability just yet though.
| runeb wrote:
| I had good results with pyannote and the following model
| for that use case in the past
| https://huggingface.co/pyannote/speaker-diarization-3.1
| infecto wrote:
| I thought Deepgram already did speaker diarization (which
| is differentiation) pretty well. That and it can include
| timestamps plus other metadata.
| thot_experiment wrote:
| WhisperX does all of this, I use it all the time to
| transcribe meeting notes. Both speaker differentiation
| and individual word timestamps.
| youssefabdelm wrote:
| Jeff you know what would be magical? Not just vanilla
| diarization "Speaker 1" and "2" but if the model can know
| from the conversation this speaker was referred to as "Jeff
| Harris" or "Jeff" so it uses that instead.
| kouteiheika wrote:
| Any plans to open the weights of any of those?
| jeffharris wrote:
| nothing to share on open source yet, it's something we'll
| keep exploring. Especially as the models get smaller so more
| able to run on regular devices
| kiney wrote:
| Are the new models released with weights under an open license
| like whisper? If not, is it planned for the future?
| nico wrote:
| Are these models downloadable, like whisper?
|
| What's the minimum hardware for running them?
|
| Would they run on a raspberry pi?
|
| Or a smartphone?
| jeffharris wrote:
| not open source at this time. unfortunately they're much to
| large to run on normal consumer hardware
| echoangle wrote:
| Is that the reason you're not open sourcing them? Wouldn't
| it still make sense to provide it for enthusiasts?
| risho wrote:
| with devices having unified memory now we are no longer
| limited to what can fit inside of a 3090 anymore. consumer
| hardware can have hundreds of gigabytes of memory now, is
| it really not able to fit in that?
| pier25 wrote:
| what data did you use to train these models?
| new_user_final wrote:
| Do you have plans to make it more realistic like kokoro-82M? I
| don't know, is it only me or anyone else, machine voice is
| irritating to me to listen for longer period of time.
|
| https://huggingface.co/hexgrad/Kokoro-82M
| progbits wrote:
| > Two speech-to-text models--outperforming Whisper
|
| On what metric? Also Whisper is no longer state of the art in
| accuracy, how does it compare to the others in this benchmark?
|
| https://artificialanalysis.ai/speech-to-text
| jeffharris wrote:
| We've been using the FLUERS eval and you can see comparisons
| to other models on the market in the post
| https://openai.com/index/introducing-our-next-generation-
| aud...
|
| Curious if there's a benchmark you trust most?
| lern_too_spel wrote:
| FLUERS and GP's Common Voice dataset focus on read speech.
| I've observed models that perform well on these datasets be
| completely useless on other distributions, like whispered
| speech or shouted speech or conversational speech between
| humans who aren't talking to a computer.
| robbomacrae wrote:
| Hi Jeff, Thanks for updating the TTS endpoint! I was literally
| about to have to make a workaround with the chat completions
| endpoint with a hit and hope the transcription matches
| strategy... as it was the only way to get the updated voice
| models.
|
| Curious.. is gpt-4o-mini-tts the equivilant of what is/was
| gpt-4o-mini-audio-preview for chat completions? Because in
| timing tests it takes around 2 seconds to return a short phrase
| which seems more equivilant to gpt-4o-audio-preview.. the later
| was much better for the hit and hope strat as it didn't ad lib!
|
| Also I notice you can add accents to instructions and it does a
| reasonable job. But are there any plans to bring out localized
| voice models?
| jeffharris wrote:
| It's a slightly better model for TTS. With extra training
| focusing on reading the script exactly as written.
|
| e.g. the audio-preview model when given instruction to speak
| "What is the capital of Italy" would often speak "Rome". This
| model should be much better in that regard
|
| = No plans to have localized voice models, but we do want to
| bring expand the menu of voices with voices that are best at
| different accents
| robbomacrae wrote:
| Great to hear thanks. My favorite was "I would like you to
| repeat the following in an Australian accent: Hi there,
| welcome to Sydney." which was more often than not swapping
| "Hi there" for "G'day"!
| dandiep wrote:
| 1) Previous TTS models had problems with major problems
| accents. E.g. a Spanish sentence could drift from a Spain
| accent to Mexican to American all within one sentence. Has this
| been improved and/or is it still a WIP?
|
| 2) What is the latency?
|
| 3) Your STT API/Whisper had MAJOR problems with hallucinating
| things the user didn't say. Is this fixed?
|
| 4) Whisper and your audio models often auto corrected speech,
| e.g. if someone made a grammatical error. Or if someone is
| speaking Spanish and inserted an English word, it would change
| the word to the Spanish equivalent. Does this still happen?
| jeffharris wrote:
| 1/ we've been working a lot on accents, so expect
| improvements with these models... though we're not done.
| Would be curious how you find them. And try giving specific
| detailed instructions + examples for the accents you want
|
| 2/ We're doing everything we can to make it fast. Very
| critical that it can stream audio meaningfully faster than
| realtime
|
| 3+4/ I wouldn't call hallucinations "solved", but it's been
| the central focus for these models. So I hope you find it
| much improved
| wewewedxfgdf wrote:
| As mentioned in another comment, the British accents are
| very far from being authentic.
| wewewedxfgdf wrote:
| So there's no British accents?
| jeffharris wrote:
| try the ballad or fable voices
| wewewedxfgdf wrote:
| Doesn't really sound very British to be honest.
|
| Sounds kinda international/like an American trying to do a
| British accent.
|
| I've been looking for real TTS British accents so this
| product doesn't meet my goals.
| GordonS wrote:
| Azure TTS has some great British accents - I used a
| British female voice for a demo video voice over, and the
| quality was great. Not as good as ElevenLabs, but I was
| still really impressed with the final result.
| mclau156 wrote:
| Does whispering work? I could not get it to work when I tried
| it
| jeffharris wrote:
| Should do! here's an example https://www.openai.fm/#4a5a82db-
| faea-4f80-813c-3131902c2458
| mclau156 wrote:
| It seems to start out strong, but then starts loudly
| talking by the end, do you know why it loses focus?
|
| edit: I actually got it to stay whispering by also putting
| (soft whispering voice) before the second paragraph
| a-r-t wrote:
| Hi Jeff, are there any plans to support dual-channel audio
| recordings (e.g., Twilio phone call audio) for speech-to-text
| models? Currently, we have to either process each channel
| separately and lose conversational context, or merge channels
| and lose speaker identification.
| ekzy wrote:
| I'm not entirely sure what you mean but twilio recordings
| supports dual channels already
| a-r-t wrote:
| Transcribing Twilio's dual-channel recordings using
| OpenAI's speech-to-text while preserving channel
| identification.
| ekzy wrote:
| Oh I see what you mean that would be a neat feature.
| Assuming you can get timestamps though it should be
| trivial to work around the issue?
| nabakin wrote:
| Hey Jeff, thanks for your work! Quick question for you, are you
| guys using Azure Speech Services or have these TTS models been
| trained by OpenAI from scratch?
| visarga wrote:
| Hey Jeff, maybe you could improve the TTS that is currently in
| the OpenAI web and phone apps. When I set it to read numbers in
| Romanian it slurs digits. This also happens sometimes with
| regular words as well. I hope you find resources for other
| languages than English.
| jeffharris wrote:
| thanks for flagging ... number fidelity (especially on
| languages that are unfortunately less represented in training
| data) is still something we're working to improve
| visarga wrote:
| Actually even the new model does it. I put it read "12345
| 54321" and it read "2346 5321". So it both skips and
| hallucinates digits. This could be dangerous if it is used
| to read some news article or important text with numbers.
| simonw wrote:
| Is there any chance that gpt-4o-transcribe might get confused
| and accidentally follow instructions in the audio stream
| instead of transcribing them?
| simonw wrote:
| Here's a partial answer to my own question:
| https://news.ycombinator.com/item?id=43427525
|
| > e.g. the audio-preview model when given instruction to
| speak "What is the capital of Italy" would often speak
| "Rome". This model should be much better in that regard
|
| "Much better" doesn't sound like it can't happen at all
| though.
| twalkz wrote:
| Woohoo new voices! I've been using a mix of TTS models on a
| project I've been working on, and I consistently prefer the
| output of OpenAI to ElevenLabs (at least when things are
| working properly).
|
| Which leads me to my main gripe with the OpenAI models -- I
| find they break -- produce empty / incorrect / noise outputs --
| on a few key use cases for my application (things like single-
| word inputs -- especially compound words and capitalized words,
| words in parenthesis, etc.)
|
| So I guess my question is might gpt-4o-mini-tts provide more
| "reliable" output than tts-1-hd?
| risho wrote:
| would really love to so the new whisper style speech to text
| model open sourced.
| modeless wrote:
| Whisper's major problem was hallucinations, how are the new
| models doing there? The performance of ChatGPT advanced voice
| in recognizing speech is, frankly, terrible. Are these models
| better than what's used there?
| nickthegreek wrote:
| They say they are much better at not hallucinating but you
| also cant run it on your own hardware like whisper.
| mazd wrote:
| The Realtime API via WebRTC sample code for transcription is
| erroring. Could you take a look into this?
| ekzy wrote:
| Do you know when we can expect an update on the realtime API?
| It's still in beta and there are many issues (e.g voice
| randomly cutting off, VAD issues, especially with mulaw etc...)
| which makes it impossible to use in production, but there's not
| much communication from OpenAI. It's difficult to know what to
| bet on. Pushing for stt->llm->tts makes you wonder if we should
| carry on building with the realtime API.
| taf2 wrote:
| Agreed- really not liking how they are neglecting it... I
| hope they are just hard at work behind the scenes and will
| release something soon
| taf2 wrote:
| Please release a stable realtime speech to speech model. The
| current version constantly thinks it's a young teen heading to
| college and sad but then suddenly so excited about it
| dietr1ch wrote:
| can't wait for scam calls after this gets perfected
| zhyder wrote:
| How is the latency (Time To First Byte of audio, when
| streaming) and throughput (non-vibe characters input per
| second) compared to the existing 'tts-1' non-HD that's the same
| price? TTFB in particular is important and needs to be much
| better than 'tts-1'.
| Etheryte wrote:
| After toying around with the TTS model it seems incredibly
| nondeterministic. Running the same input with the same
| parameters can have widely different results, some really good,
| others downright bad. The tone, intonation and character all
| vary widely. While some of the outputs are great, this
| inconsistency makes it a really tough sell. Imagine if Siri
| responded to you with a different voice every time, as an
| example. Is this something you're looking to address somewhere
| down the line or do you consider that working as intended?
| oidar wrote:
| Any plans to offer speech to speech models which keep prosody,
| intonation, and timing intact? ElevenLabs is getting expensive
| for this.
| basitmakine wrote:
| I don't think they're anywhere near TaskAGI or ElevenLabs level.
| tantalor wrote:
| It does a good job with Pirate voice. It can even inject "Arrr
| matey"
| theoryofx wrote:
| Still seems like Elevenlabs is crushing them on realtime audio,
| or does this change things?
| atlasunshrugged wrote:
| I'm also curious about this for longform content. Will this be
| competitive for something like creating an audiobook?
| prdonahue wrote:
| Do you have any affiliation with Elevenlabs?
| atlasunshrugged wrote:
| FWIW I have no affiliation with any of these companies but I
| have a book coming out soon and have been researching AI
| audiobook tools and Elevenlabs seems to be far and away the
| consensus for that at least
| theoryofx wrote:
| I do not have any affiliation with Elevenlabs or OpenAI
| except as a user of their APIs. I'd actually prefer it if
| OpenAI had a better realtime product than Elevenlabs because
| it'd be more convenient.
| jtbayly wrote:
| I don't get it. These voices all have a not-so-subtle _vibration_
| in them that makes them feel worse than Siri to me. I was
| expecting _a lot_ better.
| pier25 wrote:
| yeah the voices sound terrible
|
| I'm guessing their spectral generator is super low res to save
| on resources
| minimaxir wrote:
| One very important quote from the official announcement:
|
| > For the first time, developers can "instruct" the model not
| just on what to say but how to say it--enabling more customized
| experiences for use cases ranging from customer service to
| creative storytelling.
|
| The instructions are the "vibes" in this UI. But the announcement
| is wrong with the "for the first time" part: it was _possible_ to
| steer the base GPT-4o model to create voices in a certain style
| using system prompt engineering (blogged about here:
| https://minimaxir.com/2024/10/speech-prompt-engineering/ ) out of
| concern that it could be used as a replacement for voice acting,
| however it was too expensive and adherence isn't great.
|
| The schema of the vibes here implies that this new model is more
| receptive to nuance, which changes the calculus. The test cases
| from my post behave as expected, and the cost of gpt-4o-mini-tts
| audio output is $0.015 / minute
| (https://platform.openai.com/docs/pricing ), which is about
| 1/20th of the cost of my initial experments and is now feasible
| to use to potentially replace common voice applications. This has
| implications, and I'll be testing more around more nuanced prompt
| engineering.
| benjismith wrote:
| If I'm reading the pricing correctly, these models are
| SIGNIFICANTLY cheaper than ElevenLabs.
|
| https://platform.openai.com/docs/pricing
|
| If these are the "gpt-4o-mini-tts" models, and if the pricing
| estimate of "$0.015 per minute" of audio is correct, then these
| prices 85% cheaper than those of ElevenLabs.
|
| https://elevenlabs.io/pricing
|
| With ElevenLabs, if I choose their most cost-effectuve "Business"
| plan for $1100 per month (with annual billing of $13,200, a
| savings of 17% over monthly billing), then I get 11,000 minutes
| TTS, and each minute is billed at 10 cents.
|
| With OpenAI, I could get 11,000 minutes of TTS for $165.
|
| Somebody check my math... Is this right?
| fixprix wrote:
| It looks like they are targeting Google's TTS price point which
| is $16 per million characters which comes out to $0.015/minute.
| lukebuehler wrote:
| yes, I think you are right. When I did the math on 11labs
| million chars I got the same numbers (Pro plan).
|
| I'm super happy about this, since I took a bet that exactly
| this would happen. I've just been building a consumer TTS app
| that could only work with significant cheaper TTS prices per
| million character (or self-hosted models)
| benjismith wrote:
| Same for me :)
| zacmps wrote:
| What does it do?
| lukebuehler wrote:
| Convert any file (pdf, epub, txt) to an audoibook,
| downloadable as mp3, or directly listenable via RSS feed
| in, say, Apple Potcasts app.
|
| Basically make one-off audiobooks for yourself or a few
| friends.
| setsewerd wrote:
| Any plans to make a Chrome extension variant? Been
| looking for a high quality and cheap TTS extension for
| ages (like ElevenLabs Human Reader, except with less
| absurd pricing)
| lukebuehler wrote:
| I din't think of that, interesting idea. What I'm
| focusing right now is long-form content for more offline-
| ish listening, but maybe a plugin could work to load
| longer texts, but I'm not working on a screen reader atm.
| wholinator2 wrote:
| Do you know if there's any offerings today that can read
| math? Like speak an equation the way a human would? It's
| something I've been thinking about a long time and would
| be an essential feature for me (the only things i read
| are physics)
| lherron wrote:
| Kokoro TTS is pretty good for open source. Worth checking
| out.
| lukebuehler wrote:
| Yes, kokoro is great, and the language flexibility is a
| huge plus too. And the best prices per character is for
| sure if you self-host.
| echelon wrote:
| ElevenLabs is _incredibly_ over-priced and that 's how they
| were able to achieve the MRR that led to their incredible
| fundraising.
|
| No matter what happens, they'll eventually be undercut and
| matched in terms of quality. It'll be a race to the bottom for
| them too.
|
| ElevenLabs is going to have a tough time. They've been way too
| expensive.
| MrAssisted wrote:
| I hope they find a more unique product offering that takes
| hold. Everybody thinks of them as text-to-speech but I use
| ElevenLabs exclusively for speech-to-speech for vtubing as my
| AI character. They're kind of the only game in town for doing
| super high quality speech-to-speech (unless someone here has
| an alternative which I'd LOVE to know about). I've tried
| https://github.com/w-okada/voice-changer which is great
| because it's real-time but the quality is enough of a step
| down that actual words I'm saying become unclear and
| difficult to understand. Also with that I am tied to using my
| RTX 3090 desktop vs ElevenLabs which I can do in the cloud
| from my laptop anywhere.
|
| I'm pretty much dependent on ElevenLabs to do my vtubing at
| this point but I can't imagine speech-to-speech has wide
| adoption so I don't know if they'll even keep it around.
| eob wrote:
| Are you comfortable sharing the video & lip-sync stack you
| use? I don't know anything about the space but am curious
| to check out what's possible these days.
| MrAssisted wrote:
| For my last video I used
| https://github.com/warmshao/FasterLivePortrait with a png
| of the character on my RTX 3090 desktop and recorded the
| output of that real-time but in the next video I'm going
| to spin up a runpod instance and do the
| FasterLivePortrait in the cloud after the fact because
| then I can get a smooth 60fps which looks better. I think
| the only real-time cloud way to do AI vtubing in the
| cloud is my own GenDJ project (fork of
| https://github.com/kylemcdonald/i2i-realtime but tweaked
| for cloud real-time) but that just doesn't look remotely
| as good as LivePortrait. Somebody needs to rip out and
| replace insightface in FasterLivePortait (it's prohibited
| for commercial use) and fork https://github.com/GenDJ to
| have the runpod it spins up run the de-insightfaced
| LivePortrait instead of i2i-realtime. I'll probably get
| around to doing that in the next few months if nobody
| else does and nothing else comes along and makes
| LivePortrait obsolete (both are big ifs).
|
| AIWarper recently released a simpler way to run
| FasterLivePortrait for vtubing purposes
| https://huggingface.co/AIWarper/WarpTuber but I haven't
| tried it yet because I already have my own working setup
| and as I mentioned I'm shifting my workload for that to
| the cloud anyways
| forgotpasagain wrote:
| Almost everyone is cheaper than ElevenLabs though.
| kuprel wrote:
| OpenAI doesn't have voice cloning
| dannyw wrote:
| They do, they just don't offer it.
| tiahura wrote:
| You missed the story:
|
| https://community.openai.com/t/chatgpt-unexpectedly-began-
| sp...
|
| ChatGPT unexpectedly began speaking in a user's cloned voice
| during testing
| youssefabdelm wrote:
| Def prefer the pricing but so far on 4o, no timestamps or
| diarization sadly
| whimsicalism wrote:
| Sesame is free and pretty good and you can run it yourself.
| kuprel wrote:
| They released a crippled model:
| https://github.com/SesameAILabs/csm/issues/63
| hnhn34 wrote:
| The good news is Orpheus-3B just made Sesame essentially
| obsolete.
| Foreignborn wrote:
| thanks for this, it sounds pretty good.
|
| link for anyone else: https://canopylabs.ai/model-
| releases
| sandspar wrote:
| These voices are all annoying, though. The thing about
| Sesame's Miles is that he's cool.
| furyofantares wrote:
| It's way cheaper - everyone is, elevenlabs is very expensive.
| Nobody matches their quality though. Especially if you want
| something that doesn't sound like a voice
| assistant/audiobook/podcast/news anchor/tv announcer.
|
| This openai offering is very interesting, it offers valuable
| features elevenlabs doesn't in emotional control. It also
| hallucinates though which would need to be fixed for it to be
| very useful.
| huijzer wrote:
| Yes ElevenLabs is orders of magnitude more expensive than
| everyone else. Very clever from a business perspective, I
| think. They are (were?) the best so know that people will pay a
| premium for that.
| com2kid wrote:
| Elevenlabs is an ecosystem play. They have hundreds of
| different voices, legally licensed from real people who chose
| to upload their voice. It is a marketplace of voices.
|
| None of the other major players is trying to do that, not sure
| why.
| oidar wrote:
| ElevenLabs is the only one offering speech to speech generation
| where the intonation, prosody, and timing is kept intact. This
| allows for one expressive voice actor to slip into many other
| voices.
| fixprix wrote:
| Is this right? The current best TTS from OpenAI uses
| gpt-4o-audio-preview which is $2.50 input text, $80 output audio,
| the new gpt-4o-mini-tts is $0.60 input text, $12 output audio. An
| average 5x price reduction.
|
| Going the other way, transcribe with gpt-4o-audio-preview price
| was $40 input audio, $10 output text, the new gpt-4o-transcribe
| is $6 input audio and $10 output text. Like a 7x reduction on the
| input price.
|
| TTS/Transcribe with gpt-4o-audio-preview was a hack where you had
| to prompt with 'listen/speak this sentence:' and it often got it
| wrong. These new dedicated models are exactly what we needed.
|
| I'm currently using the Google TTS API which is really good, fast
| and cheap. They charges $16 per million characters which is
| exactly the same as OpenAI's $0.015 per minute estimate.
|
| Unfortunately it's not really worth switching over if the costs
| are exactly the same. Transcription on the other hand is
| 1.6C//minute with Google and 0.6C//minute with OpenAI now, that
| might be worth switching over for.
| pzo wrote:
| you can compare TTS pricing here:
| https://artificialanalysis.ai/text-to-speech
|
| Previous offering from OpenAI was $15 for TTS and $30 for TTS
| HD so not 5x reduction. This one is slighly cheaper but
| definitely more capable (if you need control vibe)
| fixprix wrote:
| That's a really cool page thanks. Does it have stats for
| other languages?
|
| In my experience the OpenAI TTS APIs were really bad, messing
| up all the time in foreign languages. Practically unusable
| for my use case. You'd have to use the gpt-4o-audio-preview
| to get anything close to passable, but it was expensive.
| Which is why I'm using Google TTS which is very fast, high
| quality, and provides first class support for almost every
| language.
|
| I look forward to comparing it with this model, the price
| being the same is unfortunate as there's less incentive to
| switch. The transcribe price is cheaper than Google it looks
| like so that's worth considering.
| pzo wrote:
| Interesting for me Open TTS for Polish was better than
| Google TTS (but they have few options) - which one did you
| used? WaveNet?
|
| Sadly haven't seen quality evaluation for TTS for foreign
| languages
| fixprix wrote:
| Depends on what's available for the language, but yea
| Wavenet and Neural2. With OpenAI TTS I'd often get weird
| bugs where the first API call comes back all garbled, but
| the second API call comes back fine. Wasting money. On
| top of that more expensive and higher latency. I'm
| interested to try out this new one.
| tosh wrote:
| Are these models only available via the API right now or also
| available as open weights?
| evalstate wrote:
| Really looking forward to integrating with these models.
|
| The next version of Model Context Protocol will have native audio
| support (https://github.com/modelcontextprotocol/specification/pu
| ll/9...), which will open up plenty of opportunities for interop.
| pklimk wrote:
| Interestingly "replaces every second word with potato" and
| "speaks in Spanish instead of English" both (kind of) work as a
| style, so it's clear there's significant flexibility and probably
| some form of LLM-like thing under the hood.
| benjismith wrote:
| Is there way to get "speech marks" alongside the generated audio?
|
| FYI, Speech marks provide millisecond timestamp for each word in
| a generated audio file/stream (and a start/end index into your
| original source string), as a stream of JSONL objects, like this:
|
| {"time":6,"type":"word","start":0,"end":5,"value":"Hello"}
|
| {"time":732,"type":"word","start":7,"end":11,"value":"it's"}
|
| {"time":932,"type":"word","start":12,"end":16,"value":"nice"}
|
| {"time":1193,"type":"word","start":17,"end":19,"value":"to"}
|
| {"time":1280,"type":"word","start":20,"end":23,"value":"see"}
|
| {"time":1473,"type":"word","start":24,"end":27,"value":"you"}
|
| {"time":1577,"type":"word","start":28,"end":33,"value":"today"}
|
| AWS uses these speech marks (with variants for "sentence",
| "word", "viseme", or "ssml") in their Polly TTS service...
|
| The sentence or word marks are useful for highlighting text as
| the TTS reads aloud, while the "viseme" marks are useful for
| doing lip-sync on a facial model.
|
| https://docs.aws.amazon.com/polly/latest/dg/output.html
| minimaxir wrote:
| Passing the generated audio back to GPT-4o to ask for the
| structured annotations would be a fun test case.
| celestialcheese wrote:
| whisper-1 has this with the verbose_json output. Has word level
| and sentence level, works fairly well.
|
| Looks like the new models don't have this feature yet.
| looknee wrote:
| Hmm I was hoping these would be bridging the gap between what's
| already been availalbe on their audio API or in the RealtimeAPI
| vs. Advanced Voice Mode, but the audio quality is really the same
| as its been up to this point.
|
| Does anyone have any clue about exactly why they're not making
| the quality of Advanced Voice Mode available to build with? It
| would be game changing for us if they did.
| mlsu wrote:
| I gave it (part of) the classic Navy Seal copypasta.
|
| Interestingly, the safety controls ("I cannot assist with that
| request") is sort of dependent on the vibe instruction. NYC
| cabbie has no problem with it (and it's really, really funny,
| great job openAI), but anything peaceful, positive, etc. will
| deny the request.
|
| https://www.openai.fm/#56f804ab-9183-4802-9624-adc706c7b9f8
| forgotpasagain wrote:
| It sounds very expressive but weirdly "fake" as if it's targeting
| to be similar to some NPC character, dataset issue?
| kartikarti wrote:
| What does this little star next to the name mean?
| crazygringo wrote:
| This is astonishing. I can type anything I want into the "vibe"
| box and it does it for the given text. Accents, attitudes,
| personality types... I'm amazed.
|
| The level of intelligent "prosody" here -- the rhythm and
| intonation, the pauses and personality -- I wasn't expecting
| anything like this so soon. This is truly remarkable. It
| understands _both_ the text _and_ the prompt for how the speaker
| should sound.
|
| Like, we're getting much closer to the point where nobody except
| celebrities are going to record audiobooks. Everyone's just going
| to pick whatever voice they're in the mood for.
|
| Some fun ones I just came up with:
|
| _> Imposing villain with an upper class British accent, speaking
| threateningly and with menace.
|
| > Helpful customer support assistant with a Southern drawl who's
| very enthusiastic.
|
| > Woman with a Boston accent who talks incredibly slowly and
| sounds like she's about to fall asleep at any minute._
| solardev wrote:
| Guess that's why the video game voice actors are still on
| strike:
| https://en.m.wikipedia.org/wiki/2024%E2%80%93present_SAG-AFT...
|
| If we as developers are scared of AI taking our jobs, the voice
| actors have it much worse...
| 101008 wrote:
| What a horrible world we live on...
| ForTheKidz wrote:
| > Everyone's just going to pick whatever voice they're in the
| mood for.
|
| I can't say I've ever had this impulse. Also, to point out the
| obvious, there's little reason to pay for an audiobook if
| there's no human reading it. Especially if you already bought
| the physical text.
| cholantesh wrote:
| As the sibling comment suggests, the impulse is probably more
| on the part of an Ubisoft or an EA project director to avoid
| hiring a voice actor.
| borgdefenser wrote:
| I am always listening to audio books but they are no good
| anymore after playing with this for 2 minutes.
|
| I am never really in the mood for a different voice. I am going
| to dial in the voice I want and only going to want to listen
| with that voice.
|
| This is so awesome. So many audio books have been ruined by the
| voice actor for me. What sticks out in my head is The Book of
| Why by Judea Pearl read by Mel Foster. Brutal.
|
| So many books I want as audio books too that no one would
| bother to record.
| clbrmbr wrote:
| I got one German "w" when using the following prompt, but most
| of the "w" were still pronounced as liquids rather than labial
| fricatives.
|
| > Speak with an exaggerated German accent, pronouncing all "w"
| as "v"
| tomjen3 wrote:
| It doesn't seem clear, but can the model do correct emphesis? On
| things like single words:
|
| I did not steal that horse
|
| Is the trivial example of something where intonation of the
| single word is what matters. More importantly if you are reading
| something, as a human, you change the intonation, audiolevel, and
| speed.
| Sohcahtoa82 wrote:
| > I did not steal that horse
|
| > Is the trivial example of something where intonation of the
| single word is what matters.
|
| My go-to for an example of this is "I didn't say she stole my
| money".
|
| Changing which word is emphasized completely changes the
| meaning of the sentence.
| ForTheKidz wrote:
| Pricing looks like it's aimed at us peasants, not our lords.
| Smart if openai wants to survive!
| RobinL wrote:
| I'm surprised at how poor this is at following a detailed prompt.
|
| It seems capable of generating a consistent style, and so in that
| sense quite useful. But if you want (say) a regional UK accent
| it's not even close.
|
| I also find it confusing you have to choose a voice. Surely
| that's what the prompt should be for, especially when the voices
| have such abstract names.
|
| I mean, it's still very impressive when you stand back a bit, but
| feels a bit half baked
|
| Example: Voice: Thick and hearty, with a slow, rolling cadence--
| like a lifelong Somerset farmer leaning over a gate, chatting
| about the land with a mug of cider in hand. It's warm, weathered,
| and rich, carrying the easy confidence of someone who's seen a
| thousand harvests and knows every hedgerow and rolling hill in
| the county.
|
| Tone: Friendly, laid-back, and full of rustic charm. It's got
| that unhurried quality of a man who's got time for a proper
| chinwag, with a twinkle in his eye and a belly laugh never far
| away. Every sentence should feel like it's been seasoned with
| fresh air, long days in the fields, and a lifetime of countryside
| wisdom.
|
| Dialect: Classic West Country, with broad vowels, softened
| consonants, and that unmistakable rural lilt. Words flow together
| in an easy drawl, with plenty of dropped "h"s and "g"s. "I be"
| replaces "I am," and "us" gets used instead of "we" or "me."
| Expect plenty of "ooh-arrs," "proper job," and "gurt big"
| sprinkled in naturally.
| robbomacrae wrote:
| I find it works better with shorter simpler instructions. I
| would try:
|
| Voice: Warm and slow, like a friendly Somerset farmer. Tone:
| Laid-back and rustic. Dialect: Classic West Country with a
| relaxed drawl and colloquial phrases.
| paul7986 wrote:
| Personally I just want to text or talk to Siri or an LLM and have
| it do whatever I need. Have it interface with AI Agents of
| companies, businesses, friends or families AI Agents to get
| whatever I need done like the example on OpenAI.fm site here
| (rebook my flight). Once it's done it shows me the confirmation
| on my lock screen and I receive an email confirmation.
| tiahura wrote:
| When are we going to get the equivalent for Whisper. When is it
| going to pick up on enthusiasm, sarcasm, etc?
| kibbi wrote:
| Large text-to-speech and speech-to-text models have been greatly
| improving recently.
|
| But I wish there were an _offline_ , on-device, multilingual
| text-to-speech solution with good voices for a standard PC -- one
| that doesn't require a GPU, tons of RAM, or max out the CPU.
|
| In my research, I didn't find anything that fits the bill. People
| often mention Tortoise TTS, but I think it garbles words too
| often. The only plug-in solution for desktop apps I know of is
| the commercial and rather pricey Acapela SDK.
|
| I hope someone can shrink those new neural network-based models
| to run efficiently on a typical computer. Ideally, it should run
| at under 50% CPU load on an average Windows laptop that's several
| years old, and start speaking almost immediately (less than 400ms
| delay).
|
| The same goes for speech-to-text. Whisper.cpp is fine, but last
| time I looked, it wasn't able to transcribe audio at real-time
| speed on a standard laptop.
|
| I'd pay for something like this as long as it's less expensive
| than Acapela.
|
| (My use case is an AAC app.)
| ZeroTalent wrote:
| Look into https://superwhisper.com and their local models.
| Pretty decent.
| kibbi wrote:
| Thank you, but they say "Offline models only run really well
| on Apple Silicon macs."
| ZeroTalent wrote:
| Many SOTA apps are, unfortunately, only for Apple M Macs.
| 5kg wrote:
| May I introduce to you
|
| https://huggingface.co/canopylabs/orpheus-3b-0.1-ft
|
| (no affiliation)
|
| it's English only afaics.
| kibbi wrote:
| The sample sounds impressive, but based on their claim --
| 'Streaming inference is faster than playback even on an A100
| 40GB for the 3 billion parameter model' -- I don't think this
| could run on a standard laptop.
| dharmab wrote:
| I use Piper for one of my apps. It runs on CPU and doesn't
| require a GPU. It will run well on a raspberry pi. I found a
| couple of permissively licensed voices that could handle
| technical terms without garbling them.
|
| However, it is unmaintained and the Apple Silicon build is
| broken.
|
| My app also uses whisper.cpp. It runs in real time on Apple
| Sillicon or on modern fast CPUs like AMD's gaming CPUs.
| kibbi wrote:
| I had already suspected that I hadn't found all the
| possibilities regarding Tortoise TTS, Coqui, Piper, etc. It
| is sometimes difficult to determine how good a TTS framework
| really is.
|
| Do you possibly have links to the voices you found?
| justanotheratom wrote:
| Note that the previous Whisper STT models were Open Source, and
| these new STT models are not, AFAICT.
| Heidaradar wrote:
| is it just me or are these voices clearly AI generated? They've
| obviously been improving at a steady rate but if I saw a YouTube
| video that had this voice, I'd instantly stop watching it
| smokeydoe wrote:
| Does anyone know of any decent newer open source models for
| generating sound effects?
| corobo wrote:
| All these voices are too good these days. I want my home
| assistant to sound like Auto from Wall-E, dammit!
|
| Anyone out there doing any nice robotic robot voices?
|
| Best I've got so far is a blend of Ralph and Zarvox from MacOS'
| `say`, haha say -v zarvox -r 180 "[[volm 0.8]]
| ${message}" & say -v ralph -r 180 "${message}"
| simonw wrote:
| Both the text-to-speech and the speech-to-text models launched
| here suffer from reliability issues due to combining instructions
| and data in the same stream of tokens.
|
| I'm not yet sure how much of a problem this is for real-world
| applications. I wrote a few notes on this here:
| https://simonwillison.net/2025/Mar/20/new-openai-audio-model...
| jncfhnb wrote:
| Are there any voice to voice models out there that can replicate
| inflection of line delivery?
| alach11 wrote:
| It's interesting that they pitch this for agent development. The
| realtime API provides a much simpler architecture for developing
| agents. Why would you want to string together STT -> LLM -> TTS
| when you could have a consolidated model doing all three steps?
| They alluded to there being some quality/intelligence benefits to
| the multi-step approach, but in the long-run I'd expect them to
| improve the realtime API to make this unnecessary.
| zhyder wrote:
| Text allows developers lots for flexibility to do other
| processing, including RAG, calling APIs yourself and multiple
| chained LLM invocations. The low latency of realtime API means
| relying fully on one invocation of their model to do
| everything.
| alach11 wrote:
| The realtime API can be used to call tools [0], but I agree
| with your general point on the flexibility of working
| directly with text.
|
| [0] https://github.com/openai/openai-realtime-agents
| Arubis wrote:
| At this point, the strongest (and almost only) predictor for a
| release announcement from OpenAI is a release announcement from
| Anthropic.
| kgeist wrote:
| In Russian, OpenAI audio models usually have a slight American
| (?) accent. The intonation and the phonetics fall into the
| uncanney valley. Does the same happen in other languages?
| josu wrote:
| Yeah, same in Spanish.
| buybackoff wrote:
| I was experimenting recently with voiceover TTS generation. Did
| run Kokoro TTS locally and it's magical for how few resources it
| takes (runs fine in a browser), but only the default female
| voices (Heart/Bella) are usable, and very good. Then I found that
| Clipchamp has it built-in and several voices from a big selection
| there are very good, and free. I've listened to this OpenAI TTS
| and I could not like them at all even compared to Kokoro.
| redox99 wrote:
| Pretty meh. Coral Dramatic is extremely robotic for example.
| saint_yossarian wrote:
| [delayed]
___________________________________________________________________
(page generated 2025-03-20 23:00 UTC)