[HN Gopher] Show HN: Affordable text-to-speech for long-form con...
___________________________________________________________________
Show HN: Affordable text-to-speech for long-form content
Hi HN, I'm Michael, creator of AudiowaveAI. I started this project
out of frustration when I couldn't find an audiobook version of
_Make_ by Pieter Levels. The available text-to-speech options were
either too robotic, overly complex, or simply too costly. It works
really well for non-fiction long-form content (i.e. hours of
audio). It's early days for AudiowaveAI, and I'm looking for
feedback to improve the product. Try it out and share your
thoughts: [AudiowaveAI](https://audiowaveai.com). Thanks!
Author : yagudaev
Score : 45 points
Date : 2024-05-23 09:18 UTC (3 days ago)
(HTM) web link (www.audiowaveai.com)
(TXT) w3m dump (www.audiowaveai.com)
| andrewinardeer wrote:
| I've been looking for something like this. Thank you.
|
| A couple of questions:
|
| How do I delete projects?
|
| I must have tapped three times after submitting a Wikipedia
| article and it created three projects that apparently cannot be
| deleted.
|
| How do I delete my account?
|
| And for $15 I get credits. How many credits do I get foe $15? Is
| each credit a word translate? 1 credit == 1 word translated to
| audio?
| yagudaev wrote:
| Hi Andrew , these are fantastic questions. Let me answer them
| one at a time:
|
| > How do I delete projects?
|
| Three dots on the side of the project, you can delete it
|
| > I must have tapped three times after submitting a Wikipedia
| article and it created three projects that apparently cannot be
| deleted.
|
| > How do I delete my account?
|
| Just email me support@audiowaveai.com with form that email and
| I'll delete it of you. Still MVP no functionality for that yet.
|
| > And for $15 I get credits. How many credits do I get foe $15?
| Is each credit a word translate? 1 credit == 1 word translated
| to audio?
|
| 1 credit = 1 character. You are right I need to be more clear
| on it. $15 would give your about 10hrs of audio or 100 articles
| (~5-6mins). ElevenLabs will cost you $99 for the same audio.
| nprateem wrote:
| I looked into using Google cloud TTS or Azure for this but it
| was too expensive.
|
| How have you got the costs so low? Also the GCP voices don't
| have as natural intonation. How did you do that?
|
| I really didn't think there would be a market for this
| either.
| DreaminDani wrote:
| This is really cool! One quick note about your marketing copy,
| though: > Audio for humans, not robots
|
| There are plenty of blind folks who use traditional text to
| speech for navigating our devices. We prefer the robot text at
| ridiculously high speeds. We're humans too.
|
| I would love the option to switch to a more natural voice for
| more literary text (or even a fan fic) so I'll definitely be
| checking this out
| qchris wrote:
| > I would love the option to switch to a more natural voice for
| more literary text (or even a fan fic) so I'll definitely be
| checking this out
|
| I'm curious if it would be possible to do some kind of analysis
| to determine the number of individual characters in the text
| who are speaking, and then assign an appropriate voice to each
| of the characters. So if you had something like descriptive
| language interspersed with a conversation between two
| characters, that you'd have three voices (a narrator, Character
| A, and Character B) that are consistent across the text.
|
| For more complex writing with many characters, you'd probably
| need a wide library of possible voices, and the analysis piece
| would need to spot-on, since it would be very confusing to have
| one characters' lines spoken by the wrong voice.
|
| Regarding fanfics, many authors give (or withhold) permissions
| around creating derivative versions of their work via avenues
| like ficbinding. Before using a tool like this to create an
| audio version of their writing, I'd suggest reaching out to a
| fic's author to see if they'd be okay with that. For personal-
| only use, though, and especially if it's in context of
| accessibility for visually-impaired folks, I imagine that many
| of them would probably be okay with it.
| fortydegrees wrote:
| Is this a custom trained TTS model or is it an implementation of
| something like StyleTTSv2?
| maddynator wrote:
| I looked into this problem a while back and haven't looked at
| since.
|
| The base ai model sounded like whisper ai from meta. Did you
| train the voice yourself or is it one of defaults?
|
| I am always curious as to what copyright issues products like
| this run into. Also whats the stack like to build something like
| this?
| lyu07282 wrote:
| isn't whisper the speech to text model by openai? which model
| did you mean?
| 101008 wrote:
| Hey. I have published a non-fiction book, and i would like to
| publish the audiobook on Amazon (Audible, etc). Do you know if
| the output is accepted by them? What format should I provide my
| book to AudiowaveAI to receive a good audio? Does it understand
| chapter titles, quotations, etc?
| Aelius wrote:
| I have a use case for a niche audience:
|
| The videogame Final Fantasy XIV has a lot of text. A LOT of text.
|
| Someone has made a plugin to pipe text to external tts services,
| or a websocket. You talk to characters in game and hear the
| dialog read by the tts.
|
| https://github.com/karashiiro/TextToTalk
|
| For whatever reason, amazon poly only exposes middling quality
| voices to the plugin. And I'd rather not have an active AWS
| account for just this use case.
|
| ElevenLabs is supported by the plugin, but their service isn't
| really about tts and I'd have to pay the $220/yr tier to unlock
| further "pay as you go (per character)" with a budget of 100,000
| characters per month. A bit steep for using it only for in this
| one game.
|
| If someone could help plumb AudiowaveAI to this plugin, I'd
| gladly turn off AWS for this!
| smeej wrote:
| Is it possible to switch back and forth between the written text
| and the audio, like Amazon's Whispersync? I prefer reading with
| my eyes when I can (especially on my ereader, so with pagination
| instead of scrolling), but I would love to be able to flip
| narration on when I need to set the book down to do something
| like wash my dishes, then pick the book back up when I'm done.
|
| I've been looking for something that would let me synchronize
| Librivox recordings with Project Gutenberg epub files, but as
| much as I love the Librivox volunteers for their contributions, a
| lot of the recordings are such low audio quality that they're not
| fun to listen to. This would be a big step up, and there's no
| copyright worries for this use case because the works are in the
| public domain!
| steffenhk wrote:
| I've used Storyteller to create an epub book with Media Overlay
| but not sure it works in all ebook readers. It worked in
| Calibre.
|
| https://smoores.gitlab.io/storyteller/docs/what-is-this/
| smeej wrote:
| My ereader is an Android tablet under the hood. I don't know
| of any apps that can do this on Android, but I can go
| hunting!
| sphars wrote:
| Storyteller itself does have an app, just requires you to
| be hosting the service:
| https://smoores.gitlab.io/storyteller/docs/reading-your-
| book...
|
| It also says that BookFusion can read the files it
| produces:
| https://smoores.gitlab.io/storyteller/docs/reading-your-
| book...
| satvikpendem wrote:
| I use the app Moon+ Reader that can do this. It uses the built-
| in text-to-speech engine so if someone makes another engine
| with more natural speech, it can plug in seamlessly.
| afrederico wrote:
| Love this tool; please include the pricing on the front page
| before one has to sign up. Thanks!
| NayamAmarshe wrote:
| This looks nice a really nice product.
| hereme888 wrote:
| Good for you!
|
| So similar to my app. But I'm not a real programmer, so of course
| your is more refined.
|
| I almost launched the same exact online business.
|
| Here's my version (my github version is a bit less refined than
| my local code):
|
| https://github.com/sm18lr88/OpenAI_TTS_GUI
| lupusreal wrote:
| I've been using Piper for this. The quality is (in my subjective
| opinion) as good as the TTS built into MacOS is, it's open
| source, and it's so fast that you can run it in real time on a
| raspberry pi. On a real computer I can generate a whole audiobook
| in about 20 minutes.
|
| What I do is I split the book up into sentences, generate speech
| for each sentence and at the same time turn that sentence into
| subtitles. Then I combine the two and stitch them all together
| into a mp4 container with audio and a subtitle track using
| ffmpeg. mpv (and think VLC) can display subtitles synced to audio
| playback even when there is no video track.
| prox wrote:
| Thats genius! Was it a lot of work to set up?
| icev wrote:
| Interesting, I read epubs on Android using aiTTS as TTS engine
| using Google cloud voices.
|
| What I would really like is an option to download the whole book
| as mp3 for offline playback, and different voices for each
| character.
| satvikpendem wrote:
| Moon+ Reader has offline playback although the voices aren't as
| good, but on the bright side, if someone makes a local AI text-
| to-speech engine, then that can plug into the app and it'll
| work fully offline.
| anonu wrote:
| Curious what the technical implementation looks like. What kind
| of TTS are you using? How do you scale it? What are the costs
| involved?
| pmg101 wrote:
| You put "medicore content" in place of (I assume) "mediocre
| content".
___________________________________________________________________
(page generated 2024-05-26 23:02 UTC)