[HN Gopher] Show HN: Affordable text-to-speech for long-form con...
       ___________________________________________________________________
        
       Show HN: Affordable text-to-speech for long-form content
        
       Hi HN, I'm Michael, creator of AudiowaveAI. I started this project
       out of frustration when I couldn't find an audiobook version of
       _Make_ by Pieter Levels. The available text-to-speech options were
       either too robotic, overly complex, or simply too costly.  It works
       really well for non-fiction long-form content (i.e. hours of
       audio).  It's early days for AudiowaveAI, and I'm looking for
       feedback to improve the product. Try it out and share your
       thoughts: [AudiowaveAI](https://audiowaveai.com). Thanks!
        
       Author : yagudaev
       Score  : 45 points
       Date   : 2024-05-23 09:18 UTC (3 days ago)
        
 (HTM) web link (www.audiowaveai.com)
 (TXT) w3m dump (www.audiowaveai.com)
        
       | andrewinardeer wrote:
       | I've been looking for something like this. Thank you.
       | 
       | A couple of questions:
       | 
       | How do I delete projects?
       | 
       | I must have tapped three times after submitting a Wikipedia
       | article and it created three projects that apparently cannot be
       | deleted.
       | 
       | How do I delete my account?
       | 
       | And for $15 I get credits. How many credits do I get foe $15? Is
       | each credit a word translate? 1 credit == 1 word translated to
       | audio?
        
         | yagudaev wrote:
         | Hi Andrew , these are fantastic questions. Let me answer them
         | one at a time:
         | 
         | > How do I delete projects?
         | 
         | Three dots on the side of the project, you can delete it
         | 
         | > I must have tapped three times after submitting a Wikipedia
         | article and it created three projects that apparently cannot be
         | deleted.
         | 
         | > How do I delete my account?
         | 
         | Just email me support@audiowaveai.com with form that email and
         | I'll delete it of you. Still MVP no functionality for that yet.
         | 
         | > And for $15 I get credits. How many credits do I get foe $15?
         | Is each credit a word translate? 1 credit == 1 word translated
         | to audio?
         | 
         | 1 credit = 1 character. You are right I need to be more clear
         | on it. $15 would give your about 10hrs of audio or 100 articles
         | (~5-6mins). ElevenLabs will cost you $99 for the same audio.
        
           | nprateem wrote:
           | I looked into using Google cloud TTS or Azure for this but it
           | was too expensive.
           | 
           | How have you got the costs so low? Also the GCP voices don't
           | have as natural intonation. How did you do that?
           | 
           | I really didn't think there would be a market for this
           | either.
        
       | DreaminDani wrote:
       | This is really cool! One quick note about your marketing copy,
       | though: > Audio for humans, not robots
       | 
       | There are plenty of blind folks who use traditional text to
       | speech for navigating our devices. We prefer the robot text at
       | ridiculously high speeds. We're humans too.
       | 
       | I would love the option to switch to a more natural voice for
       | more literary text (or even a fan fic) so I'll definitely be
       | checking this out
        
         | qchris wrote:
         | > I would love the option to switch to a more natural voice for
         | more literary text (or even a fan fic) so I'll definitely be
         | checking this out
         | 
         | I'm curious if it would be possible to do some kind of analysis
         | to determine the number of individual characters in the text
         | who are speaking, and then assign an appropriate voice to each
         | of the characters. So if you had something like descriptive
         | language interspersed with a conversation between two
         | characters, that you'd have three voices (a narrator, Character
         | A, and Character B) that are consistent across the text.
         | 
         | For more complex writing with many characters, you'd probably
         | need a wide library of possible voices, and the analysis piece
         | would need to spot-on, since it would be very confusing to have
         | one characters' lines spoken by the wrong voice.
         | 
         | Regarding fanfics, many authors give (or withhold) permissions
         | around creating derivative versions of their work via avenues
         | like ficbinding. Before using a tool like this to create an
         | audio version of their writing, I'd suggest reaching out to a
         | fic's author to see if they'd be okay with that. For personal-
         | only use, though, and especially if it's in context of
         | accessibility for visually-impaired folks, I imagine that many
         | of them would probably be okay with it.
        
       | fortydegrees wrote:
       | Is this a custom trained TTS model or is it an implementation of
       | something like StyleTTSv2?
        
       | maddynator wrote:
       | I looked into this problem a while back and haven't looked at
       | since.
       | 
       | The base ai model sounded like whisper ai from meta. Did you
       | train the voice yourself or is it one of defaults?
       | 
       | I am always curious as to what copyright issues products like
       | this run into. Also whats the stack like to build something like
       | this?
        
         | lyu07282 wrote:
         | isn't whisper the speech to text model by openai? which model
         | did you mean?
        
       | 101008 wrote:
       | Hey. I have published a non-fiction book, and i would like to
       | publish the audiobook on Amazon (Audible, etc). Do you know if
       | the output is accepted by them? What format should I provide my
       | book to AudiowaveAI to receive a good audio? Does it understand
       | chapter titles, quotations, etc?
        
       | Aelius wrote:
       | I have a use case for a niche audience:
       | 
       | The videogame Final Fantasy XIV has a lot of text. A LOT of text.
       | 
       | Someone has made a plugin to pipe text to external tts services,
       | or a websocket. You talk to characters in game and hear the
       | dialog read by the tts.
       | 
       | https://github.com/karashiiro/TextToTalk
       | 
       | For whatever reason, amazon poly only exposes middling quality
       | voices to the plugin. And I'd rather not have an active AWS
       | account for just this use case.
       | 
       | ElevenLabs is supported by the plugin, but their service isn't
       | really about tts and I'd have to pay the $220/yr tier to unlock
       | further "pay as you go (per character)" with a budget of 100,000
       | characters per month. A bit steep for using it only for in this
       | one game.
       | 
       | If someone could help plumb AudiowaveAI to this plugin, I'd
       | gladly turn off AWS for this!
        
       | smeej wrote:
       | Is it possible to switch back and forth between the written text
       | and the audio, like Amazon's Whispersync? I prefer reading with
       | my eyes when I can (especially on my ereader, so with pagination
       | instead of scrolling), but I would love to be able to flip
       | narration on when I need to set the book down to do something
       | like wash my dishes, then pick the book back up when I'm done.
       | 
       | I've been looking for something that would let me synchronize
       | Librivox recordings with Project Gutenberg epub files, but as
       | much as I love the Librivox volunteers for their contributions, a
       | lot of the recordings are such low audio quality that they're not
       | fun to listen to. This would be a big step up, and there's no
       | copyright worries for this use case because the works are in the
       | public domain!
        
         | steffenhk wrote:
         | I've used Storyteller to create an epub book with Media Overlay
         | but not sure it works in all ebook readers. It worked in
         | Calibre.
         | 
         | https://smoores.gitlab.io/storyteller/docs/what-is-this/
        
           | smeej wrote:
           | My ereader is an Android tablet under the hood. I don't know
           | of any apps that can do this on Android, but I can go
           | hunting!
        
             | sphars wrote:
             | Storyteller itself does have an app, just requires you to
             | be hosting the service:
             | https://smoores.gitlab.io/storyteller/docs/reading-your-
             | book...
             | 
             | It also says that BookFusion can read the files it
             | produces:
             | https://smoores.gitlab.io/storyteller/docs/reading-your-
             | book...
        
         | satvikpendem wrote:
         | I use the app Moon+ Reader that can do this. It uses the built-
         | in text-to-speech engine so if someone makes another engine
         | with more natural speech, it can plug in seamlessly.
        
       | afrederico wrote:
       | Love this tool; please include the pricing on the front page
       | before one has to sign up. Thanks!
        
       | NayamAmarshe wrote:
       | This looks nice a really nice product.
        
       | hereme888 wrote:
       | Good for you!
       | 
       | So similar to my app. But I'm not a real programmer, so of course
       | your is more refined.
       | 
       | I almost launched the same exact online business.
       | 
       | Here's my version (my github version is a bit less refined than
       | my local code):
       | 
       | https://github.com/sm18lr88/OpenAI_TTS_GUI
        
       | lupusreal wrote:
       | I've been using Piper for this. The quality is (in my subjective
       | opinion) as good as the TTS built into MacOS is, it's open
       | source, and it's so fast that you can run it in real time on a
       | raspberry pi. On a real computer I can generate a whole audiobook
       | in about 20 minutes.
       | 
       | What I do is I split the book up into sentences, generate speech
       | for each sentence and at the same time turn that sentence into
       | subtitles. Then I combine the two and stitch them all together
       | into a mp4 container with audio and a subtitle track using
       | ffmpeg. mpv (and think VLC) can display subtitles synced to audio
       | playback even when there is no video track.
        
         | prox wrote:
         | Thats genius! Was it a lot of work to set up?
        
       | icev wrote:
       | Interesting, I read epubs on Android using aiTTS as TTS engine
       | using Google cloud voices.
       | 
       | What I would really like is an option to download the whole book
       | as mp3 for offline playback, and different voices for each
       | character.
        
         | satvikpendem wrote:
         | Moon+ Reader has offline playback although the voices aren't as
         | good, but on the bright side, if someone makes a local AI text-
         | to-speech engine, then that can plug into the app and it'll
         | work fully offline.
        
       | anonu wrote:
       | Curious what the technical implementation looks like. What kind
       | of TTS are you using? How do you scale it? What are the costs
       | involved?
        
       | pmg101 wrote:
       | You put "medicore content" in place of (I assume) "mediocre
       | content".
        
       ___________________________________________________________________
       (page generated 2024-05-26 23:02 UTC)