[HN Gopher] Neutts-air - Open-source, on device TTS
       ___________________________________________________________________
        
       Neutts-air - Open-source, on device TTS
        
       Author : nopelynopington
       Score  : 23 points
       Date   : 2025-10-06 09:06 UTC (3 days ago)
        
 (HTM) web link (github.com)
 (TXT) w3m dump (github.com)
        
       | nopelynopington wrote:
       | If this lives up to the demo it's a huge development for anyone
       | looking to do realistic tts without paying to use an API
        
       | gardnr wrote:
       | The model weighs 1.5GB [1] (the q4 quant is ~500MB)
       | 
       | The demo is impressive. It uses reference audio at inference
       | time, and it looks like the training code is mostly available
       | [2][3] with a reference dataset [4] as well.
       | 
       | From the README:
       | 
       | > NeuTTS Air is built off Qwen 0.5B
       | 
       | 1. https://huggingface.co/neuphonic/neutts-air/tree/main
       | 
       | 2. https://github.com/neuphonic/neutts-air/issues/7
       | 
       | 3. https://github.com/neuphonic/neutts-air/blob/feat/example-
       | fi...
       | 
       | 4. https://huggingface.co/datasets/neuphonic/emilia-yodas-
       | engli...
        
       | curioussquirrel wrote:
       | Could we finally get a decent opensource TTS app for Android?
       | This project is very cool.
        
       | ks2048 wrote:
       | Every couple of weeks I see a new TTS model showcased here and
       | it's always difficult to see how they differ from one another.
       | Why don't they describe the architecture and details of the
       | trailing data?
       | 
       | My cynical side thinks people just take the state-of-the-art open
       | source model, use an LLM to alter the source, minimal fine tuning
       | to change the weights and they are able to claim "we built our
       | own state of the art tts".
       | 
       | I know it's open source, so I can dig into the details myself,
       | but are they any good high-level overviews of modern TTS,
       | comparing/contrasting the top models?
        
       | joshstrange wrote:
       | This is really neat. I cloned my voice and can generate text, but
       | I can't seem to generate longer clips. The README.md says:
       | 
       | > Context Window: 2048 tokens, enough for processing ~30 seconds
       | of audio (including prompt duration)
       | 
       | But it's cutting off for me before even that point. I fed it a
       | paragraph of text and it gets part of the way through it before
       | skipping a few words ahead, saying a few words more, then cutting
       | off at 17 seconds. Another test just cut off after 21 seconds (no
       | skipping).
       | 
       | Lastly, I'm on a MBP M3 Max with 128GB running Sequoia. I'm
       | following all the "Guidelines for minimizing Latency" but
       | generating a 4.16 second clip takes 16.51s for me. Not sure what
       | I'm doing wrong or how you would use this in practice since it's
       | not realtime and the limit is so low (and unclear). Maybe you are
       | supposed to cut your text into smaller chunks and run them in
       | parallel/sequence to get around the limit?
        
       ___________________________________________________________________
       (page generated 2025-10-09 23:00 UTC)