[HN Gopher] Show HN: I built a sub-500ms latency voice agent fro...
       ___________________________________________________________________
        
       Show HN: I built a sub-500ms latency voice agent from scratch
        
       I built a voice agent from scratch that averages ~400ms end-to-end
       latency (phone stop - first syllable). That's with full STT - LLM -
       TTS in the loop, clean barge-ins, and no precomputed responses.
       What moved the needle:  Voice is a turn-taking problem, not a
       transcription problem. VAD alone fails; you need semantic end-of-
       turn detection.  The system reduces to one loop: speaking vs
       listening. The two transitions - cancel instantly on barge-in,
       respond instantly on end-of-turn - define the experience.  STT -
       LLM - TTS must stream. Sequential pipelines are dead on arrival for
       natural conversation.  TTFT dominates everything. In voice, the
       first token is the critical path. Groq's ~80ms TTFT was the single
       biggest win.  Geography matters more than prompts. Colocate
       everything or you lose before you start.
        
       Author : nicktikhonov
       Score  : 37 points
       Date   : 2026-03-02 21:23 UTC (1 hours ago)
        
 (HTM) web link (www.ntik.me)
 (TXT) w3m dump (www.ntik.me)
        
       | MbBrainz wrote:
       | Love it! Solving the latency problem is essential to making voice
       | ai usable and comfortable. Your point on VAD is interesting -
       | hadn't thought about that.
        
       | NickNaraghi wrote:
       | Pretty exciting breakthrough. This actually mirrors the early
       | days of game engine netcode evolution. Since latency is an
       | orchestration problem (not a model problem) you can beat general-
       | purpose frameworks by co-locating and pipelining aggressively.
       | 
       | Carmack's 2013 "Latency Mitigation Strategies" paper[0] made the
       | same point for VR too: every millisecond hides in a different
       | stage of the pipeline, and you only find them by tracing the full
       | path yourself. Great find with the warm TTS websocket pool saving
       | ~300ms, perfect example of this.
       | 
       | [0]: https://danluu.com/latency-mitigation/
        
       | jangletown wrote:
       | impressive
        
       | lukax wrote:
       | Or you could use Soniox Real-time (supports 60 languages) which
       | natively supports endpoint detection - the model is trained to
       | figure out when a user's turn ended. This always works better
       | than VAD.
       | 
       | https://soniox.com/docs/stt/rt/endpoint-detection
       | 
       | Soniox also wins the independent benchmarks done by Daily, the
       | company behind Pipecat.
       | 
       | https://www.daily.co/blog/benchmarking-stt-for-voice-agents/
       | 
       | You can try a demo on the home page:
       | 
       | https://soniox.com/
       | 
       | Disclaimer: I used to work for Soniox
       | 
       | Edit: I commented too soon. I only saw VAD and immediately
       | thought of Soniox which was the first service to implement real
       | time endpoint detection last year.
        
         | nicktikhonov wrote:
         | If you read the post, you'll see that I used Deepgram's Flux.
         | It also does endpointing and is a higher-level abstraction than
         | VAD.
        
           | lukax wrote:
           | Sorry, I commented too soon. Did you also try Soniox? Why did
           | you decide to use Deepgram's Flux (English only)?
        
             | nicktikhonov wrote:
             | I didn't try Soniox, but I made a note to check it out! I
             | chose Flux because I was already using Deepgram for STT and
             | just happened to discover it when I was doing research. It
             | would definitely be a good follow-up to try out all the
             | different endpointing solutions to see what would shave off
             | additional latency and feel most natural.
             | 
             | Another good follow-up would be to try PersonaPlex,
             | Nvidia's new model that would completely replace this
             | architecture with a single model that does everything:
             | 
             | https://research.nvidia.com/labs/adlr/personaplex/
        
       | loevborg wrote:
       | Nice write-up, thanks for sharing. How does your hand-vibed
       | python program compare to frameworks like pipecat or livekit
       | agents? Both are also written in python.
        
         | nicktikhonov wrote:
         | I'm sure LiveKit or similar would be best to use in production.
         | I'm sure these libraries handle a lot of edge cases, or at
         | least let you configure things quite well out of the box.
         | Though maybe that argument will become less and less potent
         | over time. The results I got were genuinely impressive, and of
         | course most of the credit goes to the LLM. I think it's worth
         | building this stuff from scratch, just so that you can be sure
         | you understand what you'll actually be running. I now know how
         | every piece works and can configure/tune things more
         | confidently.
        
       | perelin wrote:
       | Great writeup! For VAD did you use heaphone/mic combo, or an open
       | mic? If open, how did you deal with the agent interupting itself?
        
         | nicktikhonov wrote:
         | I was using Twilio, and as far as I'm aware they handle any
         | echos that may arise. I'm actually not sure where in the
         | telephony stack this is handled, but I didn't see any issues or
         | have to solve this problem myself luckily.
        
       | boznz wrote:
       | "Voice is an orchestration problem" is basically correct. The two
       | takeaways from this for me are
       | 
       | 1. I wonder if it could be optimised more by just having a single
       | language, and
       | 
       | 2. How do we get around the problem of interference, humans are
       | good at conversation discrimination ie listing while multiple
       | conversations, TV, music, etc are going on in the background,
       | I've not had too much success with voice in noisy environments.
        
       | modeless wrote:
       | IMO STT -> LLM -> TTS is a dead end. The future is end-to-end. I
       | played with this two years ago and even made a demo you can
       | install locally on a gaming GPU:
       | https://github.com/jdarpinian/chirpy, but concluded that making
       | something worth using for real tasks would require training of
       | end-to-end models. A really interesting problem I would love to
       | tackle, but out of my budget for a side project.
        
         | nicktikhonov wrote:
         | If you're of that opinion, you'll enjoy the new stuff coming
         | out from nvidia:
         | 
         | https://research.nvidia.com/labs/adlr/personaplex/
        
       | age123456gpg wrote:
       | Hi all! Check out this Handy app https://github.com/cjpais/Handy
       | - a free, open source, and extensible speech-to-text application
       | that works completely offline.
       | 
       | I am using it daily to drive Claude and it works really-well for
       | me (much better than macOS dictation mode).
        
       | armcat wrote:
       | This is an outstanding write up, thank you! Regarding LLM
       | latency, OpenAI introduced web sockets in their Responses client
       | recently so it should be a bit faster. An alternative is to have
       | a super small LLM running locally on your device. I built my own
       | pipeline fully local and it was sub second RTT, with no streaming
       | nor optimisations https://github.com/acatovic/ova
        
       ___________________________________________________________________
       (page generated 2026-03-02 23:00 UTC)