[HN Gopher] Show HN: Ghost Pepper - Local hold-to-talk speech-to...
___________________________________________________________________
Show HN: Ghost Pepper - Local hold-to-talk speech-to-text for macOS
I built this because I wanted to see how far I could get with a
voice-to-text app that used 100% local models so no data left my
computer. I've been using a ton for coding and emails.
Experimenting with using it as a voice interface for my other
agents too. 100% open-source MIT license, would love feedback, PRs,
and ideas on where to take it.
Author : MattHart88
Score : 151 points
Date : 2026-04-06 19:50 UTC (3 hours ago)
(HTM) web link (github.com)
(TXT) w3m dump (github.com)
| charlietran wrote:
| Thank you for sharing, I appreciate the emphasis on local speed
| and privacy. As a current user of Hex
| (https://github.com/kitlangton/Hex), which has similar goals,
| what are your thoughts on how they compare?
| ipsum2 wrote:
| Parakeet is significantly more accurate and faster than Whisper
| if it supports your language.
| yeutterg wrote:
| Are you running Parakeet with VoiceInk[0]?
|
| [0]: https://github.com/beingpax/VoiceInk
| zackify wrote:
| i am, working great for a long time now
| ipsum2 wrote:
| I'm using https://github.com/senstella/parakeet-mlx library.
| treetalker wrote:
| I have been using Parakeet with MacWhisper's hold-to-talk on a
| MacBook Neo and it's been awesome.
| rahimnathwani wrote:
| Right, and if you're on MacOS you can use it for free with Hex:
| https://github.com/kitlangton/Hex
| lloyd-christmas wrote:
| Or write your own custom one with the library that backs it:
| https://github.com/FluidInference/FluidAudio
|
| I did that so that I could record my own inputs and finetune
| parakeet to make it accurate enough to skip post-processing.
| obrajesse wrote:
| And indeed, Ghost Pepper supports parakeet v3
| goodroot wrote:
| Nice one! For Linux folks, I developed
| https://github.com/goodroot/hyprwhspr.
|
| On Linux, there's access to the latest Cohere Transcribe model
| and it works very, very well. Requires a GPU though. Larger local
| models generally shouldn't require a subordinate model for clean
| up.
|
| Have you compared WhisperKit to faster-whisper or similar? You
| might be able to run turbov3 successfully and negate the need for
| cleanup.
|
| Incidentally, waiting for Apple to blow this all up with native
| STT any day now. :)
| hephaes7us wrote:
| Thanks for sharing! I was literally getting ready to build,
| essentially, this. Now it looks like I don't have to!
|
| Have you ever considered using a foot-pedal for PTT?
|
| Apple incidentally already has native STT, but for some reason
| they just don't use a decent model yet.
| goodroot wrote:
| They do, and they even have that nice microphone F5 key for
| it, and an ideal OS level API making the input experience
| >perfect<.
|
| Apparently they do have a better model, they just haven't
| exposed it in their own OS yet!
|
| https://developer.apple.com/documentation/speech/bringing-
| ad...
|
| Wonder what's the hold up...
|
| For footpedal:
|
| Yes, conceptually it's just another evdev-trigger source,
| assuming the pedal exposes usable key/button events.
|
| Otherwise we'd bridge it into the existing external control
| interface. Either way, hooks are there. :)
| jiehong wrote:
| The only issue with Apple models is that they do not detect
| languages automatically, nor switch if you do between
| sentences.
|
| Parakeet does both just fine.
| chrisweekly wrote:
| sorry, PTT?
| serf wrote:
| push-to-talk.
| LuxBennu wrote:
| I've been running whisper large-v3 on an m2 max through a self-
| hosted endpoint and honestly the accuracy is good enough that i
| stopped bothering with cleanup models. The bigger annoyance for
| me was latency on longer chunks, like anything over 30 seconds
| starts feeling sluggish even with metal acceleration. Haven't
| tried whisperkit specifically but curious how it handles longer
| audio compared to the full model.
| goodroot wrote:
| Ah yeah, longform is interesting.
|
| Not sure how you're running it, via whichever "app thing",
| but...
|
| On resource limited machines: "Continuous recording" mode
| outputs when silence is detected via a configurable
| threshold.
|
| This outputs as you speak in more reasonable chunks; in
| aggregate "the same output" just chunked efficiently.
|
| Maybe you can try hackin' that up?
| LuxBennu wrote:
| Yeah that makes sense, chunking on silence would sidestep
| the latency issue pretty cleanly. I've been running it
| through a basic fastapi wrapper so it just takes whatever
| audio blob gets thrown at it, no chunking logic on the
| server side. Might be worth adding a vad pass before
| sending to whisper though, would cut down on processing
| dead air too.
| VorpalWay wrote:
| How does it compare to the more well established
| https://github.com/cjpais/handy? Are there any stand out
| features (for either option)? What was the reason for writing
| your own rather than using or improving existing software?
| goodroot wrote:
| Not sure I know what you mean by IR...
|
| But in this case I built hyprwhspr for Linux (Arch at first).
|
| The goal was (is) the absolute best performance, in both
| accuracy & speed.
|
| Python, via CUDA, on a NVIDIA GPU, is where that exists.
|
| For example:
|
| The #1 model on the ASR (automatic speech recognition)
| hugging face board is Cohere Transcribe and it is not yet 2
| weeks old.
|
| The ecosystem choices allowed me to hook it up in a night.
|
| Other hardware types also work great on Linux due to its
| adaptability.
|
| In short, the local stt peak is Linux/Wayland.
| VorpalWay wrote:
| IR was a typo, meant "it" (fixed it). I blame the phone
| keyboard plus insufficient proof reading on my part.
|
| If this needs nvidia CPU acceleration for good performance
| it is not useful to me, I have Intel graphics and handy
| works fine.
| goodroot wrote:
| It works well with anything. :)
|
| That said: If handy works, no need whatsoever to change.
| konaraddi wrote:
| That's awesome! Do you know how it compares to Handy? Handy is
| open source and local only too. It's been around a while and what
| I've been using.
|
| https://github.com/cjpais/handy
| youniverse wrote:
| I love and have been using handy for a while too, what we need
| is this for mobile apps I don't think there's any free apps and
| native dictation is not always fully local and not as good.
| swaptr wrote:
| Handy is awesome! I used it for quite a while before Claude
| Code added voice support. Solid software, very good linux and
| mac integration. Shoutout to Parakeet models as well, extremely
| fast and solid models for their relatively modest memory
| requirements.
| stavros wrote:
| Handy is fantastic.
| vunderba wrote:
| I'd also be interested to know what the impetus was for
| developing ghost-pepper, which looks relatively recent, given
| that Handy exists and has been pretty well received.
|
| Extra bonus is that Handy lets add an automatic LLM post-
| processor. This is very handy for the Parakeet V3 model, which
| can sometimes have issues where it repeats words or makes
| recognition errors for example, duplicating the recognition of
| a single word a dozen dozen dozen dozen dozen dozen dozen dozen
| times.
| rob wrote:
| Yep. Using Handy with Parakeet v3 + a custom coding-tailored
| prompt to post-process on my 2019 Intel Mac and it's been
| working great.
|
| Once in a while it will only output a literal space instead
| of the actual translation, but if I go into the 'history'
| page the translation is there for me to copy and paste
| manually. Maybe some pasting bug.
| olup wrote:
| I use handy all day long as a software engineer, and
| recommended it to all of my team members. I love it.
| JohnPDickerson wrote:
| Handy is an awesome project, highly recommended - many of our
| engineers and PMs use it! CJ, Handy's creator, recently joined
| us as a Builder in Residence at Mozilla.ai. So for those
| interested in deploying a more raw/lightweight approach to
| local speech-to-text (or other multimodal) models, feel free to
| check out llamafile - which includes whisperfile, a single-file
| whisper.cpp + cosmopolitan framework-based executable. We're
| hoping to build some bridges between the two projects as well.
| https://github.com/mozilla-ai/llamafile
| mathis wrote:
| If you don't feel like downloading a large model, you can also
| use `yap dictate`. Yap leverages the built-in models exposed
| though Speech.framework on macOS 26 (Tahoe).
|
| Project repo: https://github.com/finnvoor/yap
| hyperhello wrote:
| Feature request or beg: let me play a speech video and transcribe
| it for me.
| MattHart88 wrote:
| I like this idea and it should work -- whatever microphone you
| have on should be able to hear the speaker. LMK if not (e.g.,
| are you wearing headphones? if so, the mic can't hear the
| speaker)
| aristech wrote:
| Great job. How about the supported languages? System languages
| gets recognised?
| MattHart88 wrote:
| Thanks! We currently have 2 multi-lingual options available: -
| Whisper small (multilingual) (~466 MB, supports many languages)
| - Parakeet v3 (25 languages) (~1.4 GB, supports 25 languages
| via FluidAudio)
| guzik wrote:
| Sadly the app doesn't work. There is no popup asking for
| microphone permission.
|
| EDIT: I see there is an open issue for that on github
| ttul wrote:
| And many people are mailing in Codex and Claude Code generated
| PRs - myself included. Fingers crossed, I suppose.
| MattHart88 wrote:
| Thanks to everyone who submitted PRs! The fix is merged, new
| version is up.
| parhamn wrote:
| I see a lot of whisper stuff out there. Are these the same old
| OpenAI whispers or have they been updated heavily?
|
| I've been using parakeet v3 which is fantastic (and tiny).
| Confused why we're still seeing whisper out there, there's been a
| lot of development.
| zackify wrote:
| same, even have kokoro for speech back to text for home
| assistant and parakeet on mac os through voice ink.
|
| Also vibe coded a way to use parakeet from the same parakeet
| piper server on my grapheneos phone https://zach.codes/p/vibe-
| coding-a-wispr-clone-in-20-minutes
| daemonologist wrote:
| Whisper is still old reliable - I find that it's less prone to
| hallucinations than newer models, easier to run (on AMD GPU,
| via whisper.cpp), and only ~2x slower than parakeet. I even
| bothered to "port" Parakeet to Nemo-less pytorch to run it on
| my GPU, and still went back to Whisper after a couple of days.
| goodroot wrote:
| Whisper is very good in many languages.
|
| It's also in many flavours, from tiny to turbo, and so can fit
| many system profiles.
|
| That's what makes it unique and hard to replace.
| gegtik wrote:
| how does this compare to macos built in siri TTS, in quality and
| in privacy?
| realityfactchex wrote:
| Exactly my question. I double-tap the control button and macOS
| does native, local TTS dictation pretty well. (Similar to
| Keyboard > Enable Dictation setting on iOS.)
|
| The macOS built-in TTS (dictation) seems better than all the
| 3rd party, local apps I tried in the past that people raved
| about. I have tried several.
|
| Is this better somehow?
|
| If the 3rd party apps did streaming with typing in place and
| corrections within a reasonable window when they understand
| things better given more context, that would be cool.
| Theoretically, a custom model or UX could be "better" than what
| comes free built into macOS (more accurate or customizable).
|
| But when I contacted the developer of my favorite one they said
| that would be pretty hard to implement due to having to go back
| and make corrections in the active field, etc.
|
| I assume streaming STT in these utilities for Mac will get
| better at some point, but I haven't seen it yet (been waiting).
| It seems these tools generally are not streaming, e.g. they
| want you to finish speaking first before showing you anything.
| Which doesn't work for me when I'm dictating. I want to see
| what I've been saying lately, to jog my memory about what I've
| just said and help guide the next thing I'm about to say. I
| certainly don't want to split my attention by manually toggling
| the control (whether PTT or not) periodically to indicate "ok,
| you can render what I just said now".
|
| I guess "hold-to-talk" tools are for delivering discrete, fully
| formed messages, not for longer, running dictation.
|
| AFAICT, TFA is focused on hold-to-talk as the differentiator,
| over double-tap to begin speaking and double-tap to end
| speaking?
| Supercompressor wrote:
| I've been looking for the opposite - wanting to dump text and it
| be read to me, coherently. Anyone have good recommendations?
| realityfactchex wrote:
| Sure, Chatterbox TTS Server is rather high quality:
| https://github.com/devnen/Chatterbox-TTS-Server
|
| You could hook it up to some workflow over the local API
| depending on how you want to dump the text, but the web UI is
| good too.
|
| The Show HN by the author was at:
| https://news.ycombinator.com/item?id=44145564
| Supercompressor wrote:
| Appreciated - thank you.
| ericmcer wrote:
| I see quite a few of these, the killer feature to me will be one
| that fine tunes the model based on your own voice.
|
| E.G. if your name is `Donold` (pronounced like Donald) there is
| not a transcription model in existence that will transcribe your
| name correctly. That means forget inputting your name or email
| ever, it will never output it correctly.
|
| Combine that with any subtleties of speech you have, or industry
| jargon you frequently use and you will have a much more useful
| tool.
|
| We have a ton of options for "predict the most common word that
| matches this audio data" but I haven't found any "predict MY most
| common word" setups.
| MattHart88 wrote:
| I've found the "corrections" feature works well for most of the
| jargon and misspelling use cases. Can you give it a try and let
| me know edge cases?
| sorenjan wrote:
| Whisper supports a prompt, you can put your "Donold" there.
|
| https://developers.openai.com/cookbook/examples/whisper_prom...
| bonkler59 wrote:
| My experience is that Aqua voice does a good job of this with
| custom dictionary and replacements.
| __mharrison__ wrote:
| Cool, I've been doing a lot of "coding" (and other typing tasks)
| recently by tapping a button on my Stream Deck. It starts
| recording me until I tap it again. At which point, it transcribes
| the recording and plops it into the paste buffer.
|
| The button next to it pastes when I press it. If I press it
| again, it hits the enter command.
|
| You can get a lot done with two buttons.
| purplehat_ wrote:
| Hi Matt, there's lots of speech-to-text programs out there with
| varying levels of quality. 100% local is admirable but it's
| always a tradeoff and users have to decide for themselves what's
| worth it.
|
| Would you consider making available a video showing someone using
| the app?
| semiquaver wrote:
| Slop
| douglaswlance wrote:
| does it input the text as soon as it hears it? or does it wait
| until the end?
| romeroej wrote:
| always mac. when windows? why can you just make things multios
| primaprashant wrote:
| Speech-to-text has become integral part of my dev flow especially
| for dictating detailed prompts to LLMs and coding agents.
|
| I have collected the best open-source voice typing tools
| categorized by platform in this awesome-style GitHub repo. Hope
| you all find this useful!
|
| https://github.com/primaprashant/awesome-voice-typing
| rcarmo wrote:
| Not sure why I should use this instead of the baked-in OS
| dictation features (which I use almost daily--just double-tap the
| world key, and you're there). What's the advantage?
| qq66 wrote:
| I haven't used this one but WisprFlow is vastly better than the
| built-in functionality on MacOS. Apple is way behind even
| startups, even for fundamental AI functionality like
| transcribing speech
| ibero wrote:
| WisprFlow has a lot of good recommendations behind it but the
| fact they used Delve for SOC2 compliance gives me major
| pause.
| janalsncm wrote:
| The fact that a company could suck up all of your data and
| then use Delve for their SOC2 is a great reason to use
| local models.
| jonwinstanley wrote:
| I use the baked in Apple transcription and haven't had any
| issues. But what I do is usually pretty simple.
|
| What makes the others vastly better?
| MattDamonSpace wrote:
| I've rarely had macOS TTS produce a sentence I didn't have
| to edit
|
| Whisper models I barely bother checking anymore
| atlgator wrote:
| This thread is a support group for people who have each
| independently built the same macOS speech-to-text app.
| brcmthrowaway wrote:
| Oh to be 20-something and do a bunch of free work for your
| portfolio again
| obrajesse wrote:
| I'll have you know that I'm Matt's top contributor to Ghost
| Pepper and I'm nearly _fifty_
|
| But I did it because I wanted it to work exactly the way I
| wanted it.
|
| Also, for kicks, I (codex) ported it to Linux. But because my
| Linux laptop isn't as fast, I've had to use a few tricks to
| make it fast. https://github.com/obra/pepper-x
| tito wrote:
| This is great. I'm typing this message now using Ghost Pepper.
| What benefits have you seen from the OCR screen sharing step?
| janalsncm wrote:
| I think the jab at the bottom of the readme is referring to
| whispr flow?
|
| https://wisprflow.ai/new-funding
___________________________________________________________________
(page generated 2026-04-06 23:00 UTC)