[HN Gopher] Show HN: Aqua Voice 2 - Fast Voice Input for Mac and...
       ___________________________________________________________________
        
       Show HN: Aqua Voice 2 - Fast Voice Input for Mac and Windows
        
       Hey HN - It's Finn and Jack from Aqua Voice (https://withaqua.com).
       Aqua is fast AI dictation for your desktop and our attempt to make
       voice a first-class input method.  Video:
       https://withaqua.com/watch  Try it here:
       https://withaqua.com/sandbox  Finn is uber dyslexic and has been
       using dictation software since sixth grade. For over a decade, he's
       been chasing a dream that never quite worked -- using your voice
       instead of a keyboard.  Our last post
       (https://news.ycombinator.com/item?id=39828686) about this seemed
       to resonate with the community - though it turned out that version
       of Aqua was a better demo than product. But it gave us (and others)
       a lot of good ideas about what should come next.  Since then, we've
       remade Aqua from scratch for speed and usability. It now lives on
       your desktop, and it lets you talk into any text field -- Cursor,
       Gmail, Slack, even your terminal.  It starts up in under 50ms,
       inserts text in about a second (sometimes as fast as 450ms), and
       has state-of-the-art accuracy. It does a lot more, but that's the
       core. We'd love your feedback -- and if you've got ideas for what
       voice should do next, let's hear them!
        
       Author : the_king
       Score  : 60 points
       Date   : 2025-04-09 16:31 UTC (6 hours ago)
        
 (HTM) web link (withaqua.com)
 (TXT) w3m dump (withaqua.com)
        
       | oulipo wrote:
       | Interesting!
       | 
       | A nice open-source alternative is VoiceInk, check it out:
       | https://github.com/Beingpax/VoiceInk
       | 
       | do you also plan to open-source part of your platform?
        
         | pablopeniche wrote:
         | Just tried it and it crashed
        
         | fintechie wrote:
         | I've been using this one with Cursor the past few months...
         | 
         | https://github.com/foges/whisper-dictation
        
       | bklyn11201 wrote:
       | Music playing on Youtube in Chrome, Airpods in, the desktop and
       | the sandbox/demo just don't work.
        
         | the_king wrote:
         | Shoot, might be due to AirPods mic init latency. AirPods work
         | well on Desktop (though their mic quality isn't the best).
        
       | rkagerer wrote:
       | You mentioned it "lives on your desktop". How does licensing
       | work, and can you install and use it on a machine without
       | internet access?
        
       | gnfedhjmm2 wrote:
       | Broken on mobile
        
         | pablopeniche wrote:
         | fixing. ty
        
       | niel wrote:
       | Real-time text output a la Apple Dictation with the accuracy of
       | Whisper is something I've been looking for recently - I'll
       | definitely give Aqua a spin.
       | 
       | MacWhisper [0] (the app I settled on) is conspicuously missing
       | from your benchmarks [1]. How does it compare?
       | 
       | [0]: https://goodsnooze.gumroad.com/l/macwhisper
       | 
       | [1]: https://withaqua.com/blog/benchmark-nov-2024
        
         | the_king wrote:
         | We're more accurate and much faster than Mac Whisper, even
         | their strongest model (Whisper Cpp Large V3).
         | 
         | For that benchmarking table, you can use Whisper Large V3 as a
         | stand-in for Mac Whisper and Super Whisper accuracy.
        
       | fxtentacle wrote:
       | This looks like it'll slurp up all your data and upload it into a
       | cloud. Thanks, no. I want privacy, offline mode and source code
       | for something as crucial to system security as an input method.
       | 
       | "we also collect and process your voice inputs [..] We leverage
       | this data for improvements and development [..] Sharing of your
       | information [..] service providers [..] OpenAI"
       | https://withaqua.com/privacy
        
         | FloatArtifact wrote:
         | Local inference only is an absolute requirement. It's not even
         | really all that accessible if it's online only. I can say this
         | as someone that's used over 20000 hours worth of voice
         | dictation and computer control.
        
         | pokstad wrote:
         | This should be on the FAQ. I was trying to find out if it was
         | 100% processed locally.
        
         | jackthetab wrote:
         | Agreed.
         | 
         | This is where I bounce (out of this discussion).
        
         | thmsmlr wrote:
         | I totally agree, I created BetterDictation (.com) exactly
         | because of that. Offline was a super important requirement for
         | me.
        
       | hu3 wrote:
       | How does it compare to https://wisprflow.ai ?
       | 
       | btw, grats!
        
         | the_king wrote:
         | Thanks!
         | 
         | We're faster, more accurate, and have a streaming option. Aqua
         | can go from key-up to paste in as little as 450ms. Flow was
         | closer to 1000 in our tests.
         | 
         | Overall, you'll notice we make a few more tweaks to the output
         | than Wisprflow.
         | 
         | For example, Aqua + Cursor is very powerful - we syntax
         | highlight your transcript. The easiest way to see this is to
         | use streaming mode (double press Fn) + deep context + cursor
         | and try asking it to change something.
         | 
         | This also works in other "context rich" environments.
        
           | hu3 wrote:
           | Thanks! I'm tempted to try Speech to Text + Cursor/Copilot
           | for development. It's probably the future since most people
           | can speak faster than they can type.
        
             | jmcintire1 wrote:
             | This use case is great, even for people who haven't been
             | interested in dictation before
        
           | redcanvas wrote:
           | Hey, love the what you are building in this category. I've
           | been using a competing product which you know very well
           | about. They advertised about how you can improve your work
           | per minute by dictation, which was the main draw for me
           | because I do a founder. There's a lot of managerial work that
           | I'm doing.
           | 
           | It has been a godsend in terms of increasing my productivity
           | because I no longer have to type. I think your product's
           | accuracy and latency shortening just make this even better. I
           | often use it and then find out, "Hey, I need to make some
           | changes," and I need to re-edit some of the stuff, which
           | reduces the WPM productivity amount. So I think accuracy is
           | definitely key here. Key metric to differentiate a product.
           | 
           | I am pushing this to other colleagues to get them to adopt.
           | One challenge people are saying is that. One is that some
           | people may not be as organized (you know, they might be a lot
           | more organizationally structured in their mind). So for them,
           | they're having trouble - they'd like to write things out, and
           | by the time if things go out of their mouth, you know it's
           | already formulated logical thought. Whereas you know people
           | like me are a lot more verbal vomit type of person. For me
           | it's huge because I say a lot of um like in all the other
           | things I just dump stuff out and then organize it later.
           | 
           | Whereas other people organize stuff in their brain and then
           | dump the information out. So people who do a lot of
           | coordination and just you know so I feel like this could be
           | two different segments to take into account.
           | 
           | Another one that's been fantastic is that we have
           | multilingual colleagues who are speaking in Mandarin or
           | something else and then they speak it and then ask Flow to be
           | sent to translate it to a different language. That part I
           | think has been fantastic.
           | 
           | I think the ability to edit what you wrote with AI is going
           | to be the next key feature. Providing the context in the
           | window is all wiithin the conversation right? For example,
           | you just ask after because what you write out is not the
           | final and you need to do a lot of editing and formatting.
           | Sometimes when you say too much stuff, it's just like a huge
           | jumble paragraph with a lot of fluff words. Make it clear,
           | concise, trim non-effective words. I think those are a key
           | feature because it's not about your productivity, it's about
           | other people being able to ingest your information
           | efficiently. At least that's what I look at from a managerial
           | perspective.
           | 
           | To give you an example, everything I laid out above came from
           | dictation. You can see how this is inefficient. There's a lot
           | of inefficiencies here.
        
             | redcanvas wrote:
             | A feature that would be great is similar to how you can
             | write snippets in all the other tools where you can say
             | "calendar" or "cal" and then it gives you the link. If this
             | is something possible, I think that would make this
             | fantastic.
             | 
             | Another feature that would be great is actually being able
             | to have a conversation with an AI model first and then
             | refine the output iteratively until you're ready and then
             | pipe all that over. The ability to have a chat is very good
             | or do this all through voice.
        
       | SCdF wrote:
       | I currently use Talon, which I note is not in your benchmarks.
       | 
       | I can't find any documentation on how Aqua works, or how it
       | compares, so I'm not sure it's meant to be a replacement /
       | competitor to Talon? What are you configuring? How are you
       | telling it that you like "genz" style in Slack? Can I create
       | custom configurations / macros?
       | 
       | One thing I like about Talon is it's not magic. Which maybe is
       | not what you're going for. But I am giving it explicit commands
       | that I know it will understand (if it understands my accent
       | obvs), as opposed to guessing and constructing a human language
       | vague sentence and hope that an llm will work it out. Which means
       | it feels like something I can actually become fast with, and
       | build up muscle memory of.
       | 
       | Also that it's completely offline, so I can actually run it on a
       | work computer without my security folks freaking out.
        
         | willwade wrote:
         | Aqua voice is nothing like talon. I wouldn't bother trying to
         | compare. It's a dictation tool. Just entry. Not commands. But
         | it's bloody impressive. You don't need to learn anything - you
         | just talk like you would talk to someone across the way from
         | you
        
           | SCdF wrote:
           | Oh, from the video I got the impression it was more than
           | that, based on it recognising app contexts and the like. I
           | guess that's mostly just icing on the cake for the core
           | dictation part.
        
             | pablopeniche wrote:
             | >recognizing app contexts
             | 
             | Users have different preferences on the text format they
             | input into different apps. Aqua is able to pick up on these
             | explicit and implicit preferences across apps - but no
             | "open XYZ app" commands, yea
        
         | the_king wrote:
         | We're building something different, but there is some overlap.
         | Aqua is built for max speed, while keeping accuracy high. To
         | achieve that, inference runs in a datacenter (for now).
         | 
         | You can customize Aqua using custom instructions, similar to
         | ChatGPT custom instructions, and get some Talon functionality
         | from it:
         | 
         | In my own, I have:
         | 
         | 1. Breaking the paragraphs with three or four sentences.
         | 
         | 2. Don't start a sentence with "and".
         | 
         | 3. Use lowercase in Slack and iMessage.
         | 
         | 4. Here are some common terminal commands...
        
       | tomblomfield wrote:
       | I recently started using Aqua and it's great. The team really
       | improved the latency in the last few weeks.
        
       | willwade wrote:
       | You're real market you need to go hard on is the assistive tech
       | market. You know the biggest companies in this space are those
       | solving problems for dyslexia where govt grants in eg UK fund
       | pretty much all their work? I had an access to work assessment
       | and they recommend like sweets stuff from texthelp. It's then
       | paid for by the government following these assessments. But it's
       | crap. It literally is a crap tool for adhd or dyslexia because
       | these users literally CANT remember or deal with barriers like
       | learning how to dictate correctly. Aqua voice solves this. I'm
       | your biggest fan. I recommend it in my AT assessments all the
       | time :)
        
       | idk1 wrote:
       | I've been using this for some time and I have to say it is
       | fantastic. I'm intentionally not writing this with Aqua but by
       | hand and it is taking so much longer. This to me feels like what
       | Apple Intelligence could be, it is so much better than stuff all
       | of the big tech is doing. For example, if you tell Siri voice
       | dictation to go back and delete something what Siri will do is
       | just write out "go back and delete something" also if you tell
       | Siri to go back and spell a name differently all Siri will do is
       | write out the letters that you said to go back and type out.
       | Honestly, for voice dictation software it feels like travelling
       | to another planet in terms of improvement.
        
       | aminsadeghi wrote:
       | Is there going to be Linux support at some point?
        
       | TylerE wrote:
       | I will have to look into this. I am currently in the process of
       | going on disability as I cannot work due to (amongst other
       | things) carpal and cubical tunnel in both arms.
        
       ___________________________________________________________________
       (page generated 2025-04-09 23:00 UTC)