[HN Gopher] Ask HN: How to boost Gemini transcription accuracy f...
       ___________________________________________________________________
        
       Ask HN: How to boost Gemini transcription accuracy for company
       names?
        
       I'm using Gemini for speech-to-text and it often misrecognizes
       company names and acronyms.  Is there any way to use a custom
       lexicon or vocabulary with Gemini to improve recognition accuracy?
       If not directly supported, what are practical workarounds people
       use -- e.g. preprocessing prompts, fine-tuning, or combining Gemini
       with another ASR that supports phrase boosting?
        
       Author : bingwu1995
       Score  : 28 points
       Date   : 2025-10-20 04:41 UTC (6 days ago)
        
       | gallexme wrote:
       | Adding it to the instructions worked well for me with specific
       | terms
        
       | semessier wrote:
       | adding to the question, ruling out fine-tuning for practicality,
       | what about injecting names towards the embedding but not into the
       | context?
        
       | lysecret wrote:
       | I generally found 4o-transcribe to be more performant than gemini
       | fyi.
        
       | Reubend wrote:
       | Any company names or special acronyms should be added to your
       | prompt.
        
       | tifa2up wrote:
       | Don't solve it on the STT level. Get the raw transcription from
       | Gemini then pass the output to an LLM to fix company names and
       | other modifications.
       | 
       | Happy to share more details if helpful.
        
         | remus wrote:
         | I've had some luck with this in other contexts. Get the initial
         | transcript from STT (e.g. whisper), then feed that in to gemini
         | with a prompt giving it as much extra context as possible. For
         | example "This is a transcript from a youtube video. It's a
         | conversation between x people, where they talk about y and z.
         | Please clean up the transcript, paying particular attention to
         | company names and acronyms."
        
           | flyinglizard wrote:
           | I've done the same, it works very well.
        
         | idopmstuff wrote:
         | Yeah, I've done it with industry-specific acronyms and this
         | works well. Generate a list of company names and other terms it
         | gets wrong, and give it definitions and any other useful
         | context. For industry jargon, example sentences are good, but
         | that's probably not relevant for company names.
         | 
         | Feed it that list and the transcript along with a simple prompt
         | along the lines of "Attached is a transcript of a conversation
         | created from an audio file. The model doing the transcription
         | has trouble with company names/industry terms/acronyms/whatever
         | else and will have made errors with those. I have also attached
         | a list of company names/etc. that may have been spoken in the
         | transcribed audio. Please review the transcription, and output
         | a corrected version, along with a list of all corrections that
         | you made. The list of corrections should include the original
         | version of the word that you fixed, what you updated it to, and
         | where it is in the document." If it's getting things wrong, you
         | can also ask it to give an explanation of why it made each
         | change that it did and use that to iterate on your prompt and
         | the context you're giving it with your list of words.
        
           | dotancohen wrote:
           | Which specific model do you use?
        
       | simonw wrote:
       | Have you tried feeding it a system prompt with a list of custom
       | vocabulary? I would expect that to work really well.
       | 
       | "Transcribe this audio. Be careful to spell the following names
       | and acronyms right: list-goes-here"
        
       | gearhart wrote:
       | We use openwhisper for transcription which accepts a list of
       | "words to look out for" which we populate with a short list of
       | the names of all the people and companies most likely to be
       | mentioned in the text, and then we do a spell checking pass at
       | the end using Gemini with a much longer list, telling it to look
       | out for anything that might be a misspelling.
       | 
       | It's not perfect, but it's taken it from being an issue that made
       | all our transcripts look terrible, to an issue I no longer think
       | about.
       | 
       | I imagine just using the second spellchecking pass with Gemini
       | would be almost as effective.
        
       | vayup wrote:
       | Something along these lines, as part of the prompt, has worked
       | for me.                              # User-Defined Dictionary
       | Always use the following exact terms if they sound similar in the
       | audio:                          ```json
       | {{jsonDictionary}}                     ```
        
       | another_twist wrote:
       | Use any proper ASR service that supports custom vocabulary ?
       | Transcribe and Deepgram definitely support it and if you want to
       | go fancy Nemo with custom vocabulary.
       | 
       | Are there constraints where you have to use Gemini ?
        
       | rancar2 wrote:
       | The business edition of Wispr Flow does this well, and includes
       | sharing among teams so you can make sure that the company wide
       | vocabulary is consistent and well recognized.
       | 
       | https://wisprflow.ai/business
        
         | e1g wrote:
         | +1 from another happy Whispr Flow power user. I tried 4-5
         | similar apps and even built one with Assembly AI, but Whispr is
         | a significant upgrade above the rest for correctly recognizing
         | my accent and jargon. Having the custom vocabulary helps.
        
       | alex-skobe wrote:
       | We have used markdown and list of vocabulary at the end like
       | 
       | Return company name only from dictionary
       | 
       | #dictionary 1:Apple 2:..
       | 
       | And than Vercel AI sdk + Zod Schema + Gemini 2.5 pro and it
       | pretty accurate
        
       | meerab wrote:
       | I use a two-pass approach - first pass with ASR (OpenAI Whisper)
       | and second pass with an LLM. I ask users to provide context
       | upfront and use that as the "initial_prompt" parameter in
       | Whisper:
       | https://github.com/openai/whisper/discussions/963#discussion...
       | 
       | Gemini might have similar capabilities for custom vocabulary,
       | though I'm not certain about their specific implementation. The
       | two-pass ASR+LLM approach could work with Gemini's output as
       | well.
        
       | bbarnett wrote:
       | Give it a database backend with lots and lots of facts. Things
       | verified by humans. There, AI 'fixed'.
        
         | brokensegue wrote:
         | I don't get your suggestion. How does the database tie into
         | speech to text?
        
       ___________________________________________________________________
       (page generated 2025-10-26 23:00 UTC)