[HN Gopher] Voice recognition
___________________________________________________________________
Voice recognition
Author : ingve
Score : 36 points
Date : 2023-07-30 07:03 UTC (1 days ago)
(HTM) web link (www.imperialviolet.org)
(TXT) w3m dump (www.imperialviolet.org)
| fudged71 wrote:
| It just so happens I'm in the middle of scraping numeric data out
| of a proprietary application using dictation. It took a bit of
| experimentation to get good results, but it's faster than any
| other scraping method I could find for this specific use case.
| I'm using the dictation tool in the Drafts app, and used ChatGPT
| to write a Drafts Action script to clean up and format the data
| after dictation.
| ianbicking wrote:
| Lately I've been using Google Recorder, and then when I have text
| I _actually_ want transcribed I take the audio and put it through
| Whisper.
|
| From there it would probably be helpful to send it through GPT,
| but I haven't gotten that far. I think GPT is more likely able to
| fix formatting than text at that point. But overcorrection here
| is an issue, the formatting isn't obvious, and I haven't explored
| the prompting in depth.
|
| For actual correction I think it might be superior to get GPT to
| highlight issues and provide alternatives (using some made-up
| markup language), but not actually rewrite.
|
| To get even _more_ fancy, I can imagine giving GPT the low-
| quality Google transcription, having it identify terms and
| subject, and then using that as a prompt to Whisper.
| sebastiennight wrote:
| I tried something like that GPT trick and it is pretty hard to
| get it to not over-correct and over-index the vocabulary words.
|
| For instance if I tell it that "OneTake AI" (my company) or
| "OneTake Chat" (one of our features) belong to the vocabulary
| words, it will twist correctly-transcribed sentences to make it
| appear in there.
|
| > Here is the biggest opportunity in artificial intelligence
| today if you're using Chat in your app
|
| (which is already grammatically correct!)
|
| would suddenly become
|
| > Here is the biggest opportunity in OneTake AI today if you're
| using OneTake Chat in your app
|
| This sycophantism of the LLM causes more trouble than it fixes.
| It's even worse if the vocabulary provided includes names,
| because then the model will randomly twist some sentences (mostly
| the benefit-driven ones, e.g. "so you can get faster results
| today") to add "... thanks to EXPERT_NAME!"
|
| It is also very difficult to get the model to edit structured
| text (like a JSON of the transcription) without breaking the
| schema or the IDs.
|
| I am still iterating to find a better method, and I think the
| "Functions" API can also help.
| sebastiennight wrote:
| If this can be valuable to anybody else, here is an improved
| version of OP's prompt that will perform better than the one in
| the original article (it fixes every mistake plus punctuation
| and capitalization), with chatgpt-3.5-turbo-16k at a
| temperature of 0.7.
|
| I find that using the wording in the OP's article seems to
| create fewer sycophantic corrections than my previous approach,
| so that's a win. However the structure of my prompt seems to
| work better. So here's a combined result.
|
| System prompt:
|
| ``` The following English transcription is the
| output of a poor speech recognition system. Please correct its
| errors and reply with the corrected version without any
| preamble. Some of the uncommon words that the speech
| recognition system probably misspelled are: ["U2F",
| "CTAP", "FIDO", "CTAP2", "JSON", "FIDO2", "Web Authn", "bearer
| token", "hash", "phishing", "RP ID", "passkey", "caBLE"]
| If those words are present and misspelled, please fix their
| spelling. Please fix the spelling, grammar, any
| inconsistencies and transcription errors in each sentence.
|
| ```
|
| Note: to improve performance, I also recommend:
|
| - providing the model with a before-and-after example, by using
| the USER and ASSISTANT roles to roleplay the whole thing (not
| putting the example in SYSTEM!)
|
| - providing the source text as a CSV or JSON, and asking the
| model to provide JSON as the output.
|
| - prefixing the Assistant response with an Assistant reminder
| of the instruction, like this:
|
| ASSISTANT
|
| ``` Return a corrected version with correct
| English grammar, punctuation and spelling.
| Correction (as JSON):
|
| ```
| kleiba wrote:
| For those who are wondering:
|
| This article is about (automatic) _speech recogniton_ , also
| known by its acronym ASR or as _speech-to-text_ (STT).
|
| It has nothing to do with automatic _speaker identification_ ,
| which in turn is also different from _speaker diarization_. The
| latter refers to the process of automatically identifying from a
| group of copresent speakers who is the one speaking at any point
| in time without caring about the actual identity of that person.
| DarthNebo wrote:
| I'm building a tool for transcriptions where your brand, product
| or any other technical jargon or heck even your own name does not
| need a fine-tuned model all the time. Both as a free native Mac
| app & a SaaS tool for those who need to process in bulk. Hit me
| up at nebo@minusgreed.com to check it out, will launch as
| FortuneSpeech.com once out of beta.
|
| Thanks to this post, I kind of have another idea for a
| 'corrections to transcript' feature that Llama2-7B even on CPU
| can help with.
| arjvik wrote:
| See sebastiennight's parent comment above - I think they meant
| to reply to you.
|
| Curious to see your response.
| sebastiennight wrote:
| Actually I wasn't, but if this can be useful to GP I shared
| my actual prompt in a follow-up comment.
|
| I would be interested to see if LLama2 can perform a similar
| task.
| voiper1 wrote:
| You can also try Whisper, or try google or amazon's ASR and give
| it hints/dictionary of common words.
___________________________________________________________________
(page generated 2023-07-31 23:02 UTC)