[HN Gopher] After All Is Said and Indexed - Unlocking Informatio...
___________________________________________________________________
After All Is Said and Indexed - Unlocking Information in Recorded
Speech
Author : jeadie
Score : 47 points
Date : 2023-04-12 07:01 UTC (3 days ago)
(HTM) web link (github.com)
(TXT) w3m dump (github.com)
| moneywoes wrote:
| How does this compare to using Whisper and feeding that into a
| vector DB and querying with a LLM
|
| Pardon the dumb question I only have an elementary understanding
| jeadie wrote:
| A really interesting blog post I found using LLMs for audio
| search which I think is a pretty nifty/new idea.
|
| I've found it cumbersome using some of the new vector DBs
| (chroma, faiss, etc) to make end to end systems, but with Marqo
| it doesn't seem too hard.
| thomasahle wrote:
| > I've found it cumbersome using some of the new vector DBs
| (chroma, faiss, etc) to make end to end systems
|
| What parts are cumbersome?
| rektide wrote:
| Hadn't heard of the thing they were putting their data into,
| Marqo, a "tensor search for humans" , https://github.com/marqo-
| ai/marqo
| notjulianjaynes wrote:
| This is interesting but what problem does it solve better than
| CTRL+F-ing a transcript? It seems like this would be a worse
| solution for when the precise way someone says something could be
| important (ex. journalists parsing an interview, students
| studying their recorded lectures) and that it would be most
| useful if you were working with a large volume of recorded audio,
| such as customer service calls. This makes me somewhat
| uncomfortable, but perhaps I am not fully understanding how it
| works.
|
| Edit: wording
| UncleEntity wrote:
| > This is interesting but what problem does it solve better
| than CTRL+F-ing a transcript?
|
| Producing the transcript?
|
| Being able to classify and search data seems like a pretty big
| deal these days too.
| password4321 wrote:
| Both speaker and speech recognition are done in the article using
| huggingface.
|
| Is there anything as good ready to use on-prem for the
| diarization (speaker recognition)?
|
| I've heard good things about whisper(.cpp) for speech recognition
| and vosk used to be king of that hill...
| boredemployee wrote:
| vosk is really good, but also a good example of an open source
| project with great potential, but doesn't scale up because the
| person behind it is a douchebag.
|
| documentation is poor, and what you find is sparsed outdated
| shit on the web, so it's really hard to find help.
| rolisz wrote:
| Diarization can be done on premise using pyannote (what they
| use in the article). Huggingface offers a library to run things
| locally and an API to run things on their cloud. Pyannote is
| available under an MIT licence
___________________________________________________________________
(page generated 2023-04-15 23:01 UTC)