[HN Gopher] Speaking without vocal cords, thanks to a new AI-ass...
___________________________________________________________________
Speaking without vocal cords, thanks to a new AI-assisted wearable
device
Author : geox
Score : 116 points
Date : 2024-03-24 00:39 UTC (22 hours ago)
(HTM) web link (newsroom.ucla.edu)
(TXT) w3m dump (newsroom.ucla.edu)
| khimaros wrote:
| i'm really excited to see progress being made in this space.
| subvocal speech recognition seems to be an underfunded area of
| research.
|
| my sense is that it has the potential to make hands free
| interaction with our devices in public spaces less obnoxious and,
| consequently, more socially acceptable.
|
| however, i notice that the article doesn't mention anything about
| dictionary size, which is a very important consideration for a
| tool of this kind.
| graphe wrote:
| https://en.m.wikipedia.org/wiki/Basic_english this communicates
| English efficiently with 850 words. I don't think it's basic
| English is any good but I can see them making simplified
| English the lingua franca to boost 'literacy rates' in the
| future.
| thorum wrote:
| > The research team demonstrated the system's accuracy by
| having the participants pronounce five sentences -- both aloud
| and voicelessly -- including "Hi, Rachel, how are you doing
| today?" and "I love you!" (...) Going forward, the research
| team plans to continue enlarging the vocabulary of the device
| through machine learning and to test it in people with speech
| disorders.
|
| It's a proof of concept at this stage but very cool.
| dontreact wrote:
| While subvocal is cool and would allow for speech in more
| places, something that's earlier on the tech tree and that I
| would like to see is just robust lipreading.
|
| I already am comfortable talking to my phone quietly using my
| AirPods while looking at my screen, but it seems like in loud
| public places the accuracy becomes unusable. I imagine it could
| be easily recovered by the additional signal of lipreading.
| zharknado wrote:
| Very cool! This is an insanely impressive sensor, but the
| proposed application is still in dream phase.
|
| > Going forward, the research team plans to continue enlarging
| the vocabulary of the device through machine learning and to test
| it in people with speech disorders.
|
| They haven't tried giving it to a person with a voice disorder.
| So it just might not work in that application at all. That will
| likely depend on the degree to which laryngeal muscles are
| implicated in a given person's disorder.
|
| That's certainly a valid starting place for research purposes,
| but it's very early days.
|
| And I imagine you'll need some very interesting cabling attached
| to a somewhat beefy device to actually run live inference from
| this data, plus to drive the speech synthesis.
| anonylizard wrote:
| This seems only useful to people who once had a voice, then lost
| their voice. Because only this way, would they have a unified
| mapping of voice cord movements to actual voices. Deaf and mutes
| can't really use this.
|
| It also basically mandates a patch to your throat, because no way
| of detecting vibrations otherwise.
|
| I wonder if there are visual based ways, like sign->text,
| expression->text, that would benefit from the larger developments
| in LLMs. Like an LLM that has access to your conversation
| history, so when you give your smartphone camera a hand sign and
| a smile, it can guess and output an entire intended speech.
| tbenst wrote:
| This is a super cool device. Note that the decoding is highly
| limited: they decode into one of five different sentences. This
| is easier than five words for example as there is more
| information to distinguish.
|
| Unfortunately the media is blowing this way of out proportion as
| the larynx alone does not contain sufficient information to
| decode silent speech.
|
| If you also sense the lips, tongue articulators, and jaw, then
| general English decoding becomes possible with high accuracy (eg
| see our recent work here:
| https://x.com/tbenst/status/1767952614157848859). It's not in the
| preprint but I've done experiments with only the larynx recorded
| and performance is pretty abysmal on even a 10 word vocabulary---
| hence why they did a five sentence task.
| irviss wrote:
| > If you also sense the lips, tongue articulators, and jaw,
| then general English decoding becomes possible with high
| accuracy
|
| A bit OT but I see this frequently and I'm curious. Why do you
| English speakers (or just a US phenomenon?) tend to use the
| word "English" instead of "language", "linguistic" or one of
| its related words to refer to a general concept?
| atopal wrote:
| There are about 6000 spoken languages around the world with
| an extreme variety in how they produce meaning. How could you
| make sweeping statements about all of them?
| x1798DE wrote:
| Not OP, but as a native English speaker and former scientist
| (though not in this area), I would interpret "x does y on
| English tasks" to mean "we tested this in English and don't
| know if the effect generalizes to other languages".
| thaumasiotes wrote:
| In this case we do know if the effect generalizes to other
| languages. It cannot fail to; the larynx, lips, tongue, and
| jaw are almost all there is. For example, vowels are
| conventionally defined by jaw position ("height"), tongue
| position ("frontness"), and lip configuration ("rounded" or
| not).
|
| You might miss some things like creaky voice or ejectives,
| you'll probably miss aspiration, but all that does is give
| you a worst-case scenario analogous to a native speaker
| trying to understand someone with a foreign accent.
| Extremely high accuracy will be possible.
| AlecSchueler wrote:
| This is a reasonable hypothesis but if only English has
| been studied then it would be unscientific to extrapolate
| at this time.
| thaumasiotes wrote:
| Sure, in the same sense that it would be "unscientific"
| to conclude that someone's amputated leg didn't
| regenerate by chance, because the sample size is only 1.
|
| If you know how you're recognizing English, and you know
| that other languages do not differ from English in
| relevant ways, then you know you can recognize those
| other languages. Pretending you don't know something you
| do know is not scientific.
| metabagel wrote:
| Other languages have different sounds which aren't
| present in English.
| thaumasiotes wrote:
| So? They don't have sounds that are produced in a manner
| other than arranging the lips, tongue, and jaw.
|
| (Actually, they do. So does English; I already mentioned
| aspiration. But those are minor elements.)
| wizzwizz4 wrote:
| They're minor elements _in English_ - and even then, you
| can construct sentences where the meaning changes based
| on aspiration.
| brookst wrote:
| This seems like damned-either-way. If they had only
| tested English and asserted that it was universally
| applicable to all languages, it's likely you (or someone
| else) would rightfully object that it's annoying when
| English speakers assume that's all there is.
| thaumasiotes wrote:
| That's not a similar claim. Anyone can be annoyed by
| anything; the idea that it's "unscientific" to state that
| a method of recognizing English by measuring the
| positions of the lips, tongue, and jaw alongside the
| activity of the larynx will apply to every other spoken
| language in the world is ludicrous on its face. It will,
| because those measurements capture nearly every dimension
| of phonetic variation that exists. No one could believe
| otherwise, except apparently for metabagel.
| tbenst wrote:
| x1798DE captured my intent well. For example, tonal
| languages like Mandarin or Cantonese may be more
| difficult to decode if vocal cords aren't vibrating, and
| languages with more phonemes that have both a voiced and
| unvoiced version might be more difficult. I still think
| decoding will be possible for general language, but
| that's a hypothesis whereas I know it's true for English.
| khazhoux wrote:
| This is your misperception.
|
| In the instances where a person says "English" in this kind
| of context, it catches your attention and you infer that the
| person is an English-speaker, and possibly American.
|
| But when a person uses the generic word "language", you don't
| notice it.
|
| This leads you to believe that English speakers "tend to use
| the word English," when that's not the case necessarily.
|
| I don't know what this perceptual fallacy is called, but
| there's probably a word. In English :-)
| johnisgood wrote:
| I have not noticed this. I just assume that they are
| specifically talking about a language, in this case: English.
| roenxi wrote:
| I'd speculate English speakers are used to being part of a
| society where non-English speakers are present and
| politically important. It is polite not to assume that
| English = language. Even on the British Isles English isn't a
| universal thing. Let alone somewhere like America where it
| isn't even native.
|
| "Language" just doesn't mean "English". In Australia if
| someone is talking about "language" on its own I'd assume
| they're Aboriginal advocates.
| AlecSchueler wrote:
| > Even on the British Isles English isn't a universal
| thing. Let alone somewhere like America where it isn't even
| native.
|
| English isn't native to all of those isles either only
| Great Britain.
| ImHereToVote wrote:
| I bet if you listened to the feedback you could teach yourself
| to talk using the larynx and surrounding muscles.
| jvanderbot wrote:
| Why can't the muscles of the larnex and perhaps chest /
| diaphragm, be monitored and mapped to vocal chord noises,
| rather than full speech? Just put the noise in the throat and
| let the rest of the body make it work.
| ImHereToVote wrote:
| I'll take the whole lot.
| feverzsj wrote:
| How it compares to electrolarynx, which give you robotic sound.
| croemer wrote:
| "Speaking" is a hyperbole. It allows you to say exactly 5 phrases
| with 95% accuracy, after repeating each sentence 100 times. In
| other words, it's totally useless. The sentences are so different
| that they can be distinguished almost entirely by length. I'm
| very surprised anyone thinks this is useful.
|
| Excerpt: "A brief demonstration was made with five sentences that
| we had selected for training the algorithm (S1: "Hi Rachel, how
| you are doing today?", S2: "Hope your experiments are going
| well!", S3: "Merry Christmas!", S4: "I love you!", S5: "I don't
| trust you."). Each participant repeated each sentence 100 times
| for data collection."
|
| I never read a press release from a university, it's always
| exaggerated.
|
| Original study:
| https://www.nature.com/articles/s41467-024-45915-7
| light_hue_1 wrote:
| Exactly. They took a neat device and wrapped it in a BS story
| that's wildly unscientific.
|
| Nature and Science don't mind if people outright lie about what
| their research means as long as it gets hits. The paper is
| pretty much just as bad.
|
| This is how mistrust for science slowly builds up when people
| publish obvious falsehoods.
| joshspankit wrote:
| Is it just me, or does anyone else think it would be amazing to
| use upcoming voice assistants with something like this letting
| you "talk" silently?
___________________________________________________________________
(page generated 2024-03-24 23:02 UTC)