[HN Gopher] Speaking without vocal cords, thanks to a new AI-ass...
       ___________________________________________________________________
        
       Speaking without vocal cords, thanks to a new AI-assisted wearable
       device
        
       Author : geox
       Score  : 116 points
       Date   : 2024-03-24 00:39 UTC (22 hours ago)
        
 (HTM) web link (newsroom.ucla.edu)
 (TXT) w3m dump (newsroom.ucla.edu)
        
       | khimaros wrote:
       | i'm really excited to see progress being made in this space.
       | subvocal speech recognition seems to be an underfunded area of
       | research.
       | 
       | my sense is that it has the potential to make hands free
       | interaction with our devices in public spaces less obnoxious and,
       | consequently, more socially acceptable.
       | 
       | however, i notice that the article doesn't mention anything about
       | dictionary size, which is a very important consideration for a
       | tool of this kind.
        
         | graphe wrote:
         | https://en.m.wikipedia.org/wiki/Basic_english this communicates
         | English efficiently with 850 words. I don't think it's basic
         | English is any good but I can see them making simplified
         | English the lingua franca to boost 'literacy rates' in the
         | future.
        
         | thorum wrote:
         | > The research team demonstrated the system's accuracy by
         | having the participants pronounce five sentences -- both aloud
         | and voicelessly -- including "Hi, Rachel, how are you doing
         | today?" and "I love you!" (...) Going forward, the research
         | team plans to continue enlarging the vocabulary of the device
         | through machine learning and to test it in people with speech
         | disorders.
         | 
         | It's a proof of concept at this stage but very cool.
        
         | dontreact wrote:
         | While subvocal is cool and would allow for speech in more
         | places, something that's earlier on the tech tree and that I
         | would like to see is just robust lipreading.
         | 
         | I already am comfortable talking to my phone quietly using my
         | AirPods while looking at my screen, but it seems like in loud
         | public places the accuracy becomes unusable. I imagine it could
         | be easily recovered by the additional signal of lipreading.
        
       | zharknado wrote:
       | Very cool! This is an insanely impressive sensor, but the
       | proposed application is still in dream phase.
       | 
       | > Going forward, the research team plans to continue enlarging
       | the vocabulary of the device through machine learning and to test
       | it in people with speech disorders.
       | 
       | They haven't tried giving it to a person with a voice disorder.
       | So it just might not work in that application at all. That will
       | likely depend on the degree to which laryngeal muscles are
       | implicated in a given person's disorder.
       | 
       | That's certainly a valid starting place for research purposes,
       | but it's very early days.
       | 
       | And I imagine you'll need some very interesting cabling attached
       | to a somewhat beefy device to actually run live inference from
       | this data, plus to drive the speech synthesis.
        
       | anonylizard wrote:
       | This seems only useful to people who once had a voice, then lost
       | their voice. Because only this way, would they have a unified
       | mapping of voice cord movements to actual voices. Deaf and mutes
       | can't really use this.
       | 
       | It also basically mandates a patch to your throat, because no way
       | of detecting vibrations otherwise.
       | 
       | I wonder if there are visual based ways, like sign->text,
       | expression->text, that would benefit from the larger developments
       | in LLMs. Like an LLM that has access to your conversation
       | history, so when you give your smartphone camera a hand sign and
       | a smile, it can guess and output an entire intended speech.
        
       | tbenst wrote:
       | This is a super cool device. Note that the decoding is highly
       | limited: they decode into one of five different sentences. This
       | is easier than five words for example as there is more
       | information to distinguish.
       | 
       | Unfortunately the media is blowing this way of out proportion as
       | the larynx alone does not contain sufficient information to
       | decode silent speech.
       | 
       | If you also sense the lips, tongue articulators, and jaw, then
       | general English decoding becomes possible with high accuracy (eg
       | see our recent work here:
       | https://x.com/tbenst/status/1767952614157848859). It's not in the
       | preprint but I've done experiments with only the larynx recorded
       | and performance is pretty abysmal on even a 10 word vocabulary---
       | hence why they did a five sentence task.
        
         | irviss wrote:
         | > If you also sense the lips, tongue articulators, and jaw,
         | then general English decoding becomes possible with high
         | accuracy
         | 
         | A bit OT but I see this frequently and I'm curious. Why do you
         | English speakers (or just a US phenomenon?) tend to use the
         | word "English" instead of "language", "linguistic" or one of
         | its related words to refer to a general concept?
        
           | atopal wrote:
           | There are about 6000 spoken languages around the world with
           | an extreme variety in how they produce meaning. How could you
           | make sweeping statements about all of them?
        
           | x1798DE wrote:
           | Not OP, but as a native English speaker and former scientist
           | (though not in this area), I would interpret "x does y on
           | English tasks" to mean "we tested this in English and don't
           | know if the effect generalizes to other languages".
        
             | thaumasiotes wrote:
             | In this case we do know if the effect generalizes to other
             | languages. It cannot fail to; the larynx, lips, tongue, and
             | jaw are almost all there is. For example, vowels are
             | conventionally defined by jaw position ("height"), tongue
             | position ("frontness"), and lip configuration ("rounded" or
             | not).
             | 
             | You might miss some things like creaky voice or ejectives,
             | you'll probably miss aspiration, but all that does is give
             | you a worst-case scenario analogous to a native speaker
             | trying to understand someone with a foreign accent.
             | Extremely high accuracy will be possible.
        
               | AlecSchueler wrote:
               | This is a reasonable hypothesis but if only English has
               | been studied then it would be unscientific to extrapolate
               | at this time.
        
               | thaumasiotes wrote:
               | Sure, in the same sense that it would be "unscientific"
               | to conclude that someone's amputated leg didn't
               | regenerate by chance, because the sample size is only 1.
               | 
               | If you know how you're recognizing English, and you know
               | that other languages do not differ from English in
               | relevant ways, then you know you can recognize those
               | other languages. Pretending you don't know something you
               | do know is not scientific.
        
               | metabagel wrote:
               | Other languages have different sounds which aren't
               | present in English.
        
               | thaumasiotes wrote:
               | So? They don't have sounds that are produced in a manner
               | other than arranging the lips, tongue, and jaw.
               | 
               | (Actually, they do. So does English; I already mentioned
               | aspiration. But those are minor elements.)
        
               | wizzwizz4 wrote:
               | They're minor elements _in English_ - and even then, you
               | can construct sentences where the meaning changes based
               | on aspiration.
        
               | brookst wrote:
               | This seems like damned-either-way. If they had only
               | tested English and asserted that it was universally
               | applicable to all languages, it's likely you (or someone
               | else) would rightfully object that it's annoying when
               | English speakers assume that's all there is.
        
               | thaumasiotes wrote:
               | That's not a similar claim. Anyone can be annoyed by
               | anything; the idea that it's "unscientific" to state that
               | a method of recognizing English by measuring the
               | positions of the lips, tongue, and jaw alongside the
               | activity of the larynx will apply to every other spoken
               | language in the world is ludicrous on its face. It will,
               | because those measurements capture nearly every dimension
               | of phonetic variation that exists. No one could believe
               | otherwise, except apparently for metabagel.
        
               | tbenst wrote:
               | x1798DE captured my intent well. For example, tonal
               | languages like Mandarin or Cantonese may be more
               | difficult to decode if vocal cords aren't vibrating, and
               | languages with more phonemes that have both a voiced and
               | unvoiced version might be more difficult. I still think
               | decoding will be possible for general language, but
               | that's a hypothesis whereas I know it's true for English.
        
           | khazhoux wrote:
           | This is your misperception.
           | 
           | In the instances where a person says "English" in this kind
           | of context, it catches your attention and you infer that the
           | person is an English-speaker, and possibly American.
           | 
           | But when a person uses the generic word "language", you don't
           | notice it.
           | 
           | This leads you to believe that English speakers "tend to use
           | the word English," when that's not the case necessarily.
           | 
           | I don't know what this perceptual fallacy is called, but
           | there's probably a word. In English :-)
        
           | johnisgood wrote:
           | I have not noticed this. I just assume that they are
           | specifically talking about a language, in this case: English.
        
           | roenxi wrote:
           | I'd speculate English speakers are used to being part of a
           | society where non-English speakers are present and
           | politically important. It is polite not to assume that
           | English = language. Even on the British Isles English isn't a
           | universal thing. Let alone somewhere like America where it
           | isn't even native.
           | 
           | "Language" just doesn't mean "English". In Australia if
           | someone is talking about "language" on its own I'd assume
           | they're Aboriginal advocates.
        
             | AlecSchueler wrote:
             | > Even on the British Isles English isn't a universal
             | thing. Let alone somewhere like America where it isn't even
             | native.
             | 
             | English isn't native to all of those isles either only
             | Great Britain.
        
         | ImHereToVote wrote:
         | I bet if you listened to the feedback you could teach yourself
         | to talk using the larynx and surrounding muscles.
        
         | jvanderbot wrote:
         | Why can't the muscles of the larnex and perhaps chest /
         | diaphragm, be monitored and mapped to vocal chord noises,
         | rather than full speech? Just put the noise in the throat and
         | let the rest of the body make it work.
        
       | ImHereToVote wrote:
       | I'll take the whole lot.
        
       | feverzsj wrote:
       | How it compares to electrolarynx, which give you robotic sound.
        
       | croemer wrote:
       | "Speaking" is a hyperbole. It allows you to say exactly 5 phrases
       | with 95% accuracy, after repeating each sentence 100 times. In
       | other words, it's totally useless. The sentences are so different
       | that they can be distinguished almost entirely by length. I'm
       | very surprised anyone thinks this is useful.
       | 
       | Excerpt: "A brief demonstration was made with five sentences that
       | we had selected for training the algorithm (S1: "Hi Rachel, how
       | you are doing today?", S2: "Hope your experiments are going
       | well!", S3: "Merry Christmas!", S4: "I love you!", S5: "I don't
       | trust you."). Each participant repeated each sentence 100 times
       | for data collection."
       | 
       | I never read a press release from a university, it's always
       | exaggerated.
       | 
       | Original study:
       | https://www.nature.com/articles/s41467-024-45915-7
        
         | light_hue_1 wrote:
         | Exactly. They took a neat device and wrapped it in a BS story
         | that's wildly unscientific.
         | 
         | Nature and Science don't mind if people outright lie about what
         | their research means as long as it gets hits. The paper is
         | pretty much just as bad.
         | 
         | This is how mistrust for science slowly builds up when people
         | publish obvious falsehoods.
        
       | joshspankit wrote:
       | Is it just me, or does anyone else think it would be amazing to
       | use upcoming voice assistants with something like this letting
       | you "talk" silently?
        
       ___________________________________________________________________
       (page generated 2024-03-24 23:02 UTC)