[HN Gopher] Omnilingual ASR: Advancing automatic speech recognit...
       ___________________________________________________________________
        
       Omnilingual ASR: Advancing automatic speech recognition for 1600
       languages
        
       HF Demo: https://huggingface.co/spaces/facebook/omniasr-
       transcription...  GitHub:
       https://github.com/facebookresearch/omnilingual-asr
        
       Author : jean-
       Score  : 50 points
       Date   : 2025-11-10 18:10 UTC (4 hours ago)
        
 (HTM) web link (ai.meta.com)
 (TXT) w3m dump (ai.meta.com)
        
       | meetpateltech wrote:
       | HF Demo: https://huggingface.co/spaces/facebook/omniasr-
       | transcription...
       | 
       | GitHub: https://github.com/facebookresearch/omnilingual-asr
        
         | dang wrote:
         | Thanks! I've added those links to the toptext as well.
        
       | tschellenbach wrote:
       | any insights on latency?
        
       | samat wrote:
       | How hard is it to make TTS out of this? A few independent
       | journalists from Belarus asked for TTS in their language, but I
       | am no expert, was thinking about re-using Mozilla's work. What's
       | the easiest way to get working TTS for a language?
        
         | kulahan wrote:
         | From TFA, it says that it's extremely easy to add new languages
         | with just a few examples. I didn't see specifics on how "few"
         | it really is, though.
        
           | nl wrote:
           | This is ASR not TTS though.
        
         | woodson wrote:
         | You can use the OmniASR SSL models instead of their older MMS
         | models to create TTS models:
         | https://github.com/ylacombe/finetune-hf-vits
        
       | stuffoverflow wrote:
       | This seems like a massive improvement for openly available local
       | ASR. Even the 300M model outperforms whisper-large-v3 according
       | to the paper's benchmarks.
        
         | lostmsu wrote:
         | Not sure, I recorded 3 seconds of voice (a single sentence) and
         | the hf demo misrecognized about half of the words.
        
       | AIorNot wrote:
       | the global language explorer is fascinating -great work guys
       | 
       | https://aidemos.atmeta.com/omnilingualasr/language-globe
       | 
       | - we are getting closer to BabelFish.. at least for the Earth!
        
       | cadamsdotcom wrote:
       | Only a few gb of weights will recognize speech in 1600+
       | languages.
       | 
       | Freely downloadable and usable by anyone for almost anything.
       | 
       | We truly live in the future.
        
       ___________________________________________________________________
       (page generated 2025-11-10 23:00 UTC)