[HN Gopher] Exploring JEPA for real-time speech translation
___________________________________________________________________
Exploring JEPA for real-time speech translation
Author : christiansafka
Score : 13 points
Date : 2026-03-11 08:14 UTC (2 days ago)
(HTM) web link (www.startpinch.com)
(TXT) w3m dump (www.startpinch.com)
| numpad0 wrote:
| > You'd want parallel speech data: the same utterance spoken in
| English, Portuguese, Japanese, Arabic, Mandarin, and dozens more
| languages.
|
| There is no such things as parallel speech data. The idea that
| parallel text is a thing is dubious in the first place, like
| there's "translation tone" in Japanese that refers to the voice-
| of-text distinct to translated Western texts. The entire concept
| of _translation_ between distinct human languages is a thing born
| out of practical necessity rather than something with a concrete
| theoretical basis.
|
| The Babel fish in _The Hitchhiker 's Guide to the Galaxy_ is
| supposed to be mind-reading. They don't rely on spoken utterances
| at all, but they read the minds of creatures within its
| telepathic range and feed the language portions of it into its
| host's brain, thereby achieving zero-lag realtime translation.
| This may or may not mean the author Douglas Adams knew that
| general universal translation is impossible, but it makes the
| fictional fish not contradictory to the reality that zero-lag
| interpretation through audio is basically impossible.
|
| You CAN probably do a parallel speech voice to voice if you're
| okay with something like 30s delay. But if you want a voice-to-
| voice no-pause zero-delay, I mean, people sometimes think as they
| speak, everyone can do, yet not everyone speaks languages with
| same word orders, you literally need a crystal ball that reads
| dices before they're even rolled.
| brandonb wrote:
| Very cool. I learned something new about why EMA (exponential
| moving average) is needed:
|
| > EMA-based training dynamics like JEPA's don't optimize any
| smooth mathematical function, yet they provably converge to
| useful, non-collapsed representations.
|
| All the papers say EMA avoids "representation collapse" without
| justifying it. Didn't realize there were any theoretical results
| here.
___________________________________________________________________
(page generated 2026-03-13 23:00 UTC)