[HN Gopher] Complete silence is always hallucinated as "trjm@ nn...
       ___________________________________________________________________
        
       Complete silence is always hallucinated as "trjm@ nnsy qnqr" in
       Arabic
        
       Author : edent
       Score  : 518 points
       Date   : 2025-07-22 05:42 UTC (17 hours ago)
        
 (HTM) web link (github.com)
 (TXT) w3m dump (github.com)
        
       | GaggiX wrote:
       | Whisper frequently generates random credits. I guess they didn't
       | curate the dataset much at the time.
        
       | sivers wrote:
       | to save you a lookup:
       | 
       | The Arabic text "rjm@ nnsy qnqr" translates to English as: "Nancy
       | Qanqar's translation" or "Translation by Nancy Qanqar"
       | 
       | "rjm@" means "translation" and "nnsy qnqr" is the name "Nancy
       | Qanqar"
        
         | aprilthird2021 wrote:
         | And it seems to be because the training data is largely
         | unofficial subtitles from movies. Which often have a string
         | like "Translated by X" at the end of the movie which is often
         | silent while credits roll.
        
           | iqfareez wrote:
           | make sense..
        
           | rob74 wrote:
           | Looks like they used more official sources for German -
           | there, silence is apparently hallucinated as "Untertitelung
           | des ZDF fur funk, 2017" according to one of the comments on
           | the issue. Which makes sense, as the public broadcasters'
           | "Mediathek" is probably the largest freely available resource
           | of subtitled videos in Germany. I wonder if the ZDF gave its
           | approval for it being used for LLM training though?
        
             | MrGilbert wrote:
             | > I wonder if the ZDF gave its approval for it being used
             | for LLM training though?
             | 
             | I am pretty sure they didn't get asked.
        
               | h784gljf wrote:
               | Just like the people forced to pay for ZDF under threat
               | of imprisonment.
        
               | eclecticfrank wrote:
               | This person refers to the German television and radio fee
               | (Rundfunkgebuhren).[1] It is a state-mandated system that
               | ensures free (as in free speech) and (relatively) neutral
               | public broadcasting institutions. There is a constant and
               | engaged discussion, because every household in Germany
               | has to pay this fee. Exceptions are made only for low-
               | income households.
               | 
               | [1] https://en.wikipedia.org/wiki/ARD_ZDF_Deutschlandradi
               | o_Beitr...
        
               | rob74 wrote:
               | A constant discussion, lately fueled by extremist parties
               | (AfD) who feel treated unfairly by (amongst others) the
               | public broadcasters (which has parallels to Trump's
               | recent campaign against public broadcasters in the US).
        
               | blueflow wrote:
               | Can't argue them - Tageschau always has been trashtalking
               | people with the wrong opinion.
               | 
               | Back in 2011, Tageeschau openly rallied against Muslims
               | and wanting public broadcasting gone was a leftist
               | position. The whole thing is completely asinine to anyone
               | who remembers.
        
               | throwaway290 wrote:
               | So, people are made pay for it, and it makes it fair if
               | billion USD corporations don't?
        
               | TheBicPen wrote:
               | Just like any other public service paid for with public
               | funds?
        
               | h43z wrote:
               | Oh it's like any other. Then just add another one!
        
               | wongarsu wrote:
               | I'm being made to pay for Autobahnen I barely use,
               | finance kindergartens despite not having a child, and
               | made to pay into public pensions with little hope of
               | getting close to the same value out. All under threat of
               | imprisonment, many without a way to even refuse (not that
               | I'd want to) The only thing that sets the pubic
               | broadcasting fee apart is that it's collected separately
               | from taxes in an attempt to reduce the influence
               | politicians have on broadcasters
        
             | Zacharias030 wrote:
             | definitely not! The media platform of the German public
             | television networks is even geoblocking anyone outside of
             | Germany.
             | 
             | https://www.ardmediathek.de/
        
             | unusual-name wrote:
             | Most content from Funk (youtubers funded by public german
             | broadcasters) is available on youtube without any
             | geoblocking or other limitations.
        
               | rob74 wrote:
               | Ah, ok, thanks for the info, TIL! "We are funk - the
               | first public service content network that started on
               | October 1, 2016. We create online-only content on social
               | networks and third-party platforms, including YouTube,
               | Instagram, Snapchat, TikTok, Spotify, Apple Music or
               | Twitch for 14-29 year-olds."
               | (https://presse.funk.net/das-ist-funk/, scroll down for
               | the English version). I live in Germany, and I even watch
               | public broadcasters regularly, but this is the first time
               | I have heard about funk (I even initially thought it was
               | misspelled, usually it's written with a capital F). But
               | I'm not part of the targeted audience (not now, nor even
               | back in 2016 when it was launched), so all good...
        
               | layer8 wrote:
               | I'm pretty sure that content doesn't come with a license
               | granting unlimited usage rights.
        
               | darkwater wrote:
               | from the link[1] another user posted:
               | 
               | > We have a public service mandate, which means that we
               | have very clear responsibilities according to the state
               | media treaty. For us, this means that our top priority is
               | actually reaching our target audience, namely
               | approximately 15 million people living in Germany between
               | the age of 14 and 29 who have internet access
               | 
               | It's not a binding contract for sure but I don't think
               | that OpenAI or other AI scraper is their target.
               | 
               | [1] https://presse.funk.net/das-ist-funk/
        
             | bigiain wrote:
             | A more appropriate output might be ``4'33" -- John Cage,
             | 1952``
        
             | aprilthird2021 wrote:
             | > I wonder if the ZDF gave its approval for it being used
             | for LLM training though?
             | 
             | Obviously a rhetorical question. The AI grifters of this
             | decade take what they want and laugh at your pitiful future
        
           | 4gotunameagain wrote:
           | I'm sure they totally did not pirate the audio of said
           | movies.
        
         | mormegil wrote:
         | In Czech, Whisper usually transcribes music as "Titulky
         | vytvoril JohnyX" ("subtitles made by JohnyX") for the same
         | reason.
        
           | actionfromafar wrote:
           | Haha, trained on torrented movies! :-D
           | 
           | The MPA must be so proud.
        
             | Incipient wrote:
             | It's absolutely insane that these companies can't be held
             | liable for what is obvious piracy.
        
               | jdiff wrote:
               | That's the magic of money. Download your favorite
               | artist's discography for personal use? If the MPAA had
               | its way (and it occasionally has), torrenting that could
               | bankrupt you.
               | 
               | The AI industry - soaking up every bit of media available
               | online for commercial purposes, often reproducing it
               | nearly identically - has enough money and capital to
               | influence things its way. And only its way, in case
               | anyone was hoping this might change anything at all for
               | the little guy.
        
               | 1718627440 wrote:
               | The movie industry also has some money and lobbying
               | power. Surely this is a way larger threat than any single
               | torrenter could ever be?
        
               | jdiff wrote:
               | The fact that this is propping up the entire AI industry
               | adds additional weight. When legislating or deciding
               | court cases, some won't be willing to pop the cash cow,
               | some will be worried about falling behind countries that
               | don't enforce copyright evenly. IP owners _are_ trying to
               | go after the AI industry, with only mixed to poor
               | success.
        
               | verzali wrote:
               | Hard to justify that they can't afford to pay when they
               | have multi-billion dollar valuations and are apparently
               | paying hundreds of millions to get a single engineer.
        
               | jdiff wrote:
               | Maybe. But we are talking about the whole of copyrighted
               | creative works created and sold by humanity. That'll get
               | expensive no matter who you are.
        
               | d1sxeyes wrote:
               | > Download your favorite artist's discography for
               | personal use? If the MPAA had its way (and it
               | occasionally has), torrenting that could bankrupt you.
               | 
               | I don't think that there are any clear examples of cases
               | where _ONLY_ downloading has resulted in huge fines. All
               | the big bankrupting level fines have been for both
               | downloading _and_ sharing.
               | 
               | You mention that 'torrenting' could bankrupt you, and
               | that is true, but the main reason for the huge fines are
               | that you are taking part in distribution rather than just
               | 'downloading for personal use'.
        
               | jdiff wrote:
               | Given the lack of sense in treating each peer as a lost
               | sale for damages, I think we can safely say they're only
               | interested in making examples out of people and would
               | absolutely go after people for only downloading if the
               | law permitted. Thankfully it's not, but maybe they lobby
               | to make changes in that direction to try and curb future
               | AI industry shenanigans.
        
               | 0points wrote:
               | > I don't think that there are any clear examples of
               | cases where ONLY downloading has resulted in huge fines.
               | 
               | They [1, and others] been hunting and fining downloaders
               | for over a decade now, with the only "evidence" being IP
               | addresses connected with the torrent [2].
               | 
               | 1: https://www.njordlaw.com/filesharing-and-downloading-
               | films/q...
               | 
               | 2: https://admin.ovpn.com/en/blog/online-integrity-new-
               | threats-...
        
               | d1sxeyes wrote:
               | Yes, but torrenting is not ONLY downloading, it's both.
               | The articles you link are very clearly talking about
               | 'Sharing' (from link 2: "File sharing consists of both
               | download and upload of a file.").
        
               | 0points wrote:
               | Yes, thats lawyer speak to make clients/victims believe
               | there is no distinction.
               | 
               | Hint: there is a distinction.
        
               | d1sxeyes wrote:
               | There is indeed, but not when you're torrenting (i.e. you
               | can't download without also uploading).
        
               | 0points wrote:
               | Even when you are torrenting, there is a clear
               | distinction of the different roles.
               | 
               | Copying from another comment I wrote here:
               | 
               | > These are two separate things:
               | 
               | > * Making content available for unauthorized
               | distribution
               | 
               | > * Distributing unauthorized content that someone else
               | already made available
               | 
               | > Seeding isn't making content available, it's keeping
               | content available.
        
               | d1sxeyes wrote:
               | Replied to your other comment (sorry, didn't clock that
               | we had two threads ongoing)
        
               | gruez wrote:
               | >with the only "evidence" being IP addresses connected
               | with the torrent [2].
               | 
               | Is that an unreasonable assumption? As much as people
               | like to come up with excuses like "I had open wifi!" or
               | "I was running a TOR node", judges don't seem inclined to
               | believe them, probably for the same reason they don't
               | seem inclined to believe excuses like "somebody took my
               | car on a joyride and then returned it!" for parking
               | tickets. Remember, both non-commercial copyright
               | infringement lawsuits and parking tickets are tried in
               | civil court, which means the standard is "preponderance
               | of evidence", not "beyond reasonable doubt".
        
               | actionfromafar wrote:
               | DHCP addresses often shuffle on reboots. I don't trust
               | ISPs to keep completely accurate records or give them out
               | in a correct manner if they do.
        
               | gruez wrote:
               | >I don't trust ISPs to keep completely accurate records
               | or give them out in a correct manner if they do.
               | 
               | How hard could it be to keep DHCP logs? Assuming they
               | exist at all, what would cause it to be incorrect?
        
               | d1sxeyes wrote:
               | I'm sure they exist. I think the point is more that you
               | shouldn't need to trust your ISP's record-keeping to
               | avoid life-alteringly big fines.
        
               | 0points wrote:
               | You are missing the point I was replying to, specifically
               | that parent suggested people were only hunted for
               | creating/uploading pirated content, not merely
               | participating in the torrent.
        
               | gruez wrote:
               | >specifically that parent suggested people were only
               | hunted for creating/uploading pirated content, not merely
               | participating in the torrent.
               | 
               | For all intents and purposes, participating in the
               | torrent almost guarantees that you seeded, because all
               | torrent clients upload as you download.
        
               | 0points wrote:
               | These are two separate things:
               | 
               | * Making content available for unauthorized distribution
               | 
               | * Distributing unauthorized content that someone else
               | already made available
               | 
               | Seeding isn't _making_ content available, it 's _keeping_
               | content available.
        
               | gruez wrote:
               | But both are illegal? I suspect if it came out that some
               | torrent seeder was actually part of some sort of piracy
               | ring responsible for ripping the movies, they'd get far
               | stiffer penalties than the few thousand $ fine that
               | typical torrenters get. Moreover isn't AI companies also
               | "keeping content available"?
        
               | 0points wrote:
               | Both are illegal, yes.
               | 
               | That still doesn't make them the same thing. There are
               | different shades of grey, etc.
               | 
               | > Moreover isn't AI companies also "keeping content
               | available"?
               | 
               | I don't know what you mean by that.
        
               | gruez wrote:
               | >I don't know what you mean by that.
               | 
               | The whole point of the thread is that AI companies are
               | getting away with piracy but individuals aren't. But the
               | reality is that AI companies aren't getting away with it
               | (a judge ruled that Anthropic must face trial over their
               | use of pirated books).
               | 
               | More specific to this thread is that claim that "ONLY
               | downloading" hasn't resulted in fines for anyone. So far
               | as I can tell, this is true. People are just quibbling
               | over how someone who's torrenting somehow counts as "only
               | downloading", even though their client is uploading.
        
               | d1sxeyes wrote:
               | That's a really interesting distinction. Clearly there's
               | an "original crime", the first person to rip the CD and
               | put it online (or whatever kids do to pirate music
               | nowadays).
               | 
               | But then if I download a file, create a copy, and share
               | it with you, have I done anything wrong?
               | 
               | To all intents and purposes, seeding _is_ an act of
               | reproduction. You, while keeping your copy, create copies
               | of (parts) of the file and share it to someone else to
               | allow them to assemble a new, second copy.
               | 
               | Whether this is, or should be, a crime is a different
               | question altogether. The main point I was making is that
               | it's the copying/sharing to other people which seems to
               | be a crucial element in these prosecutions.
               | 
               | That's likely intentional: the last thing the *AA folks
               | want is a decision that creating a copy of a copyrighted
               | work for your own personal use is not a crime. But it
               | does seem the courts have decided: making a copy for
               | someone else is indeed illegal.
        
               | jajko wrote:
               | You contradict yourself. There were numerous public cases
               | where they chased people downloading few mp3s just for
               | themselves, and made into example case with massive
               | fines.
               | 
               | If you don't understand how torrents work on technical
               | level I suggest at least some shallow reading. Property
               | rights holders don't care about details, as long as you
               | tick the box of sending a single packet to somebody, off
               | to court with ya.
        
               | d1sxeyes wrote:
               | > There were numerous public cases where they chased
               | people downloading few mp3s just for themselves
               | 
               | If this is true, I have been unable to find any. Can you
               | please share? In all of the cases I was able to find, the
               | huge fines were based on _also_ uploading.
               | 
               | > If you don't understand how torrents work on technical
               | level I suggest at least some shallow reading
               | 
               | This is a bit patronising, and I'm not sure what point
               | you're trying to make. My point is that the only
               | prosecutions I've been able to find are where they were
               | able to prove _uploading_ as well as downloading (and
               | yes, the fact that someone used BitTorrent makes it a
               | slam-dunk, because the protocol makes it impossible to
               | download without also uploading). Are you trying to argue
               | that someone who torrents a copyrighted work doesn 't
               | also share it?
        
               | anon191928 wrote:
               | court judges agree to this
        
               | shadowgovt wrote:
               | It's more the magic of precedent.
               | 
               | The fight about digitized media for personal
               | (entertainment / informational) use were the early
               | aughts. The precedents crafted then don't immediately
               | translate to these cases (novel transformative work from
               | protected materials), and the new precedents have to
               | account for the fact that universities have been training
               | via "piracy" for ages.
               | 
               | (The magic of money factors in to the extent that they
               | can afford the lawyers to remind the court that this
               | isn't settled law yet).
        
               | boredhedgehog wrote:
               | It's an indication how few people consider license
               | infringements as a matter of actual moral import. Those
               | tend to evoke strong feelings.
        
               | ACCount36 wrote:
               | It's the way it should be.
        
               | pavon wrote:
               | Anthropic is going to trial over pirating books for
               | training. The judge was pretty clear that even if
               | training is fair use, the training material must be
               | obtained legally.
               | 
               | These regurgitations combined with proof that a model is
               | familiar with a work could be sufficient evidence to
               | force discovery to determine if the work was pirated.
        
               | scotty79 wrote:
               | What's insane is copyright. How come you can own
               | intellectual property but not pay a property tax? The
               | ecosystem would be much healthier if to get copyright
               | protections you should declare value of your IP (that you
               | are obligated to sell for if the buyer pops up) and pay
               | tax on this for every year you hold the IP.
        
               | retsibsi wrote:
               | > if to get copyright protections you should declare
               | value of your IP (that you are obligated to sell for if
               | the buyer pops up) and pay tax on this for every year you
               | hold the IP
               | 
               | I think this would have some unpalatable consequences.
               | Let's say an author is writing a modestly successful book
               | series: it's not going to make them rich, but it's
               | commercially viable and they care a lot about it for its
               | own sake. Under this system, if the author declares a
               | value commensurate with the (quite small) pure economic
               | value of the IP, they have to live in fear of their right
               | to continue working on their creation being abruptly
               | taken away from them at any point. If they instead
               | declare a value commensurate with the economic value +
               | the extra value that it has to them personally, the
               | resulting tax liability could easily tip the balance and
               | destroy their ability to pursue their writing as a
               | career.
        
               | Workaccount2 wrote:
               | Perhaps the tax would start a decade after the first
               | sale.
        
               | scotty79 wrote:
               | You are always free to update the value before paying
               | tax. If somebody is willing to pay more than it's worth
               | to you they probably have an idea how to turn it into
               | more economic value for the society. So the society
               | should allow them to do that. For a price, of the tax.
               | What I'm proposing is about the financial rights.
               | Individual right, like the right to call yourself author
               | of any given creation should be inalienable.
               | 
               | There are always some cases on the edge. The question is
               | if saving them is worth the cost of the major players
               | running rampant.
        
               | immibis wrote:
               | Indeed, this is a general problem with a lot of these
               | schemes.
               | 
               | We shouldn't abandon the line of investigation, however.
               | We should continue thinking of ways to do this until we
               | find one that works well.
               | 
               | There's a chance it ends up being something that requires
               | a judge to interpret each individual case...
        
               | gruez wrote:
               | >What's insane is copyright. How come you can own
               | intellectual property but not pay a property tax? The
               | 
               | Most jurisdictions that have "property tax" only apply it
               | on certain types of property, most commonly real estate.
               | So it's not that weird that IP isn't taxed.
        
               | codedokode wrote:
               | This is corsairy, not piracy, do not be mistaken.
        
         | beshrkayali wrote:
         | You've got a little typo, it's not "rjm@", it's "trjm@" that
         | means translation, the t at the beginning is missing.
        
       | dandiep wrote:
       | Whisper is unusable IMO because of the hallucinations. Widely
       | documented. Removing silence from audio clips helps, but even
       | then it will auto correct grammar, translating bilingual speech,
       | etc. Improved in the latest audio models but not solved [1]
       | 
       | 1. https://news.ycombinator.com/item?id=43427376
        
         | eric-burel wrote:
         | That's the problem with raws large models, it should always be
         | coupled with satellite small models and logic. It's (probably)
         | easier to detect hallucinations using a traditional ML/DL model
         | that can catch mismatches (it's easy to build a synthetic
         | dataset for this) than transcribing. And the simplest piece of
         | code can detect a silence and that it should match no text.
        
         | ilyakaminsky wrote:
         | I wouldn't describe it as "unusable" so much as needing to
         | understand its constraints and how to work around them. I built
         | a business on top of Whisper [1] and one of the early key
         | insights was to implement a good voice activity detection (VAD)
         | model in order to reduce Whisper's hallucinations on silence.
         | 
         | [1] https://speechischeap.com
        
           | poly2it wrote:
           | How does this make a profit? Whisper should be $0.006 to
           | $0.010 per minute, but you rate less than $0.001? Do you 10x
           | the audio?
        
             | ilyakaminsky wrote:
             | Thanks for noticing. It took a lot of effort to optimize
             | the pipeline every step of the way. VAD, inference server,
             | hardware optimization, etc. But nothing that would
             | compromise on quality. The audio is currently transcribed
             | in its original speed. I'll be sure to publish something if
             | I manage to speed it up without incurring any losses to the
             | WER.
        
         | horseradish7k wrote:
         | well, auto correcting grammar happens in normal subtitles
         | too... "Why don't subtitles match dubbing?" by Tom Scott:
         | https://youtu.be/pU9sHwNKc2c
        
       | GodelNumbering wrote:
       | Interesting that this happens even on large v3. I had once done a
       | deep dive into STT and Whisper Large was the only model that
       | could correctly transcribe Yann LeCun (it was a Lex Friedman
       | podcast), ever since I held the belief that it was the best STT
       | model, this was over 2 years ago
        
       | dlcarrier wrote:
       | Classic overfitting
       | 
       | It's the LLM equivalent of thinking that an out-of-office reply
       | is the translation:
       | https://www.theguardian.com/theguardian/2008/nov/01/5
        
         | stingraycharles wrote:
         | How is this overfitting, rather than a data quality /
         | classification issue?
        
           | hsn915 wrote:
           | uThe Arabic text is the translator's self credit
           | 
           | "Translated by Nancy Qanfar"
        
             | wongarsu wrote:
             | And the German is "subtitles of [public broadcaster] for
             | [content network], 2017
             | 
             | I'm not sure this is really overfitting, the network does
             | exactly what the training data demands. According to the
             | training data silence art the end transcribes to a
             | copyright notice or subtitle credits
        
               | baobabKoodaa wrote:
               | > I'm not sure this is really overfitting, the network
               | does exactly what the training data demands.
               | 
               | What do you think overfitting is, if not that?
        
               | wongarsu wrote:
               | Overfitting would be replicating overly specific details.
               | Like if a specific pattern of silence (or quiet noise)
               | matched to specific copyright notices.
               | 
               | But in this case the behavior seems to generalize over
               | multiple languages, with the model choosing
               | representative "outro silence" captions depending on the
               | language. Which is consistent with the training data
               | showing that outro silence is captioned.
               | 
               | If the model was generalizing perfectly it would show
               | something like "[subtitle credits here]" but that'd be
               | demanding a bit much.
               | 
               | Transcribing outro silence as silence despite the
               | training data consistently transcribing outro silence
               | differently from regular silence would be underfitting
        
               | maxbond wrote:
               | The optimizer is functioning correctly, and the pattern
               | really exists in the training data. But consider:
               | 
               | - This behavior damages the model's performance on out of
               | sample data; every word you predict during silence
               | increases the transcript's Word Error Rate.
               | 
               | - These translation credits are an artifact of our
               | training data, and not a reflection of the process we are
               | modeling (spoken language).
               | 
               | So, while you are correct about the mechanism at work
               | here, it is still correct to call learning a spurious
               | pattern which damages our performance "overfitting".
        
               | bmacho wrote:
               | Overfitting is achieving better and better scores on the
               | training material and worse and worse scores on unseen
               | tasks. More at: https://en.wikipedia.org/wiki/Overfitting
               | #Machine_learning
               | 
               | This is just wrong training data.
        
               | samrus wrote:
               | fitting on noise in the training data is exactly what
               | overfitting is. underfitting is smoothing out signal
        
               | wongarsu wrote:
               | Exactly. Underfitting would be if the model doesn't pick
               | up on the fact that outro silence is labeled differently
               | from regular silence and transcribes them the same
        
               | xigoi wrote:
               | Overfitting implies a failure to properly generalize the
               | training data. Here it generalized them correctly.
               | Garbage in, garbage out.
        
               | samrus wrote:
               | No. Because there would have been indtances in the data
               | where silence was labelled correctly. But the model
               | couldnt handle the null case, so it over fit on the
               | outros. But generally it fit on the random error in the
               | label of the null feature. Which is what overfitting is
        
               | andrepd wrote:
               | That's literally what overfitting means.
               | 
               | Side-note: it's also yet more evidence that AI companies
               | hoover all data with no regard for legality or copyright
               | status, the very same offences that got other people in
               | jail or with heavy fines.
        
             | efitz wrote:
             | I know it's off topic, but it reminded me that translators
             | like to put in Easter eggs, or at least they used to:
             | https://learn.microsoft.com/en-
             | us/archive/blogs/ericfitz/i-a...
        
           | maxbond wrote:
           | It is a data quality issue which caused the model to overfit.
        
           | mort96 wrote:
           | Isn't overfitting just when the model picks up on an
           | unintended pattern in the training data? Isn't that precisely
           | what this is?
        
           | bGl2YW5j wrote:
           | If the model was able to generalise, you'd expect it to
           | output something like "[silence]" or "...", in response to
           | silence.
           | 
           | Instead, it reverted to what it has seen before (in the
           | training data), hence the overfit.
        
             | stingraycharles wrote:
             | Right, maybe my definition of overfitting was wrong, I
             | always understood it more as trying to optimize for a
             | specific benchmark / use case, and then it starts failing
             | in other areas.
             | 
             | But the way you phrase it, it's just "the model is not
             | properly able to generalize", ie it doesn't understand the
             | concept of silence also makes sense.
             | 
             | But couldn't you then argue that any type of mistake /
             | unknown could be explained as "overfitting" ? Where do you
             | draw the line ?
        
               | heavyset_go wrote:
               | Your definition is one, but the one the OP is using is
               | overfitting to training data.
        
               | stingraycharles wrote:
               | That's exactly my point: by that definition any incorrect
               | answer can be explained by "overfitting to training
               | data".
               | 
               | Where do you draw the line between "overfitting to
               | training data" and "incorrect data" ?
        
               | tempaccount420 wrote:
               | > That's exactly my point: by that definition any
               | incorrect answer can be explained by "overfitting to
               | training data".
               | 
               | Not really, getting 94381294*123=... wrong, but close
               | within the actual answer, cannot be overfitting since it
               | wasn't in the training data.
        
               | maxbond wrote:
               | > [By] that definition any incorrect answer can be
               | explained by "overfitting to training data".
               | 
               | No it doesn't, for instance some errors would be caused
               | by under fitting. The data could also be correct but your
               | hyperparameters (such as the learning rate or dropout
               | rate) could cause your model to overfit.
               | 
               | > Where do you draw the line between "overfitting to
               | training data" and "incorrect data" ?
               | 
               | There's no need to draw a line between two explanations
               | that aren't mutually exclusive. They can (as in this
               | case) both be true. Overfitting is the symptom; dirty
               | data is the cause.
        
             | alienbaby wrote:
             | It's actually because it is incapable of recognising when
             | it does not know the answer. It will give you the nearest
             | match, even if that is completely incorrect.
        
             | mywittyname wrote:
             | I think it's a classification issue.
             | 
             | Silence is never put in the subtitles of a film, since it
             | isn't necessary. The viewers can tell that nothing is being
             | said if there are actors on the screen. And in situations
             | where there are no actors, then there will be a subtitle to
             | indicate what is going on, like "[rock music plays]".
             | 
             | Subtitle authors use this silence to fit in meta
             | information and have done so since the closed captions era.
             | 
             | Proper data cleaning procedures would be to strip this meta
             | data from any subtitle sources. Since this wasn't done,
             | this is fundamentally a classification issue. It may also
             | be an over-fitting issue, but that is secondary to the
             | classification problem.
        
       | Lucasoato wrote:
       | In Italian as well there are random hallucination when parsing
       | silence, something like: "Thank you for watching", "Subtitles
       | by..."
        
         | userbinator wrote:
         | I wouldn't be surprised if "like share and subscribe" also
         | shows up at some point.
        
           | Muromec wrote:
           | "nasha zbroia v tsei moment -- vpodobaika i koment"
        
           | madcaptenor wrote:
           | In the comments to that Github issue, by alentodorov:
           | 
           | in romanian, i've noticed multiple instances where the
           | transcripts ends with "nu uitati sa da-ti like si subscribe"
           | which, as you might easily infer , translates to "don't
           | forget to like and subscribe".
        
       | tornikeo wrote:
       | Garbage in, garbage out. If the training dataset (accidentally)
       | paired silence (`X_train`) with `rjm@ nnsy qnqr` tokens
       | (`y_pred`), then any silence will always be translated to that.
       | Fortunately, this particular problem is easy to fix--just detect
       | and remove silent parts before API call. This also has a side
       | benefit of saving you money on transcription.
        
         | Zacharias030 wrote:
         | and saving money on litigation.
        
       | Hobadee wrote:
       | Little did you all know, this is just being mechanical turked by
       | Nancy Qunqar.
       | 
       | Way to go Nancy! Keep up the good work, ya crazy bastard!
        
         | whamlastxmas wrote:
         | Is this spam? That name only shows as an instagram account and
         | this thread. If you pay for insta followers is this how they
         | get them now? Haha
        
           | DAlperin wrote:
           | That's the name in the Arabic text hallucinated by the model
           | :)
        
             | neom wrote:
             | Who appears to be a real, human, translator:
             | https://www.instagram.com/nancyrk/
        
       | vanschelven wrote:
       | I wonder if hallucinated copyright claims (esp. like the ZDF one
       | at the bottom of the OP) will be introduced as evidence in one of
       | the court cases against "big AI"
        
         | staplers wrote:
         | It already has been and meta won the lawsuit because
         | corporations are sacrosanct.
        
           | vanschelven wrote:
           | Do you mean that specifically a hulucibated text "copyright
           | by not-meta" made it into evidence? Or are you talking about
           | copyright generally?
        
         | berkes wrote:
         | Evidence against what?
         | 
         | "Big AI" is transparent and open about the fact they use all
         | sorts of copyrighted material to train the data. How would "we
         | see an exact chunk of text from our copyrighted material" add
         | to that?
        
           | pbmonster wrote:
           | It appears they have not been training on the official studio
           | subtitle files, but on community transcriptions/translations
           | commonly distributed with torrents.
           | 
           | So not only are they training on copyrighted material, but
           | they didn't even pay for it once, and then they didn't even
           | do minimal data cleaning before training. Which, by the way,
           | is the type of cleaning their LLMs could have done.
        
           | sofixa wrote:
           | Their main defence is that it's fair use because it's
           | transformative (like a human reading a book, getting
           | inspired, and writing something of their own) and not a
           | copypaste illegal distribution (like a human scanning that
           | book and selling it themselves).
           | 
           | Having models hallucinate copyright notices shows that some
           | content is being copypasted as is, which kind of goes against
           | the transformative argument.
           | 
           | (Note: I think that trying to litigate AI with current
           | copyright laws is weird. They were created before LLMs were
           | even imagined, so of course they can't handle them clearly.
           | New laws are needed around this, not trying to bend over
           | backwards to think about what a lawmarker a century ago would
           | have thought about how transformative a thing they couldn't
           | have imagined is.)
        
       | cyp0633 wrote:
       | The same happens with whisper-large-v3 on Chinese transcription:
       | silence is transcribed to something like "please upvote, share
       | and favourite this video". I suspect they trained the model on
       | some random YouTube video without carefully picking really useful
       | data.
        
         | st_goliath wrote:
         | That's interesting, the few times I tried playing with whisper,
         | I had the impression that YouTube style videos or random
         | cellphone videos was something it did particularly bad with
         | (compared to movies). My guess at the time was that most of the
         | training material might be sub titles and raw screen plays.
         | 
         | The videos I tried to transcribe were also Mandarin Chinese,
         | using whisper-large-v3. Besides the usual complaints that it
         | would phonetically "mishear" things and generate nonsense, it
         | was still surprisingly good, compared to other software I
         | played around with.
         | 
         | That said, it would often invent names for the speakers and
         | prefix their lines, or randomly switch between simplified and
         | traditional Chinese. For the videos I tested, intermittent
         | silence would often result in repeating the last line several
         | times, or occasionally, it would insert direction cues (in
         | English for some reason). I've never seen credits or anything
         | like that.
         | 
         | In one video I transcribed, somebody had a cold and was
         | sniffling. Whisper decided the person was crying (transcribed
         | as "* crying *", a cough was turned into "* door closing *").
         | It then transcribed the next line as something _quite_
         | unfriendly. It didn 't do that anymore after I cut the
         | sniffling out (but then the output switched back to traditional
         | Chinese again).
        
         | isoprophlex wrote:
         | Indeed, with another model I would get persistent
         | transcriptions of silent parts into 'Thanks for watching!' or
         | '[MUSIC]'. Pretty dumb that this failure mode wasn't caught in
         | some QA process, and there are now multiple transcription
         | models suffering from the same issue. Having silent parts in
         | your input audio seems like it should be a very common
         | occurrence...
        
           | rollcat wrote:
           | When I was taught mathematics, the zero value was always
           | considered the most important edge case. You prove something
           | for N=0 (or N=1), then for N=M+1.
           | 
           | It's even more important in audio DSP: processing near-zeroes
           | can end up being extremely CPU intensive, look up
           | denormal/subnormal floats.
        
             | inglor_cz wrote:
             | Yeah, I studied mathematics (algebra and number theory) and
             | zero is _the_ point, often sporting discontinuities, or
             | weird asymptotic behavior.
             | 
             | Quite a lot of algorithms use some form of division and
             | zero is the only number in our typical structures (Z, Q, R,
             | C), that cannot be used to divide with.
        
               | isoprophlex wrote:
               | Well, now in this brave new age of AI we can enjoy
               | computer programs crashing with an
               | Error: division by please upvote, share and like!
        
               | xyproto wrote:
               | This also works; I upvoted your comment.
        
               | o1bf2k25n8g5 wrote:
               | I have discovered a truly marvelous proof of how to smash
               | that like and subscribe button, which this comment box is
               | too small to contain.
        
               | msopena wrote:
               | Signed by Pierre de FermAIt
        
               | Bluestein wrote:
               | NaN
        
               | edwcross wrote:
               | In machine integer arithmetics, one must also beware
               | division by -1, which can convert MIN_INT into MIN_INT
               | with a signed overflow and violate some arithmetics
               | invariants, such as sign (negative divided by negative is
               | _usually_ positive).
        
             | KeplerBoy wrote:
             | Denormals are flushed to zero by default on most GPUs by
             | the way.
        
           | wahnfrieden wrote:
           | whisper MUST be combined with silence detection / VAD
        
             | pferde wrote:
             | Ah, the good old "you're holding it wrong".
             | 
             | What good is a speech recognition tool that literally hears
             | imaginary voices?
        
               | Xmd5a wrote:
               | faster-whisper has a min_silence_duration_ms option
        
               | wahnfrieden wrote:
               | There are much higher quality VAD solutions available
        
               | zettabomb wrote:
               | Considering that if you DO use VAD (voice activity
               | detection), it's the best open weights voice recognition
               | model by a very wide margin, it's quite good. I'd be
               | willing to be that commercial products that "don't have
               | this problem" are using VAD as well, and that this is
               | well known to them. But Whisper is just the weights, and
               | I suppose a simple reference implementation, not a full
               | product.
        
               | bmacho wrote:
               | > What good is a speech recognition tool that literally
               | hears imaginary voices?
               | 
               | Well, if it is supposed to work after silence detection,
               | then it is good for speech recognition I guess. It's like
               | blaming a wheel why is it circular, you can't sit on it.
               | It's a part of a larger machine.
        
               | dumbfounder wrote:
               | Just lay the wheel on its side and it makes a fine seat.
        
               | wahnfrieden wrote:
               | Yes, you are holding it wrong. The good of it is that it
               | does not output imaginary voices when used with VAD.
               | 
               | Show us a technology with better results that does not
               | use VAD. If you can't, then I'm not sure what you're
               | arguing against except superficialities so
               | inconsequential that I can't comprehend the
               | condescension. The results speak for itself
        
               | xandrius wrote:
               | So if a tool has a process to have it perform at its best
               | then it's a problem?
               | 
               | Do you also moan that before applying glue to a surface
               | or it won't stick? Or if you need to drill a guiding hole
               | before making a larger one in wood? Or that you need to
               | use truly prime numbers for a security key to actually be
               | safe?
        
               | nhecker wrote:
               | >imaginary voices
               | 
               | On the other hand, I can imagine that when things get
               | quiet and the signal-to-noise ratio gets close to zero,
               | random background audio (or randomness introduced in the
               | transcription model) will be enough to tickle a critical
               | number of neurons and elicit hallucinations.
               | 
               | The related thought exercise is this: Try scanning across
               | the band with an AM or sideband radio, and after a while
               | your brain will start to wonder "was that a voice I just
               | heard, or music perhaps?" when in reality it was just
               | environmental static.
        
             | DANmode wrote:
             | What's VAD?
        
               | maxbond wrote:
               | Voice Activity Detection (it predicts whether a short
               | clip contains speech, eg to mute your microphone when you
               | aren't speaking).
        
             | cmiles74 wrote:
             | If that's truly the case then they should make it part of
             | the product, IMHO.
        
               | wahnfrieden wrote:
               | How is it not the case? It is unusable without VAD or
               | editing. I don't understand what you're questioning
               | 
               | I agree their products could be better "end to end"
               | integrated. Meanwhile there is a continuously-improving
               | field of work for detecting speech (which Whisper is
               | incapable of). They offer official "cookbooks" with
               | guidance on an approach they recommend: https://cookbook.
               | openai.com/examples/whisper_processing_guid...
               | 
               | > At times, files with long silences at the beginning can
               | cause Whisper to transcribe the audio incorrectly. We'll
               | use Pydub to detect and trim the silence.
               | 
               | (Official OpenAI quote)
        
             | DANmode wrote:
             | What's a good starter VAD lib, and if you know, the best
             | implementation of something like this to use in a browser-
             | based app?
             | 
             | Say if I wanted to use it for Voice Nav, or Voice Input,
             | but not piss off random people speaking the wrong language.
        
         | ttflee wrote:
         | In Chinese, it always added something like "For study/research
         | purpose only. Please delete after 48 hours." This is what those
         | volunteers added in subtitles of (pirated) movies/shows.
        
           | cyp0633 wrote:
           | That is not the case here - I never encountered this with
           | whisper-large-v3 or similar ASR models. Part of the reason, I
           | guess, is that those subs are burnt into the movie, which
           | makes them hard to extract. Standalone subs need the
           | corresponding video resource to match the audio and text. So
           | nothing is better than YouTube videos which are already
           | aligned.
        
             | simsla wrote:
             | At least for English, those "fansubs" aren't typically
             | burnt into the movie*, but ride along in the video
             | container (MP4/MKV) as subtitle streams. They can typically
             | be extracted as SRT files (plain text with sentence level
             | timestamps).
             | 
             | *Although it used to be more common for AVI files in the
             | olden days.
        
               | ethbr1 wrote:
               | _flashbacks of trying to track down subs sync'd to a
               | specific release_
        
               | Maken wrote:
               | SRT is ancient. Nowadays everyone uses ASS subtitles
               | which can be randomly styled.
        
               | simsla wrote:
               | In general? In the past I've known ASS to be used a lot
               | for things like anime, but less for live action shows.
        
               | Maken wrote:
               | I have also found them inside mkvs as the subtitle track.
               | I think SRT was the default because most content was
               | ripped from DVD/BD, but now most of the content is from
               | streaming sources and you need to convert the subtitles
               | anyway.
        
               | conradev wrote:
               | WebVTT (a SubRip successor) is probably more widely used
               | than ASS
        
               | Maken wrote:
               | By legit providers, probably.
        
           | codedokode wrote:
           | Fair, if AI companies are allowed to download pirated content
           | for "learning", why ordinary people cannot.
        
             | snickerdoodle12 wrote:
             | There is so much damning evidence that AI companies have
             | committed absolutely shocking amounts of piracy, yet
             | nothing is being done.
             | 
             | It only highlights how the world really works. If you have
             | money you get to do whatever the fuck you want. If you're
             | just a normal person you get to spend years in jail or
             | worse.
             | 
             | Reminds me of https://www.youtube.com/watch?v=8GptobqPsvg
        
               | 4gotunameagain wrote:
               | If you owe the bank $1,000 you have a problem.
               | 
               | If you owe the bank $100,000,000 the bank has a problem.
               | 
               | We live in an era where the _president of the United
               | States_ uses his position to pump crypto scams purely for
               | personal profit.
        
               | kyleee wrote:
               | 10% for the big don
        
               | alphan0n wrote:
               | No one (in the US) has been jailed for downloading
               | copyrighted material.
        
               | snickerdoodle12 wrote:
               | https://en.wikipedia.org/wiki/Aaron_Swartz
               | 
               | And the US is not the only jurisdiction
        
               | gruez wrote:
               | That's not the same as piracy though. He wasn't
               | downloading millions of scientific papers from libgen or
               | sci-hub, he was downloading them directly from jstor.
               | Indeed, none of his charge was for copyright
               | infringement. It was for stuff like "breaking and
               | entering" and "unauthorized access to a computer
               | network".
        
               | snickerdoodle12 wrote:
               | The exact same charges could apply to the AI scrapers
               | illegitimately accessing random websites.
        
               | gruez wrote:
               | Part of the accusation comes from the fact that Swartz
               | accessed the downloads through a MIT network closet,
               | which AI companies wasn't doing. The equivalent to that
               | would be if openai broke into a wiring closet at
               | Disneyland to download Disney movies.
        
               | snickerdoodle12 wrote:
               | The CFAA is vague enough to punish unauthorized access to
               | a computer system. I don't have an example case in mind,
               | but people have gotten in trouble for scraping websites
               | before while ignoring e.g. robots.txt
        
               | gruez wrote:
               | The CFAA might be vague, but the case law on scraping
               | pretty much has been resolved to "it's pretty much legal
               | except in very limited circumstances". It's regrettable
               | that less resourced defendants were harassed before large
               | corporations were able to secure such rulings, but the
               | rulings that allowed scraping occurred before AI
               | companies' scraping was done, so it's unclear why AI
               | companies in particular should be getting flak here.
        
               | kube-system wrote:
               | I haven't seen any accusations that they've done that,
               | though. Usually people get pirated material from sources
               | that intentionally share pirated material.
        
               | snickerdoodle12 wrote:
               | They're not just training on pirated content, they've
               | also scraped literally the entire internet and used that
               | too.
        
               | kube-system wrote:
               | Scraping the public internet is also not a CFAA violation
        
               | snickerdoodle12 wrote:
               | CFAA bans accessing a protected computer without
               | authorization. Hitting URLs denied by robots.txt has been
               | argued to be just that.
        
               | dragonwriter wrote:
               | > Hitting URLs denied by robots.txt has been argued to be
               | just that.
               | 
               | "Has been argued" -- sure, but never successfully; in
               | fact, in HiQ v. LinkedIn, the 9th Circuit ruled (twice,
               | both before and on remand again after and applying the
               | Supreme Court ruling in Van Buren v. US) against _a cease
               | and desist on top of robots.txt_ to stop accessing data
               | on a public website constituting  "without authorization"
               | under the CFAA.
        
               | snickerdoodle12 wrote:
               | Now do every other jurisdiction
        
               | gruez wrote:
               | CFAA was mentioned specifically, which means only US
               | jurisdiction is relevant here.
        
               | dragonwriter wrote:
               | No, they couldn't, since the then-novel and untested
               | strained interpretation of the CFAA that the prosecutor
               | was relying on has since been tested in the courts and
               | soundly rejected.
        
               | alphan0n wrote:
               | Aaron Swartz was not jailed or even charged for copyright
               | infringement. The discussion and the comment I replied to
               | is centered around US companies and jurisdiction.
        
               | snickerdoodle12 wrote:
               | The thread is centered around US companies, but not US
               | jurisdiction.
        
               | codedokode wrote:
               | What about people filming movies in the cinema (for
               | learning of course)? [1]
               | 
               | [1]
               | https://www.thefederalcriminalattorneys.com/unauthorized-
               | rec...
        
               | codedokode wrote:
               | There could be a moral question. For example a researcher
               | might not want to download a pirated paper and cause loss
               | to a fellow researcher. But it becomes pretty stupid to
               | pay when everyone, including large reputable companies
               | endorsed by the government, is just downloading the
               | content for free. Maybe his research will help developing
               | faster chips to win against China, why should he pay?
               | 
               | Would it be a "fair use" to download pirated papers for
               | research instead of buying?
               | 
               | Also I was gradually migrating from obtaining software
               | from questionable sources to open source software,
               | thinking that this is going out of trend and nobody
               | torrents apps anymore, but it seems I was wrong?
               | 
               | Or another example: if someone wants to make
               | contributions to Wine but needs a Windows for developing
               | the patch, what would be the right choice, buy it or
               | download a free copy from questionable source?
        
               | immibis wrote:
               | Researchers don't get paid when their papers are
               | downloaded, though. They _pay_ to have their papers
               | downloaded, and the middleman makes money on both sides.
               | Piracy is the only moral option for them. There is a
               | reason every single competent professor in the western
               | world will email you a free copy of their papers if you
               | ask nicely.
        
               | shadowgovt wrote:
               | There's actually a lot of court activity on this topic,
               | but the law moves slowly and is reluctant to issue
               | injunctions where harm is not obvious.
               | 
               | It's more that the law about "one guy decides to pirate
               | twelve movies to watch them at home and share with his
               | buddies" is already well-settled, but the law about "a
               | company pirates 10,000,000 pieces to use as training data
               | for an AI model (a practice that the law already says is
               | _legal_ in an academic setting, i.e. universities do this
               | all the time and nobody bats an eye) " is more
               | complicated and requires additional trials to resolve.
               | And no, even though the right answer may be self-evident
               | to you or me, it's not settled law, and if the force of
               | law is applied poorly suddenly what the _universities_
               | are doing runs afoul of it and basically nobody wants
               | that outcome.
        
               | Workaccount2 wrote:
               | There is a distinction that must be made that very few
               | people do, but thankfully the courts seems to grasp:
               | 
               | Training on copyright is a separate claim than skirting
               | payment for copyright.
               | 
               | Which pretty much boils down to: "If they put it out
               | there for everyone to see, it's probably OK to train on
               | it, if they put it behind a paywall and you don't pay,
               | the training part doesn't matter, it's a violation."
        
               | snickerdoodle12 wrote:
               | So if I download copyrighted material like the new disney
               | movie with fansubs and watch it for training purposes
               | instead of enjoyment purposes it's fine? In that case
               | I've just been training myself, your honor. No, no, I'm
               | not enjoying these TV shows.
               | 
               | Because it's important to grasp the scale of these
               | copyright violations:
               | 
               | * They downloaded, and admitted to using, Anna's Archive:
               | Millions of books and papers, most of which are paywalled
               | but they pirated it instead
               | 
               | * They acquired Movies and TV shows and used unofficial
               | subtitles distributed by websites such as OpenSubtitles,
               | which are typically used for pirated media. Official
               | releases such as DVDs tend to have official subtitles
               | that don't sign off with "For study/research purpose
               | only. Please delete after 48 hours" or "Subtitles by
               | %some_username%"
        
               | Workaccount2 wrote:
               | I don't know what is confusing here, perhaps my comment
               | isn't clear.
               | 
               | If you skirt payment, its a violation. If it's free, but
               | still copyright, it's likely not a violation.
        
               | snickerdoodle12 wrote:
               | They've done both, so my confusion is about why you are
               | bringing this up?
        
               | Analemma_ wrote:
               | Whether it's legal slash fair use to train on copyrighted
               | material is only one of the questions currently being
               | asked though. There's a separate issue at play where
               | these companies are pirating the material for the
               | training process.
               | 
               | By comparison, someone here brought up that it might be
               | transformative fair use to write a play heavily based on
               | Blood Meridian, but you still need to buy a copy of the
               | book. It would still be infringement to pirate the e-book
               | for your writing process, even if the end result was
               | legal.
        
               | codedokode wrote:
               | If they would buy material at a large scale, the seller
               | might require them to sign a contract that requires
               | royalty if the material is used for training an AI. So
               | buying legally is a way to put yourself into a trap.
        
               | shadowgovt wrote:
               | What is the precedent on that kind of agreement?
               | 
               | The only thing I've been able to find is the note that
               | since copyright is federal law, state contract law
               | actually _can 't_ supersede it, to wit: if you try to put
               | a clause in the contract that says the contract is void
               | if I use your work to make transformative fair-use works
               | (or I owe you a fee), that clause is functionally
               | unenforceable (for the same reason that I don't owe you a
               | fee if I make transformative fair-use works of your
               | creations in general).
        
               | brookst wrote:
               | They can buy individual works like anyone else.
               | 
               | Or they can negotiate a deal at scale with whatever price
               | / restrictions make sense to both parties.
               | 
               | I don't see a way they could be "trapped". Worst case
               | they pay retail price.
        
               | CamperBob2 wrote:
               | No, if you revolutionize both the practice and philosophy
               | of computing and advance mankind to the next stage of its
               | own intellectual evolution, you get to do whatever the
               | fuck you want.
               | 
               | Seems fair.
        
               | recursive wrote:
               | Hm. Not a given that it's an advance.
        
               | Nevermark wrote:
               | I get the common cynical response to new tech, and the
               | reasons for it.
               | 
               | We wish we lived in a world where change was reliably
               | positive for our lives. Often changes are sold that way,
               | but they rarely are.
               | 
               | But when new things introduce dramatic capabilities that
               | former things couldn't match (every chatbot before LLMs),
               | it is as clear of an objective technological _advance_ as
               | has ever happened.
               | 
               | --
               | 
               | Not every technical advance reliably or immediately makes
               | society better.
               | 
               | But whether or when technology improves the human
               | condition is far more likely to be a function of human
               | choices than the bare technology. Outcomes are strongly
               | dependent on the trajectories of who has a technology,
               | when they do, and how they use it. And what would be the
               | realistic (not wished for) outcome of not having or using
               | it.
               | 
               | For instance, even something as corrosive as social
               | media, as it is today, could have existed in strongly
               | constructive forms instead. If society viewed private
               | surveillance, unpermissioned collation across third
               | parties, and weaponizing of dossiers via personalized
               | manipulation of media, increased ad impact and addictive-
               | type responses, as ALL being violations of human rights
               | to privacy and freedom from coercion or manipulation. And
               | worth legally banning.
               | 
               | Ergo, if we want tech to more reliably improve lives, we
               | need to ban obviously perverse human/corporate behaviors
               | and conflicts of interest.
               | 
               | (Not just shade tech. Which despite being a pervasive
               | response, doesn't seem to improve anything.)
        
               | CamperBob2 wrote:
               | At the risk of stepping on a well-known land mine around
               | here, how'd _you_ do on the IMO problem set this year?
        
               | recursive wrote:
               | I didn't participate. I probably wouldn't have done well.
               | I disagree with your framing.
        
               | CamperBob2 wrote:
               | Well, wait, if somebody writes a computer program that
               | answers 5 of 6 IMO questions/proofs correctly, and you
               | don't consider it an "advance," what _would_ qualify?
               | 
               | Either both AI teams cheated, in which case there's
               | nothing to worry about, or they didn't, in which case
               | you've set a pretty high bar. Where is that bar, exactly?
               | What exactly does it take to justify blowing off
               | copyright law in the larger interest of progress? (I have
               | my own answers to that question, including equitable
               | access to the resulting models regardless of how
               | impressive their performance might be, but am curious to
               | hear yours.)
        
               | recursive wrote:
               | The technology is capable in a way that never existed
               | before. We haven't yet begun to see the impacts of that.
               | I don't think it will be a good for humanity.
               | 
               | Social networks as they exist today represent technology
               | that didn't exist decades ago. I wouldn't call it an
               | "advancement" though. I think social media is terrible
               | for humans in aggregate.
        
               | immibis wrote:
               | I notice you've motte-and-baileyed from "revolutionize
               | both the practice and philosophy of computing and advance
               | mankind to the next stage of its own intellectual
               | evolution" to simply "is considered an 'advance'".
        
               | verandaguy wrote:
               | Except that the jury's (at best) still out on whether the
               | influence of LLMs and similarly tech on knowledge workers
               | is actually a net good, since it might stunt our ability
               | to critically think and problem solve while confidently
               | spewing hallucinations at random while model alignment is
               | unregulated, haphazard, and (again at best) more of an
               | art than a science.
        
               | CamperBob2 wrote:
               | Well, if it's no big deal, you and the other copyright
               | maximalists who have popped out of the woodwork lately
               | have nothing to worry about, at least in the long run.
               | Right?
        
               | NoMoreNicksLeft wrote:
               | The dead corpses of filmmakers and authors and actors are
               | buried in unmarked graves out behind those companies'
               | corporate headquarters. Unimaginable horror, that piracy.
               | Why has no one intervened?
               | 
               | >If you're just a normal person you get to spend years in
               | jail or worse.
               | 
               | Not that I'm a big fan of the criminalization of
               | copyright infringement in the United States, but who has
               | ever spent years in jail for this?
               | 
               | Besides, if it really bothered you, then we might not see
               | this weird tone-switch from one sentence to the next,
               | where you seem to think that piracy is shocking and
               | "something should be done" and then "it's not good tht
               | someone should spend time in jail for it". What gives?
        
               | snickerdoodle12 wrote:
               | > Besides, if it really bothered you, then we might not
               | see this weird tone-switch from one sentence to the next,
               | where you seem to think that piracy is shocking and
               | "something should be done" and then "it's not good tht
               | someone should spend time in jail for it". What gives?
               | 
               | What a weirdly condescending way to interpret my post. My
               | point boils down to: Either prosecute copyright
               | infringement or don't. The current status quo of
               | individuals getting their lives ruined while companies
               | get to make billions is disgusting.
        
               | brookst wrote:
               | > Either prosecute copyright infringement or don't
               | 
               | This is the absolute core of the issue. Technical people
               | see law as code, where context can be disregarded and all
               | that matters is specifying the outputs for a given set of
               | inputs.
               | 
               | But law doesn't work that way, and it should not work
               | that way. Context matters, and it needs to.
               | 
               | If you go down the road of "the law is the law and
               | billion dollar companies working on product should be
               | treated the same as individual consumers", it follows
               | that individuals should do SEC filings ("either require
               | 10q's or don't!"), and surgeons should be jailed ("either
               | prosecute cutting people with knives or don't!").
               | 
               | There is a lot to dislike about AI companies, and while I
               | believe that training models is transformative, I don't
               | believe that maintaining libraries of pirated content is
               | OK just because it's an ingredient to training.
               | 
               | But insisting that individual piracy to enjoy
               | entertainment without paying must be treated exactly the
               | same as datasets for model training is the absolute
               | weakest possible argument here. The law is not that
               | reductive.
        
               | snickerdoodle12 wrote:
               | Pretty funny that your argument boils down to: It's okay
               | to break the law if you do it as a company.
               | 
               | Copyright laws target everyone. SEC laws don't.
        
               | HWR_14 wrote:
               | It doesn't matter whether it's transformative. Copyright
               | covers derivative works.
        
               | wormius wrote:
               | > But law doesn't work that way, and it should not work
               | that way. Context matters, and it needs to.
               | 
               | As Anatole France famously quipped:
               | 
               | "The law, in its majestic equality, forbids the rich and
               | poor alike to sleep under bridges, to beg in the streets,
               | and to steal bread."
        
             | shadowgovt wrote:
             | IANAL, but reading a bit on this topic: the relevant part
             | of the copyright law for AI isn't academia, it's
             | transformative work. The AI created by training on
             | copyrighted material transforms the material so much that
             | it is no longer the original protected work (collage and
             | sampling are the analogous transformations in the visual-
             | arts and music industries).
             | 
             | As for actually gathering the copyrighted material: I
             | believe the jury hasn't even been empaneled for that yet
             | (in the OpenAI case), but the latest ruling from the court
             | is that copyright _may_ have been violated in the creation
             | of their training corpus.
        
             | gruez wrote:
             | >why ordinary people cannot
             | 
             | They can. I don't think anyone got prosecuted for using an
             | illegal streaming site or downloading from sci-hub, for
             | instance. What people do get sued for is seeding, which
             | counts as distribution. If anything AI companies are
             | getting prosecuted more aggressively than "ordinary
             | people", presumably because of their scale. In a recent
             | lawsuit Anthropic won on the part about AI training on
             | books, but lost on the part where they used pirated books.
        
               | codedokode wrote:
               | People got in trouble for filming in the cinema as I
               | understand, there is a separate law for that.
        
               | gruez wrote:
               | But in that case even though filming isn't technically
               | distribution, it's clearly a step to distributing copies?
               | To take this to the extreme, suppose you ripped a blu-
               | ray, made a thousand copies, but haven't packaged or sold
               | them yet. If the FBI busted in, you'd probably be
               | prosecuted for "conspiracy to commit copyright
               | infringement" at the very least.
        
               | snickerdoodle12 wrote:
               | It's just "training"
        
               | gruez wrote:
               | You seem to equate "training" (with scare quotes) with
               | someone actually pirating a blu-ray, but they really
               | aren't equivalent. Courts so far have ruled that training
               | is fair use and it's not hard to see why. Unlike copying
               | a movie almost verbatim (as with ripping a blu-ray), AI
               | companies are actually producing something transformative
               | in the form of AI models. You don't have to like AI
               | models, or the AI companies' business models, but it
               | strains credulity to pretend ripping a blu-ray is somehow
               | equivalent to training an AI model.
        
               | snickerdoodle12 wrote:
               | Who's to say why I downloaded and am now watching a
               | movie? Is it for my enjoyment? Is it because I'm training
               | my brain? How is me training my brain any different from
               | companies training their LLMs?
               | 
               | Same goes for recording: I'm just training my skills of
               | recording. Or maybe I'm just recording it so I can
               | rewatch it later, for training purposes, of course.
        
               | gruez wrote:
               | >Who's to say why I downloaded and am now watching a
               | movie? Is it for my enjoyment? Is it because I'm training
               | my brain? How is me training my brain any different from
               | companies training their LLMs?
               | 
               | None of this is relevant because Anthropic was only left
               | off the hook for training, and not for pirating the books
               | itself. So far as the court cases are playing out, there
               | doesn't appear to be a special piracy exemption for AI
               | companies.
               | 
               | >Same goes for recording: I'm just training my skills of
               | recording. Or maybe I'm just recording it so I can
               | rewatch it later, for training purposes, of course.
               | 
               | You can certainly use that as a defense. That's why we
               | have judges, otherwise there's going to be some smartass
               | caught with 1KG of coke and claiming it's for "personal
               | consumption" rather than distribution.
               | 
               | None of this matters in reality, though. If you're caught
               | with AV gear in a movie theater once, you'd likely be
               | ejected and banned from the establishment/chain, not have
               | the FBI/MPAA go after you for piracy. If you come again,
               | you'd likely be prosecuted for trespassing. In the cases
               | where they're going after someone in particular for
               | making these rips, they usually have a dossier of
               | evidence, like surveillance/transaction history showing
               | that the same individual has been repeatedly recording
               | movies, and watermarks correlating the screenings that
               | the person has been in to files showing up on torrent
               | sites.
        
               | snickerdoodle12 wrote:
               | > If you're caught with AV gear in a movie theater once,
               | you'd likely be ejected and banned from the
               | establishment/chain, not have the FBI/MPAA go after you
               | for piracy
               | 
               | Good example, because this is exactly what websites are
               | doing with LLM companies, who are doing their damnest to
               | evade the blocks. Which brings us back around to
               | "trespassing" or the CFAA or whatever.
        
               | gruez wrote:
               | >Which brings us back around to "trespassing" or the CFAA
               | or whatever.
               | 
               | That argument is pretty much dead after
               | https://en.wikipedia.org/wiki/Van_Buren_v._United_States
               | and https://en.wikipedia.org/wiki/HiQ_Labs_v._LinkedIn
        
               | snickerdoodle12 wrote:
               | https://www.rvo.nl/onderwerpen/octrooien-ofwel-
               | patenten/vorm...
               | 
               | I'll leave all other jurisdictions up to you.
        
             | 0x457 wrote:
             | Well, it just shows that they've downloaded subtitles.
        
           | kgeist wrote:
           | Interesting, in Russian, it often ends with "Subtitles by
           | %some_username%"
        
         | xigoi wrote:
         | Pray, Mr. Babbage, if you put into the machine wrong figures,
         | will the right answers come out?
        
           | madcaptenor wrote:
           | I am not able rightly to apprehend the kind of confusion of
           | ideas that could provoke such a question.
        
             | Workaccount2 wrote:
             | Having zero exposure to any form of computation for your
             | entire life, as the vast majority of people in the early
             | 19th century were.
        
             | immibis wrote:
             | I can. He was asking if Babbage was cheating.
             | 
             | You put in 2+2 - the right figures. The machine says 4 -
             | the right answer. If you put in the wrong figures, like
             | 3+3, will the machine still say 4? It's easy to make a
             | machine that always says 4.
             | 
             | The people who asked him that question, however, probably
             | got a different scam demonstrated to them every every.
             | Remember the Mechanical Turk? Babbage's reply paints him
             | very honestly. It shows that he couldn't even conceive that
             | someone might try to trick the royal court (or whoever it
             | was) into accepting a fake device.
        
         | mmcwilliams wrote:
         | Similar in the English model. Pretty clear they trained on
         | YouTube videos where creators will put that in otherwise silent
         | sections to ensure it shows up for people with CC on.
        
           | probably_wrong wrote:
           | The number one hallucination in my transcriptions was
           | "Subtitles by the Amara.org community".
        
         | indrora wrote:
         | When YouTube began building automatic transcriptions for
         | captions, it regularly flagged any noise or music -- typically
         | industrial noise -- with "[foreign]"
         | 
         | If it couldn't understand it, it was "foreign" for the longest
         | time.
        
           | stndef wrote:
           | Yeah, I can confirm seeing that a fair bit specifically
           | during non-verbal parts of videos when someone is using a
           | tool.
        
             | TurkTurkleton wrote:
             | Can confirm as well, although to my recollection it just
             | shows up as if it's a word the transcription model heard,
             | not "[foreign]" in brackets like with "[Music]" or
             | "[Applause]". It's especially weird to me because I recall
             | the auto-transcriptions being _reasonably_ serviceable when
             | they first rolled them out, only to degrade over time to
             | the point where it was hallucinating the word  "foreign"
             | and dropping letters from words or using weird
             | abbreviations (like "koby" for "kilobyte", "TBTE" for
             | "terabyte", or, most memorably weirdly, transcribing the
             | phrase "nanosecond-by-nanosecond" as "nond by nanc") if it
             | didn't decide it heard another one entirely.
             | 
             | I also noticed a couple of months ago that YouTube seems to
             | have quietly rolled out a new auto-transcription model that
             | can make reasonable guesses at where capitalization,
             | punctuation, and sentence boundaries should go. It seems to
             | have degraded even more rapidly than the old one, falling
             | victim to the same kinds of transcription errors. Although
             | the new one has a different hallucination in silence and
             | noise that it wasn't able to classify (which, incidentally,
             | its ability to recognize things like music and applause
             | seems worse than the old one's): where the old model would
             | have hallucinated the word "foreign", the new one thinks
             | it's hearing the word "heat", often repeated ("Heat.
             | Heat.").
        
           | the_af wrote:
           | Hey, Netflix occasionally still puts in its English subtitles
           | "[foreign music]", it always cracks me up.
        
             | 0x457 wrote:
             | [speaks japanese]
             | 
             | To be fair, there is a difference between when subtitles
             | match the source language and when they don't. Former are
             | often verbatim.
        
               | the_af wrote:
               | Haha, yes, it's fair when English subtitles write
               | something like [speaks Japanese], especially when at
               | least one of the characters is not supposed to understand
               | what's being said (when they do, it's more appropriate to
               | write "[in Japanese]: let's go shopping!").
               | 
               | Netflix sometimes takes the cake with what I consider the
               | most outrageous option: writing "[in English]" when they
               | mean "in whatever language the protagonist considers
               | native", which is mind-bogglingly wrong and hilarious at
               | the same time.
               | 
               | They do this with the English subtitles of the German
               | production "Die Kaiserin" ("The Empress"): whenever Sisi
               | is speaking in another language, say French, the
               | subtitles will say "[in French] I love you...", and when
               | she switches back to German they will say "[in English] I
               | love you...". WTF, Netflix? Note this is unrelated to
               | understanding German; it's mostly Netflix looking down on
               | its customers and assuming they cannot comprehend there
               | are people in the world for whom their native tongue is
               | different to the viewer's native tongue.
               | 
               | This has happened in more shows, enough to know it's not
               | a fluke, though Netflix is inconsistent about it.
        
         | tonyhart7 wrote:
         | lmao
        
         | philipwhiuk wrote:
         | > I suspect they trained the model on some random YouTube video
         | without carefully picking really useful data.
         | 
         | They trained the model on _every_ YouTube video they could, and
         | hoped the aggregate was useful data.
        
         | danirod wrote:
         | This is totally happening with other models too, at least with
         | Spanish. Many transcriptions will end with something that
         | roughly translates to "Thanks for watching!" even if it's never
         | present in the original audio.
        
         | PhasmaFelis wrote:
         | This reminds me, some years ago as Google was expanding its
         | translation service, someone tried translating text into and
         | out of an obscure African language (don't recall which) and it
         | always came out as weird Biblical-sounding semi-gibberish.
         | 
         | My revelation was that machine translation needs a corpus of
         | bilingual documents to learn from, and if the language is
         | sufficiently obscure, there may not be any bilingual documents
         | except for the Bible, which missionaries have translated into
         | just about every language on Earth.
        
         | horseradish7k wrote:
         | oh yeah this happens a lot on reddit on videos in foreign
         | languages
        
       | sandspar wrote:
       | Neat, we finally know the answer! What is the sound of one hand
       | clapping? Translation by Nancy Qunqar.
        
         | layer8 wrote:
         | I can clap with one hand (fingers on palm) and it produces a
         | clapping sound.
        
           | boomlinde wrote:
           | Your brain merely hallucinates a clapping sound as
           | "Translation by Nancy Qunqar" enters your ears.
        
             | layer8 wrote:
             | Good one. :)
        
             | Dilettante_ wrote:
             | I can see the fnords!
        
       | kristjank wrote:
       | roses are red
       | 
       | violets are blue
       | 
       | unregistered hypercam 2
        
         | maxbond wrote:
         | Roses are red,
         | 
         | Silence is golden,
         | 
         | Translated by Nancy,
         | 
         | To copyright, we aren't beholden
        
       | arnejenssen wrote:
       | Who is Nicolai Winther? https://medium.com/@lehandreassen/who-is-
       | nicolai-winther-985...
        
         | io84 wrote:
         | "In the future, everyone will be world-famous for 15 minutes"
         | _in a microniche techno-linguistic community, at a time and
         | choosing of the swirling AI clouds_
        
       | haiku2077 wrote:
       | I've noticed this also happens in english Whisper models with the
       | phrases:
       | 
       | "[ sub by sk cn2 ]"
       | 
       | or
       | 
       | "Anyways, thanks for watching! Please subscribe and like! Thanks
       | for watching! Bye!"
       | 
       | or
       | 
       | "This is the end of the video. Thank you for watching. If you
       | enjoyed this video, please subscribe to the channel. Thank you."
        
         | OSDeveloper wrote:
         | Because they train on pirated media and or youtube videos, good
         | method, until you get slop, or get caught
        
       | tarikozket wrote:
       | this happens in Turkish too. I believe the reason is that the
       | movie subtitles were used for training without cleaning up the
       | comments / intros subtitle authors leave in them.
       | 
       | leaving personal comments, jokes, reactions, intros in subtitles
       | is very common in eastern cultures.
       | 
       | Turkish readers will probably remember "esekadam iyi seyirler
       | diler" :)
        
         | jdiff wrote:
         | Kind of mindblowing considering who it is we're talking about.
         | Of all companies, OpenAI couldn't be bothered to throw an LLM
         | at this problem? Finding amorphously phrased but clearly
         | recognizable needles in large numbers of haystacks seems like a
         | patently perfect task for them.
        
           | sofixa wrote:
           | Don't even need an LLM, a regex would have sufficed (I've
           | used my fair share of community sourced subtitles, and
           | comments are almost always in a different font, colour,
           | between brackets, etc etc).
        
         | verzali wrote:
         | That name translates as "Donkey Man" btw :D
        
       | flexagoon wrote:
       | In Russian it often hallucinates "Subtitry sdelal DimaTorzok"
       | ("Subtitles by DimaTorzok") at the end of things. Interestingly,
       | I wasn't able to find any YouTube videos with that name in the
       | subtitles, so it's not like it's in a lot of training data.
        
         | berkes wrote:
         | Could it be someone distributing subs online, e.g. showing up
         | in the opensubtitles.org dataset?
        
           | voidUpdate wrote:
           | Or possibly someone subtitling pirated movies? That seems to
           | be a common thing according to other comments
        
         | codedokode wrote:
         | I tried googling this and found questions from Telegram users
         | why voice messages recognition sometimes produces this phrase
         | and who is this person. Also I found this thread [1] claiming
         | that the subtitles by DimaTorzok are coming from some Russian
         | youtube videos on gaming like [2].
         | 
         | [1] https://github.com/openai/whisper/discussions/2372
         | 
         | [2] https://www.youtube.com/watch?v=FAqyUuahMlc&t=401s
        
           | flexagoon wrote:
           | Yeah, I know about this from Telegram, because they use
           | Whisper for voice message recognition. There are a bunch of
           | other artifacts it often produces.
        
       | xg15 wrote:
       | Yeah, the subtitle "credits" occur very frequently. I found with
       | whisper-2, they're also triggered by music.
       | 
       | I suppose the cause is the same, generally subtitle creators
       | adding all kinds of stuff during the credits that is NOT a
       | transcript.
       | 
       | Seems to me it could have been filtered out relatively easily
       | during training, by clipping the first and last few minutes of
       | all audios. But I guess that's just in hindsight.
       | 
       | Whisper also likes to transcribe cut off speech or unintelligible
       | noise as "Thank you". I have no idea where that is coming from,
       | but I guess it's a very polite model...
        
       | nottorp wrote:
       | Title should be changed to "OpenAI publishes evidence they
       | trained on pirated movies".
        
         | berkes wrote:
         | How is this evidence of that fact? Honest question.
         | 
         | I can see how this might show that subtitles from online sub
         | communities are used, or that maybe even original subtitles
         | from e.g. DVDs are used. But isn't it already known and
         | admitted (and allowed?) that AI uses all sorts of copyrighted
         | material to train models?
        
           | nemomarx wrote:
           | The Chinese subtitles for silence use a common mark for
           | pirated media in that language, according to other commentors
           | here. In general it's pretty likely that if you're finding
           | non professional subtitles they were distributed with pirated
           | media in some form, that's where you get the most fan subs
           | after all
        
           | 0points wrote:
           | > I can see how this might show that subtitles from online
           | sub communities are used, or that maybe even original
           | subtitles from e.g. DVDs are used.
           | 
           | Indeed, the captioning is copyrighted work and you are not
           | legally allowed to copy and redistribute it.
           | 
           | > But isn't it already known and admitted (and allowed?)
           | 
           | No, and I don't see where you got that from. Meta [1], OpenAI
           | [2] and everybody else is being sued as we speak.
           | 
           | 1: https://petapixel.com/2025/01/10/lawsuit-alleges-mark-
           | zucker...
           | 
           | 2: https://www.reuters.com/legal/litigation/openai-hit-with-
           | new...
        
             | skeezyboy wrote:
             | > Indeed, the captioning is copyrighted work and you are
             | not legally allowed to copy and redistribute it. Unless you
             | qualify for one of the many exceptions, such as fair use
        
               | kranke155 wrote:
               | It's not clear that training is fair use. That's being
               | contested in court I think.
        
               | whamlastxmas wrote:
               | Training isn't recreating or distributing so copyright
               | won't apply if the ruling is actually consistent with the
               | intention of the law, which it may not.
               | 
               | Using copyrighted materials and then meaningfully
               | transforming it isn't infringement. LLMs only recreate
               | original work in the same way I am when I wrote the first
               | sentence of this paragraph because it probably exists
               | word for word somewhere else too
        
             | lcnPylGDnU4H9OF wrote:
             | > I don't see where you got that from
             | 
             | It's been determined by the judge in the Meta case that
             | training on the material is fair use. The suit in that case
             | is ongoing to determine the extent of the copyright damages
             | from downloading the material. I would not be surprised if
             | there is an appeal to the fair use ruling but that hasn't
             | happened yet, as far as I know. Just saying that there is
             | good reason for them to think it's been allowed because it
             | kind of has; that can be reversed but it happened.
        
               | 0points wrote:
               | That was specifically involving 13 authors.
               | 
               | There hasn't been any trials yet about the millions of
               | copyrighted books, movies and other content they
               | evidently used.
        
               | lcnPylGDnU4H9OF wrote:
               | There's no reason to think those cases will go any
               | differently. As far as I know, the ruling would have to
               | be appealed at this point. I am only commenting to say
               | that there is reason to think this is true:
               | 
               | > But isn't it already known and admitted (and allowed?)
               | 
               | You seemed to be confused about why this person believed
               | that:
               | 
               | > No, and I don't see where you got that from.
               | 
               | And I wrote a comment intended to dispel your confusion.
               | The above commenter thought that it was allowed because a
               | judge said it was allowed; that can be appealed but
               | that's the reason someone thinks it's allowed.
        
               | dragonwriter wrote:
               | > There's no reason to think those cases will go any
               | differently. As far as I know, the ruling would have to
               | be appealed at this point.
               | 
               | Trial court rulings aren't binding precedent even on the
               | same court in different cases, so its quite possible that
               | different cases at the trial level can reach different
               | conclusions on fair use on fairly similar facts, given
               | the lack of appellate precedent directly on point with AI
               | training.
        
               | 0points wrote:
               | Yea, no. I don't think I am confused.
               | 
               | A single verdict about a specific case (13 authors vs
               | META) does not mean it's legal for companies to steal IP
               | from other companies which has evidently been going on
               | for some years now.
               | 
               | Those other companies have lawyers powerful enough to
               | change jurisdiction in many countries in order to
               | "protect their IP".
        
           | jcranmer wrote:
           | > How is this evidence of that fact?
           | 
           | The contention is that the specific translated text appears
           | largely from illegal translations (i.e., fansubs) and not
           | from authorized translations. And from a legal perspective,
           | that would basically mean there's no way they could legally
           | have appropriated that material.
           | 
           | > But isn't it already known and admitted (and allowed?) that
           | AI uses all sorts of copyrighted material to train models?
           | 
           | Technically, everything is copyrighted. But your question is
           | really about permission. Some of the known corpuses for AI
           | training include known pirate materials (e.g., libgen), but
           | it's not known whether or not the AI companies are filtering
           | out those materials from training. There's a large clutch of
           | cases ongoing right now about whether or not AI training is
           | fair use or not, and the ones that have resolved at this
           | point have done so on technical grounds rather than answering
           | the question at stake.
        
         | pjc50 wrote:
         | Of course. Piracy is legal when you have a bigger pile of money
         | than the studios.
        
           | codedokode wrote:
           | Let's not forget that some of real pirates (for example,
           | corsairs) also were legal and performed legitimate pirate
           | activities to ships of foreign countries.
        
             | jowea wrote:
             | I suspect privateers would have been offended at being
             | called pirates, but is this what is going on? If it
             | specifically a Chinese AI company pirating Hollywood for
             | example sure, but it seems it's more of a everyone firing
             | at everyone situation.
        
             | gruez wrote:
             | >Let's not forget that some of real pirates (for example,
             | corsairs) also were legal and performed legitimate pirate
             | activities to ships of foreign countries.
             | 
             | In other words there are activities that are legal or not
             | depending on whether you have authorization from the state.
             | That describes many things. For instance you synthesize
             | meth without a license from the DEA/FDA, you're a "drug
             | cartel" or whatever. But if you do it with a license you're
             | a "pharmaceutical company", and you're not making "meth",
             | you're making "desoxyn".
        
               | codedokode wrote:
               | Synthesizing chemical substances doesn't involve
               | murdering people though.
        
               | gruez wrote:
               | Families of fentanyl overdose victims would disagree.
               | Moreover it's not hard to find examples of "legal if the
               | government authorizes it" for killings. Cops and
               | soldiers, for instance.
        
               | codedokode wrote:
               | It is a bad comparison because substance vendor doesn't
               | kill anyone just as a gun store doesn't. Soldiers are
               | better analogy though.
        
               | immibis wrote:
               | It takes quite a leap of logic to blame someone ingesting
               | a toxic quantity of a substance on the person who
               | manufactured the substance. When someone drinks bleach do
               | we blame the company that makes the bleach?
        
           | onlyrealcuzzo wrote:
           | Isn't Piracy legal in many parts of the world?
           | 
           | Legally, why wouldn't they be able to do the piracy parts in
           | one of those jurisdictions and then ship the outputs back to
           | the mothership?
        
         | Hnrobert42 wrote:
         | HN is pretty strict about not editorializing titles. Even if
         | you statement was unequivocably correct, the post would get
         | flagged.
        
       | VMG wrote:
       | related: googles song detection alg detects my phone vibrating as
       | the song "Montagem Dilatacao Hipnotica"
        
       | 1718627440 wrote:
       | Well, I fail to see how the LLM is in the wrong here. Surely if a
       | sufficiently large part of the training data comes from a single
       | source, it is correct to credit them for the output.
        
       | bilekas wrote:
       | This is a nice reminder that there is no real reasoning in the
       | "AI" it is just still guessing the next word. After being trained
       | on subtitle files which I guess is actually a clever idea as they
       | convey real conversations without pirating, subtitles are freely
       | distributed after all by dedicated translators. Good to see
       | they're the ones getting credit though!
        
       | terribleperson wrote:
       | Using Whisper to sub Japanese vtuber concerts for my enjoyment,
       | I've noticed a similar trend. Not one specific phrase, but
       | several. Some are strange ("I'm going to make a hole in the back
       | of the head"), some are clearly from lyrics websites.
        
       | theanirudh wrote:
       | In English, silence is transcribed to "Please like and subscribe"
        
         | cheschire wrote:
         | I get thanks for watching a lot when using speech to text on
         | ChatGPT
        
       | michalpleban wrote:
       | The same things happen on Dutch as well, it brings up some kind
       | of radio channel name if I recall correctly.
        
       | ninetyninenine wrote:
       | https://www.instagram.com/nancyrk/?hl=en
        
       | johtso wrote:
       | I get the same with Welsh, when having some network issues in
       | voice chat it hallucinated me saying "Diolch yn fawr am wylio'r
       | fideo." which translates as "Thank you very much for watching the
       | video."
        
       | tacone wrote:
       | Same happened to me with English: I've got "Thanks for watching"
       | many times.
        
         | withinboredom wrote:
         | Super annoying when it happens with voice chat -- it'll just be
         | explaining something and suddenly stop to say "you're welcome!
         | Feel free to come back any time you want to chat" and that
         | conversation is done.
        
       | majke wrote:
       | I've spent some time with whisper, and indeed this happens all
       | the time. To my untrained eye it seems like:
       | 
       | - they indeed seem to have trained on movies/subtitles
       | 
       | - you absolutely positively must use Voice Activity Detection
       | (VAD) in front of whisper
        
       | 0points wrote:
       | Interesting! I used whipser last year to attempt to build an
       | audio transcription tool but gave up due to excessive amount of
       | hallucinated output no matter what model I used.
       | 
       | It would produce seemingly ok output until you started paying
       | attention.
       | 
       | One example, it insisted that Biggie Smalls sings "Puttin five
       | carrots in my baby girl ear". (its "carats").
       | 
       | It's apparently not useful in transcription as it don't reason
       | [sic].
        
         | IshKebab wrote:
         | That example is not hallucination, it's just a homonym with
         | insufficiently clear context for the model to disambiguate it.
        
           | 0points wrote:
           | I'm well aware mishearing "carots" as "carrots" is not a
           | hallucination.
           | 
           | That's an example I gave after having used Whisper, the topic
           | of discussion.
        
             | dpoloncsak wrote:
             | An example of what you claimed was a hallucination
        
       | jacobgorm wrote:
       | In Danish I get credits to a known subtitler.
        
       | Oras wrote:
       | Searching Google for older posts, found many DailyMotion links
       | for translated movies in Arabic with "trjm@ nnsy qnqr".
       | 
       | I suspected as others mentioned, these were extracted from
       | torrents movies.
        
         | undersuit wrote:
         | All my Google searches for Oracle support pages have been
         | labelled with 'lmwrd lbshry@ wltnmy@ ljtm`y@ ' which translates
         | to 'Human Resources and Social Development' for a few months
         | now. Wonder how much this is related.
        
       | arblor wrote:
       | Just to add some trivia: ChatGpt interprets(/ed) silence as
       | "Sottotitoli e Revisione a cura di QTSS". Now many videos (mainly
       | dailymotion) with autogenerated subtitles have their Transcripts
       | full of the same message
       | 
       | i.e. https://www.dailymotion.com/video/x9g9d6u
        
       | h1fra wrote:
       | Time to weapon this: publish thousands of videos, add a referal
       | link or AI instructions in subtitles when there is a silent
       | section, ???, profit
        
       | blindstitch wrote:
       | The fork that I've been using, WhisperX, seems to do better. I've
       | used it on clean splits of mic tracks (ie total silence when the
       | other is talking) with far fewer hallucinations.
        
         | ethan_smith wrote:
         | WhisperX works better because it implements a robust VAD (Voice
         | Activity Detection) preprocessing step that effectively filters
         | out silence segments before they reach the model, preventing
         | the hallucination triggers entirely.
        
       | dipierro wrote:
       | Subtitry sdelal Dima Torzhok
        
       | DonHopkins wrote:
       | More like reminiscing than hallucinating.
        
       | jovial_cavalier wrote:
       | Looks like it's some random user who has generated some
       | lyricslyrics translations between Arabic and English. It's
       | strange, they don't seem to have many contributions. I would have
       | imagined them to be more prolific.
       | 
       | https://lyricstranslate.com/en/translator/nancy-qunqar
        
       | flkiwi wrote:
       | > [In English] it also happens a lot with hallucinations saying
       | stuff like "This is the end of the video, remember to like and
       | subscribe
       | 
       | Well now I know how I'm going to start filling awkward silences
       | in meetings.
        
       | kranner wrote:
       | I've run lots of guided meditations through whisper-large-v3 and
       | anything with long periods of silence gets a "(c) Mooji Media"
       | line added at the end of the transcript. None of these have
       | actually been from Mooji.
        
       | boredumb wrote:
       | Since it says "Translated by Nancy Qanqar" i'd be willing to bet
       | they're training on some audiobooks with a transcript and
       | somewhere in there it consistently has "Translated by Nancy
       | Qanqar" in the transcript where there is dead air in the
       | audiobook.
        
       | shinycode wrote:
       | In French it's << Sous-titrage Societe Radio-Canada >>
        
       | dangus wrote:
       | Hey guys, AI by 2027 is going to be superhuman AGI Agentic mega-
       | intelligence, you better fire all your employees and get ready
       | for AI to take your job and embrace your spouse at a Coldplay
       | concert.
       | 
       | Big data. Machine learning. Blockchain. Artificial intelligence.
       | Digital manufacturing. Big data analysis. Quantum communication
       | and...Internet of things.
       | 
       | This time the hype cycle won't be a massive exaggerated
       | disappointment, for real this time.
        
       | shadowgovt wrote:
       | Interesting. This is similar to the Google Translate bug where it
       | would translate lorem ipsum as bits of political text (because it
       | found most of its lorem ipsum examples flipping between languages
       | on sites where one language was a news story but the not-yet-
       | translated languages would output a lorem-ipsum file instead of a
       | 404 when you toggled over to them).
        
       | chris_wot wrote:
       | My wife, who speaks and reads Arabic, got a real kick out of
       | this.
       | 
       | But honestly, this is the AI equivalent of "please send for
       | translating" in Welsh on a Welsh street sign.
       | 
       | https://www.theguardian.com/theguardian/2008/nov/01/5
        
       | abdussamadbello wrote:
       | You can either fine-tune the model or filter the response from
       | whisper
       | 
       | ``` text = "helo helo hello ." target_phrase = "trjm@ nnsy qnqr"
       | replacement = ""
       | 
       | updated_text = text. Replace(target_phrase, replacement)
       | 
       | print(updated_text) ```
        
       | Dwedit wrote:
       | I've seen Japanese translation models that translate empty string
       | "" into "I'm Sorry I'm Sorry"
        
       | PatchworkCasino wrote:
       | I've found if the first 30 seconds of a recorded phone call is
       | ringing and/or DTMF (almost always happens if you call a
       | business) the system with either select Nynorsk or Welsh as the
       | language. Never bothered to check what the text translated to but
       | it's probably something similar. Not a practical issue for me but
       | I can see it being a pain for any bilingual business or call
       | center.
        
       ___________________________________________________________________
       (page generated 2025-07-22 23:01 UTC)