[HN Gopher] Anti-Piracy Group Takes AI Training Dataset 'Books3'...
       ___________________________________________________________________
        
       Anti-Piracy Group Takes AI Training Dataset 'Books3' Offline
        
       Author : YeGoblynQueenne
       Score  : 95 points
       Date   : 2023-10-03 12:55 UTC (10 hours ago)
        
 (HTM) web link (gizmodo.com)
 (TXT) w3m dump (gizmodo.com)
        
       | nektro wrote:
       | Anti-Piracy: bad
       | 
       | Group Takes AI Training Dataset: good good good
        
       | spandextwins wrote:
       | Now that it's already been used for training they don't want it
       | hanging around?
        
       | SirMaster wrote:
       | Looks like you can get books3 and other stuff here:
       | 
       | https://www.thenose.cc/wiki/Overview
        
         | gglon wrote:
         | Also magnet for the pile works well:
         | magnet:?xt=urn:btih:0d366035664fdf51cfbe9f733953ba325776e667
        
         | artninja1988 wrote:
         | Good man
        
       | delecti wrote:
       | This seems like the inevitable outcome. Hosting pirated media for
       | _any_ purpose is going to draw moneyed attention, and doing so
       | for a purpose as controversial as AI training doesn 't help.
       | Multiple authors I follow have expressed their disapproval of
       | their works being included in that dataset.
        
         | startupsfail wrote:
         | It is a very good question, what should be protected by
         | copyright and what shouldn't be.
         | 
         | I understand that in Japan, for example, it is legal to apply
         | deep analysis techniques, including deep learning and
         | generative models to any data, regardless the copyright.
         | 
         | I would guess that in Japan, they've looked at China, where
         | copyrights are also not an issue. And decided that it's a
         | really bad idea to stop research because of copyrights and let
         | China to be first at the AGI race.
         | 
         | If a lot of time will get wasted on the GDPR and copyrights
         | discussions in Europe and United States, instead of actually
         | doing research, the democratic world will get behind.
         | 
         | Giant corporations that are scrambling to keep R&D going are
         | quite law-abiding actually. And are currently struggling, as
         | most of the datasets out there can't be used, due to unclear
         | copyright and licensing regulations.
        
           | mistrial9 wrote:
           | "race to the bottom" concept comes to mind here
        
           | dbtc wrote:
           | It's really too bad that human conflict in the form of an
           | arms race is what's driving the development of these things.
        
           | __loam wrote:
           | Looking to China for ethical business practice seems pretty
           | bad to me. I don't think some imaginary race to agi is a good
           | excuse for violating the rights of thousands of people. If
           | you're getting value out of someone's data and they're not
           | being compensated, some kind of shitty externality is
           | happening. Hurting authors and artists to the point that they
           | make fewer works that can be fed to the AI also hurts AI
           | companies in the long run, and also it sucks a lot.
        
             | firecall wrote:
             | No one is looking to China for ethical business practice.
             | 
             | Copyright ethics are something we have invented.
             | 
             | China chooses not to follow western copyright regulations.
             | That's not unethical, to them anyway.
             | 
             | Indeed, Copyright law was originally intended to allow
             | people to legally copy things, not protect corporate
             | interests for an eternity!
        
         | almatabata wrote:
         | I agree that this database looked blatantly illegal as the ones
         | using it never actually bought the books. You can clearly sue
         | them for counterfeiting.
         | 
         | I still would like to see if legally a judge would consider the
         | training illegal. I do not think this got tested and i would
         | not consider the answer as clear cut. If a company buys a
         | million ebooks, can it simply train an AI on it? If no employee
         | reads the books and only the AI "reads" it you could argue that
         | the company bought it only for a single reader/entity. They
         | will probably need to update the copyright system to take this
         | case into account. We will see in a few years how this plays
         | out.
         | 
         | In the end these authors will eventually lose their copyright
         | after their death. And nothing will stop the AI training at
         | that point. This only delays the inevitable by a few years. The
         | public domain already has a lot of amazing books and resources.
         | Google even has access to very rare books nobody else has
         | access to digitally i think:
         | 
         | https://blog.google/outreach-initiatives/arts-culture/in-beg...
        
         | toomuchtodo wrote:
         | There are enormous digital libraries being used for training
         | where authors will have no recourse in preventing this from
         | happening. To continue, those performing model generation just
         | have to either avoid certain jurisdictions or not announce what
         | they're doing. I am not unsympathetic to author feelings on the
         | topic, but regulation in this regard moving faster than bits
         | and compute will be challenging to say the least. For these
         | small dataset amounts currently discussed (<1PB), the cost to
         | geographically relocate is obviously trivial.
         | 
         | TLDR Model training will go "dark" or underground, with models
         | distributed like illicit content.
        
           | AnthonyMouse wrote:
           | Would you even need to distribute the models that way? How
           | would anyone know what you trained them on? Being able to
           | occasionally emit a direct quote from a literary work doesn't
           | tell you whether it was trained on the work itself or on
           | various internet posts that have quoted from it, nor how the
           | work was obtained. Is it illegal to train a model on a copy
           | of a book that you own?
        
             | [deleted]
        
             | __loam wrote:
             | If you get sued, the data you trained on could be subject
             | to the discovery process.
        
               | AnthonyMouse wrote:
               | That's assuming you kept records of what you used, and
               | that the person who trained the model is the same person
               | distributing it, or that the former is even in your
               | jurisdiction.
        
               | __loam wrote:
               | We're good as long as we outsource our abuses I guess.
        
               | AnthonyMouse wrote:
               | It not obvious that this even _should_ be illegal, but a
               | decent factor in determining that is whether making it
               | illegal would even do any good. Unenforceable laws are
               | all cost and no benefit.
        
               | Retric wrote:
               | The models themselves record the training data. ChatGPT
               | is happy to spit out Albus Dumbledore when asked about a
               | hypothetical magic school. If you don't have the exact
               | training dataset that's not going to be an actual
               | defense.
               | 
               | The only protection is specific jurisdictions ruling this
               | stuff is fair use, which is risky when you could be found
               | liable in literally any country.
        
               | AnthonyMouse wrote:
               | > ChatGPT is happy to spit out Albus Dumbledore when
               | asked about a hypothetical magic school.
               | 
               | How are you supposed to know if that's because it
               | ingested the text of Harry Potter or a bunch of fan blogs
               | talking about Harry Potter?
               | 
               | > The only protection is specific jurisdictions ruling
               | this stuff is fair use, which is risky when you could be
               | found liable in literally any country.
               | 
               | Saudi Arabia has quite strict blasphemy laws but people
               | don't seem to be bothered much by it when hosting quite
               | obvious violations of them on their servers in North
               | America.
        
               | Retric wrote:
               | > or a bunch of fan blogs talking about Harry Potter?
               | 
               | Doesn't actually matter here because those fan blogs are
               | also derivative works. I may have read "Call me Ishmael"
               | in a blog, but the quote is from Moby Dick.
               | 
               | ChatGPT can argue for a fair use exception, but combining
               | lots of derivative works can be copyright infringement
               | even if none of the things you directly copied where. IE:
               | If you copy a lot of excerpts from a poem and recreate
               | the poem that doesn't mean you can now use the poem.
               | 
               | Being able to say Harry Potter isn't directly in their
               | training data is at best useful to arguing something was
               | unintentional copyright infringement rather than not
               | actually being copyright infringement.
               | 
               | > Blasphemy laws
               | 
               | That's a criminal not civil issue and there aren't a huge
               | number of treaties on the subject. If a US company is
               | sued in Australia for copyright infringement they don't
               | get to move the case to the US.
        
               | __loam wrote:
               | Worth noting that Fair Use is more than just
               | transformative work. Your work can still be found to be
               | infringing even if it's transformative if it egregiously
               | violates other tenets of fair use, like disrupting the
               | market for the original.
        
               | TrueDuality wrote:
               | That kind of information is present in fan fiction,
               | Wikipedia, reviews. Plenty of other sources. You're also
               | incredibly wrong on being able to assume a book is
               | present in the dataset. It is up to the complainant to
               | prove the case, you are not guilty because you didn't
               | keep sufficient records.
        
               | Retric wrote:
               | Copyright is like a shit sandwich, just those two words
               | alone prove it's a derivative work. This is true even if
               | you copied from a copy. Once something is demonstrated as
               | a derivative work, proving your use is fair use is now on
               | the person who created it. That's normally fairly easy
               | with a book, but harder with LLM's if they easily spit
               | out large chunks of a clearly copied work.
               | 
               | Remember lawsuits aren't beyond a reasonable doubt the
               | standard is much lower.
        
               | diogenes4 wrote:
               | Surely this opens a door where, because an agent is
               | unlikely to reproduce work _verbatim_ but rather in a
               | more compressed or decompressed wording, humans are held
               | to a different standard? Otherwise humans would remain an
               | effective way to launder the output of the agent.
        
           | artninja1988 wrote:
           | I mean the pile, which books3 is only a subset of, is only
           | about 850GB big. Raw text doesn't take up much space at all
        
             | truckerbill wrote:
             | got a mirror?
        
               | justinclift wrote:
               | This seems like it?
               | 
               | https://news.ycombinator.com/item?id=37386478
        
               | crtasm wrote:
               | Or this direct download linked on the same thread:
               | https://news.ycombinator.com/item?id=37386854
        
               | toomuchtodo wrote:
               | One should be mindful to, in general, not touch anything
               | with a public IP that leads back to them.
        
           | NewbornKittens wrote:
           | I agree, realistically this is just going to make training
           | sets less findable in mainstream knowledge, but I think
           | writers being against it regardless makes sense on principle.
           | I, too, would be pissed af if someone took say, my
           | fingerprints, and started putting my fingerprints on random
           | stuff and places where I wasn't.
        
         | rg111 wrote:
         | This is bad for the small guys, though.
         | 
         | Do you think that DeepMind, OpenAI, etc. don't have this
         | dataset copied 10 times over? Think again!
        
           | delecti wrote:
           | Your argument seems to be: why stop the little guys from
           | doing something illegal if the big guys are doing it too? We
           | should ideally stop them all, but it isn't surprising that
           | the _blatant_ examples of illegality are stopped first.
        
             | artninja1988 wrote:
             | If datasets are not shared externally, like how deepmind
             | etc. does it, it is much more of a grey area. Training may
             | well end up being fair use. So yes, the efforts of Shawn
             | compiling books3 just levels the playing field a bit.
        
           | justapassenger wrote:
           | It's a valid question, how can smaller companies be
           | competitive in this business.
           | 
           | But saying "we steal to stay relevant, because we don't have
           | as much funds as our competitors" is not the answer.
        
           | quonn wrote:
           | How is this different from running pirated software for
           | example? Big business doesn't do it. And if they do they get
           | fined heavily. Same should be true for illegal AI training
           | data. Surely this is a problem that can be solved.
        
             | michaelt wrote:
             | _> How is this different from running pirated software for
             | example? Big business doesn't do it._
             | 
             | All the current LLMs are trained on data scraped from
             | random websites, with no regard for the website's
             | copyright, aren't they?
             | 
             | Presumably because these big businesses have a theory
             | training an LLM is 'fair use'.
             | 
             | Why wouldn't they treat books the same way they treat web
             | pages?
        
               | __loam wrote:
               | I'm tired of tech companies abusing the public then
               | asking for forgiveness.
        
               | disposition2 wrote:
               | I agree and feel the same way.
               | 
               | But it also feels like standing at the edge of the sea
               | complaining about the tide coming in...I'm not sure
               | there's really much that can / will be done about it.
        
               | rpd9803 wrote:
               | We've tried nothing and are all out of ideas!
        
               | MPSimmons wrote:
               | I actually think this kind of use is purer to the
               | original intent of the web, where everything on the
               | internet was freely available and consumption was
               | encouraged. "Information wants to be free" used to be the
               | rallying cry.
               | 
               | That being said, it feels like there's also a shade of
               | perspective from the old quote:
               | 
               | "In its majestic equality, the law forbids rich and poor
               | alike to sleep under bridges, beg in the streets and
               | steal loaves of bread." - assuming everything in public
               | is fair game, then everyone is welcome to build a multi-
               | petabyte database of text and use millions of dollars
               | worth of GPUs to train an AI on it.
        
               | __loam wrote:
               | That last point is great. I've definitely seen a lot of
               | people talking about how we need to let the little guy
               | develop their own AI, with very little attention paid to
               | the actual realistic costs of doing so. GPT-4 I believe
               | cost $100 million ish to train.
        
               | oooyay wrote:
               | Bold of you to think they're asking for forgiveness at
               | all. I've seen mostly apologists.
        
             | simbolit wrote:
             | Except big business does it all the time. Your parent
             | comment mentioned OpenAI, who does it.
        
       | wolverine876 wrote:
       | Using the term 'anti-piracy' makes the speaker a shill for 'pro-
       | censorhip' or 'anti-knowledge' or 'elitist' or whatever groups.
       | Look at the effectiveness of well-crafted public communications:
       | The people you don't like are 'pirates' - criminals; threats to
       | order; dangerous to civilians; wild, drunken and low status. And
       | 'our' side is against piracy!
       | 
       | We need neutral terminology for both sides. Even 'intellectual
       | property' tries to transform a time-limited monopolistic license
       | from government for intangible goods into eternal real estate.
       | Any ideas?
        
         | thomastjeffery wrote:
         | Here's the problem: we could potentially deescalate this debate
         | by using more tactful language. But do we really want to?
         | 
         | Is the answer to this situation a neutral position? I don't
         | think so.
         | 
         | Intellectual monopoly is bullshit. I see no value in
         | deescalating the narrative.
         | 
         | What we need is _change_. Change can 't be made by taking a
         | step back and deescalating the narrative.
        
         | armchairhacker wrote:
         | I'm at the point where "piracy" isn't even a Bad Word anymore,
         | because it always refers to pirating from big companies where
         | there's either no way to legitimately buy, buying is
         | unreasonably expensive (e.g. due to different seasons being on
         | different platforms), or it's research where the scientists
         | don't care about their work being pirated because they don't
         | get paid for it anyways.
         | 
         | You'd have to reword it to something like "anti-digital theft"
         | to bring back the negative connotation. Or have more articles
         | about how piracy is destroying small businesses, which are
         | suspiciously rare.
        
         | nonrandomstring wrote:
         | This is a really good question. Can language be de-weaponised?
         | 
         | Wittgenstein and Marcuse thought not - that words stand in as
         | tools of intent. It is our hostile times that make weapons of
         | words.
         | 
         | But it's frustrating because neutral prose is hard to reach. We
         | seem in an age where it's hard to say anything succinctly, in a
         | fashion that will pique the wretched attention span of the
         | ordinary reader, that isn't either drenched in political bias
         | or ungainly.
         | 
         | I think the media industry absolutely embarrassed themselves by
         | misappropriating the word "piracy". Even after 30 years it
         | still jars as awful, childish cringe. When I see a grown man in
         | a suit talking of "piracy" I can't help but see a pitiful clown
         | before me.
         | 
         | But one cannot police other's language. They must be allowed to
         | make themselves known through the words they choose.
        
           | [deleted]
        
         | __loam wrote:
         | You're giving a lot of credit to the "pro-piracy side", which
         | in this case are massive corporate surveillance platforms
         | generating huge amounts of revenue with data they had no part
         | in creating. There's gotta be a middle ground here that doesn't
         | turn all artists and writers into paste for the AI machine.
        
         | barrysteve wrote:
         | No, there's no ideas. It's already trivial to devour the works
         | of many good writers for the cost of an internet connection and
         | the capital outlay of 'a laptop'.
         | 
         | Piracy is killing the motive to work. I've seen it happen
         | countless times and nobody says a thing about it.
         | 
         | How about copyright becomes so pervasive and all enforced that
         | massive corporations like Disney cannot bully the small fry
         | into legally moving into their turf.
         | 
         | Copyright is an hellish nightmare because it used by big
         | corporations to devour, rather than used to emancipate every
         | single person.
         | 
         | ChatGPT ravaging and pillaging the internet was the last straw
         | for "faith in freedom".
         | 
         | I'm putting all my faith and work into walking on water,
         | establishing the cleansing task of getting out from under the
         | sawblade.
         | 
         | People think to write is to compile and package in on itself
         | like a computer program or a spirit-sucking business seminar.
         | Typing on keyboards does that to a man.
        
           | visarga wrote:
           | It is known that LLMs are possible, so there is no stoping
           | it. It's too late to think about if we like it. All ideas and
           | skills are fair game now, but LLMs lack in depth in all of
           | them. No LLM system is better than a domain expert. Maybe
           | with better training data we will have better models. There
           | are some models trained only on public domain and synthetic
           | (generated) data. So maybe LLMs won't learn directly from our
           | works. They need better data than internet stuff, human text
           | is not good enough. See the Microsoft model Phi-1.5, trained
           | on diverse synthetic text of high quality.
        
         | rpd9803 wrote:
         | I mean, except if they are clearly copying works they don't
         | have copyright to, they are violating the law and effectively
         | taking some amount of money out of the pocket of authors and
         | rights holders. You can _try_ and demand people use more
         | "neutral" language to disucss copyright violations, but good
         | luck.
         | 
         | I hope you never know the misfortune of having assets devalued
         | by people that feel they don't need to operate within the
         | boundaries of the rules.
        
       | achrono wrote:
       | "Is this the right thing to have done or not?" is an important
       | question obviously but the more impactful question for our times
       | is: "is the right thing being done _consistently_? "
       | 
       | So, why is similar action not being taken against the actual big
       | offenders -- Google, OpenAI, rest of big-tech?
        
       | artninja1988 wrote:
       | Hope someone hosts it can't be taken down by imaginary "property"
       | DMCA requests
        
         | [deleted]
        
         | h0p3 wrote:
         | Okay, comrade. `/salute`. It is discoverable over filesharing
         | networks designed for the job.
        
       | [deleted]
        
       | YeGoblynQueenne wrote:
       | Note well:
       | 
       | https://twitter.com/theshawwn/status/1320282155457646592
       | 
       | Might it be that a frank conversation between adults is what is
       | needed instead of puerile posturing?
        
         | boomboomsubban wrote:
         | From their site
         | 
         | >We are and always have been DMCA compliant, this is our
         | official policy. https://the-eye.eu/dmca.html However we're
         | often hounded by false claimants to which this(mp4 the Twitter
         | post links to) is our unofficial policy, popularised as a meme
         | in our community that time a church tried to sue us for
         | $22,000,000.
        
           | YeGoblynQueenne wrote:
           | Yes, but the tweet links directly to that video calling it
           | the "DMCA policy" of the-eye.eu. My comment is about the
           | tweet.
        
         | justinclift wrote:
         | https://nitter.net/theshawwn/status/1320282155457646592
        
       | awestroke wrote:
       | Now it will be mirrored 1000x more due to the Streisand effect.
       | Good job I guess?
        
       | crtasm wrote:
       | https://news.ycombinator.com/item?id=37379297 The Battle over
       | Books3
       | 
       | 122 points | 29 days ago | 130 comments
        
       | ttt3ts wrote:
       | It is part of the pile. So, ya don't know why this matters. You
       | can easily get it.
       | 
       | Also, dupe, also old
       | https://news.ycombinator.com/item?id=37154633
        
       ___________________________________________________________________
       (page generated 2023-10-03 23:01 UTC)