[HN Gopher] Anti-Piracy Group Takes AI Training Dataset 'Books3'...
___________________________________________________________________
Anti-Piracy Group Takes AI Training Dataset 'Books3' Offline
Author : YeGoblynQueenne
Score : 95 points
Date : 2023-10-03 12:55 UTC (10 hours ago)
(HTM) web link (gizmodo.com)
(TXT) w3m dump (gizmodo.com)
| nektro wrote:
| Anti-Piracy: bad
|
| Group Takes AI Training Dataset: good good good
| spandextwins wrote:
| Now that it's already been used for training they don't want it
| hanging around?
| SirMaster wrote:
| Looks like you can get books3 and other stuff here:
|
| https://www.thenose.cc/wiki/Overview
| gglon wrote:
| Also magnet for the pile works well:
| magnet:?xt=urn:btih:0d366035664fdf51cfbe9f733953ba325776e667
| artninja1988 wrote:
| Good man
| delecti wrote:
| This seems like the inevitable outcome. Hosting pirated media for
| _any_ purpose is going to draw moneyed attention, and doing so
| for a purpose as controversial as AI training doesn 't help.
| Multiple authors I follow have expressed their disapproval of
| their works being included in that dataset.
| startupsfail wrote:
| It is a very good question, what should be protected by
| copyright and what shouldn't be.
|
| I understand that in Japan, for example, it is legal to apply
| deep analysis techniques, including deep learning and
| generative models to any data, regardless the copyright.
|
| I would guess that in Japan, they've looked at China, where
| copyrights are also not an issue. And decided that it's a
| really bad idea to stop research because of copyrights and let
| China to be first at the AGI race.
|
| If a lot of time will get wasted on the GDPR and copyrights
| discussions in Europe and United States, instead of actually
| doing research, the democratic world will get behind.
|
| Giant corporations that are scrambling to keep R&D going are
| quite law-abiding actually. And are currently struggling, as
| most of the datasets out there can't be used, due to unclear
| copyright and licensing regulations.
| mistrial9 wrote:
| "race to the bottom" concept comes to mind here
| dbtc wrote:
| It's really too bad that human conflict in the form of an
| arms race is what's driving the development of these things.
| __loam wrote:
| Looking to China for ethical business practice seems pretty
| bad to me. I don't think some imaginary race to agi is a good
| excuse for violating the rights of thousands of people. If
| you're getting value out of someone's data and they're not
| being compensated, some kind of shitty externality is
| happening. Hurting authors and artists to the point that they
| make fewer works that can be fed to the AI also hurts AI
| companies in the long run, and also it sucks a lot.
| firecall wrote:
| No one is looking to China for ethical business practice.
|
| Copyright ethics are something we have invented.
|
| China chooses not to follow western copyright regulations.
| That's not unethical, to them anyway.
|
| Indeed, Copyright law was originally intended to allow
| people to legally copy things, not protect corporate
| interests for an eternity!
| almatabata wrote:
| I agree that this database looked blatantly illegal as the ones
| using it never actually bought the books. You can clearly sue
| them for counterfeiting.
|
| I still would like to see if legally a judge would consider the
| training illegal. I do not think this got tested and i would
| not consider the answer as clear cut. If a company buys a
| million ebooks, can it simply train an AI on it? If no employee
| reads the books and only the AI "reads" it you could argue that
| the company bought it only for a single reader/entity. They
| will probably need to update the copyright system to take this
| case into account. We will see in a few years how this plays
| out.
|
| In the end these authors will eventually lose their copyright
| after their death. And nothing will stop the AI training at
| that point. This only delays the inevitable by a few years. The
| public domain already has a lot of amazing books and resources.
| Google even has access to very rare books nobody else has
| access to digitally i think:
|
| https://blog.google/outreach-initiatives/arts-culture/in-beg...
| toomuchtodo wrote:
| There are enormous digital libraries being used for training
| where authors will have no recourse in preventing this from
| happening. To continue, those performing model generation just
| have to either avoid certain jurisdictions or not announce what
| they're doing. I am not unsympathetic to author feelings on the
| topic, but regulation in this regard moving faster than bits
| and compute will be challenging to say the least. For these
| small dataset amounts currently discussed (<1PB), the cost to
| geographically relocate is obviously trivial.
|
| TLDR Model training will go "dark" or underground, with models
| distributed like illicit content.
| AnthonyMouse wrote:
| Would you even need to distribute the models that way? How
| would anyone know what you trained them on? Being able to
| occasionally emit a direct quote from a literary work doesn't
| tell you whether it was trained on the work itself or on
| various internet posts that have quoted from it, nor how the
| work was obtained. Is it illegal to train a model on a copy
| of a book that you own?
| [deleted]
| __loam wrote:
| If you get sued, the data you trained on could be subject
| to the discovery process.
| AnthonyMouse wrote:
| That's assuming you kept records of what you used, and
| that the person who trained the model is the same person
| distributing it, or that the former is even in your
| jurisdiction.
| __loam wrote:
| We're good as long as we outsource our abuses I guess.
| AnthonyMouse wrote:
| It not obvious that this even _should_ be illegal, but a
| decent factor in determining that is whether making it
| illegal would even do any good. Unenforceable laws are
| all cost and no benefit.
| Retric wrote:
| The models themselves record the training data. ChatGPT
| is happy to spit out Albus Dumbledore when asked about a
| hypothetical magic school. If you don't have the exact
| training dataset that's not going to be an actual
| defense.
|
| The only protection is specific jurisdictions ruling this
| stuff is fair use, which is risky when you could be found
| liable in literally any country.
| AnthonyMouse wrote:
| > ChatGPT is happy to spit out Albus Dumbledore when
| asked about a hypothetical magic school.
|
| How are you supposed to know if that's because it
| ingested the text of Harry Potter or a bunch of fan blogs
| talking about Harry Potter?
|
| > The only protection is specific jurisdictions ruling
| this stuff is fair use, which is risky when you could be
| found liable in literally any country.
|
| Saudi Arabia has quite strict blasphemy laws but people
| don't seem to be bothered much by it when hosting quite
| obvious violations of them on their servers in North
| America.
| Retric wrote:
| > or a bunch of fan blogs talking about Harry Potter?
|
| Doesn't actually matter here because those fan blogs are
| also derivative works. I may have read "Call me Ishmael"
| in a blog, but the quote is from Moby Dick.
|
| ChatGPT can argue for a fair use exception, but combining
| lots of derivative works can be copyright infringement
| even if none of the things you directly copied where. IE:
| If you copy a lot of excerpts from a poem and recreate
| the poem that doesn't mean you can now use the poem.
|
| Being able to say Harry Potter isn't directly in their
| training data is at best useful to arguing something was
| unintentional copyright infringement rather than not
| actually being copyright infringement.
|
| > Blasphemy laws
|
| That's a criminal not civil issue and there aren't a huge
| number of treaties on the subject. If a US company is
| sued in Australia for copyright infringement they don't
| get to move the case to the US.
| __loam wrote:
| Worth noting that Fair Use is more than just
| transformative work. Your work can still be found to be
| infringing even if it's transformative if it egregiously
| violates other tenets of fair use, like disrupting the
| market for the original.
| TrueDuality wrote:
| That kind of information is present in fan fiction,
| Wikipedia, reviews. Plenty of other sources. You're also
| incredibly wrong on being able to assume a book is
| present in the dataset. It is up to the complainant to
| prove the case, you are not guilty because you didn't
| keep sufficient records.
| Retric wrote:
| Copyright is like a shit sandwich, just those two words
| alone prove it's a derivative work. This is true even if
| you copied from a copy. Once something is demonstrated as
| a derivative work, proving your use is fair use is now on
| the person who created it. That's normally fairly easy
| with a book, but harder with LLM's if they easily spit
| out large chunks of a clearly copied work.
|
| Remember lawsuits aren't beyond a reasonable doubt the
| standard is much lower.
| diogenes4 wrote:
| Surely this opens a door where, because an agent is
| unlikely to reproduce work _verbatim_ but rather in a
| more compressed or decompressed wording, humans are held
| to a different standard? Otherwise humans would remain an
| effective way to launder the output of the agent.
| artninja1988 wrote:
| I mean the pile, which books3 is only a subset of, is only
| about 850GB big. Raw text doesn't take up much space at all
| truckerbill wrote:
| got a mirror?
| justinclift wrote:
| This seems like it?
|
| https://news.ycombinator.com/item?id=37386478
| crtasm wrote:
| Or this direct download linked on the same thread:
| https://news.ycombinator.com/item?id=37386854
| toomuchtodo wrote:
| One should be mindful to, in general, not touch anything
| with a public IP that leads back to them.
| NewbornKittens wrote:
| I agree, realistically this is just going to make training
| sets less findable in mainstream knowledge, but I think
| writers being against it regardless makes sense on principle.
| I, too, would be pissed af if someone took say, my
| fingerprints, and started putting my fingerprints on random
| stuff and places where I wasn't.
| rg111 wrote:
| This is bad for the small guys, though.
|
| Do you think that DeepMind, OpenAI, etc. don't have this
| dataset copied 10 times over? Think again!
| delecti wrote:
| Your argument seems to be: why stop the little guys from
| doing something illegal if the big guys are doing it too? We
| should ideally stop them all, but it isn't surprising that
| the _blatant_ examples of illegality are stopped first.
| artninja1988 wrote:
| If datasets are not shared externally, like how deepmind
| etc. does it, it is much more of a grey area. Training may
| well end up being fair use. So yes, the efforts of Shawn
| compiling books3 just levels the playing field a bit.
| justapassenger wrote:
| It's a valid question, how can smaller companies be
| competitive in this business.
|
| But saying "we steal to stay relevant, because we don't have
| as much funds as our competitors" is not the answer.
| quonn wrote:
| How is this different from running pirated software for
| example? Big business doesn't do it. And if they do they get
| fined heavily. Same should be true for illegal AI training
| data. Surely this is a problem that can be solved.
| michaelt wrote:
| _> How is this different from running pirated software for
| example? Big business doesn't do it._
|
| All the current LLMs are trained on data scraped from
| random websites, with no regard for the website's
| copyright, aren't they?
|
| Presumably because these big businesses have a theory
| training an LLM is 'fair use'.
|
| Why wouldn't they treat books the same way they treat web
| pages?
| __loam wrote:
| I'm tired of tech companies abusing the public then
| asking for forgiveness.
| disposition2 wrote:
| I agree and feel the same way.
|
| But it also feels like standing at the edge of the sea
| complaining about the tide coming in...I'm not sure
| there's really much that can / will be done about it.
| rpd9803 wrote:
| We've tried nothing and are all out of ideas!
| MPSimmons wrote:
| I actually think this kind of use is purer to the
| original intent of the web, where everything on the
| internet was freely available and consumption was
| encouraged. "Information wants to be free" used to be the
| rallying cry.
|
| That being said, it feels like there's also a shade of
| perspective from the old quote:
|
| "In its majestic equality, the law forbids rich and poor
| alike to sleep under bridges, beg in the streets and
| steal loaves of bread." - assuming everything in public
| is fair game, then everyone is welcome to build a multi-
| petabyte database of text and use millions of dollars
| worth of GPUs to train an AI on it.
| __loam wrote:
| That last point is great. I've definitely seen a lot of
| people talking about how we need to let the little guy
| develop their own AI, with very little attention paid to
| the actual realistic costs of doing so. GPT-4 I believe
| cost $100 million ish to train.
| oooyay wrote:
| Bold of you to think they're asking for forgiveness at
| all. I've seen mostly apologists.
| simbolit wrote:
| Except big business does it all the time. Your parent
| comment mentioned OpenAI, who does it.
| wolverine876 wrote:
| Using the term 'anti-piracy' makes the speaker a shill for 'pro-
| censorhip' or 'anti-knowledge' or 'elitist' or whatever groups.
| Look at the effectiveness of well-crafted public communications:
| The people you don't like are 'pirates' - criminals; threats to
| order; dangerous to civilians; wild, drunken and low status. And
| 'our' side is against piracy!
|
| We need neutral terminology for both sides. Even 'intellectual
| property' tries to transform a time-limited monopolistic license
| from government for intangible goods into eternal real estate.
| Any ideas?
| thomastjeffery wrote:
| Here's the problem: we could potentially deescalate this debate
| by using more tactful language. But do we really want to?
|
| Is the answer to this situation a neutral position? I don't
| think so.
|
| Intellectual monopoly is bullshit. I see no value in
| deescalating the narrative.
|
| What we need is _change_. Change can 't be made by taking a
| step back and deescalating the narrative.
| armchairhacker wrote:
| I'm at the point where "piracy" isn't even a Bad Word anymore,
| because it always refers to pirating from big companies where
| there's either no way to legitimately buy, buying is
| unreasonably expensive (e.g. due to different seasons being on
| different platforms), or it's research where the scientists
| don't care about their work being pirated because they don't
| get paid for it anyways.
|
| You'd have to reword it to something like "anti-digital theft"
| to bring back the negative connotation. Or have more articles
| about how piracy is destroying small businesses, which are
| suspiciously rare.
| nonrandomstring wrote:
| This is a really good question. Can language be de-weaponised?
|
| Wittgenstein and Marcuse thought not - that words stand in as
| tools of intent. It is our hostile times that make weapons of
| words.
|
| But it's frustrating because neutral prose is hard to reach. We
| seem in an age where it's hard to say anything succinctly, in a
| fashion that will pique the wretched attention span of the
| ordinary reader, that isn't either drenched in political bias
| or ungainly.
|
| I think the media industry absolutely embarrassed themselves by
| misappropriating the word "piracy". Even after 30 years it
| still jars as awful, childish cringe. When I see a grown man in
| a suit talking of "piracy" I can't help but see a pitiful clown
| before me.
|
| But one cannot police other's language. They must be allowed to
| make themselves known through the words they choose.
| [deleted]
| __loam wrote:
| You're giving a lot of credit to the "pro-piracy side", which
| in this case are massive corporate surveillance platforms
| generating huge amounts of revenue with data they had no part
| in creating. There's gotta be a middle ground here that doesn't
| turn all artists and writers into paste for the AI machine.
| barrysteve wrote:
| No, there's no ideas. It's already trivial to devour the works
| of many good writers for the cost of an internet connection and
| the capital outlay of 'a laptop'.
|
| Piracy is killing the motive to work. I've seen it happen
| countless times and nobody says a thing about it.
|
| How about copyright becomes so pervasive and all enforced that
| massive corporations like Disney cannot bully the small fry
| into legally moving into their turf.
|
| Copyright is an hellish nightmare because it used by big
| corporations to devour, rather than used to emancipate every
| single person.
|
| ChatGPT ravaging and pillaging the internet was the last straw
| for "faith in freedom".
|
| I'm putting all my faith and work into walking on water,
| establishing the cleansing task of getting out from under the
| sawblade.
|
| People think to write is to compile and package in on itself
| like a computer program or a spirit-sucking business seminar.
| Typing on keyboards does that to a man.
| visarga wrote:
| It is known that LLMs are possible, so there is no stoping
| it. It's too late to think about if we like it. All ideas and
| skills are fair game now, but LLMs lack in depth in all of
| them. No LLM system is better than a domain expert. Maybe
| with better training data we will have better models. There
| are some models trained only on public domain and synthetic
| (generated) data. So maybe LLMs won't learn directly from our
| works. They need better data than internet stuff, human text
| is not good enough. See the Microsoft model Phi-1.5, trained
| on diverse synthetic text of high quality.
| rpd9803 wrote:
| I mean, except if they are clearly copying works they don't
| have copyright to, they are violating the law and effectively
| taking some amount of money out of the pocket of authors and
| rights holders. You can _try_ and demand people use more
| "neutral" language to disucss copyright violations, but good
| luck.
|
| I hope you never know the misfortune of having assets devalued
| by people that feel they don't need to operate within the
| boundaries of the rules.
| achrono wrote:
| "Is this the right thing to have done or not?" is an important
| question obviously but the more impactful question for our times
| is: "is the right thing being done _consistently_? "
|
| So, why is similar action not being taken against the actual big
| offenders -- Google, OpenAI, rest of big-tech?
| artninja1988 wrote:
| Hope someone hosts it can't be taken down by imaginary "property"
| DMCA requests
| [deleted]
| h0p3 wrote:
| Okay, comrade. `/salute`. It is discoverable over filesharing
| networks designed for the job.
| [deleted]
| YeGoblynQueenne wrote:
| Note well:
|
| https://twitter.com/theshawwn/status/1320282155457646592
|
| Might it be that a frank conversation between adults is what is
| needed instead of puerile posturing?
| boomboomsubban wrote:
| From their site
|
| >We are and always have been DMCA compliant, this is our
| official policy. https://the-eye.eu/dmca.html However we're
| often hounded by false claimants to which this(mp4 the Twitter
| post links to) is our unofficial policy, popularised as a meme
| in our community that time a church tried to sue us for
| $22,000,000.
| YeGoblynQueenne wrote:
| Yes, but the tweet links directly to that video calling it
| the "DMCA policy" of the-eye.eu. My comment is about the
| tweet.
| justinclift wrote:
| https://nitter.net/theshawwn/status/1320282155457646592
| awestroke wrote:
| Now it will be mirrored 1000x more due to the Streisand effect.
| Good job I guess?
| crtasm wrote:
| https://news.ycombinator.com/item?id=37379297 The Battle over
| Books3
|
| 122 points | 29 days ago | 130 comments
| ttt3ts wrote:
| It is part of the pile. So, ya don't know why this matters. You
| can easily get it.
|
| Also, dupe, also old
| https://news.ycombinator.com/item?id=37154633
___________________________________________________________________
(page generated 2023-10-03 23:01 UTC)