Post B4JzUswX26pRtzf4Ou by patrick@retro.social
(DIR) More posts by patrick@retro.social
(DIR) Post #B4JzUseo60eP11Mum0 by spaceraser@polymaths.social
0 likes, 0 repeats
Ok, so I’m curious as to our opinions, polymathsians.I want to know how we feel about the process at the heart of creating an LLM, “training” it on data that already exists. My first reaction was that this was theft and copyright infringement. The best good-faith argument that I have heard for the opposite side, the side that says training is fine and not ethically problematic, is that we train human language models with large amounts of the written word. It’s not “copyright infringement” for me to borrow a book from the library, read the entire thing, and integrate that data in to the language model that exists in my brain. That’s the analogy.I want to see if we can focus on this specific issue, and not the energy use and the awful companies and the job displacement etc etc. Is training a language model with material from the open web, library archives, basically any info that is free for a human to access for free, is that by itself theft or fair use? If you feel strongly one way, can you articulate an argument that it is definitely not the other?
(DIR) Post #B4JzUswX26pRtzf4Ou by patrick@retro.social
0 likes, 0 repeats
@spaceraser "Theft or fair use": it depends ;-)And that's true even for "human language models", that's why the difference exists in the first place! And the workarounds exist for the same reason, too: the clean room reverse engineering process is designed specifically to avoid transfer of copyrighted material "trained into a human brain" into the output of the process.A botched clean room reverse engineering process pollutes one artifact, and that can be followed up as needed. A sampling situation can be discussed in court (e.g. the base lines of Queen - Under Pressure v. Vanilla Ice - Ice, Ice, Baby).With LLMs, nobody knows - not the creator of the system, not its user - if the output contains enough "original" material to not count as fair use anymore. And it gets worse: The process is industrialised - good luck discussing every single case, there are just too many.
(DIR) Post #B4JzUtEFyD0UmxxE1o by kabel42@polymaths.social
0 likes, 0 repeats
@patrick @spaceraser does that mean training is fine and only the output is bad?
(DIR) Post #B4JzUtVGwwcNdjuoYC by patrick@retro.social
0 likes, 0 repeats
@spaceraser @kabel42 Assume that I have a weird affliction that by merely seeing a book, only its spine, I know and digest its _entire_ content. Now I walk through the archives of the Library of Congress or Deutsche Nationalbibliothek, any one of those reference libraries that carry ~everything in their area of responsibility.By the time I get out, I haven't committed copyright infringement, I merely "read" all those books. If I can generate an original thought once I get out, that isn't copyright infringement, because it's my thought.The problem I see: how to discern if my utterance is still an original thought? how can I make sure that I can still cite every source properly? (and I struggled with these questions on a _much_ smaller scale when writing my theses: the entire body of "you just know" facts, how could I possibly attribute them properly?)So yes, training is okay as far as authors' rights are concerned (there are other issues, e.g. provenance of data, energy investment cost/benefit).The regurgitation machines would have to provide an "original thought" as output, though. I'm not seeing that.
(DIR) Post #B4JzUtfuJP7oAitJ7w by kabel42@polymaths.social
0 likes, 0 repeats
@patrick @spaceraser but there is some difference, if you looked at images and then painted something, you would not recreate the watermark, you would understand that that is another image that happens to be in the same space
(DIR) Post #B4JzUtolmSDKcD2NwO by patrick@retro.social
0 likes, 0 repeats
@spaceraser @kabel42 That's training, though: Give me a sufficiently abstract image, put abstract watermarks on them without telling me, and I might just copy them with the rest of the mess, thinking it's one and the same.Send me to a course about abstract arts first, and I might not._Somebody_ might have broken copyright to build the corpus, but the training step itself is not the problem: If I read the result of blatant plagiarism, I'm still in the clear, unlike the author of the treatise.
(DIR) Post #B4JzUtxzEBaR4nLkJ6 by kabel42@polymaths.social
0 likes, 0 repeats
@patrick @spaceraser but the claim is, it learns like humans and reading copyrighted material and then creating works is fine and not a derived work. But if it reproduces the watermark it doesn't understand and transform the input, it just outputs an interpolation of the inputs and that is not a new work.If the license of the output matched the input that could be fine, like training on GPL compatible code and then outputting GPL code with all the copyrights from the training data.
(DIR) Post #B4JzUu8GbxoHag9xKa by spaceraser@polymaths.social
0 likes, 0 repeats
@kabel42 @patrick One of the things that niggles at me a bit is that the same communities that get extremely righteous and picky about licenses and what's legal fair use and illegal piracy when it comes to a giant silicon valley AI company are (typically, on balance) pretty likely to either dismiss or outright defend the right of, like, a 14 year old learner to do all of the same things and more.That tells me that we (myself included in this statement) are not really that fussed over piracy, we're fussed over a bigass company making a shitton of money by ignoring rules that are inconvenient to them. That's not unique to AI. I sincerely doubt any internet connected device was made without similarly underhanded tactics. That doesn't make it right but that suggests that the LLM as a category is nothing special.
(DIR) Post #B4JzUuJbvmss9rT10q by mirabilos@toot.mirbsd.org
0 likes, 0 repeats
@spaceraser @kabel42 @patrick a program is not a 14-year-old learner. A program is not a human. A program is a fucking program, running on a deterministic machine (PRNG use notwithstanding). LLM output is a deterministic machine translation of its inputs (which includes the “training data”).I’ll even go so far as to say, any attempted comparisons of LLM reproduction of works with human learning is fascist TESCREAL idelogoy.
(DIR) Post #B4JzZ0pHfL8rVeEte4 by mirabilos@toot.mirbsd.org
0 likes, 0 repeats
@kabel42 @patrick @spaceraser training for output is bad, the copyright exception only covers training for analytical purposes (and even then must honour an opt-out).
(DIR) Post #B4JzdPRe4lRf1OHNCa by mirabilos@toot.mirbsd.org
0 likes, 0 repeats
@spaceraser that analogy you chose is fascist TESCREAL ideology (see comment further downthread).
(DIR) Post #B4LcQOp7ALVIhS0J1M by spaceraser@polymaths.social
0 likes, 0 repeats
@mirabilos let me take a second and Wikipedia whatever tf a tescreal is real quick…So, I’ll just say that I don’t agree. I have a pretty high view of anthropology (due to a combination of philosophical and religious positions) and I really don’t need to be convinced that programs and humans are not equivalent. I’m just trying to get at why people are so laisses Faire with people pirating things and so furiously righteous about a company pirating things to develop a program that simulates plausible human speech.
(DIR) Post #B4LcQP2aMGHNNEJ41A by mirabilos@toot.mirbsd.org
0 likes, 0 repeats
@spaceraser ok.(Depends, I guess, on what, when, what for, etc. but I haven’t seen a detailled scenario (but don’t recall the whole thread from yesterday) and cannot type well on smartphone anwyay)