Post B4o4b3jy87LndkLXUW by pwithnall@mastodon.social
(DIR) More posts by pwithnall@mastodon.social
(DIR) Post #B4mHHpxU3bjITvudTk by hbons@mastodon.social
0 likes, 0 repeats
RE: https://sfba.social/@drahardja/116311524946860153the recent “clean room implementation” tools really hammered this home for me.generated images always felt icky because visual. code felt more diffuse and less emotional for me. even though the process is the same.but now #FOSS projects can be targeted and stripped of author, copyright, and effort with a single command.RT: https://sfba.social/users/drahardja/statuses/116311524946860153
(DIR) Post #B4mHHqNMVOQhmC1Ioa by hbons@mastodon.social
0 likes, 0 repeats
not that I expected much from them lately, but it sure would be nice if the #FSF had an opinion on the end of copyleft.
(DIR) Post #B4o4b2s5MSFUwd8VVI by ebassi@mastodon.social
0 likes, 0 repeats
@nirbheek @slomo @pwithnall @hbons we're talking about genAI-based "clean room" reimplementations; those should not be allowed, unless you can demonstrate that the training data set is actually clean. This should be encoded in the licensing terms and copyright law.
(DIR) Post #B4o4b39oIYQXpbQf8C by slomo@toot.cat
0 likes, 0 repeats
@ebassi @nirbheek @pwithnall @hbons Whether something is a "clean room" implementation or not is also not that easy. Maybe the LLM saw other implementations of the same thing during training, but is that different from you having read some other implementation of something some time before in your life and then writing your own? In either case it's not like the LLM or you can recite* any of the originals but you have an abstract model of it and anything else you ever learned that you work from.* Give it a try: let a recent LLM write you the FreeBSD implementation of /bin/yes that was surely in its training data and is small enough, and then compare to the original. Or maybe more interesting: let an LLM write Rust/GStreamer code, and you'll clearly see that it learned from code I have written. Just like most humans did if you look over github/etc. Unlike what some humans do, you don't see 1:1 copy&paste of whole little helper functions though.Very different to that is the case when you (or an LLM) actively look, compare, copy another implementation during development. (Which is also probably more common than we'd like to pretend based on all the code I've seen over the years)I'm sure we're going to have lots of interesting philosophical discussions between lawyers and courts in the future, ideally with outcomes that don't backfire at us.
(DIR) Post #B4o4b3W8xWI8wrsUwS by pwithnall@mastodon.social
0 likes, 0 repeats
@slomo @ebassi @nirbheek @hbons Because humans are bad at license and copyright attribution on a small scale, does not mean that LLMs should be allowed to get away with bad license and copyright attribution on a vast scale.1/3
(DIR) Post #B4o4b3jy87LndkLXUW by pwithnall@mastodon.social
0 likes, 1 repeats
@slomo @ebassi @nirbheek @hbons As a thought experiment: if there was an LLM which had been trained purely on (say) LGPL-2.1+ code, had low environmental impact, was not funded by VCs who are counting down the time until they turn on the monetisation switch, and which ran local-only and didn’t exfiltrate your stuff to the cloud, and someone used it to rewrite my project, I think I would still be massively pissed off.Why? Because it’s a social problem.2/3
(DIR) Post #B4o4b59Att6i0D9BQ0 by pwithnall@mastodon.social
0 likes, 0 repeats
@slomo @ebassi @nirbheek @hbons Why rewrite my code rather than contributing to it? Why relicense someone’s project from GPL to something more permissive? What’s the motive?Both from the point of view of the people who have released FOSS code which has been put in training datasets, and from the point of view of people who receive unexpected LLM contributions, LLMs enable this kind of anti-community/anti-social behaviour at a huge scale.That’s my current thought-in-progress, anyway.3/3