[HN Gopher] Is legal the same as legitimate: AI reimplementation...
       ___________________________________________________________________
        
       Is legal the same as legitimate: AI reimplementation and the
       erosion of copyleft
        
       Author : dahlia
       Score  : 262 points
       Date   : 2026-03-09 15:12 UTC (7 hours ago)
        
 (HTM) web link (writings.hongminhee.org)
 (TXT) w3m dump (writings.hongminhee.org)
        
       | wccrawford wrote:
       | "Antirez closes his careful legal analysis as though it settles
       | the matter. Ronacher acknowledges that "there is an obvious moral
       | question here, but that isn't necessarily what I'm interested
       | in." Both pieces treat legal permissibility as a proxy for social
       | legitimacy. "
       | 
       | This whole article is just complaining that other people didn't
       | have the discussion he wanted.
       | 
       | Ronacher even acknowledged that it's a different discussion, and
       | not one they were trying to have at the moment.
       | 
       | If you want to have it, have it. Don't blast others for _not_
       | having it for you.
        
         | wizzwizz4 wrote:
         | Having this discussion _involves_ blasting others for not
         | considering it. Consider the rest of the paragraph you quoted:
         | 
         | > But law only says what conduct it will not prevent--it does
         | not certify that conduct as right. Aggressive tax minimization
         | that never crosses into illegality may still be widely regarded
         | as antisocial. A pharmaceutical company that legally acquires a
         | patent on a long-generic drug and raises the price a
         | hundredfold has not done something legal and therefore fine.
         | Legality is a necessary condition; it is not a sufficient one.
        
           | amarant wrote:
           | If the discussion inherently cannot be had without blasting
           | innocent bystanders, I don't think it's a discussion worth
           | having.
           | 
           | It might even be morally abhorrent to have such a discussion
           | in the first place!
        
       | ordu wrote:
       | I believe it is a narrow view of the situation. If we take a look
       | into the history, into the reasons for inventing GPL, we'll see
       | that it was an attempt to fight copyrights with copyrights. The
       | very name 'copyleft' is trying to convey the idea.
       | 
       | What AI are eroding is copyright. You can re-implement not just a
       | GPL program, but to reverse engineer and re-implement a closed
       | source program too, people have demonstrated it already, there
       | were stories here on HN about it.
       | 
       | AI is eroding copyright, so there may no longer be a need for the
       | GPL. GNU should stop and rethink its stance, chuck away the GPL
       | as the main tool to fight evil software corporations and embrace
       | LLM as the main weapon.
        
         | stebalien wrote:
         | Copyleft is a mirror of copyright, not a way to fight
         | copyright. It grants rights to the consumer where copyright
         | grants rights to the creator. Importantly, it gives the end-
         | user the right to modify the software running on their devices.
         | 
         | Unfortunately, there are cases where you simply can't just "re-
         | implement" something. E.g., because doing so requires access to
         | restricted tools, keys, or proprietary specifications.
        
           | rileymat2 wrote:
           | > It grants rights to the consumer where copyright grants
           | rights to the creator.
           | 
           | It also grants one major right/feature to the creator, the
           | ability to spread their work while keeping it as open as they
           | intend.
        
           | ordu wrote:
           | These are words of Stallman:
           | 
           |  _" So, I looked for a way to stop that from happening. The
           | method I came up with is called "copyleft." It's called
           | copyleft because it's sort of like taking copyright and
           | flipping it over. [Laughter] Legally, copyleft works based on
           | copyright. We use the existing copyright law, but we use it
           | to achieve a very different goal."_
           | 
           | https://writings.hongminhee.org/2026/03/legal-vs-legitimate/
        
             | sarchertech wrote:
             | That's not a rebuttal of the OP's point. None of that says
             | anything about fighting copyright. It literally says he
             | flipped it which is wha the OP said when they said it's a
             | mirror.
        
             | dathinab wrote:
             | > flipping it over.
             | 
             | i.e. mirroring it
             | 
             | > use it to achieve a very different goal."
             | 
             | "very different goal" isn't the same as "fundamentally
             | destroying copyright"
             | 
             | the very different goal include to protect public code to
             | stay public, be properly attributed, prevent companies from
             | just "sizing" , motivate other to make their code public
             | too etc.
             | 
             | and even if his goals where not like that, it wouldn't make
             | a difference as this is what many people try to archive
             | with using such licenses
             | 
             | this kind of AI usage is very much not in line with this
             | goals,
             | 
             | and in general way cheaper to do software cloning isn't
             | sufficient to fix many of the issues the FOSS movement
             | tried to fix, especially not when looking at the current
             | ecosystem most people are interacting with (i.e. Phones)
             | 
             | ---
             | 
             | ("sizing"): As in the typical MS embrace, extend and
             | extinguish strategy of first embracing the code then giving
             | it proprietary but available extensions/changes/bug
             | fixes/security patches to then make them no longer
             | available if you don't pay them/play by their rules.
             | 
             | ---
             | 
             | Through in the end using AI as a "fancy complicated"
             | photocopier for code is as much removing copyright as using
             | a photocopier for code would. It doesn't matter if you use
             | the photocopier blind folded and never looked at the thing
             | you copied.
        
             | sjunot wrote:
             | > We use the existing copyright law, but we use it to
             | achieve a very different goal.
             | 
             | For the right goal, he should have called it "rightcopy".
        
         | davidw wrote:
         | > LLM as the main weapon
         | 
         | LLM's - to date - seem to require massive capital expenditures
         | to have the highest quality ones, which is a monumental shift
         | in power towards mega corporations and away from the world of
         | open source where you could do innovative work on your own
         | computer running Linux or FreeBSD or some other open OS.
         | 
         | I don't think that's an exciting idea for the Free Software
         | Foundation.
         | 
         | Perhaps with time we'll be able to run local ones that are
         | 'good enough', but we're not there yet.
         | 
         | There's also an ethical/moral question that these things have
         | been trained on millions of hours of people's volunteer work
         | and the benefits of that are going to accrue to the mega
         | corporations.
         | 
         | Edit: I guess the conclusion I come to is that LLM's are good
         | for 'getting things done', but the context in which they are
         | operating is one where the balance of power is heavily tilted
         | towards capital, and open source is perhaps less interesting to
         | participate in if the machines are just going to slurp it up
         | and people don't have to respect the license or even
         | acknowledge your work.
        
           | ordu wrote:
           | _> LLM 's - to date - seem to require massive capital
           | expenditures to have the highest quality ones, which is a
           | monumental shift in power towards mega corporations and away
           | from the world of open source_
           | 
           | Yeah, a bit of a conundrum. But I don't think that fighting
           | for copyright now can bring any benefits for FOSS. GNU should
           | bring Stallman back and see whether he can come with any new
           | ideas and a new strategy. Alternatively they could try
           | without Stallman. But the point is: they should stop and
           | think again. Maybe they will find a way forward, maybe they
           | won't but it means that either they could continue their
           | fight for a freedom meaningfully, or they could just stop
           | fighting and find some other things to do. Both options are
           | better then fighting for copyright.
           | 
           |  _> There 's also an ethical/moral question that these things
           | have been trained on millions of hours of people's volunteer
           | work and the benefits of that are going to accrue to the mega
           | corporations._
           | 
           | I want a clarify this statement a bit. The thing with LLM
           | relying on work of others are not against GPU philosophy as I
           | understand it: algorithms have to be free. Nothing wrong with
           | training LLMs on them or on programs implementing them.
           | Nothing wrong with using these LLMs to write new (free)
           | programs. What is wrong are corporations reaping all the
           | benefits now and locking down new algorithms later.
           | 
           | I think it is important, because copyright is deemed to be an
           | ethical thing by many (I think for most people it is just a
           | deduction: abiding the law is ethical, therefore copyright is
           | ethical), but not for GNU.
        
             | balamatom wrote:
             | >Yeah, a bit of a conundrum.
             | 
             | IMO the primary significant trend in AI. Doesn't get talked
             | about nearly enough. Means the AI is working, I guess.
             | 
             | >GNU should bring Stallman back ... Alternatively they
             | could try without Stallman.
             | 
             | Leave Britney alone >:(
             | 
             | >copyright is deemed to be an ethical thing by many (I
             | think for most people it is just a deduction: abiding the
             | law is ethical, therefore copyright is ethical)
             | 
             | I've busted out "intellectual property is a crime against
             | humanity" at layfolk to see if that shortcuts through that
             | entire little politico-philosophical minefield. They emote
             | the requisite mild shock when such things as crimes against
             | humanity are mentioned; as well as at someone making such a
             | radical statement which seems to come from no familiar
             | species of echo chamber; and then a moment later they begin
             | to very much look like they see where I'm coming from.
        
               | Serenacula wrote:
               | How do you even argue such a thing? I've had no such
               | luck, I've met many people who seem to view copyright and
               | a person owning their ideas and work as a sort of
               | inherent moral.
        
               | bo1024 wrote:
               | Not saying this gets through to people, but copyright is
               | purely about the legal ability to restrict what other
               | people do. Whereas property rights are about not allowing
               | others to restrict what you do (e.g. by taking your
               | stuff).
        
           | zozbot234 wrote:
           | > LLM's - to date - seem to require massive capital
           | expenditures to have the highest quality ones
           | 
           | There are near-SOTA LLM's available under permissive
           | licenses. Even running them doesn't require prohibitive
           | expenses on hardware unless you insist on realtime use.
        
             | walterbell wrote:
             | _> running them doesn 't require prohibitive expenses on
             | hardware_
             | 
             | What async tasks could a local LLM accomplish on Intel 11th
             | gen CPU with 32GB RAM?
        
           | jacquesm wrote:
           | > There's also an ethical/moral question that these things
           | have been trained on millions of hours of people's volunteer
           | work and the benefits of that are going to accrue to the mega
           | corporations.
           | 
           | This was already the case and it just got worse, not better.
        
             | davidw wrote:
             | At a certain point, I think we had reached a kind of
             | equilibrium where some corporations were decent open source
             | citizens. They understood that they could open source
             | things like infrastructure or libraries and keep their
             | 'crown jewels' closed. And while Stallman types might not
             | have been happy with that, it seemed to work out for
             | people.
             | 
             | Now they've just hoovered up all the free stuff into
             | machines that can mix it up enough to spit it out in a way
             | that doesn't even require attribution, and you have to pay
             | to use their machine.
        
               | jacquesm wrote:
               | AI essentially gatekeeps all of open source to companies
               | to pluck from to their hearts content. And individual
               | contributors using these tools and freely mixing it with
               | their own - usual minor - contributions are another step
               | of whitewashing because they're definitely not going to
               | own up to writing only 5% of the stuff they got paid for.
               | 
               | Before we had RedHat and Ubuntu, who at least were
               | contributing back, now we have Microsoft, Anthropic and
               | OpenAI who are racing to lock the barn door around their
               | new captive sheep. It's just a massive IP laundromat.
        
           | thenewnewguy wrote:
           | Is massive capital expenditure not also required to enforce
           | the GPL? If some company steals your GPLed code and doesn't
           | follow the license, you will have to sue them and somebody
           | will have to pay the lawyers.
        
             | davidw wrote:
             | > Is massive capital expenditure not also required to
             | enforce the GPL?
             | 
             | It's nowhere near the order of magnitude of the kind of
             | spending they're sinking into LLM's. The FSF and other
             | groups were reasonably successful at enforcing the GPL,
             | operating on a budget 1000's of times smaller than that of
             | AI companies.
        
               | cloverich wrote:
               | Right but LLM companies are building frontier models with
               | frontier talent while trying to sock up demand with a
               | loss leader strategy, on top of an historic
               | infrastructure build out.
               | 
               | Being able to coat efficiently run frontier models is i
               | think, not a high priced endeavor for an org (compared to
               | an individual).
               | 
               | IMO the proposition is little fishy, but its not totally
               | without merit and imo deserves investigation. If we are
               | all worried about our jobs, even via building custom for
               | sale software, there is likely something there that may
               | obviate the need at least for end user applications.
               | Again, im deeply skeptical, but it is interesting.
        
               | overfeed wrote:
               | > Being able to coat efficiently run frontier models is i
               | think, not a high priced endeavor for an org
               | 
               | Running proprietary model would make you subject to
               | whatever ToS the LLM companies choose on a particular
               | day, and what you can produce with them, which circles
               | back to the raison d'etre for the GPL and GNU.
               | 
               | Until all software copyright is dead and buried, there is
               | no need for copyleft to change tack. Otherwise there
               | rising tide may rise high enough to drown GPL, but not
               | proprietary software.
               | 
               | Open source is easier to counterfeit/license-launder/re-
               | implement using LLMs because source code is much lower-
               | hanging fruit, and is understood by more people than
               | closed-source assembly.
        
           | socalgal2 wrote:
           | Maybe a good open source idea is to "seti at home" style
           | crowd-source training, assuming that's possible.
        
           | shadowgovt wrote:
           | How close are we to good enough and who's working on that? I
           | would be interested in supporting that work; to my mind, many
           | of the real objections to LLMs are diminished if we can make
           | them small and cheap enough to run in the home (and, perhaps,
           | trained with distributed shared resources, although the
           | training problem is the harder one).
        
             | davidw wrote:
             | Good question. It seems like most of the tech world is
             | perfectly happy to be sharecroppers on the Big AI farms. I
             | guess that's not quite the right analogy, since they're
             | doing their own things with it; just that at the end of the
             | day, the tool they're building everything on is owned by
             | someone else.
        
           | tmp10423288442 wrote:
           | > LLM's - to date - seem to require massive capital
           | expenditures to have the highest quality ones, which is a
           | monumental shift in power towards mega corporations and away
           | from the world of open source where you could do innovative
           | work on your own computer running Linux or FreeBSD or some
           | other open OS.
           | 
           | When the FSF and GPL were created, I don't think this was
           | really a consideration. They were perfectly happy with
           | requiring Big Iron Unix or an esoteric Lisp Machine to use
           | the software - they just wanted to have the ability to
           | customize and distribute fixes and enhancements to it.
        
           | Aozora7 wrote:
           | >Perhaps with time we'll be able to run local ones that are
           | 'good enough', but we're not there yet.
           | 
           | Right now, we can get local models that you can run on
           | consumer hardware, that match capabilities of state of the
           | art models from two years ago. The improvements to model
           | architecture may or may not maintain the same pace in the
           | future, but we will get a local equivalent to Opus 4.6 or
           | whatever other benchmark of "good enough" you have, in the
           | foreseeable future.
        
         | cubefox wrote:
         | That's naive. Copyright doesn't just apply to software. There
         | already have been countless lawsuits about copying music long
         | before the term "open source" was invented. No, changing the
         | lyrics a bit doesn't circumvent copyright. Nor does translating
         | a Stephen King novel to German and switching the names of the
         | places and characters.
         | 
         | A court ordered the first Nosferatu movie to be destroyed
         | because it had too many similarities to Dracula. Despite the
         | fact that the movie makes rather large deviations from the
         | original.
         | 
         | If Claude was indeed asked to reimplement the existing
         | codebase, just in Rust and a bit optimized, that could well be
         | a copyright violation. Just like rephrasing _A Song ot Ice and
         | Fire_ a bit, and switching to a different language, doesn 't
         | remove its copyright.
        
           | zozbot234 wrote:
           | Claude was asked to implement a public API, not an entire
           | codebase. The definition of a public API is largely
           | functional; even in an unusually complex case like the Java
           | standard facilities (which are unusually creative even in the
           | structure and organization of the API itself) the
           | reimplementation by Google was found to be fair use.
        
             | cubefox wrote:
             | > Claude was asked to implement a public API, not an entire
             | codebase.
             | 
             | Allegedly. There have been several people who doubted this
             | story. So how to find out who is right? Well, just let
             | Claude compare the sources. Coincidentally, Claude Opus 4.6
             | doesn't just score 75.6% on SWE-bench Verified but also
             | 90.2% on BigLaw Bench.
             | 
             | It's like our copyright lawyer is conveniently also a
             | developer. And possibly identical to the AI that carried
             | out the rewrite/reimplemention in question in the first
             | place.
        
           | Marsymars wrote:
           | > Just like rephrasing A Song ot Ice and Fire a bit, and
           | switching to a different language, doesn't remove its
           | copyright.
           | 
           | There is some precedent for this, e.g. _Alchemised_ is a
           | recent best seller that had just enough changed from its
           | Harry Potter fan fiction source in order to avoid copyright
           | infringement: https://en.wikipedia.org/wiki/Alchemised
           | 
           | (I avoided the term "remove copyright" here because the new
           | work is still under copyright, just not Harry Potter -
           | related copyright.)
        
             | cubefox wrote:
             | That's apparently a different story with different plot, so
             | that's not comparable.
        
         | webstrand wrote:
         | Its purpose "if you run the software you should be able to
         | inspect and modify that software, and to share those
         | modifications with your peers" not explicitly resist copyright.
         | Yes copyright is bad in that it often prevents one from doing
         | that, but it is not the purpose of the GPL to dismantle
         | copyright.
         | 
         | Reducing it to "well you can clone the proprietary software
         | you're forced to use by LLM" is really missing the soul of the
         | GPL.
        
           | pocksuppet wrote:
           | If not for copyright, you could always do that and copyleft
           | wouldn't be needed.
        
             | webstrand wrote:
             | Just because something is copyleft doesn't mean the person
             | who gave you the binary you're using has to supply you with
             | the code the used to build it. That's what the GPL does.
        
         | dathinab wrote:
         | > we'll see that it was an attempt to fight copyrights with
         | copyrights
         | 
         | it's not that simple
         | 
         | yes, GPLs origins have the idea of "everyone should be able to
         | use"
         | 
         | but it also is about attribution the original author
         | 
         | and making sure people can't just de-facto "size public goods"
         | 
         | the kind of AI usage is removing attribution and is often
         | sizing public goods in a way far worse then most companies
         | which just ignored the license did
         | 
         | so today there is more need then ever in the last few decades
         | for GPL like licenses
        
           | amiga386 wrote:
           | You've said "size" twice in comments, did you mean "seize"?
        
         | johnofthesea wrote:
         | > AI is eroding copyright, so there may no longer be a need for
         | the GPL. GNU should stop and rethink its stance, chuck away the
         | GPL as the main tool to fight evil software corporations and
         | embrace LLM as the main weapon.
         | 
         | Is this LLM thing freely available or is it owned and
         | controlled by these companies? Are we going to rent the tools
         | to fight "evil software corporations"?
        
           | cozzyd wrote:
           | easy, we ask Claude to write an open-source freely-available
           | version of Claude with equal or better capabilities.
        
           | josephg wrote:
           | Open models do exist. They're nowhere near aa good as
           | frontier models, but they're getting better all the time.
           | 
           | It's probably only a matter of time before open models are as
           | good as Claude code is today.
        
             | Aozora7 wrote:
             | With the release of GLM-5, I would say that they are pretty
             | much almost as good. Basically 90% as good as Opus 4.6 on
             | most tasks for 20% of inference cost, and open weights.
        
           | Aozora7 wrote:
           | There already are LLMs with open weights that are better at
           | code than state of the art closed source models from a year
           | ago. For now, you most people may have to rent the hardware
           | to run those models, since it's too expensive for most people
           | to own something that can run inference on one trillion
           | parameters, but I wouldn't consider LLMs to be controlled by
           | "evil software corporations" at this point.
        
         | thomastjeffery wrote:
         | While I personally agree with you, Richard Stallman (the
         | creator of the GPL) does not. He has always advocated in favor
         | of strong copyright protection, because the foundation of the
         | GPL is the monopoly power granted by copyright. The problem
         | that the GPL is intended to solve is _proprietary software_.
         | 
         | Generative models (AI) are not really eroding copyright. They
         | are calling its bluff. The very notion of intellectual property
         | depends on a property line: some arbitrary boundary where the
         | property begins and ends. Generative models blur that line,
         | making it impractical to distinguish which property belongs to
         | whom.
         | 
         | Ironically, these models are made by giant monopolistic
         | corporations whose wealth is quite literally a market valuation
         | (stock price) of their copyrights! If generative models ever
         | become good enough to reimplement CUDA, what value will NVIDIA
         | have left?
         | 
         | The reality is that generative models are nowhere near good
         | enough to actually call the bluff. Copyright is still the
         | winning hand, and that is likely to continue, particularly
         | while IP holders are the primary authors of law.
         | 
         | ---
         | 
         | This whole situation is missing the forest for the trees.
         | Intellectual Property is _bullshit_. A system predicated on
         | monopoly power can only result in consolidated wealth driving
         | the consolidation of power; which is precisely what has
         | happened. The words  "starving artist" ring every bit as
         | familiar today as any time in history. Copyright has utterly
         | failed the very goals it was explicitly written with.
         | 
         | It isn't the GPL that needs changing. So long as a system of
         | copyright rules the land, copyleft is the best way to
         | participate. What we really need is a cohesive political
         | movement against monopoly power; one that _isn 't_ conveniently
         | ignorant of copyright as its most significant source.
        
           | pennomi wrote:
           | Right, anything that can be copied instantly for free cannot
           | be realistically owned.
        
         | re-thc wrote:
         | > What AI are eroding is copyright.
         | 
         | At the moment it's people that are eroding copyright. E.g. in
         | this case someone did something.
         | 
         | "AI" didn't have a brain, woke up and suddenly decided to do
         | it.
         | 
         | Realistically nothing to do with AI. Having a gun doesn't mean
         | you randomly shoot.
        
         | xantronix wrote:
         | So not only are we moving goalposts here, but we've decided the
         | GNU team should join the other team? I don't understand how GNU
         | would see mass model LLM training as anything but the most
         | flagrant violations of their ethos. LLM labs, in their view,
         | would be among the most evil software corporations to have ever
         | existed.
        
         | Peritract wrote:
         | > chuck away the GPL as the main tool to fight evil software
         | corporations and embrace LLM as the main weapon.
         | 
         | LLMs are one of the primary manifestations of 'evil software
         | corporations' currently.
        
         | wolvesechoes wrote:
         | > AI is eroding copyright
         | 
         | Unless it is IP of the same big corpos that consumed all
         | content available. Good luck with eroding them.
        
         | mikkupikku wrote:
         | I agree with almost all of that, except the part about GNU
         | changing their stance. I think GNU should stay true and
         | consistent, if for no other reason than to not make many of
         | their supporters who aren't on board with AI feel betrayed and
         | have GNUs legacy soured. If the cause of LLMs conquering
         | proprietary software needs an organization to champion it, let
         | that be a new organization, not GNU.
        
         | martin-t wrote:
         | This is naive. Advertisement and network effects win.
         | Individuals cannot compete with corporations on equal ground
         | here.
        
       | sharkjacobs wrote:
       | > Blanchard's account is that he never looked at the existing
       | source code directly. He fed only the API and the test suite to
       | Claude and asked it to reimplement the library from scratch
       | 
       | This feels sort of like saying "I just blindly threw paint at
       | that canvas on the wall and it came out in the shape of Mickey
       | Mouse, and so it can't be copyright infringement because it was
       | created without the use of my knowledge of Micky Mouse"
       | 
       | Blanchard is, of course, familiar with the source code, he's been
       | its maintainer for years. The premise is that he prompted Claude
       | to reimplement it, without using his own knowledge of it to
       | direct or steer.
        
         | re-thc wrote:
         | > This feels sort of like saying "I just blindly threw paint at
         | that canvas on the wall and
         | 
         | > He fed only the API and the test suite to Claude and asked it
         | 
         | Difference being Claude looked; so not blind. The equivalent is
         | more like I blindly took a photo of it and then used that to...
         | 
         | Technically did look.
        
           | amarant wrote:
           | The article is poorly written. Blanchard was a chardet
           | maintainer for years. Of course _he_ had looked at it 's
           | code!
           | 
           | What he claimed, and what was interesting, was that Claude
           | _didn 't_ look at the code, only the API and the test suite.
           | The new implementation is all Claude. And the implementation
           | is different enough to be considered original, completely
           | different structure, design, and hey, a 48x improvement in
           | performance! It's just API-compatible with the original.
           | Which as per the Google Vs oracle 2021 decision is to be
           | considered fair use.
        
             | mrgoldenbrown wrote:
             | did he claim that Claude wasn't trained on the original? Or
             | just that he didn't personally provide Claude with a copy?
        
               | amarant wrote:
               | I recon the latter, how would he know what was in
               | Claude's training data?
        
             | re-thc wrote:
             | > What he claimed, and what was interesting, was that
             | Claude didn't look at the code
             | 
             | Who opened the PR? Who co-authored the commits? It's
             | clearly on Github.
             | 
             | > Blanchard was a chardet maintainer for years. Of course
             | he had looked at its code!
             | 
             | So there you have it. If he looked, he co-authored then
             | there's that.
        
               | kjksf wrote:
               | If I put my signature on Picasso painting, it doesn't
               | make me co-author of said painting.
               | 
               | Blanchard is very clear that he didn't write a single
               | line of code. He isn't an author, he isn't a co-author.
               | 
               | Signing GitHub commit doesn't change that.
        
               | re-thc wrote:
               | > Blanchard is very clear that he didn't write a single
               | line of code
               | 
               | He used Claude to write it. Difference? The fact that I
               | write on the notepad vs printed it out = I didn't do it?
               | 
               | > Signing GitHub commit doesn't change that.
               | 
               | That's the equivalent of me saying I didn't kill anyone.
               | The fingerprints on the knife doesn't change that.
        
               | satvikpendem wrote:
               | I'll take a commit authored by someone else and then git
               | amend the author to myself, did I write that commit then?
               | By your logic I did apparently.
        
         | dathinab wrote:
         | > Blanchard is, of course, familiar with the source code, he's
         | been its maintainer for years.
         | 
         | I would argue it's irrelevant if they looked or didn't look at
         | the code. As well as weather he was or wasn't familiar with it.
         | 
         | What matters is, that they feed to original code into a tool
         | which they setup to make a copy of it. How that tool works
         | doesn't really matter. Neither does it make a difference if you
         | obfuscate that it's an copy.
         | 
         | If I blindfold myself when making copies of books with a book
         | scanner + printer I'm still engaging in copyright infringement.
         | 
         | If AI is a tool, that should hold.
         | 
         | If it isn't "just" a tool, then it did engage in copyright
         | infringement (as it created the new output side by side with
         | the original) in the same way an employee might do so on
         | command of their boss. Which still makes the boss/company
         | liable for copyright infringement and in general just because
         | you weren't the one who created an infringing product doesn't
         | mean you aren't more or less as liable of distributing it, as
         | if you had done so.
        
           | spullara wrote:
           | if the actual text of the code isn't the same or obviously
           | derivative, copyright doesn't apply at all.
        
             | sigseg1v wrote:
             | What does derivative mean here? Because IMO it means that
             | the existing work was used as input. So if you used a LLM
             | and it was trained on the existing work, that's a
             | derivative work. If you rot13 encode something as input, so
             | you can't personally read it, and then a device decides to
             | rot13 on it again and output it, that's a derivative work.
        
               | nicole_express wrote:
               | Of course, the problem with this interpretation is that
               | all modern LLMs are derivatives from huge amounts of text
               | under completely different licenses, including "All
               | rights reserved", and therefore can not be used for any
               | purpose.
               | 
               | I'm not sure how you square the circle of "it's alright
               | to use the LLM to write code, unless the code is a
               | rewrite of an open source project to change its license".
        
               | JoshTriplett wrote:
               | > Of course, the problem with this interpretation is that
               | all modern LLMs are derivatives from huge amounts of text
               | under completely different licenses, including "All
               | rights reserved", and therefore can not be used for any
               | purpose.
               | 
               | > I'm not sure how you square the circle of "it's alright
               | to use the LLM to write code
               | 
               | You seem like you're on the cusp of stating the obvious
               | correct conclusion: it isn't.
        
               | wizzwizz4 wrote:
               | See also: https://monolith.sourceforge.net/, which seeks
               | to ask the question:
               | 
               | > But how far away from direct and explicit
               | representations do we have to go before copyright no
               | longer applies?
        
               | ghostpepper wrote:
               | As a cynical person I assume all the frontier LLMs were
               | trained on datasets that include every open source
               | project, but as a thought experiment, if an LLM was
               | trained on a dataset that included every open source
               | project _execept_ chardet, do you think said LLM would
               | still be able to easily implement something very similar?
        
               | spullara wrote:
               | There is no doubt in my mind that it could still do it.
        
               | spullara wrote:
               | In order for it to be creatively derivative you would
               | need to copy the structure, logic, organization, and
               | sequence of operations not just reimplement the
               | functionality. It is pretty clear in this case that
               | wasn't done.
        
               | cubefox wrote:
               | It's not clear at all.
        
               | bmcahren wrote:
               | LLMs do not encode nor encrypt their training data. The
               | fact they can recite training data is a defect not a
               | default. You can understand this more simply by
               | calculating the model size as an inverse of a fantasy
               | compression algorithm that is 50% better than SOTA.
               | You'll find you'd still be missing 80-90% of the training
               | data even if it were as much of a stochastic parrot as
               | you may be implying. The outputs of AI are not derivative
               | just because they saw training data including the
               | original library.
               | 
               | Then onto prompting: 'He fed only the API and (his) test
               | suite to Claude'
               | 
               | This is Google v Oracle all over again - are APIs
               | copyrightable?
        
               | satvikpendem wrote:
               | > _This is Google v Oracle all over again - are APIs
               | copyrightable?_
               | 
               | Yes this is the best way to ask the question. If I take a
               | public facing API and reimplement everything, whether
               | it's by human or machine, it should be sufficient. After
               | all, that's what Google did, and it's not like their
               | engineers never read a single line of the Java source
               | code. Even in "clean room" implementations, a human might
               | still have remembered or recalled a previous
               | implementation of some function they had encountered
               | before.
        
               | thunderfork wrote:
               | I find the "compression" argument not very strong, both
               | because copyright still applies to (very) lossy codecs
               | (e.g. your 16kbps Opus file of Thriller infringes, even
               | if the original 192khz/32bit wav file was 12,000kbps),
               | and because copyright still applies to transformed
               | derivative works (a tiny midi file of Thriller might
               | still be enough for the Jackson's label to get you)
        
               | azakai wrote:
               | > LLMs do not encode nor encrypt their training data. The
               | fact they can recite training data is a defect not a
               | default.
               | 
               | About this specific point, it is unclear how much of a
               | defect memorization actually is - there are also reasons
               | to see it as necessary for effective learning. This link
               | explains it well:
               | 
               | https://infinitefaculty.substack.com/p/memorization-vs-
               | gener...
        
               | satvikpendem wrote:
               | > _Because IMO it means that the existing work was used
               | as input_
               | 
               | That's your opinion (since you said "IMO"), not the
               | actual legal definition.
        
             | NSUserDefaults wrote:
             | If you pirate a movie and reencode it, does that apply as
             | well? You can still watch the movie and it is "obviously"
             | the same movie, even though the bytes are completely
             | different. Here you can use the program and it is, to the
             | user, also the same.
        
             | yorwba wrote:
             | Copyright protects even very abstract aspects of human
             | creative expression, not just the specific form in which it
             | is originally expressed. If you translate a book into
             | another language, or turn it into a silent movie, none of
             | the actual text may survive, but the story itself remains
             | covered by the original copyright.
             | 
             | So when you clone the behavior of a program like chardet
             | without referencing the original source code except by
             | executing it to make sure your clone produces exactly the
             | same output, you may still be infringing its copyright if
             | that output reflects creative choices made in the design of
             | chardet that aren't fully determined by the functional
             | purpose of the program.
        
           | margalabargala wrote:
           | > If it isn't "just" a tool, then it did engage in copyright
           | infringement
           | 
           | Copyright infringement is a thing humans do. It's not a
           | human.
           | 
           | Just like how the photos taken by a monkey with a camera have
           | no copyright. Human law binds humans.
        
             | malicka wrote:
             | Correct. The human who shares the copy is the one who
             | engages in copyright infringement.
        
               | margalabargala wrote:
               | So, let's say that rather than actually touching any
               | copyrighted material, a human merely tells an AI about
               | how to go onto the internet and find copyrighted
               | material, download it, and ingest it for training. The
               | AI, fully autonomously, does so, and after training
               | itself on the material deletes it so no human ever
               | downloads, consumes, or shares it.
               | 
               | If we are saying AI is "more than a tool", which seems to
               | be the case courts are leaning since they've ruled AI
               | output without direct human involvement is not
               | copyrightable[0], then the above seems like it would be
               | entirely legal.
               | 
               | [0] https://www.copyright.gov/newsnet/2025/1060.html
        
               | Ekaros wrote:
               | Someone would likely get prosecuted if they instructed AI
               | agent to run say a pump and dump scheme...
               | 
               | Even if the final output doesn't have copyright
               | protection it might still be copyright violation. I think
               | it could be reasonable to have work that itself violates
               | copyright when distributed even if it does not have copy
               | right itself.
        
           | Legend2440 wrote:
           | >that they feed to original code into a tool which they setup
           | to make a copy of it
           | 
           | Well, no. They fed the spec (test cases, etc) into a tool
           | which made a new program matching the spec. This is not a
           | copy of the original code.
           | 
           | But also this feels like arguing over the color of the
           | iceberg while the titanic sinks. If you have a tool that can
           | make code to spec, what is the value in source code anymore?
           | Even if your app is closed-source, you can just tell claude
           | to write new code that does the same thing.
        
             | timeinput wrote:
             | Blanchard fed the spec to the tool, and Anthropic fed the
             | code to the tool, so Blanchard didn't do anything wrong,
             | and Anthropic didn't do anything wrong. Nothing to see
             | here.
        
               | foresto wrote:
               | > Blanchard fed the spec to the tool,
               | 
               | Yes...
               | 
               | > and Anthropic fed the code to the tool,
               | 
               | Presumably, as part of the massive amount of open-source
               | code that must have been fed in to train their model.
               | 
               | > so Blanchard didn't do anything wrong, and Anthropic
               | didn't do anything wrong. Nothing to see here.
               | 
               | This is meant as irony, right?
        
             | vbarrielle wrote:
             | Everyone writes as if he just fed the spec and tests to
             | Claude Code. Ignoring for now that the tests are under LGPL
             | as well, the commit history shows that this has been done
             | with two weeks of steering Claude Code towards the desired
             | output. At every one of these interactions, the maintainer
             | used his deep knowledge of the chardet codebase to steer
             | Claude.
        
         | logicprog wrote:
         | I just don't see how it's relevant whether he did look or
         | didn't. In my opinion, it's not just legally valid to make a
         | re-implementation of something if you've seen the code as long
         | as it doesn't copy expressive elements. I think it's also
         | ethically fine as well to use source code as a reference for
         | re-implementing something as long as it doesn't turn into an
         | exact translation.
        
           | sarchertech wrote:
           | Ignoring the legal or ethical concerns. Let's say we live in
           | a world where the cost of copying code is so close to zero
           | that it's indistinguishable from a world without copyright.
           | 
           | Anything you put out can and will be used by whatever giant
           | company wants to use it with no attribution whatsoever.
           | 
           | Doesn't that massively reduce the incentive to release the
           | source of anything ever?
        
             | pocksuppet wrote:
             | Yes, and it reduces the incentives to release binaries too.
             | Such a world will be populated by almost entirely SaaS,
             | which can still compete on freedom.
        
             | intrasight wrote:
             | Most commercial software that I've used has the model of a
             | legal moat around a pretty crappy database schema.
             | 
             | The non IP protection has largely been in the effort
             | involved in replicating an application's behavior and that
             | effort is dropping precipitously.
        
               | sarchertech wrote:
               | You must not have used much commercial software outside
               | of crappy business SaaS.
        
             | satvikpendem wrote:
             | No, because (most) people don't work on OSS for vanity,
             | they do it to help other people, whether it's individuals
             | or groups of individuals, ie corporations.
             | 
             | It's the same question as, if an AI can generate "art", or
             | photographers can capture a scene better than any
             | (realistic) painter, then will people still create art?
             | Obviously yes, and we see it of course after Stable
             | Diffusion was released three years ago, people are still
             | creating.
        
               | sarchertech wrote:
               | I don't know what a world without copyright does to
               | corporate sponsored open source. It certainly reduces it
               | because there are many corporate sponsored projects that
               | monetize through dual licensing. My guess is in a world
               | where you can't even guarantee attribution, it's much
               | harder to convince your boss to let you open source a
               | project in the first place.
               | 
               | So ignoring people who are being paid by corporations
               | directly to work on open source, in my experience the
               | vast majority of contributors expect to be able to
               | monetize their work eventually in a way that requires
               | attribution. And out of the small number who don't expect
               | a monetary return of any kind, a still smaller number
               | don't expect recognition.
               | 
               | If this weren't the case you'd see a much larger amount
               | of anonymous contributions. There are people who
               | anonymously donate to charity. The vast majority want
               | some kind of recognition.
               | 
               | Obviously we still see art, if you greatly reduce the
               | monetary benefit to producing art, you'll see a lot less
               | of it. This is especially true of non trivial open source
               | software that unlike static artwork requires continual
               | maintenance.
        
           | atomicnumber3 wrote:
           | It's actually not legally fine, or at least it's extremely
           | dangerous. Projects that re-implement APIs presented by
           | extremely litigious companies specifically do not allow
           | people who, for instance, have seen the proprietary source
           | code to then work on the project.
        
             | jpc0 wrote:
             | I don't think fear or legal action makes it illegal.
             | 
             | If I know it is legal to make a turn at a red light. And I
             | know a court will uphold that I was in the right but a
             | police officer will fine me regardless and I would need to
             | go to actually pursue some legal remedy I'm unlikely to do
             | it regardless of whether it is legal because it is
             | expensive, if not in money but time.
             | 
             | In the case of copyright lawsuits they are notoriously
             | expensive and long so even if a court would eventually deem
             | it fine, why take the chance.
        
               | atomicnumber3 wrote:
               | That's my point. It's dangerous and there are sharks in
               | the water. That sounds like you're not going to have a
               | good time if you do the described approach to someone who
               | might assert you're infringing.
        
             | sunshowers wrote:
             | My understanding is that that is a maximalist position for
             | the avoidance of risk, and is sufficient but probably not
             | necessary.
        
           | simonw wrote:
           | Right. The alternative is that we reward Dan for his 14 years
           | of volunteer maintenance of a project... by banning him from
           | working on anything similar under a different license for the
           | rest of his life.
        
         | esafak wrote:
         | If you only stick to the API and ignore the implementation, it
         | is not Mickey Mouse any more but a rodent. If it was just a
         | clone it wouldn't be 50x as fast. Nevertheless, APIs apparently
         | can be copyrightable. I generally disagree with this; it's how
         | PC compatibles took off, giving consumers better options.
        
           | amarant wrote:
           | Wait what, didn't oracle lose the case against Google? Have I
           | been living in an alternate reality where API compatibility
           | is fair use?
        
         | Aurornis wrote:
         | Can anyone find the actual quote where Blanchard said this?
         | 
         | My understanding was that his claim was that _Claude_ was not
         | looking at the existing source code while writing it.
        
           | pklausler wrote:
           | Conveniently ignoring the likelihood that Claude had been
           | trained on the freely accessible source code.
        
           | mrgoldenbrown wrote:
           | Does he have access to Claude's training data? How can he
           | claim Claude wasn't trained on the original code?
        
         | SpicyLemonZest wrote:
         | Isn't this a red herring? An API definition is fair use under
         | Google v. Oracle, but the test suite is definitely
         | copyrightable code!
        
         | babypuncher wrote:
         | What if we said that generative AI output is simply not
         | copyrightable. Anything an AI spits out would automatically be
         | public domain, except in cases where the output directly
         | infringes the rights of an existing work.
         | 
         | This would make it so relicensing with AI rewrites is
         | essentially impossible unless your goal is to transition the
         | work to be truly public domain.
         | 
         | I think this also helps somewhat with the ethical quandary of
         | these models being trained on public data while contributing
         | nothing of value back to the public, and disincentivize the
         | production of slop for profit.
        
           | kjksf wrote:
           | We did in fact say so.
           | 
           | https://www.carltonfields.com/insights/publications/2025/no-.
           | ..
           | 
           | > No Copyright Protection for AI-Assisted Creations: Thaler
           | v. Perlmutter
           | 
           | > A recent key judicial development on this topic occurred
           | when the U.S. Supreme Court declined to review the case of
           | Thaler v. Perlmutter on March 2, 2026, effectively upholding
           | lower court rulings that AI-generated works lacking human
           | authorship are not eligible for copyright protection under
           | U.S. law
        
             | pseudalopex wrote:
             | > > A recent key judicial development on this topic
             | occurred when the U.S. Supreme Court declined to review the
             | case of Thaler v. Perlmutter on March 2, 2026, effectively
             | upholding lower court rulings that AI-generated works
             | lacking human authorship are not eligible for copyright
             | protection under U.S. law
             | 
             | This was AI summary? Those words were not in the article.
             | 
             | The courts said Thaler could not have copyright because he
             | refused to list himself as an author.
        
           | idle_zealot wrote:
           | > This would make it so relicensing with AI rewrites is
           | essentially impossible unless your goal is to transition the
           | work to be truly public domain.
           | 
           | That's not true at all. Anyone could follow these steps:
           | 
           | 1. Have the LLM rewrite GPL code.
           | 
           | 2. Do _not_ publish that public domain code. You have no
           | obligation to.
           | 
           | 3. Make a few tweaks to that code.
           | 
           | 4. Publish a compiled binary/use your code to host a service
           | under a proprietary license of your choice.
        
         | axus wrote:
         | Oracle had it's day in court with Google over the Java APIs.
         | Reimplementing APIs can be done without copyright infringement,
         | but Oracle must have tried to find real infringement during
         | discovery.
         | 
         | In this case, we could theoretically prove that the new chardet
         | is a clean reimplementation. Blanchard can provide all of the
         | prompts necessary to re-implement again, and for the cost of
         | the tokens anyone can reproduce the results.
        
         | NewsaHackO wrote:
         | >This feels sort of like saying "I just blindly threw paint at
         | that canvas on the wall and it came out in the shape of Mickey
         | Mouse, and so it can't be copyright infringement because it was
         | created without the use of my knowledge of Micky Mouse"
         | 
         | IANAL, but that analogy wouldn't work because Mickey Mouse is a
         | trademark, so it doesn't matter how it is created.
        
       | throwaway2027 wrote:
       | I think we're going one step too far even, AI itself is a gray
       | area and how can they guarantee it was trained legally or if it's
       | even legal what they're doing and how can they assert that the
       | input training data didn't contain any copyrighted data.
        
         | observationist wrote:
         | Google already spent billions of dollars and decades of lawyer
         | hours proving it out as fair use. The legal challenges we see
         | now are the dying convulsions of an already broken system of
         | publishers and IP hoarders using every resource at their
         | disposal to manipulate authors and creators and the public into
         | thinking that there's any legitimacy or value underlying modern
         | copyright law.
         | 
         | AI will destroy the current paradigm, completely and utterly,
         | and there's nothing they can do to stop it. It's unclear if
         | they can even slow it, and that's a good thing.
         | 
         | We will be forced to legislate a modern, digital oriented
         | copyright system that's fair and compatible with AI. If
         | producing any software becomes a matter of asking a machine to
         | produce it - if things like AI native operating systems come
         | about, where apps and media are generated on demand, with
         | protocols as backbone, and each device is just generating its
         | own scaffolding around the protocols - then nearly none of
         | modern licensing, copyright, software patents, or IP
         | conventions make any sense whatsoever.
         | 
         | You can't have horse and buggy traffic conventions for
         | airplanes. We're moving in to a whole new paradigm, and maybe
         | we can get legislation that actually benefits society and
         | individuals, instead of propping up massive corporations and
         | making lawyers rich.
        
           | casey2 wrote:
           | Google has cut out some very specific ruling that have
           | nothing to do with modern AI. These systems are just a really
           | slow/lossy git clone, current law has no trouble with it,
           | it's broadly illegal.
           | 
           | If corporations are allowed to launder someone else work as
           | their own people will simply stop working and just start
           | endlessly remixing a la popular music.
        
       | throawayonthe wrote:
       | shall we now have to think about the tradeoffs in adopting
       | 
       | - proprietary
       | 
       | - free
       | 
       | - slop-licensed
       | 
       | software?
        
         | megous wrote:
         | We should just use LLMs to free more software and HW. Make it
         | work against the system.
        
       | mfabbri77 wrote:
       | What if someone doesn't declare that it has been reimplemented
       | using an LLM? Isn't it enough to simply declare that you have
       | reimplemented the software without using an LLM? Good luck
       | proving that in court...
       | 
       | One thing is certain, however: copyleft licenses will disappear:
       | If I can't control the redistribution of my code (through a GPL
       | or similar license), I choose to develop it in closed source.
        
         | bigyabai wrote:
         | Arguably, the GPL has always been the wrong choice if you want
         | to authoritatively control redistribution.
        
       | dwroberts wrote:
       | One of the things that irks me about this whole thing is, if it's
       | so clean room and distinct, why make the changes to the existing
       | project? Why not make an entirely new library?
       | 
       | The answer to that, I think, is that the authors wanted to squat
       | an existing successful project and gain a platform from it. Hence
       | we have news cycle discussing it.
       | 
       | Nobody cares about a new library using AI, but squash an existing
       | one with this stuff, and you get attention. It's the reputation,
       | the GitHub stars, whatever
        
         | nicole_express wrote:
         | I mean, Blanchard was the longtime maintainer of chardet
         | already, and had wanted to relicense it for years. So I think
         | that complicates your picture of "squatting an existing
         | successful project".
         | 
         | Honestly it's a weird test case for this sort of thing. I don't
         | think you'd see an equivalent in most open source projects.
        
         | intrasight wrote:
         | I agree. But you can't copyright goodwill and reputation.
         | Trademark does provide some protection there, right?
        
       | delichon wrote:
       | Imagine if the author has his way, and when we have AI write
       | software, it becomes legally under the license of some other
       | sufficiently similar piece of software. Which may or may not be
       | proprietary. "I see you have generated a todo app very similar to
       | Todoist. So they now own it." That does not seem like a good path
       | either for open source software or for opening up the benefits of
       | AI generated software.
        
       | moi2388 wrote:
       | Perhaps we should finally admit that copyright has always been
       | nonsense, and abolish this ridiculous measure once and for all
        
         | vladms wrote:
         | Probably a wiser approach is to consider different times
         | require different measures (in general!).
         | 
         | I did not study in detail if copyright "has always been
         | nonsense", but I do agree that nowadays some of the copyright
         | regulations are nonsense (for example the very long duration of
         | life + 70 years)
        
         | intrasight wrote:
         | I think AI is very much eroding the legitimacy of copyright -
         | at least to software, which is long been questioned since it's
         | more like math than creative expression.
         | 
         | I think the industry will realize that it made a huge mistake
         | by leaning on copyright for protection rather than on patents.
        
         | joshmoody24 wrote:
         | IMO the core idea of copyright isn't nonsense, but I do think
         | the current implementation (70+ years after death) is
         | egregiously overpowered. I've always thought the current laws
         | were too deeply entrenched to ever change, but I'm tentatively
         | optimistic AI will shock the system hard enough to trigger
         | actual reform.
        
         | mbgerring wrote:
         | Actually I think the last 20 years of the Internet demonstrates
         | that copyright is more important than ever, because unless it's
         | enforced, people with more capital than the copyright owner
         | will simply steal creative works and profit from them.
         | 
         | The idea that "information wants to be free" was always a lie,
         | meant to transfer value from creators to platform owners. The
         | result of that has been disastrous, and it's long past time to
         | push the pendulum in the other direction.
        
       | logicprog wrote:
       | > Ronacher notes this as an irony and moves on. But the irony
       | cuts deeper than he lets on. Next.js is MIT licensed.
       | Cloudflare's vinext did not violate any license--it did exactly
       | what Ronacher calls a contribution to the culture of openness,
       | applied to a permissively licensed codebase. Vercel's reaction
       | had nothing to do with license infringement; it was purely
       | competitive and territorial. The implicit position is:
       | reimplementing GPL software as MIT is a victory for sharing, but
       | having our own MIT software reimplemented by a competitor is
       | cause for outrage. This is what the claim that permissive
       | licensing is "more share-friendly" than copyleft looks like in
       | practice. The spirit of sharing, it turns out, runs in one
       | direction only: outward from oneself.
       | 
       | This argument makes no sense. Are they arguing that because
       | Vercel, specifically, had this attitude, this is an attitude
       | necessitated by AI, reimplementation, and those who are in favor
       | of it towards more permissive licenses? That certainly doesn't
       | seem to be an accurate way to summarize what antirez or Ronacher
       | believe. In fact, under the legal and ethical frameworks
       | (respectively) that those two put forward, Vercel has no right to
       | claim that position and no way to enforce it, so it seems very
       | strange to me to even assert that this sort of thing would be the
       | _practical_ result of AI reimplementations. This seems to just be
       | pointing towards the hypocrisy of one particular company, and
       | assuming that this would be the inevitable universal, attitude,
       | and result when there 's no evidence to think so.
       | 
       | It's ironic, because antirez actually literally addresses this
       | specific argument. They completely miss the fact that a lot of
       | his blog post is not actually just about legal but also about
       | ethical matters. Specifically, the idea he puts forward is that
       | yes, corporations can do these kinds of rewrites now, but they
       | always had the resources and manpower to do so anyway. What's
       | different now is that individuals can do this kind of rewrites
       | when they never have the ability to do so before, and the vector
       | of such a rewrite can be from a permissive to copyleft or even
       | from decompile the proprietary to permissive or copyleft. The
       | fact that it hasn't been so far is a more a factor of the fact
       | that most people really hate copyleft and find an annoying and
       | it's been losing traction and developer mind share for decades,
       | not that this tactic can't be used that way. I think that's
       | actually one of the big points he's trying to make with his GNU
       | comparison -- not just that if it was legal for GNU to do it,
       | then it's legal for you to do with AI, and not even just the
       | fundamental libertarian ethical axiom (that I agree with for the
       | most part) that it should remain legal to do such a rewrite in
       | either direction because in terms of the fundamental axioms that
       | we enforce with violence in our society, there should be a level
       | playing field where we look at the action itself and not just
       | whether we like or dislike the consequences, but specifically the
       | fact that if GNU did it once with the ability to rewrite things,
       | it can be done again, even in the same direction, it now even
       | more easily using AI.
        
         | antirez wrote:
         | > They completely miss the fact that a lot of his blog post is
         | not actually just about legal but also about ethical matters.
         | 
         | Honestly I was confused about the summarization of my blog post
         | into just a legal matter as well. I hope my blog post will be
         | able to flash at least a short time in the HN front page so
         | that the actual arguments it contain will get a bit more
         | exposure.
        
         | Talanes wrote:
         | I'm failing to see what in the quoted text you took to be about
         | AI rewrites specifically? It just reads as a slightly catty
         | aside about the social reaction of rewrites in general (by
         | implying the one example is generalizable.)
        
       | throwaway2027 wrote:
       | Perhaps software patents may play an even bigger role in the
       | future.
        
         | intrasight wrote:
         | Or, hopefully, even less of a role.
        
       | drnick1 wrote:
       | It should be noted that the Rust community is also guilty of
       | something similar. That is, porting old GPL programs, typically
       | written in C, to Rust and relicensing them as MIT.
        
         | wolvesechoes wrote:
         | > porting old GPL programs, typically written in C, to Rust and
         | relicensing them as MIT
         | 
         | Everything for memory safety.
        
       | nicole_express wrote:
       | Not a lawyer, but my understanding is: In theory, copyright only
       | protects the creative expression of source code; this is the
       | point of the "clean room" dance, that you're keeping only the
       | functional behavior (not protected by copyright). Patents are, of
       | course, an entirely different can of worms. So using an LLM to
       | strip all of the "creative expression" out of source code but
       | create the same functionality feels like it could be equivalent
       | enough.
       | 
       | I like the article's point of legal vs. legitimate here, though;
       | copyright is actually something of a strange animal to use to
       | protect source code, it was just the most convenient pre-existing
       | framework to shove it in.
        
         | dathinab wrote:
         | > this is the point of the "clean room" dance
         | 
         | which is the actual relevant part: they didn't do that dance
         | AFIK
         | 
         | AI is a tool, they set it up to make a non-verbatim copy of a
         | program.
         | 
         | Then they feed it the original software (AFIK).
         | 
         | Which makes it a side by side copy, as in the original source
         | was used as reference to create the new program. Which tend to
         | be seen as derived work even if very different.
         | 
         | IMHO They would have to:
         | 
         | 1. create a specification of the software _without looking at
         | the source code_, i.e. by behavior observation (and an
         | interface description). I.e. you give the AI access to running
         | the program, but not to looking into the insides of it. I
         | really don't think they did it as even with AI it's a huge pain
         | as you normally can't just brute force all combinations of
         | inputs and instead need to have a scientific
         | model=>test=>refine loop (which AI can do, but can take long
         | and get stuck, so you want it human assisted, and the human
         | can't have inside knowledge about the program).
         | 
         | 2. then generate a new program from specification, And only
         | from it. No git history, no original source code access, no
         | program access, no shared AI state or anything like that.
         | 
         | Also for the extra mile of legal risk avoidance do both human
         | assisted and use unrelated 3rd parties without inside knowledge
         | for both steps.
         | 
         | While this does majorly cut cost of a clean room approach, it
         | still isn't cost free. And still is a legal mine field if done
         | by a single person, especially if they have enough familiarity
         | to potentially remember specific peaces of code verbatim.
        
           | nicole_express wrote:
           | Well sure they didn't do the dance, but you don't have to do
           | the dance. The reason to do it is that it's a good defense in
           | a lawsuit. Like you say, all of this is a legal minefield.
           | 
           | So my understanding was that the original code was
           | specifically not fed into Claude. But was almost certainly
           | part of its training data, which complicates things, but if
           | that's fair use then it's not relevant? If training's not
           | fair use and taints the output, then new-chardet is a
           | derivative of a lot of things, not just old-chardet...
           | 
           | This is all new legal ground. I'm not sure if anyone will go
           | to court over chardet, though, but something that's an actual
           | money-maker or an FSF flagship project like readline, on the
           | other hand, well that's a lot more likely.
        
           | RaffaelCH wrote:
           | > Then they feed it the original software (AFIK).
           | 
           | My understanding is they did do the dance. From the article:
           | "He fed only the API and the test suite to Claude and asked
           | it to reimplement the library from scratch."
           | 
           | One could still make the argument that using the test suite
           | was a critical contributing factor, but it is not a part of
           | the resulting library. So in my uninformed opinion, it seems
           | to me like the clean room argument does apply.
        
       | grahamlee wrote:
       | It's clear that we're entering a new era of copyright
       | _expectations_ (whether we get new _legislation_ is different),
       | but for now realise this: the people like me who like copyleft
       | can do this too. We can take software we like, point an agent at
       | it, and tell it to make a new version with the AGPL3.0-or-later
       | badge on the front.
        
         | anonymous_sorry wrote:
         | But the LLM contributions would likely be ruled public domain,
         | so AGPL may not be enforceable on these.
        
         | armchairhacker wrote:
         | The point of GPL is to restrict distribution. If there's
         | already an MIT version, it's useless.
        
           | grahamlee wrote:
           | but the point of an EULA is to restrict distribution, so
           | AGPL3 can help there.
        
           | enriquto wrote:
           | > The point of GPL is to restrict distribution.
           | 
           | no, it isn't. The point of the GPL is to grant users of the
           | software four basic freedoms (run, study, modify and
           | _redistribute_ ). There's no restriction to distribution per
           | se, other than disallowing the removal of these freedoms to
           | other users.
        
       | largbae wrote:
       | This is only worth arguing about because software has value.
       | Putting this in context of a world where the cost of writing code
       | is trending to 0, there are two obvious futures:
       | 
       | 1. The cost continues to trend to 0, and _all_ software loses
       | value and becomes immediately replaceable. In this world,
       | proprietary, copyleft and permissive licenses do not matter, as I
       | can simply have my AI reimplement whatever I want and not
       | distribute it at all.
       | 
       | 2. The coding cost reduction is all some temporary mirage, to be
       | ended soon by drying VC money/rising inference costs, regulatory
       | barriers, etc. In that world we should be reimplementing
       | everything we can as copyleft while the inferencing is good.
        
         | anonymous_sorry wrote:
         | There was a recent ruling that LLM output is inherently public
         | domain (presumably unless it infringes some existing
         | copyright). In which case it's not possible to use them to
         | "reimplement everything we can as copyleft".
        
           | dathinab wrote:
           | it's more complicated, the ruling was that AI can't be an
           | author and the thing in question is (de-facto) public domain
           | because it has no author in context of the "dev" claim it was
           | fully build by AI
           | 
           | but AI assisted code has an author and claiming it's AI
           | assisted even if it is fully AI build is trivial (if you
           | don't make it public that you didn't do anything)
           | 
           | also some countries have laws which treat it like a tool in
           | the sense that the one who used it is the author by default
           | AFIK
        
         | sarchertech wrote:
         | There's an other option. The cost of copying existing software
         | trends to 0, but the cost of writing new software stays far
         | enough above 0 that it is still relatively expensive.
        
         | casey2 wrote:
         | The value of software has never been tied to the cost of
         | writing it, even if you don't distribute it your still breaking
         | the law.
        
           | largbae wrote:
           | The article is proceeding from the premise that a
           | reimplementation is legal (but evil). To help my
           | understanding of your comment, do you mean:
           | 
           | 1. An LLM recreating a piece of software violates its
           | copyright and is illegal, in which case LLM output can never
           | be legally used because someone somewhere probably has a
           | copyright on some portion of any software that an LLM could
           | write.
           | 
           | 2. You read my example as "copying a project without
           | distributing it", vs. "having an LLM write the same
           | functionality just for me"
        
         | beepbooptheory wrote:
         | There will always be cost though. Even if perfect code is
         | getting one-shotted out, that is constantly maintained and
         | adapted to changing conditions and technology, it simply can't
         | _stay_ at 0 forever because one day the power is surely going
         | to go out!
         | 
         | More and more I am drawn to these kinds of ideas lately,
         | perhaps as a kind of ethical sidestep, but still:
         | 
         | - https://wiki.xxiivv.com/site/permacomputing.html
         | 
         | - https://permacomputing.net/
         | 
         | It's not going to solve any general issue here, but the one
         | thing these freaks _need_ that can 't be generated by their
         | models is energy, tons of it. So, the one thing I can do as an
         | individual and in my (digital) community is work to be, in a
         | word, self-sustainable. And depending on my company I guess, if
         | I was a CEO I would hope I was wise enough to be thinking on
         | the same lines.
         | 
         | Everyone is making beautiful mountains from paper and wire. I
         | will just be happy to make a small dollhouse of stone, I think
         | it will be worth it. How can we see not just at least some
         | small-level of hubris otherwise?
        
       | t43562 wrote:
       | Why does anyone need his new library? They can do what he did and
       | make their own.
       | 
       | I'm glad we can fork things at a point and thumb our noses at
       | those who wish to cash in on other's work.
        
         | warkdarrior wrote:
         | Why would I make my own? The new library is released under MIT
         | license and faster than the old one.
        
           | t43562 wrote:
           | If you decide to improve it in any way to fit your needs you
           | can merely tell your own AI to re-implement it with your
           | changes. Then it's proprietary to you.
        
       | righthand wrote:
       | I think what is happening is the collapse of the "greater good".
       | Open source is dependent upon providing information for the
       | greater good and general benefit of its readers. However now that
       | no one is reading anything, its purpose is for the great good of
       | the most clever or most convincing or richest harvester.
        
       | sayrer wrote:
       | I don't think this part is correct: "If you distribute modified
       | code, _or offer it as a networked service_ , you must make the
       | source available under the same terms."
       | 
       | That's what something like AGPL does.
        
       | kazinator wrote:
       | You can't put a copyright and MIT license on something you
       | generated with AI. It is derived from the work of many unknown,
       | uncredited authors.
       | 
       | Think about it; the license says that copies of the work must be
       | reproduced with the copyright notice and licensing clauses
       | intact. Why would anyone obey that, knowing it came from AI?
       | 
       | Countless instances of such licenses were ignored in the training
       | data.
        
         | moralestapia wrote:
         | Courts have already ruled that AI-generated work belongs to the
         | public domain. So, even the MIT license does not apply.
        
         | harshreality wrote:
         | When learning is sufficiently atomized and recombined,
         | creations cease to be "derived from" in a legal sense.
         | 
         | A lego sculpture is copyrighted. Lego blocks are not. The
         | threshold between blocks and sculpture is not well-defined, but
         | if an AI isn't prompted specifically to attempt to mimic an
         | existing work, its output will be safely on the non-copyrighted
         | side of things.
         | 
         | A derivative work is separately copyrightable, but
         | redistribution needs permission from the original author too.
         | Since that usually won't be granted or would be uneconomical,
         | the derivative work can't usually be redistributed.
         | 
         | AI-produced material is inherently not copyrightable, but not
         | because it's a derivative work.
        
           | kazinator wrote:
           | Token prediction is a form of "learning" that is reinforced
           | by the goal of reproducing the correct next token of the
           | work, rather that acquiring ideas and concepts. For instance,
           | given the prefix "Four score and seven years", the weights
           | are adjusted until "ago" is correctly predicted, which is a
           | fancy way of saying that it was stored in the model in a
           | lossy way. The model "learned" that "ago" follows "four score
           | and seven years" exactly the way your hard drive "learns" the
           | audio and video frames of a movie when you download a .mp4
           | file.
           | 
           | I dispute the idea that token sequences reproduced from the
           | model are not derived works.
           | 
           | I predict, no pun intended, that a time is coming when the
           | idea that it's not a derived work will be challenged in
           | mainstream law.
           | 
           | The slop merchants are getting a free ride for the time
           | being.
        
             | harshreality wrote:
             | Only because that quote is famous.
             | 
             | As you said, it's lossy. Try it with any other distinctive
             | but non-famous passage, and you won't get a correct
             | prediction for the immediately following clause, much less
             | for multiple sentences or paragraphs.
             | 
             | That's the case _even when an LLM correctly identifies
             | which book the prompted text is from._ It still won 't
             | accurately continue on from some arbitrary passage. By the
             | time you ask it to reproduce hundreds of words, you're into
             | brand new book territory. Even when it's slop content, it's
             | distinct slop.
             | 
             | The exceptions are cases where a significant number of
             | humans would also know a particular quote from memory.
             | Then, chances are, a frontier LLM will too.
             | 
             | You know how else you can reproduce a quote? Search for it
             | on google, and search the resulting top hits; if it's a
             | significant quote, multiple people have probably quoted it
             | -- _legally_. You can also search a pirate library for the
             | actual book, and search the book for the quote; while
             | illegal, it 's very simple to do, so unless you propose to
             | make the free and open internet illegal, I'd suggest that
             | banning LLMs for being "derivative work" creation engines
             | is not so different from destroying the internet.
             | 
             | > I predict, no pun intended, that a time is coming when
             | the idea that it's not a derived work will be challenged in
             | mainstream law.
             | 
             | If judges have any sense whatsoever, LLM generations
             | (without specific prompt crafting to mimic existing works)
             | will be judged to not be derived works and therefore not be
             | violating copyright, in the same sense that you can live
             | and breathe Taylor Swift's music, create new music in the
             | same style, and still not be violating copyright.
             | 
             | More likely, perhaps, is that everything will be so infused
             | with LLM output that copyright ceases to be relevant, or
             | forces copyright law to be rewritten from the ground up.
        
       | skybrian wrote:
       | Broadly speaking, the "freedom of users" is often protected by
       | competition from competing alternatives. The GNU command line
       | tools were replacements for system utilities. Linux was was a
       | replacement for other Unix kernels. People chose to install them
       | instead of proprietary alternatives. Was it due to ideology or
       | lower cost or more features? All of the above. Different users
       | have different motivations.
       | 
       | Copyleft could be seen as an attempt to give Free Software an
       | edge in this competition for users, to counter the increased
       | resources that proprietary systems can often draw on. I think
       | success has been mixed. Sure, Linux won on the server. Open
       | source won for libraries downloaded by language-specific package
       | managers. But there's a long tail of GPL apps that are not really
       | all that appealing, compared to all the proprietary apps
       | available from app stores.
       | 
       | But if reimplementing software is easy, there's just going to be
       | a lot more competition from both proprietary and open source
       | software. Software that you can download for free that has better
       | features and is more user-friendly is going to have an advantage.
       | 
       | With coding agents, it's likely that you'll be able to modify
       | apps to your own needs more easily, too. Perhaps plugin systems
       | and an AI that can write plugins for you will become the norm?
        
         | jacquesm wrote:
         | > Was it due to ideology or lower cost or more features?
         | 
         | It was due to _access_.
        
       | casey2 wrote:
       | If the model wasn't trained on copyleft, if he didn't use a
       | copyleft test suite and if he wasn't the maintainer for years.
       | Clearly the intent here is copyright infringement.
       | 
       | If you have software your testsuite should be your testsuite, you
       | do dev with a testsuite and then mit without releasing one.
       | Depending on the test-suite it may break clean room rules,
       | especially for ttd codebases.
        
       | strongpigeon wrote:
       | I feel like the licenses that suffer the most isn't the GPL, but
       | the ones like SSPL. If your code can be re-implemented easily and
       | legally by AWS using an LLM, why risk publishing it?
       | 
       | It does feel like open source is about to change. My hunch is
       | that commercial open source (beyond the consultation model) risks
       | disappearing. Though I'd be happy to be proven wrong.
        
       | kccqzy wrote:
       | > When GNU reimplemented the UNIX userspace, the vector ran from
       | proprietary to free. Stallman was using the limits of copyright
       | law to turn proprietary software into free software. [...] The
       | vector in the chardet case runs the other way.
       | 
       | That's just your subjective opinion which many other people would
       | disagree. I bet Armin Ronacher would agree that an MIT licensed
       | library is even freer than an LGPL licensed library. To them, the
       | vector is running from free to freer.
        
       | bjt wrote:
       | > If source code can now be generated from a specification, the
       | specification is where the essential intellectual content of a
       | GPL project resides. Blanchard's own claim--that he worked only
       | from the test suite and API without reading the source--is,
       | paradoxically, an argument for protecting that test suite and API
       | specification under copyleft terms.
       | 
       | This is an interesting reversal in itself. If you make the
       | specification protected under copyright, then the whole practice
       | of clean room implementations is invalid.
        
       | dleslie wrote:
       | IMHO, the API and Test Suite, particularly the latter, define the
       | contract of the functional definition of the software. It almost
       | doesn't matter what that definition looks like so long as it
       | conforms to the contract.
       | 
       | There was an issue where Google did something similar with the
       | JVM, and ultimately it came down to whether or not Oracle owned
       | the copyright to the header files containing the API. It went all
       | the way to the US supreme court, and they ruled in Google's
       | favour; finding that the API wasn't the implementation, and that
       | the amount of shared code was so minimal as to be irrelevant.
       | 
       | They didn't anticipate that in less than half a decade we'd have
       | technology that could _rapidly_ reimplement software given a
       | strong functional definition and contract enforcing test suite.
        
       | mwkaufma wrote:
       | A lot of untagged IANAL takes here today.
        
       | ineedasername wrote:
       | This article is setting up a bit of a moving target. Legal vs
       | legitimate is at least only a single vague question to be defined
       | but then the target changes to "socially legitimate" defined only
       | indirectly by way of example, like aggressive tax avoidance as
       | "antisocial"-- and while I tend to agree with that
       | characterization my agreement is predicated on a layering of
       | other principals.
       | 
       | The fundamental problem is that once you take something outside
       | the realm of law and rule of law in its many facets as the
       | legitimizing principal, you have to go a whole lot further to be
       | coherent and consistent.
       | 
       | You can't just leave things floating in a few ambiguous things
       | you don't like and feel "off" to you in some way- not if you're
       | trying to bring some clarity to your own thoughts, much less
       | others. You don't _have to land on a conclusion_ either. By all
       | means chew over things, but once you try to settle, things fall
       | apart if you haven't done the harder work of replacing the
       | framework of law with that of another conceptual structure.
       | 
       | You need to at least be asking "to what ends? What purpose is
       | served by the rule?" Otherwise you're stuck in things where half
       | the time you end up arguing backwards in ways that put purpose
       | serving rules, the maintenance of the rule with justifications
       | ever further afield pulled in when the rule is questioned and
       | edge cases reached. If you're asking, essentially, "is the spirit
       | of the rule still there?" You've got to stop and fill in what
       | that spirit is or you or people that want to control you or have
       | an agenda will sweep in with their own language and fill the void
       | to their own ends.
        
       | kelseyfrog wrote:
       | In the corporate world, we've started using reimplementation as a
       | way to access tooling that security won't authorize.
       | 
       | Sec has a deny by default policy. Eng has a use-more-AI policy.
       | Any code written in-house is accepted by default. You can see
       | where this is going.
       | 
       | We've been using AI to reimplement tooling that security won't
       | approve. The incentives conspired in the worst outcome, yet here
       | we are. If you want a different outcome, you need to create
       | different incentives.
        
         | kemitchell wrote:
         | Not Invented Here's long, slow mutagenic march toward full
         | antibiotic resistance continues apace.
         | 
         | There is a fundamental corpo-cognitive dissonance, to boot. If
         | "AI" is cheap enough and good enough to implement security-
         | relevant software from `git init` repeatedly, why isn't it also
         | cheap enough and good enough to assess and approve the security
         | of third-party software at pace with internal adoption? Is
         | there some basis to believe LLMs' leverage on production
         | differs from its leverage on analysis of existing code?
        
       | ticulatedspline wrote:
       | Surprised they don't mention Google LLC v. Oracle America, Inc.
       | Seems a bit myopic to condone the general legality while arguing
       | "you can only use it how I like it".
       | 
       | It also doesn't talk about the far more interesting philosophical
       | queston. Does what Blanchard did cover ALL implementations from
       | Claude? What if anyone did exactly what he did, feed it the test
       | cases and say "re-implement from scratch", ostensibly one would
       | expect the results to be largely similar (technically under the
       | right conditions deterministically similar)
       | 
       | could you then fork the project under your own name and a
       | commercial license? when you use an LLM like this, to basically
       | do what anyone else could ask it to do how do you attach any
       | license to it? Is it first come first serve?
       | 
       | If an agent is acting mostly on its own it feels like if you
       | found a copy of Harry Potter in the fictional library of Babel,
       | you didn't write it, just found it amongst the infinite library,
       | but if you found it first could you block everyone else that
       | stumbles on a near-identical copy elsewhere in the library? or
       | does each found copy represent a "Re-implementation" that could
       | be individually copyrighted?
        
       | danbruc wrote:
       | Why are people even having problems with sharing their changes to
       | begin with? Just publishing it somewhere does not seem too
       | expensive. The risk of accidentally including stuff that is not
       | supposed to become public? Or are people regularly completely
       | changing codebases and do not want to make the effort freely
       | available, maybe especially to competitors? I would have assumed
       | that the common case is adding a missing feature here, tweaking
       | something there, if you turn the entire thing on its head, why
       | not have your own alternative solution from scratch?
        
       | PaulDavisThe1st wrote:
       | If Blanchard is claiming not to have been substantively involved
       | in the creation of the new implementation of chardet (i.e.
       | "Claude did it"), then the new implementation is machine
       | generated, and in the USA cannot be copyright and thus cannot be
       | licensed.
       | 
       | If he is claiming to have been somehow substantively "enough"
       | involved to make the code copyrightable, then his own familiarity
       | with the previous LGPL implementation makes the new one almost
       | certainly a derivative of the original.
        
         | sigmar wrote:
         | >then his own familiarity with the previous LGPL implementation
         | makes the new one almost certainly a derivative of the
         | original.
         | 
         | The "clean room rewrite" is just an extreme way to have a
         | bulletproof shield against litigation. Not doing it that way
         | doesn't automatically make all new code he writes derivative
         | solely because he saw how the code worked previously.
        
           | PaulDavisThe1st wrote:
           | If the clean room re-write was done entirely by Claude, then
           | the result cannot be copyright in the USA, and thus there is
           | no license at all.
           | 
           | And if he was in fact more involved (which he appears to
           | deny) that it's a bit weak to say that someone with huge
           | familiarity with chardet could choose to reimplement chardet
           | without the result being derivative.
        
             | serial_dev wrote:
             | So if I read any LGPL code in my life, I can never think
             | about working on something similar in my life?
        
               | vbarrielle wrote:
               | There's a difference between "I've read a LGPL code once,
               | maybe I could do something similar" and "I've been
               | reading this LGPL code for 12 years and now I'm going to
               | do exactly the same thing".
        
       | hexyl_C_gut wrote:
       | I'm less concerned about AI eroding copyleft and more exited
       | about AI eroding copy right.
        
       | AndriyKunitsyn wrote:
       | There's a Japanese version of that page, written in classical
       | text writing direction, in columns. Which is cool. Makes me
       | wonder, though - how readable is it with so many English
       | loanwords which should be rotated sideways to fit into columns?
        
         | ddellacosta wrote:
         | Total digression but yeah, that layout is stupid and the way
         | those words are dropped in using Romaji makes no sense. That's
         | not how Japanese people lay out pages on the web. In fact I
         | don't think I've ever seen a Japanese web page laid out like a
         | book like this, and in general I'd expect the English proper
         | nouns and words that don't have obvious translations to get
         | transliterated into Katakana. Smells like automatic conversion
         | added by someone not really familiar with common practices for
         | presenting Japanese on the web.
        
           | tmp10423288442 wrote:
           | He also has a Korean vertical layout that lays out Latin-
           | character words the same way. Is this common in Korea when
           | vertical layout is used? The author seems to be Korean.
           | 
           | Looks like Wikipedia has an example of Traditional Chinese
           | vertical layout with the Latin letters rotated as in TFA's
           | layout (https://en.wikipedia.org/wiki/Horizontal_and_vertical
           | _writin...)
        
       | Khaine wrote:
       | Someone be brave, and do this to ZFS. Poke the Oracle bear!
        
       | mh2266 wrote:
       | Buried in here: Mark Pilgrim suddenly reappearing after his
       | sudden disappearance years ago! Has he been up to anything since
       | then?
        
       | moralestapia wrote:
       | That's a non-sequitur. chardet v7 is GPL-derived work (currently
       | in clear violation of the GPL). If xe wanted it to be a different
       | thing xe should've published as such. Simple as.
        
       | svilen_dobrev wrote:
       | i've been following this for a while.. and the trend for
       | copyright (of any form - books code pictures music whatever)
       | being laundered by reinventing the "same" thing in-some-way.. is
       | kind-of clear.
       | 
       | But what happens with the _new_ things? Has the era of software-
       | making (or creating things at large) finished, and from now on
       | everything will be re-(gurgitated|implemented|polished) old
       | stuff?
       | 
       | Or all goes back to proprietary everything.. Babylon-tower style,
       | noone talks to noone?
       | 
       | edit: another view - is open-source from now on only for resume-
       | building? "see-what-i've-built" style
        
       | t43562 wrote:
       | It seems that this chap didn't go and implement a new library, he
       | reimplemented an existing one and became sole-controller of it.
       | i.e. he seems to have taken its reputation, brand whatever you
       | call it away from the contributors and entirely to himself. Their
       | work of establishing it as a well known solution is no longer
       | recognised.
       | 
       | So of course we feel that something wrong has happened even if
       | it's not easy to put one's finger on it.
        
       | zmmmmm wrote:
       | The really interesting question to me is if this transcends
       | copyright and unravels the whole concept of intellectual
       | property. Because all of it is premised on an assumption that
       | creativity is "hard". But LLMs are not just writing software,
       | they are rapidly being engineered to operate completely generally
       | as knowledge creation engines: solving math proofs, designing
       | drugs, etc.
       | 
       | So: once it's not "hard" any more, does IP even make sense at
       | all? Why grant monopoly rights to something that required little
       | to no investment in the first place? Even with vestigial IP law -
       | let's say, patents: it just becomes and input parameter that the
       | AI needs to work around the patents like any other constraints.
        
         | palmotea wrote:
         | > So: once it's not "hard" any more, does IP even make sense at
         | all? Why grant monopoly rights to something that required
         | little to no investment in the first place? Even with vestigial
         | IP law - let's say, patents: it just becomes and input
         | parameter that the AI needs to work around the patents like any
         | other constraints.
         | 
         | I think it still does: IIRC, the current legal situation is AI-
         | output does not qualify for IP protections (at least not
         | without substantial later human modification). IP protections
         | are solely reserved for human work.
         | 
         | And I'm fine with that: if a person put in the work, they
         | _should_ have protections so their stuff can 't be ripped off
         | for free by all the wealthy major corporations that find some
         | use for it. Otherwise: who cares about the LLMs.
        
           | robmccoll wrote:
           | I think you have a rather idealized model of IP in mind. In
           | practice, IP law tends to be an expensive weapon the wealthy
           | major corporations use against the little guy. Deep enough
           | pockets and a big enough warchest of broad parents will drain
           | the little guy every time.
        
             | palmotea wrote:
             | > In practice, IP law tends to be an expensive weapon the
             | wealthy major corporations use against the little guy. Deep
             | enough pockets and a big enough warchest of broad parents
             | will drain the little guy every time.
             | 
             | Then fix that instead of blowing it up. Because IP law is
             | also literally the _only_ thing that protects the little
             | guy 's work in many cases.
             | 
             | Arguments like yours are kinda unfathomably incomplete to
             | me, almost like they're the remnants of some propaganda
             | campaign. It's constructed to appeal to the defense of the
             | little guy, but the actual effect would be to disempower
             | him and further empower the wealthy major corporations with
             | "big enough warchest[s]."
             | 
             | I mean, one thing I think the RIAA _would love_ is to stop
             | paying royalties to every artist ever. And the only thing
             | they 'd be worried about is an even bigger fish (like
             | Amazon, Apple, or Spotify) no longer paying royalties to
             | them. But as you said, they have a big enough war chest
             | that they probably could force a deal somehow. All the
             | artists without a war chest? Left out in the cold.
        
               | _aavaa_ wrote:
               | Blowing up IP would sink the RIAA. They would no longer
               | have legal grounds to go after file sharing, and I'm
               | confident that given the same legal footing that file
               | sharing would win any day of the week.
        
           | jbergqvist wrote:
           | Does this matter in practice though? By modifying some of the
           | generated code and not taking a solution produced by an LLM
           | end-to-end but borrowing heavily from it, can't a human claim
           | full ownership of the IP even though in reality the LLM did
           | most of the relevant work?
        
             | reverius42 wrote:
             | I think as long as the human puts in substantial and
             | transformational effort, they can claim to be the copyright
             | holder of the entire work, yes.
        
           | nkmnz wrote:
           | > AI-output does not qualify for IP protections
           | 
           | I beg to differ. AI-output did not entitle the person
           | creating the prompt for IP protections, so far - but my
           | objection is not directed towards the "so far", but towards
           | your omission of "the person creating the prompt", because if
           | an AI outputs copyrighted material from the training data,
           | that material is still copyrighted. AI is not a magical
           | copyright removal machine.
        
             | reverius42 wrote:
             | The U.S. Supreme Court just declined to hear a case, thus
             | upholding a lower court precedent that LLM output are not
             | copyrightable: https://www.reuters.com/legal/government/us-
             | supreme-court-de...
             | 
             | What this means in practice is that (currently), all output
             | of an LLM is legally considered to not be copyrightable (to
             | the extent that it's an original work). If it happens to
             | regurgitate an existing copyrighted work, though, is that
             | infringement? I'm not sure we have a legal precedent on
             | that question yet.
        
           | rlpb wrote:
           | What if a person puts in the work, but the work was worthless
           | or can be trivially reproduced without effort?
           | 
           | See also: https://en.wikipedia.org/wiki/Sweat_of_the_brow
        
         | spwa4 wrote:
         | Don't worry. The courts have consistently sided with huge
         | companies on copyright. In the US. In Europe. Doesn't matter.
         | 
         | Company incorporates GPL code in their product? Never once have
         | courts decided to uphold copyright. HP did that many times.
         | Microsoft got caught doing it. And yet the GPL was never
         | applied to their products. Every time there was an excuse. An
         | inconsistent excuse.
         | 
         | Schoolkid downloads a movie? 30,000 USD per infraction PLUS
         | armed police officer goes in and enforces removal of any
         | movies.
         | 
         | Or take the very subject here. AI training WAS NOT considered
         | fair use when OpenAI violated copyright to train. Same with
         | Anthropic, Google, Microsoft, ... They incorporated harry
         | potter and the linux kernel in ChatGPT, in the model itself.
         | Undeniable. Literally. So even if you accept that it's changed
         | now, OpenAI should still be forced to redistribute the training
         | set, code, and everything needed to run the model for
         | everything they did up to 2020. Needless to say ... courts
         | refused to apply that.
         | 
         | So just apply "the law", right. Courts' judgement of using AI
         | to "remove GPL"? Approved. Using AI to "make the next Disney-
         | style movie"? SEND IN THE ARMY! Whether one or the other
         | violates the law according to rational people? Whatever excuse
         | to avoid that discussion is good enough.
        
         | js8 wrote:
         | It might unravel intellectual property, just not in a fair way.
         | When capitalism started, public land was enclosed to create
         | private property. Despite this being in many cases a quite
         | unfair process, we still respect this arrangement.
         | 
         | With AI, a similar process is happening - publicly available
         | information becomes enclosed by the model owners. We will
         | probably get a "vestigial" intellectual property in the form of
         | model ownership, and everyone will pay a rent to use it. In
         | fact, companies might start to gatekeep all the information to
         | only their own LLM flavor, which you will be required to use to
         | get to the information. For example, product documentation and
         | datasheets will be only available by talking to their AI.
        
         | nradov wrote:
         | Nothing changes for drug patents regardless of whether an LLM
         | was used in the discovery process.
        
           | reverius42 wrote:
           | Not sure why this should be true; the US Supreme Court
           | recently chose to let precedent stand that AI creations are
           | not copyrightable.
           | https://www.reuters.com/legal/government/us-supreme-court-
           | de...
           | 
           | That also seems relevant for this whole discussion, actually
           | -- if a work can't be copyrighted it certainly can't have a
           | changed license, or any license at all. (I guess it's
           | effectively public domain to the extent that it's public at
           | all?)
        
             | nradov wrote:
             | You're really missing the point in multiple ways. First,
             | precedents on copyright law are irrelevant to patent law.
             | Second, AI generated works generally can be copyrighted
             | under the human creator's name.
        
               | reverius42 wrote:
               | No, I think you are quite incorrect, at least on the
               | latter point:
               | 
               | "Lower courts upheld a U.S. Copyright Office decision
               | that the AI-crafted visual art at issue in the case was
               | ineligible for copyright protection because it did not
               | have a human creator."
               | 
               | Not eligible for copyright protection does not mean it
               | can be copyrighted "under the human creator's name". It
               | means there is no creative work at all. No copyright.
        
               | reverius42 wrote:
               | And while courts in theory aren't supposed to apply
               | copyright precedent to patent cases, in practice, they
               | apparently do a lot of the time, so it's kind of a mess! 
               | https://scholarship.kentlaw.iit.edu/ckjip/vol16/iss1/4/#:
               | ~:t...
        
           | zmmmmm wrote:
           | Even if all I have to do is tell my agent, "here is a patent
           | for a drug, analyse the patent and determine an equivalent
           | but non-infringing drug" and it chugs away for a couple of
           | hours and spits out a drug along with all the specifications
           | to manufacture it?
           | 
           | I guess the state of play will be that for new drugs the
           | original manufacturer will _already have done that_ and
           | ensured that literally anything that could be found as a
           | workaround is included in the scope of the patent. But I feel
           | like it will not be possible to keep that wartertight.
        
             | nradov wrote:
             | Yes, even so. Human drug researchers have been doing the
             | same thing for decades. As soon as one pharmaceutical
             | company launches a successful small-molecule drug everyone
             | else jumps to find a minor tweak that will hit the same
             | target (ideally with fewer side effects) while evading the
             | patent. There is already specialized software to help with
             | this process so I'm skeptical that LLM agents would be very
             | helpful for this use case.
        
             | paxys wrote:
             | The formula is what is patented, not the process to come up
             | with it.
        
         | newyankee wrote:
         | If you think about creative outcomes as n dimensional
         | 'volumes', AI expressions can cover more than humans in many
         | domains. These are precisely artistic styles, music styles etc.
         | and tbh not everyone can be a Mozart but may be a lot more with
         | AI can be Mozart lite. This begs the question how much of
         | creativity is appreciated as a shared experience
        
         | hyperman1 wrote:
         | I've always thought the opposite: IP law was created to make
         | sure creativity stays hard, and hence controllable by the
         | elites.
         | 
         | Patents came along when farmers started making city goods,
         | threatening guilds secrets. Copyright came when the printing
         | press made copying and translating the bible easy and
         | accessible to all. (Trademark admittedly does not fit this
         | view, but doesn't seem all that damaging either)
         | 
         | To Protect The Arts, and To Time Limit Trade Secrets were just
         | the Protect The Children of old times, a way to confuse people
         | who didn't look too hard at actual consequences.
         | 
         | This means that the future of IP depends on what lets the
         | powers that be pull up the ladder behind them. Long term I'd
         | expect e.g. copyright expansion and harder enforcement, just
         | because cloning by AI gets easy enough to threaten the status
         | quo.
        
           | cobbzilla wrote:
           | > Trademark admittedly does not fit this view, but doesn't
           | seem all that damaging either
           | 
           | Isn't trademark the only thing keeping a certain cartoon
           | mouse out of the public domain, despite the fact that his
           | earliest animations are out of copyright? Not sure if you'd
           | consider that damaging, or if anyone has yet tested the
           | boundaries of the House of Mouse's patience here.
        
         | satvikpendem wrote:
         | Good. Intellectual property is now a twisted concept by the
         | elite, whatever its benefits were previously. As soon as Disney
         | made Mickey popular, it was all downhill.
        
         | rfw300 wrote:
         | More likely: this is a transitional phase where our previously
         | hard problems become easy, and we will soon set our sights on
         | new and much harder problems. The pinnacle of creative
         | achievement in the universe is probably not 2010s B2B SaaS.
         | 
         | It is entirely possible, however, that human beings will not be
         | the primary drivers of progress on those problems.
        
         | paxys wrote:
         | "Hard" or "easy" has never been part of the premise.
         | 
         | A company spends a decade and billions of dollars to develop a
         | groundbreaking drug and patents it.
         | 
         | I think of a cool new character called "Mr Poop" and publish a
         | short story about him with an hour of work.
         | 
         | Both of us get the exact same protection under the law (yes yes
         | I know copyright vs patent etc., but ultimately they are all
         | about IP protection).
        
       | mbgerring wrote:
       | See also "A Declaration of the Independence of Cyberspace"
       | (https://www.eff.org/cyberspace-independence), and what a goofy,
       | naive, misguided disaster that early internet optimism turned
       | into.
       | 
       | No, AI does not mean the end of either copyright or copyleft, it
       | means that the laws need to catch up. And they should, and they
       | will.
        
       | munk-a wrote:
       | I think the missing thing here is that the license violation
       | already happened. Most of the big models trained on data in a
       | manner that violated terms of service. We'll need a court case
       | but I think it's extremely reasonable to consider any model
       | trained on GPL code to be infected with open licensing
       | requirements.
        
         | NewsaHackO wrote:
         | I agree there has to be a court case about it. I think the
         | current argument, however, is that it is transformative, and
         | therefore falls under fair use.
        
           | munk-a wrote:
           | Yea, a finding that training is transformative would be
           | pretty significant and it's likely that the precedent of
           | thumbnail creation being deemed transformative would likely
           | steer us towards such a finding. Transformative is always a
           | hard thing to bank on because it is such a nebulous and
           | judgement based call. There are excellent examples of how
           | precise and gritty this can get in audio sampling.
        
             | NewsaHackO wrote:
             | Didn't know about thumbnails being fair use. In that case,
             | I just don't see an argument that genAI training on source
             | code is less transformative than thumbnails.
        
         | crazygringo wrote:
         | You might wish that were true, but there are very strong
         | arguments it's not. Training on copyleft licensed code is not a
         | license violation. Any more than a person reading it is. In
         | copyright terms, it's such an extreme transformative use that
         | copyright no longer applies. It's fair use.
         | 
         | But agreed that we're waiting for a court case to confirm that.
         | Although really, the main questions for any court cases are not
         | going to be around the principle of fair use itself or whether
         | training is transformative enough (it obviously is), but rather
         | on the specifics:
         | 
         | 1) Was any copyrighted material _acquired_ legally (not
         | applicable here), and
         | 
         | 2) Is the LLM always providing a unique expression (e.g. not
         | regurgitating books or libraries verbatim)
         | 
         | And in this particular case, they confirmed that the new
         | implementation is 98.7% unique.
        
           | gspr wrote:
           | > Training on copyleft licensed code is not a license
           | violation. Any more than a person reading it is.
           | 
           | Some might hold that we've granted persons certain
           | exemptions, on account of them being persons. We do not have
           | to grant machines the same.
           | 
           | > In copyright terms, it's such an extreme transformative use
           | that copyright no longer applies.
           | 
           | Has the model really performed an extreme transformation if
           | it is able to produce the training data near-verbatim? Sure,
           | it can _also_ produce extremely transformed versions, but is
           | that really relevant if it holds within it enough information
           | for a (near-)verbatim reproduction?
        
             | crazygringo wrote:
             | > _We do not have to grant machines the same._
             | 
             | No we don't _have_ to, but so far we do, because that 's
             | the most legally consistent. If you want to change that,
             | you're going to need to pass new laws that may wind up
             | radically redefining intellectual property.
             | 
             | > _Has the model really performed an extreme transformation
             | if it is able to produce the training data near-verbatim?_
             | 
             | Of course it has, if the _transformation_ is extreme, as it
             | appears to be here. If I memorize the lyrics to a bunch of
             | love songs, and then write my own love song where every
             | line is new, nobody 's going to successfully sue me just
             | because I can sing a bunch of other songs from memory.
             | 
             | Also, it's not even remotely clear that the LLM _can_
             | produce the training data near-verbatim. Generally it _can
             | 't_, unless it's something that it's been trained on with
             | high levels of repetition.
        
               | munk-a wrote:
               | I want to briefly pick at this:
               | 
               | > you're going to need to pass new laws that may wind up
               | radically redefining intellectual property
               | 
               | You're correct that this is one route to resolving the
               | situation, but I think it's reasonable to lean more
               | strongly into the original intent of intellectual
               | property laws to defend creative works as a manner to
               | sustain yourself that would draw a pretty clear
               | distinction between human creativity and reuse and LLMs.
        
             | NewsaHackO wrote:
             | >Has the model really performed an extreme transformation
             | if it is able to produce the training data near-verbatim?
             | Sure, it can also produce extremely transformed versions,
             | but is that really relevant if it holds within it enough
             | information for a (near-)verbatim reproduction?
             | 
             | I feel as though, from an information-theoretic standpoint,
             | it can't be possible that an LLM (which is almost certainly
             | <1 TB big) can contain any substantial verbatim portion of
             | its training corpus, which includes audio, images, and
             | videos.
        
           | Copyrightest wrote:
           | The big difference between people reading code and LLMs
           | reading code is that people have legal liability and LLMs do
           | not. You can't sue an LLM for copyright infringement, and
           | it's almost impossible for users to tell when it happens.
           | 
           | BTW in 2023 I watched ChatGPT spit out hundreds of lines of
           | F# verbatim from my own GitHub. A lot of people had this
           | experience with GitHub Copilot. "98.7% unique" is still a lot
           | of infringement.
        
             | satvikpendem wrote:
             | You can sue the company making the LLM, which is what many
             | have done.
        
             | crazygringo wrote:
             | > _people have legal liability and LLMs do not. You can 't
             | sue an LLM for copyright infringement_
             | 
             | That's not relevant, because you can still sue the person
             | using the LLM and publishing the repository. Legal
             | liability is completely unchanged.
        
           | madeofpalk wrote:
           | A human reading a unit of work is not a "copy". I'm pretty
           | sure our legal systems agree that thought or sight is not
           | copying something.
           | 
           | Training an LLM inherently requires making a copy of the
           | work. Even the initial act of loading it from the internet
           | and copying it into memory to then train the LLM is a copy
           | that can be governed by its license and copyright law
        
             | crazygringo wrote:
             | > _Training an LLM inherently requires making a copy of the
             | work._
             | 
             | But that's not relevant here. Because the copyleft license
             | does not prohibit that (and it's not even clear that _any_
             | license _can_ prohibit it, as courts may confirm it 's fair
             | use, as most people are currently assuming). That's why I
             | noted under (1) that it's not applicable here.
        
               | munk-a wrote:
               | It's absolutely prohibited to copy and redistribute for
               | commercial purposes materials that you're unlicensed to
               | do so with. This isn't an issue when it comes to the
               | copy-left scenario (though it may potentially enforce
               | transitive licensing requirements on the copier that LLM
               | runners don't want to follow) but it is a huge issue that
               | has come up with LLM training.
               | 
               | LLM training involves ingesting works (in a potentially
               | transformative process) and partially reproduce them -
               | that's a generally restricted action when it comes to
               | licensing.
        
               | crazygringo wrote:
               | > _It 's absolutely prohibited to copy and redistribute
               | for commercial purposes materials that you're unlicensed
               | to do so with._
               | 
               | Sure, but that's not what LLM's generally do, and it's
               | certainly not what they're intended to do.
               | 
               | The LLM companies, and many other people, argue that
               | training falls under _fair use_. One element of fair use
               | is whether the purpose /character is sufficiently
               | _transformative_ , and transforming texts into weights
               | without even a remote 1-1 correspondence is the
               | transformation.
               | 
               | And this is why LLM companies ensure that partial
               | reproduction _doesn 't_ happen during LLM usage, using a
               | kind of copyrighted-text filter as a last check in case
               | anything would unintentionally get through. (And it
               | doesn't even tend to occur in the first place, except
               | when the LLM is trained on a bunch of copies of the same
               | text.)
        
             | cortesoft wrote:
             | I think you are confusing two different meanings of the
             | word 'copy'. The fact that a computer loads it into memory
             | does not make it automatically a 'copy' in the copyright
             | sense.
        
       | arjie wrote:
       | Well, the license change sounds pretty strange, but to be honest
       | if I were to use this software I would use it without adhering to
       | the MIT. It's machine-created content which is not, in general,
       | copyrightable. You can assert whatever license you want on such
       | content, but I am not going to adhere to it. For example, I
       | declare you may use the following under the Elastic License
       | The
        
       | iberator wrote:
       | Easy solution for now:
       | 
       | Add something like this to NEW gpl /bsd/mit licenses:
       | 
       | 'you are forbidden from reimplementing it with AI'
       | 
       | or just:
       | 
       | 'all clones, reimpletetions with ai etc must still be GPL'
        
       | foresto wrote:
       | From the article:
       | 
       | > He fed only the API and the test suite to Claude and asked it
       | to reimplement the library from scratch.
       | 
       | From GPL2:
       | 
       | > The source code for a work means the preferred form of the work
       | for making modifications to it. For an executable work, complete
       | source code means all the source code for all modules it
       | contains, plus any associated interface definition files, plus
       | the scripts used to control compilation and installation of the
       | executable.
       | 
       | Is a project's test suite not considered part of its source code?
       | When I make modifications to a project, its test cases are very
       | much a part of that process.
       | 
       | If the test suite is part of this library's source code, and
       | Claude was fed the test suite or interface definition files, is
       | the output not considered a _work based on the library_ under the
       | terms of LGPL 2.1?
        
         | tty456 wrote:
         | Google v. Oracle ruled that use of APIs are fair game and could
         | be argued that test cases are strictly a use of APIs and not
         | implementation.
        
           | vbarrielle wrote:
           | Google vs Oracle ruled that APIs fall under copyright (the
           | contrary was thought before). However, it was ruled that, in
           | that specific case, fair use applied, because of
           | interoperability concerns. That's the important part of this
           | case: fair use is never automatic, it is assessed case by
           | case.
           | 
           | Regarding chardet, I'm not sure "I wanted to circumvent the
           | license" is a good way to argue fair use.
        
         | crazygringo wrote:
         | It's transformative, so no.
         | 
         | Legally, using the tests to help create the reimplementation is
         | fine.
         | 
         | However, it seems possible you can't _redistribute_ the same
         | tests under the MIT license. So the reimplementation MIT
         | distribution could need to be source code only, not source code
         | plus tests. Or, the tests can be distributed in parallel but
         | still under LGPL, not MIT. It doesn 't really matter since
         | compiled software won't be including the tests anyways.
        
           | foresto wrote:
           | > It's transformative, so no.
           | 
           | I'm not following your logic there, and I don't see any
           | mention of "transformative" in the license. Can you explain
           | what you mean?
        
             | crazygringo wrote:
             | Sorry, I misspoke. Transformation is what makes the _LLM
             | itself_ legal -- its training data is sufficiently
             | transformed into weights.
             | 
             | And so, a work being sufficiently transformative is one way
             | in which copyright no longer applies, but that's not the
             | case here specifically. The specific case here is
             | essentially just a _clean-room reimplementation_ (though
             | technically less  "clean", but still presumably the same
             | legally). But the end result is still a completely
             | different expression of underlying non-copyrightable ideas.
             | 
             | And in both cases, it doesn't matter _what_ the original
             | license was. If a resulting work is sufficiently
             | transformative _or_ a reimplementation, copyright no longer
             | applies, so the license no longer applies.
        
               | foresto wrote:
               | That's interesting, but it misses my point:
               | 
               | The library's test suite and interfaces were apparently
               | used directly, not transformed. If either of those are
               | considered part of the library's source code, as the
               | license's wording seems to suggest, then I think output
               | from their use could be considered a _work based on the
               | library_ as defined in the license.
        
               | crazygringo wrote:
               | Legally that's been established as acceptable.
               | 
               | Google LLC v Oracle America assumed (though didn't
               | establish) that API's are copyrightable... BUT that
               | developing against them _falls under fair use_ , as long
               | as the function implementations are independent.
               | 
               | Test suites are again generally considered
               | copyrightable... but the _behavior being tested is not_.
               | 
               | So no, it's not considered to be a work based on the
               | library. This seems pretty clear-cut in US law by now.
               | 
               | Also, the LGPL text doesn't say "work based on the
               | library". It says "If you modify a copy of the Library",
               | and this is not a "combined work" either. And the whole
               | point is that this is not a modified copy -- it's a
               | reimplementation.
               | 
               |  _In theory_ , a license _could_ be written to prevent
               | running its tests from being run against software not
               | derived from the original, i.e. clean-room
               | reimplementations. In practice, it remains dubious
               | whether any court would uphold that. And it would also be
               | trivial to then get around it, by taking advantage of
               | fair use to re-implement the tests in e.g. plain English
               | (or any specification language), and then re-implementing
               | those back into new test code. Because again, test
               | _behaviors_ are not copyrightable.
        
               | GardenLetter27 wrote:
               | But the tests were transformed to the new language, they
               | are not copied as-is.
               | 
               | Software patents would work as you describe, but not
               | copyright.
        
       | animitronix wrote:
       | LPGL is dead, long live the AI rewrites of your barely open
       | source code
        
       | kanemcgrath wrote:
       | without discussing copyright, I don't believe any of this is
       | copied. Which I think should be the argument that actually
       | matters.
       | 
       | I downloaded both 6.0 and 7.0 and based on only a light
       | comparison of a few key files, nothing would suggest to me that
       | 7.0 was copied from 6.0, especially for a 41x faster
       | implementation. It is a lot more organized and readable in my
       | armature opinion, and the code is about 1/10th the size.
        
       | Gigachad wrote:
       | Someone should put this to the test. Take the recently leaked
       | Minecraft source code and have Copilot build an exact replica in
       | another programming language and then publish it as open source.
       | See if Microsoft believes AI is copyright infringement or not.
        
         | Aboutplants wrote:
         | I've often thought that the key to fighting this is through
         | this exact method. Turn the tool against them
        
         | amelius wrote:
         | You will probably run into design patents.
        
           | VorpalWay wrote:
           | Software patents is not a thing in EU.
        
             | IshKebab wrote:
             | They absolutely are. That's a myth.
             | 
             | But also software patents and design patents are totally
             | different things.
        
             | throw-qqqqq wrote:
             | Very much is. "Software programs, as such" are exempt in
             | the EPC article 52. However if the software program
             | interacts with the world - if it has a "further technical
             | effect" - it is patentable.
             | 
             | https://en.wikipedia.org/wiki/Software_patents_under_the_Eu
             | r...
        
             | amelius wrote:
             | I said design patents:
             | 
             | https://en.wikipedia.org/wiki/Design_patent
        
         | peacebeard wrote:
         | The big question is: if copyrighted material was used in the
         | training material, is the LLM's output copyright infringement
         | when it resembles the training material? In your example, you
         | are taking the copyrighted material and giving it to the LLM as
         | input and instructing the LLM to process it. Regardless of
         | where the legal cards fall, this is a much less ambiguous
         | scenario.
        
           | LPisGood wrote:
           | I think Disney ran into this with people generating Marvel
           | characters etc
        
         | robmccoll wrote:
         | As described, this would not be the same thing. If the AI is
         | looking at the source and effectively porting it, that is
         | likely infringement. The idea instead should be "implement
         | Minecraft from scratch" but with behavior, graphics, etc.
         | identical. Note that you'll need to have an AI generate assets
         | or something since you can't just reuse textures and models.
        
           | Gigachad wrote:
           | AI models have already looked at the source of GPL software
           | and contain it in their dataset. Adding the minecraft source
           | to the mix wouldn't seem much different. Of course art assets
           | and trade marks would have to be replaced. But an AI "clean
           | room" implementation has yet to be legally tested.
        
             | reverius42 wrote:
             | For copyright purposes I think there is an important legal
             | distinction between training data (fed in once, ahead of
             | time, and can in theory no longer be recovered as-is) and
             | context window data (stored exactly for the duration of the
             | model call).
             | 
             | I'm not sure there should be, but I think there is.
        
             | NewsaHackO wrote:
             | That's why he is saying it's not equivalent. For it to be
             | the same, the LLM would have to train on/transform
             | Minecraft's source code into its weights, then you prompt
             | the LLM to make a game using the specifications of
             | Minecraft solely through prompts. Of course it's copyright
             | infringement if you just give a tool Minecraft's source
             | code and tell it to copy it, just like it would be
             | copyright infringement if you used a copier to copy
             | Minecraft's source code into a new document and say you
             | recreated Minecraft.
        
               | paxys wrote:
               | Is there a legal distinction between training, post-
               | training, fine tuning and filling up a context window?
               | 
               | In all of these cases an AI model is taking a copyrighted
               | source, reading it, jumbling the bytes and storing it in
               | its memory as vectors.
               | 
               | Later a query reads these vectors and outputs them in a
               | form which may or may not be similar to the original.
        
               | SatvikBeri wrote:
               | Judges have previously ruled that training counts as
               | sufficiently transformative to qualify for fair use:
               | https://www.whitecase.com/insight-alert/two-california-
               | distr...
               | 
               | I don't know of any rulings on the context window, but
               | it's certainly possible judges would rule that would not
               | qualify as transformative.
        
           | NiloCK wrote:
           | A room "as clean" as the one under dispute (chardet) is very
           | easy to replicate.
           | 
           | AI 1: - (reads the source), creates a spec + acceptance
           | criteria
           | 
           | AI 2: - implements from spec
           | 
           | AI 1 is in the position of the maintainer who facilitated the
           | license swap.
        
         | martin-t wrote:
         | They might not care. Products win not by quality or features
         | but by advertisement, hype and network effects.
         | 
         | The original implementation would still have the upper hand
         | here. OTOH if I as a nobody create something cool, there's
         | nothing stopping a huge corporation from "reimplementing"
         | (=stealing) it and and using their huge advertising budget to
         | completely overshadow me.
         | 
         | And that's how they like it.
        
           | Gigachad wrote:
           | Given how hard companies like Nintendo and Microsoft have
           | been taking down leaks or fan creations, it seems they very
           | much do care about keeping this stuff locked down.
        
         | u1hcw9nx wrote:
         | _This_ was not about legality.
         | 
         | > That question is this: does legal mean legitimate?
         | 
         | Just because something is legal does not mean it's moral thing
         | to do.
        
           | larodi wrote:
           | this question should've been posed earlier when first LLMs
           | were training. many people chose to ignore the question, and
           | now, several distillation epochs later, it is not a question
           | that matters, as both yes/no are true, and not true.
           | 
           | is it legitimate for millions of people to exploit, expound
           | on knowledge that was perhaps, to begin with, not legitimate
           | to use? well they did already, who's to judge the commons
           | now?
        
             | mirashii wrote:
             | What a ridiculous take. Many people loudly raised the
             | question and objected to the practice from the beginning,
             | but a handful of companies ignored the objections and ran
             | faster than the legal system. If they were in the wrong,
             | legally or morally, they still deserve to face
             | repercussions for it.
        
       | miggol wrote:
       | Wow, it feels like this argument rewired my brain.
       | 
       | When I first read about the chardet situation, I was conflicted
       | but largely sided on the legal permissibility side of things.
       | Uncomfortably I couldn't really fault the vibers; I guess I'm
       | just liberal at heart.
       | 
       | The argument from the commons has really invoked my belief in the
       | inherent morality of a public good. Something being
       | "impermissible" sounds bad until you realize that otherwise the
       | arrow of public knowledge suddenly points backwards.
       | 
       | Seeing this example play out in real life has had retroactive
       | effects on my previously BSD-aligned brain. Even though the
       | argument itself may have been presented before, I now understand
       | the morals that a GPL license text underpins better.
        
       | lukev wrote:
       | I agree with the thrust of this article, that norms and what we
       | perceive as good or desirable extend considerably beyond the
       | minimum established by law.
       | 
       | But a point that was not made strongly, which highlights this
       | even more, is that this goes in _every direction_.
       | 
       | If this kind of reimplementation is legal, then I can take any
       | permissive OSS and rebuild it as proprietary. I can take any
       | proprietary software and rebuild it as permissive. I can take any
       | proprietary software and rebuild it as my _own_ proprietary
       | software.
       | 
       | Either the law needs to catch up and prevent this kind of
       | behavior, or we're going to enter an effectively post-copyright
       | world with respect to software. Which ISN'T GOOD, because that
       | will disincentivize any sort of open license at all, and
       | companies will start protecting/obfuscating their APIs like trade
       | secrets.
        
         | integralid wrote:
         | It goes in one direction only.
         | 
         | Companies can take open-source software and make a proprietary
         | reimplementation. You can't take a proprietary software and
         | make an open source GPL version.
         | 
         | I am absolutely certain that if you tried you would be sued to
         | oblivion. But big company screwing up open source is not even
         | news anymore. In fact I (still) believe that the fact that even
         | though LLMs were trained on tons of GPL and AGPL or even
         | unlicensed software it's considered ok to use LLM code in
         | proprietary projects is example of just that.
        
           | lukev wrote:
           | From a strictly legal perspective the two are equivalent. The
           | fact that there are structural injustices in the system is
           | true, but that's not a question that any answer to "what
           | should be legal" can fix.
        
         | martin-t wrote:
         | I've been thinking this for over two years, that's why I
         | stopped contributing to open source at that time - my work was
         | only gonna be exploited to make rich people richer regardless
         | of the license.
         | 
         | Crazy that only now we're seeing a bunch of articles coming to
         | the same conclusion now.
         | 
         | I think copyright should still apply, but if it doesn't, we
         | need new laws - ones which protect all human work, creative or
         | not. Laws should serve and protect people, not algorithms and
         | not corporations "owning" those algorithms.
         | 
         | I put owning in quotes because ownership should go to the
         | people who did the work.
         | 
         | Buying/selling ownership of both companies and people's work
         | should be illegal just like buying/selling whole humans is.
         | Even if it took thousands of years to get here.
         | 
         | Money should not buy certain things because this is the root
         | cause of inequality. Rich people are not getting richer at a
         | faster rate by being more productive than everyone else but by
         | "owning" other people's work and using it as leverage to
         | extract even more from others.
         | 
         | Maybe LLM and mass unemployment of white collar workers will be
         | the wakeup call needed for a reform. Or revolution.
         | 
         | Last time this happened was during the second industrial
         | revolution and that's how communism got popular. We should do
         | better this time because this is the last revolution which
         | might be possible.
        
       | winstonwinston wrote:
       | > Blanchard's account is that he never looked at the existing
       | source code directly.
       | 
       | That's a weird statement while releasing the new version of the
       | same project. Maybe just release it as a new project, chardet-ai
       | v1.0 or whatever.
        
       | martin-t wrote:
       | 1) Legality and morality are obviously different and unrelated
       | concepts. More people should understand that.
       | 
       | 2) Copyright was the wrong mechanism to use for code from the
       | start, LLMs just exposed the issue. The thing to protect
       | shouldn't be creativity, it should be human work - any kind of
       | work.
       | 
       | The hard part of programming isn't creativity, it's making
       | correct decisions. It's getting the information you need to make
       | them. Figuring out and understanding the problem you're trying to
       | solve, whether it's a complex mathematical problem or a
       | customer's need. And then evaluating solutions until you find the
       | right one. (One constrains being how much time you can spend on
       | it.)
       | 
       | All that work is incredibly valuable but once the solution
       | exists, it's each easier to copy without replicating or even
       | understanding the thought process which led to it. But that
       | thought process took time and effort.
       | 
       |  _The person who did the work deserved credit and compensation._
       | 
       | And he deserves it transitively, if his work is used to build
       | other works - proportional to his contribution. The hard part is
       | quantifying it, of course. But a lot of people these days benefit
       | from throwing their hands up and saying we can't quantify it
       | exactly so let's make it finders keepers. That's exploitation.
       | 
       | 3) Both LLM training and inference are derivative works by any
       | reasonable meaning of those words. If LLMs are not derivative
       | works of the training data then why is so much training data
       | needed? Why don't they just build AI from scratch? Because they
       | can't. They just claim they found a legal loophole to exploit
       | other people's work without consent.
       | 
       | I am still hoping the legal people take time to understand how
       | LLMs work, how other algorithms, such as synonym replacement or
       | c2rust work, decide that calling it "AI" doesn't magically remove
       | copyright and the huge AI companies will be forced to destroy
       | their existing models and train new ones which respect the
       | licenses.
        
       | palata wrote:
       | > an argument for protecting that test suite and API
       | specification under copyleft terms.
       | 
       | If we protect API under copyright, it makes it easier to prevent
       | interoperability. We obviously do NOT want that. It would give
       | big companies even more power.
       | 
       | Now in the US, the Supreme Court that the output of an LLM is not
       | copyrightable. So even a permissive licence doesn't work for that
       | reimplementation: it should be public domain.
       | 
       | Disclaimer: I am all for copyleft for the code I write, but
       | already without LLMs, one could rewrite a similar project and use
       | the licence they please. LLMs make them faster at that, it's just
       | a fact.
       | 
       | Now I wonder: say I vibe-code a library (so it's public domain in
       | the US), I don't publish that code but I sell it to a customer.
       | Can I prevent them from reselling it? I guess not, since it's
       | public domain?
       | 
       | And as an employee writing code for a company. If I produce
       | public domain code because it is written by an LLM, can I publish
       | it, or can the company prevent me from doing it?
        
       | josalhor wrote:
       | I think the direction we are going, the GPL is going to fade
       | away. I think people will look at this like writing a book and
       | claiming the ideas in the book cannot be copied. This debate is
       | not that different from the ones going on in the music industry.
       | I open sourced my latest software as Apache 2.0 after debating a
       | lot about this. Unless the FSF wins in court in the next <=2-3
       | years, there is no coming back from this.
        
       | makerofthings wrote:
       | If an AI can license-wash open source software like this then the
       | licenses become meaningless. Which is fascinating. Commercial
       | software cloning that is simple enough for an average person to
       | drive is next and the ultimate form of piracy, see an app for
       | $10? Don't fancy paying? Just ask ChatGPT for a clone. Future is
       | going to be wild.
        
         | paxys wrote:
         | You've just described why every SaaS stock has taken a beating
         | in the last 6 months.
        
           | makerofthings wrote:
           | How long until this thing is good enough to clone photoshop?
           | Or Skyrim? I think all bets are off for the software world.
        
         | j-bos wrote:
         | > ultimate form of piracy
         | 
         | Nothing was stolen, not even copied, lamest piracy I've heard
         | of.
        
           | makerofthings wrote:
           | I take your point, but if the re-implementation looks the
           | same, I would say it's a form of copying. (Which I don't
           | think is a problem, I don't think you should be able to own
           | sequences of numbers.)
        
       | jrochkind1 wrote:
       | > If source code can now be generated from a specification, the
       | specification is where the essential intellectual content of a
       | GPL project resides.
       | 
       | Our foreparents fought for the right to implement works-a-like to
       | corporate software packages, even if the so-called owners did not
       | like it. We're ready to throw it all away, and let intellectual
       | property owners get so much more control.
       | 
       | The implications will not end up being anti-large-corporation or
       | pro-sharing. If you can prevent someone from re-implementing a
       | spec or building a client that speaks your API or building a
       | work-a-like, it will be the large corporations that exersize this
       | power as usual.
        
         | alterom wrote:
         | _> Our foreparents fought for the right to implement works-a-
         | like to corporate software packages, even if the so-called
         | owners did not like it_
         | 
         | Our "foreparents" weren't competing with corporations with
         | unlimited access to generative AI trained _on their work_. The
         | times, they 're-a-changin'.
         | 
         | You're rehashing the argument made in one of the articles which
         | this piece criticizes and directly addresses, while ignoring
         | the entirety of what was written before the conclusion that you
         | quoted.
         | 
         | If anyone finds themselves agreeing with the comment I'm
         | responding to, please, do yourself a favor and read the linked
         | article.
         | 
         | I would do no justice to it by reiterating its points here.
        
       | randyrand wrote:
       | It doesn't matter if it's legitimate. The people that use it
       | don't care. They just find it online and click download. This is
       | the reality.
        
       ___________________________________________________________________
       (page generated 2026-03-09 23:00 UTC)