[HN Gopher] Judge dismisses DMCA copyright claim in GitHub Copil...
       ___________________________________________________________________
        
       Judge dismisses DMCA copyright claim in GitHub Copilot suit
        
       Author : samspenc
       Score  : 82 points
       Date   : 2024-07-09 18:25 UTC (4 hours ago)
        
 (HTM) web link (www.theregister.com)
 (TXT) w3m dump (www.theregister.com)
        
       | rolph wrote:
       | copilot was apparently snipping license bearing comments, and
       | applying "semantic" variations of the remaining code.
       | 
       | i would package the entire code as a series of comments, [ideally
       | this would be snipped by the pliagarists] leaving a snippet of
       | example code that no one of sound mind would allow to execute,
       | being proffered by copilot.
        
         | ChrisMarshallNY wrote:
         | _> of sound mind_
         | 
         | That's a reach, these days...
         | 
         | I'm seeing some really ... _interesting_ ... behavior, being
         | exhibited by folks that, at first blush, I think are kids, just
         | out of bootcamp, but, on further inspection, turn out to be
         | middle-aged professionals.
         | 
         | I really think Teh Internets Tubes have been rather corrosive
         | to collective mental health.
        
           | klyrs wrote:
           | The ability to think for oneself will diminish rapidly in an
           | environment that rewards one for not doing so.
           | 
           | Smart people still exist. They just aren't online.
        
             | kirth_gersen wrote:
             | Suicide by words, here?
        
               | nyc_data_geek wrote:
               | The Internet is still one of the easiest ways to find and
               | participate in communities and conversations with other
               | smart people, if you're invested in vetting and filtering
               | who/what you're engaging with.
               | 
               | That said, I expect the ease of such will continue to
               | decline as we approach a largely dead Internet, primarily
               | consisting of bots talking to bots trying to sell each
               | other herbal brain force supplements or whatever
        
             | satvikpendem wrote:
             | From Plato's dialogue Phaedrus 14, 274c-275b:
             | 
             | Socrates: I heard, then, that at Naucratis, in Egypt, was
             | one of the ancient gods of that country, the one whose
             | sacred bird is called the ibis, and the name of the god
             | himself was Theuth. He it was who invented numbers and
             | arithmetic and geometry and astronomy, also draughts and
             | dice, and, most important of all, letters.
             | 
             | Now the king of all Egypt at that time was the god Thamus,
             | who lived in the great city of the upper region, which the
             | Greeks call the Egyptian Thebes, and they call the god
             | himself Ammon. To him came Theuth to show his inventions,
             | saying that they ought to be imparted to the other
             | Egyptians. But Thamus asked what use there was in each, and
             | as Theuth enumerated their uses, expressed praise or blame,
             | according as he approved or disapproved.
             | 
             | "The story goes that Thamus said many things to Theuth in
             | praise or blame of the various arts, which it would take
             | too long to repeat; but when they came to the letters,
             | "This invention, O king," said Theuth, "will make the
             | Egyptians wiser and will improve their memories; for it is
             | an elixir of memory and wisdom that I have discovered." But
             | Thamus replied, "Most ingenious Theuth, one man has the
             | ability to beget arts, but the ability to judge of their
             | usefulness or harmfulness to their users belongs to
             | another; and now you, who are the father of letters, have
             | been led by your affection to ascribe to them a power the
             | opposite of that which they really possess.
             | 
             | "For this invention will produce forgetfulness in the minds
             | of those who learn to use it, because they will not
             | practice their memory. Their trust in writing, produced by
             | external characters which are no part of themselves, will
             | discourage the use of their own memory within them. You
             | have invented an elixir not of memory, but of reminding;
             | and you offer your pupils the appearance of wisdom, not
             | true wisdom, for they will read many things without
             | instruction and will therefore seem to know many things,
             | when they are for the most part ignorant and hard to get
             | along with, since they are not wise, but only appear wise."
        
               | ChrisMarshallNY wrote:
               | That's great!
               | 
               | They nailed us, what, four thousand years ago?
        
               | courseofaction wrote:
               | Awesome. Serves as a counter-example - would HN consider
               | literacy to be damaging to the mind, or are we similarly
               | mistaken by thinking that LLMs necessarily degrade the
               | abilities of their users?
               | 
               | Pre-writing 'texts' (such as the Iliad) were memorized by
               | poets, which is reflected in their forms which made more
               | use of memory-friendly forms like rhyming, consistent
               | meter, and close repetition.
               | 
               | Writing allowed greater complexity and more
               | complex/information dense literary forms.
               | 
               | I feel that intelligent, critical LLM usage is just
               | writing with less laboriousnes, which opens up the
               | writer's ability to explore ideas more widely rather than
               | spend their time on the technical aspects of knowledge
               | production.
        
               | klyrs wrote:
               | Does it serve as a counterexample? Or did the predicted
               | loss of memory function come to pass?
               | 
               | Worth noting that people were smoking plain old opium
               | back in those times; I'd be reluctant to apply their
               | reasoning to fentanyl.
        
               | klyrs wrote:
               | Precisely the quote I was thinking of, thank you.
        
           | rolph wrote:
           | ..that suggests there is actually a chance that someone would
           | go for such a boobytrap.
        
       | bityard wrote:
       | This is pretty interesting, and I have conflicted feelings about
       | the (seemingly obvious) outcome of this trial.
       | 
       | I wonder, if MS and OpenAI win, does that mean it will be legal
       | for anyone to take the leaked source code for a proprietary
       | product, train an LLM on it, and then ask the LLM to emit a
       | version of it that is different enough to avoid copyright
       | infringement?
       | 
       | That would be quite the double-edged sword for proprietary
       | software companies.
        
         | ChrisMarshallNY wrote:
         | I suspect that this is exactly what will happen; not just with
         | code, but also prose and artwork.
         | 
         | Someone is likely to design an LLM that is specifically trained
         | to do exactly that.
         | 
         | Lots of money to be made...
        
           | devmor wrote:
           | On the matter of artwork there's no need for suspicion - it
           | is and has been happening for a while now. There are entire
           | online databases dedicated to providing non-consenting
           | artist's "styles" as downloadable model parameters by name.
        
             | ChrisMarshallNY wrote:
             | Try getting Mickey Mouse comics.
             | 
             | That should be fun...
        
             | satvikpendem wrote:
             | Style is not copyrightable so I see nothing wrong with
             | making essentially a robot that can paint in the style of
             | someone else.
        
               | falcolas wrote:
               | In isolation, no. But the produced works can be too close
               | for fair use (as demonstrated with the Prince pieces by
               | Andy Warhol), and passing it off as a piece from the
               | original artist can open you up to forgery/fraud charges.
               | 
               | To put another way, the motivations to produce art in
               | another artist's style can still land the artist/buyer in
               | legal trouble regardless of fair use.
        
               | satvikpendem wrote:
               | Yes that is true, but I don't think the people who use
               | style transfer are actually passing it off as the
               | original, they just like it for the aesthetic value of
               | their own images. In other words, no one using the Van
               | Gogh LoRA is actually trying to forge the Starry Night.
        
               | falcolas wrote:
               | Given the value of an "authentic" painting of the Starry
               | Night (or more realistically the value of something
               | forged in, say, Samwise Didier's style) I can't agree
               | with "no one".
               | 
               | I have to imagine that it's likely quite popular to sell
               | AI generated art that mimics or copies existing works.
        
               | satvikpendem wrote:
               | Do you use AI art generators? Flaws are extremely easily
               | found out, it is only good for a rough snapshot (without
               | much fiddling and even then, artifacts remain). I can
               | guarantee you it is definitely _not_ popular to sell
               | existing works made with AI, you are better off hiring an
               | actual forger. In fact, your suggestion is even the first
               | I 've even heard of such an idea.
        
               | 8organicbits wrote:
               | I guess there's always a greater fool, but forging an oil
               | painting using AI digital images seems pretty far
               | fetched.
        
               | devmor wrote:
               | The legality of using someone's copyrighted work to train
               | a model to reproduce it without their consent is still
               | under debate - but the morality of the act at least, is
               | not related to its legality - be it positively or
               | negatively; and I personally consider it abhorrent.
        
               | satvikpendem wrote:
               | Under what morals do you consider it "abhorrent?" I bet
               | got a straight answer from those I've asked about this as
               | the counter arguments seem too easy to make.
        
             | CuriouslyC wrote:
             | I sure wish I could non-consent to people observing me in
             | the world, I'd like to move through society invisibly and
             | only show myself when it benefitted me. Unfortunately, the
             | only answer is to stay inside if I don't want people to see
             | me.
        
               | vkou wrote:
               | > I sure wish I could non-consent to people observing me
               | in the world,
               | 
               | You aren't allowed to use photos _featuring_ a non-
               | consenting person to, for instance promote a product.
               | 
               | You are allowed to use photos _including_ a non-
               | consenting person.
               | 
               | There's a lot of complicated law, differing between
               | different jurisdictions to cover this question, and to
               | balance the needs of the public with commercial desires.
               | It's not as simple as you make it sound, and there's no
               | reason we should just default to bending over backwards
               | for commercial interests.
               | 
               | Laws exist to serve society, not the other way around.
        
               | CuriouslyC wrote:
               | I'm sure that the people who are being constantly
               | victimized by paparazzi would like to know those rules
               | that you just quoted, and have them be enforced.
        
               | vkou wrote:
               | If you had done a little research into this question,
               | you'd realize that 1A use cases ('journalism') are
               | treated by law quite differently than use of likeness for
               | commercial intent.
               | 
               | This is my whole point. There isn't a single, one-size-
               | fits-all rule that a five year old can comprehend that
               | describes any particular country's legal framework around
               | the many, many different dimensions of tension between
               | public and private interests on this incredibly broad
               | question.
               | 
               | And none of the existing frameworks fit the new use cases
               | well, and we should probably have an open political
               | debate about what we want to do going forward.
        
               | CuriouslyC wrote:
               | I'll happily take your picture against your will and put
               | it on the internet with the tag "vkou mad at
               | photographer, news at 11"
        
               | vkou wrote:
               | Okay? What will that prove? That you can be an ass?
               | 
               | Being an ass is generally not illegal. _Particular_
               | behaviours might be, but no legal or social system
               | intends to censure you for every possible one, and most
               | people who are experts in law or ethics don 't believe
               | that they should.
               | 
               | If you identify particular problems with the particular
               | paparazzi laws in your country, that's an interesting
               | conversation, and maybe, if framed well, an interesting
               | data point for this discussion, but is not in itself the
               | 'last word' on it. Just because you can torture an
               | analogy, doesn't mean the analogy has a lot of power.
        
               | sweeter wrote:
               | > consent Careful... A lot of people online have
               | selective understanding when it comes to this concept.
               | It's selfishness and self-centredness taken to it's
               | extreme, and not seeing other people as humans, but as
               | tools for their consumption to be used and tossed aside
               | for pleasure or for profit. It's one of the most
               | disgusting things I've layed eyes on.
        
               | devmor wrote:
               | We are not discussing people observing people. We are
               | discussing programs observing people.
        
           | ADeerAppeared wrote:
           | > Someone is likely to design an LLM that is specifically
           | trained to do exactly that.
           | 
           | Perplexity AI.
        
             | chimeracoder wrote:
             | > Perplexity AI.
             | 
             | How does this describe Perplexity AI more than any other
             | LLM?
        
           | epolanski wrote:
           | I feel like what really matters is who has more money to
           | throw in tribunals.
           | 
           | Somehow I feel if it was "Adobe vs dev that claims his code
           | was spit by copilot" it would not end the same.
        
         | devmor wrote:
         | Following existing law and applying reasonable expectations, I
         | would point to the old adage "intent is 9/10ths of the law".
         | 
         | It would probably be legal to do this, as long as no one could
         | reasonably show that you intentionally trained the LLM on said
         | leaked source code with the intent to reproduce the product.
         | 
         | Of course, civil suits could be another matter entirely. If you
         | pick a product to rip off that's owned by a multi-billion
         | dollar company, all that can save you is the ethical limits of
         | their legal team's consciences.
        
         | spencerflem wrote:
         | Not unless a big company is the one doing it lol
        
         | pennomi wrote:
         | Or even those AI-powered decompilers people are working on...
         | you could clone virtually any software with that. Surely there
         | will be limitations.
        
           | beeboobaa3 wrote:
           | The limitation is the amount of money & political power the
           | owner of the software you're cloning has.
        
           | wongarsu wrote:
           | The source code of Windows XP is widely available. Same with
           | a ~2 year old version of Bing, Bing Maps, Cortana etc. Yet
           | that doesn't seem to have had major negative effects on those
           | products. If anything having the Windows source code
           | available seems to be a net boon for Windows development.
           | Sometimes looking at the source is just better if the
           | documentation is unclear.
        
         | Legend2440 wrote:
         | I mean you can legally do this by hand right now. That's how
         | they cloned the IBM bios back in the day. IBM sued and lost.
        
           | marcosdumay wrote:
           | No, that's not.
           | 
           | They cloned the bios by observing how it behaved and writing
           | code that behaved the same way. Nobody even looked at the
           | bios code.
        
             | wvenable wrote:
             | That's not how they did it. They had one team read the BIOS
             | source listings in the IBM PC Technical Reference Manual
             | and create a technical specification and a second team take
             | that specification and write a new BIOS [1]. The second
             | team never saw the original code so therefore they could
             | not have copied it.
             | 
             | To do something similar with AI, you really need to train
             | one AI on the source code and then have it explain that
             | code to a second AI that never saw the original code.
             | 
             | [1] https://en.wikipedia.org/wiki/Phoenix_Technologies
        
           | axus wrote:
           | I thought there was a "clean room", where the people reading
           | it and the people writing it were different; and they made a
           | written specification instead of a Vulcan mind meld.
        
         | jeroenhd wrote:
         | A Wine fork built using an LLM trained on leaked Windows code
         | might be pretty useful.
        
         | Spivak wrote:
         | No, because judges aren't robots applying the law like code.
         | Intent matters. If you do this it will be painfully obvious
         | that your intent is to duplicate a large body copywritten code.
        
           | yellowapple wrote:
           | It's painfully obvious that the intent of GitHub Copilot is
           | to duplicate a large body of copyrighted code.
        
         | yellowapple wrote:
         | That's exactly what I expect to happen with the source code to
         | Microsoft's own software products, namely Windows.
         | 
         | Hilarity will ensue :)
        
       | daedrdev wrote:
       | > The anonymous programmers have repeatedly insisted Copilot
       | could, and would, generate code identical to what they had
       | written themselves, which is a key pillar of their lawsuit since
       | there is an identicality requirement for their DMCA claim.
       | However, Judge Tigar earlier ruled the plaintiffs hadn't actually
       | demonstrated instances of this happening, which prompted a
       | dismissal of the claim with a chance to amend it.
       | 
       | It sounds fair from how the article describes it
        
         | whimsicalism wrote:
         | Huh. There have definitely been well publicized examples of
         | this happening, like the quake inverse square root
        
           | polishTar wrote:
           | Fast inverse square root is now part of the public domain.
           | 
           | Also, even if this weren't the case you can't sue for damages
           | to other people (they'd need to bring their own suit)
        
             | anonymoushn wrote:
             | Is the particular implementation that the model spits out
             | 70+ years old?
        
             | immibis wrote:
             | Has it really already been 70 years since John Carmack
             | died?
        
               | polishTar wrote:
               | Ah, you're right. I was wrong to say "public domain".
               | 
               | It would be more correct to say Quake III Arena was
               | released to the public as free software under the GPLv2
               | license.
        
               | KnightHawk3 wrote:
               | There is a large gap between public domain and GPL. For
               | starters if Copilot is emitting GPL code for closed
               | source projects... that's copyright infringement.
        
           | voxic11 wrote:
           | You can't copyright a mathematical operation. Only a
           | particular implementation of it, and even then it may not be
           | copyrightable if its a straightforward and obvious
           | implementation.
           | 
           | That said the implementation doesn't appear to be totally
           | trivial and copilot apparently even copies the comments which
           | are almost certainly copyrightable in themselves.
           | 
           | https://x.com/StefanKarpinski/status/1410971061181681674
           | https://github.com/id-Software/Quake-III-
           | Arena/blob/dbe4ddb1...
           | 
           | However a twitter post on its own isn't evidence a court will
           | accept. You would need the original poster to testify that
           | what is seen in the post is actually what he got from copilot
           | and not just a meme or joke that he made.
           | 
           | Also the plaintiffs in this case don't include id-Software
           | and there is some evidence that id-Software actually stole
           | the fast inverse sqrt code from 3dfx so they might not want
           | to bring a claim here anyways.
        
             | beeboobaa3 wrote:
             | https://en.wikipedia.org/wiki/Illegal_number
        
             | whimsicalism wrote:
             | Not sure where you thought I said you could copyright a
             | mathematical operation, I was clearly referring to the
             | implementation due to the mention of "quake".
             | 
             | When it was reported, I was able to reproduce it myself.
        
           | wongarsu wrote:
           | It reads like the judge required them to show it happened to
           | their code, not to any code in general. That's a much higher
           | bar. There are thousands of instances of fast inverse square
           | root in the training data but only one copy of your random
           | github repositories. Getting to model to reproduce your code
           | verbatim might be possible for all we know, but it isn't
           | trivial.
        
             | whimsicalism wrote:
             | of course for standing. but it seems like with the right
             | plaintiffs this could have gone forward
        
           | daedrdev wrote:
           | The article mentions that GitHub copilot has been trained to
           | avoid directly copying specific cases it knows, and that
           | although you can get it to spit out copyright code by
           | prefixing the copyrighted code as a starting point, in normal
           | us cases its quite rare.
        
         | ADeerAppeared wrote:
         | Where it gets ethnically dubious is that:
         | 
         | 1. The copilot team rushed to slap a copyright filter on top to
         | keep these verbatim examples from showing up, and now claims
         | they never happen.
         | 
         | 2. LLMs are prone to paraphrasing. Just because you filter out
         | verbatim copies doesn't mean there isn't still copyright
         | infringement/plagiarism/whatever you want to call it. The
         | copyright filter is only a legal protection, not a practical
         | protection against the issue of copyright infringement.
         | 
         | Everyone who knows how these systems work understand this. The
         | copilot FAQ to this day claims that you should run copyright
         | scanning tools on your codebase because your developers might
         | "copy code from an online source or library".
         | 
         | Github has it's own research from 2021 showing that these tools
         | do indeed copy their training data occasionally:
         | https://github.blog/2021-06-30-github-copilot-research-recit...
         | 
         | They clearly know the problem is real. Their own research
         | agreed, their FAQs and legal documents are carefully phrased to
         | avoid admitting it. But rather than owning up to the problem,
         | it's "Ner ner ner ner ner, you can't prove it to a boomer
         | judge".
        
           | squarefoot wrote:
           | > 1.
           | 
           | Isn't that akin to destruction of evidence?
        
             | ADeerAppeared wrote:
             | Legally? No.
             | 
             | In spirit? ... Probably?
             | 
             | Unlike most LLMs, Github copilot can trivially solve their
             | copyright problem by just using only code they have the
             | right to reproduce.
             | 
             | They have a giant corpus of code tagged with license,
             | SELECT BY license MIT/Equivalent and you're done, problem
             | solved because those licenses explicitly grant permission
             | for this kind of reuse.
             | 
             | (It's still not very cash money to take open source work
             | for commercial gain without paying the original authors,
             | and there's a humorous question if MIT-copilot would need
             | to come with a multi-gigabyte attribution file, but
             | everyone widely agrees it's legal and permitted.)
             | 
             | The only reason you'd hack a filter on top rather than
             | doing the above is if you'd want to hide the copyright
             | problem. It's an objectively worse solution.
        
               | spencerflem wrote:
               | No, you misinterpret the MIT licence. It still requires
               | attribution, it is not legal to copy MIT code without
               | notice.
        
             | abigail95 wrote:
             | Not in any way I'm aware of - and would be required if they
             | were served a DMCA notification/Cease and Desist against a
             | specific prompt.
             | 
             | The people that think Copilot is infringng their copyright
             | would be happy with that I would think? Unless they take a
             | much stricter definition of fair use than current courts
             | do.
        
       | mvdtnz wrote:
       | What were the plaintiffs even thinking when they submitted a
       | claim based on identicality without being able to produce a
       | single instance of copilot generating a verbatim copy. Even the
       | research they submitted was unable to make a claim any stronger
       | than "it's possibly in theory but we've never seen it".
        
       | loceng wrote:
       | This kind of argument makes me feel like it also supports the
       | abolition of patents: eventually multiple other people will come
       | up with the same obvious solution, which becomes obvious once a
       | person spends enough time looking at a problem.
        
         | CodeWriter23 wrote:
         | The Patent System is not intended to be a test of exclusive
         | original thought.
         | 
         | The function of the Patent System is to incentivize search for
         | solutions by temporarily securing exclusive right to market
         | novel devices and processes for the discoverer.
        
           | loceng wrote:
           | Of non-obvious inventions. My argument being all inventions
           | are obvious once attention is applied to that area and scope.
        
         | erik_seaberg wrote:
         | Unfortunately USPTO takes "non-obvious" to mean that it wasn't
         | already suggested by combining patents or other written work,
         | so if you are the first to work a problem you can claim easy
         | solutions that anyone with a clue would have quickly reached.
         | Land rushes to fence off new fields seem inevitable.
        
       | pledess wrote:
       | I thought "the Copilot coding assistant was trained on open
       | source software hosted on GitHub and as such would suggest
       | snippets from those public projects to other programmers without
       | care for licenses" was explicitly allowed by the GitHub Terms of
       | Service: https://docs.github.com/en/site-policy/github-
       | terms/github-t... "If you set your pages and repositories to be
       | viewed publicly, you grant each User of GitHub a nonexclusive,
       | worldwide license to use, display, and perform Your Content
       | through the GitHub Service." In other words, in addition to
       | what's allowed by the LICENSE file in your repo, you are also
       | separately licensing your code "to use ... through the GitHub
       | Service" and this would (in my interpretation) include use by
       | Copilot for training, and use by Copilot to deliver snippets to
       | any other GitHub user.
        
         | dmitrygr wrote:
         | Lots of my code is on github (eg
         | https://github.com/syuu1228/uARM), uploaded by others. I gave
         | no license for its use in training. What now?
        
           | zdragnar wrote:
           | If the person didn't have your permission or permission from
           | the license to agree to github's terms, then you sue the
           | person who uploaded it to GitHub.
           | 
           | You don't get to go after GitHub because you have no
           | contractual relationship with them. At best, you can get an
           | injunction forcing them to take it down, though getting them
           | to un-train copilot may not be feasible. At best you'd get a
           | small cash offer, since you're unlikely to be able to justify
           | any damages in a suit.
        
             | dredmorbius wrote:
             | 17 USC SS504 says otherwise:
             | 
             |  _... the copyright owner may elect, at any time before
             | final judgment is rendered, to recover, instead of actual
             | damages and profits, an award of statutory damages for all
             | infringements ... in a sum of not less than $750 or more
             | than $30,000. ... in a case where the copyright owner
             | sustains the burden of proving, and the court finds, that
             | infringement was committed willfully, the court in its
             | discretion may increase the award of statutory damages to a
             | sum of not more than $150,000._
             | 
             | <https://www.law.cornell.edu/uscode/text/17/504>
             | 
             | The issue isn't contract. It's copyright infringement.
        
             | 201984 wrote:
             | So hypothetically, if a developer publishes GPL software on
             | Codeberg, and someone uploads it to GitHub, could the
             | original developer file takedowns against the Github copy?
             | 
             | I'm curious if Github's ToS make uploading GPL software you
             | don't own a copyright violation.
        
               | votepaunchy wrote:
               | No, because the GPL is already more permissive than the
               | GitHub TOS.
        
             | pton_xd wrote:
             | > then you sue the person who uploaded it to GitHub.
             | 
             | > You don't get to go after GitHub because you have no
             | contractual relationship with them
             | 
             | What makes you say that? If someone eg uploads my
             | copyrighted work to YouTube, I file a DMCA notice with
             | YouTube to stop distributing my work. If YT ignores the
             | notice then I can pursue them with a lawsuit.
             | 
             | How is this situation different?
        
               | singleshot_ wrote:
               | DMCA explicitly gives you a cause of action against the
               | party who does not properly comply with your request. GP
               | asserts that you lack a cause of action against GitHub
               | before they fail to comply with DMCA but I'm not certain
               | I agree.
        
               | stefan_ wrote:
               | DMCA is a narrow protection for operators of public
               | websites like GitHub. I don't see what it has to do with
               | GitHub taking the data submitted to it with dubious
               | sourcing and developing their CoPilot whatever based on
               | it. That has nothing to do with the privileges in DMCA.
        
               | singleshot_ wrote:
               | That's right. You have lost the thread of what we are
               | talking about: causes of action based on privity vs those
               | created by statute.
        
         | simion314 wrote:
         | That will work if I upload only my code, but there are many
         | open source projects where there are more then one author and
         | GithHub did not acquired the rights from all the authors, the
         | uploader to GitHub might not even be the author too.
        
         | Brian_K_White wrote:
         | That just means github can display the code, and you can see
         | the code, but that does not mean you can then profit from or
         | redistribute (profit or no) the code without attribution.
         | 
         | Amazon has the rights to publish a book, and you have the right
         | to receive a copy of the book, but neither of those gives you
         | the right to re-publish the book under your own name.
        
           | rurcliped wrote:
           | "use, display, and perform Your Content through the GitHub
           | Service" might allow a wide range of uses on GitHub Pages
           | websites, even if https://example.github.io is monetized
           | (monetization is permitted by
           | https://docs.github.com/en/site-policy/github-
           | terms/github-t... in a few cases)
        
       | purpleblue wrote:
       | Can you insist or put instructions that AIs do not train on your
       | code? If they train on your code but don't produce the exact same
       | output, is there any protection you can have from that?
        
         | archontes wrote:
         | When are people going to get that this isn't a right folks
         | have?
         | 
         | If your code is readable, the public can learn from it.
         | 
         | Copyright doesn't extend to function.
        
           | ADeerAppeared wrote:
           | People aren't going to get it, because you don't get them.
           | 
           | People have the right to learn _non-copyrightable elements_
           | from your code.
           | 
           | The claim is that AI learns _copyrightable elements_.
        
             | archontes wrote:
             | The comment chain you are replying to includes a request to
             | not train an AI on one's code.
             | 
             | I agree it's certainly possible for AI to produce
             | infringing output.
             | 
             | Nevertheless, people don't have the right to enforce a
             | limitation on training.
        
       | munificent wrote:
       | _> Indeed, last year GitHub was said to have tuned its
       | programming assistant to generate slight variations of ingested
       | training code to prevent its output from being accused of being
       | an exact copy of licensed software._
       | 
       | If I, a human, were to:
       | 
       | 1. Carefully read and memorize some copyrighted code.
       | 
       | 2. Produce new code that is textually identical to that. But in
       | the process of typing it up, I randomly mechanically tweak a few
       | identifiers or something to produce code that has the exact same
       | semantics but isn't character-wise identical.
       | 
       | 3. Claim that as new original code without the original
       | copyright.
       | 
       | I assume that I would get my ass kicked legally speaking. That
       | reads to me exactly like deliberate copyright infringement with
       | willful obfuscation of my infringement.
       | 
       | How is it any different when a machine does the same thing?
        
         | singleshot_ wrote:
         | The guy who owns the machine is really rich, while you are more
         | or less (all due respect of course) not worth suing.
         | 
         | That's why I think the opposite of what you claim is true: if
         | you were to do this, absolutely nothing would happen. When they
         | do it, they will get sued over and over until the law changes
         | and they can't be sued, or they enter some mutually-beneficial
         | relationship with the parties who keep suing.
        
           | beeboobaa3 wrote:
           | > if you were to do this, absolutely nothing would happen
           | 
           | Read up on the DMCA and the impact it has on e.g. nintendo
           | emulators and the developers thereof
        
             | dmix wrote:
             | Those emulators are very popular though to the point of
             | potentially impacting another business's bottom line. Where
             | an individual putting it out a small block of code isn't
             | exactly going to attract expensive lawyers.
             | 
             | I'm skeptical Github Copilot reproducing a couple functions
             | potentially used by some random Github project is going to
             | be a threat to another party's livelihood.
             | 
             | When AI gets good enough to make full duplicates of apps
             | I'd be more concerned about the source. Thousands of
             | smaller pieces drawn from a million sources and being
             | combined in novel ways is less worrying though.
        
               | BadHumans wrote:
               | There is no impact to a company's bottom line when you
               | are emulating a product they do not sell.
        
               | lcouturi wrote:
               | Yuzu, the emulator that was sued by Nintendo, was
               | emulating the Nintendo Switch, which is a product
               | Nintendo does sell.
        
               | BadHumans wrote:
               | Yuzu is not the only emulator taken down by Nintendo and
               | Nintendo is not the only company that has gone after
               | emulators.
        
         | beeboobaa3 wrote:
         | Rules for thee but not for me (rich companies). Think of the
         | shareholders!
        
         | Analemma_ wrote:
         | > How is it any different when a machine does the same thing?
         | 
         | Because intent matters in the law. If you intended to reproduce
         | copyrighted code verbatim but tried to hide your activity with
         | a few tweaks, that's a very different thing from using a tool
         | which _occasionally_ reproduces copyrighted code by accident
         | but clearly was not designed for that purpose, and much more
         | often than not outputs transformative works.
        
           | archontes wrote:
           | Not in copyright. The work speaks for itself, and the
           | function of code is not a copyrightable aspect.
        
             | bawolff wrote:
             | The intent of the work can matter when determining if de
             | minimis applies as well as fair use.
        
           | olliej wrote:
           | Um, the entire intent of these "AI" systems is explicitly to
           | reproduce copyrighted work with mechanical changes to make it
           | not appear to be a verbatim copy.
           | 
           | That is the whole purpose and mechanism by which they
           | operate.
           | 
           | Also the intent does not matter under law - not intending to
           | break the law is not a defense if you break the law. Not
           | intending to take someone's property doesn't mean it becomes
           | your property. You _might_ get less penalties and /or
           | charges, due to intent (the obvious examples being murder vs
           | manslaughter, etc).
           | 
           | But here we have an entire ecosystem where the model is "scan
           | copyrighted material" followed by "regurgitate that material
           | with mechanical changes to fit the surrounding context and to
           | appear to be 'new' content".
           | 
           | Moreover given that this 'new' code is just a regurgitation
           | of existing code with mutations to make it appear to fit the
           | context and not directly identical to the existing code, then
           | that 'new' code cannot be subject to copyright (you can't
           | claim copyright to something you did not create, copyright
           | does not protect output of mechanical or automatic
           | transformations of other copyrighted content, and copyright
           | does not protect the result of "natural processes", e.g 'I
           | asked a statistical model to give me a statically plausible
           | sequence of tokens and it did'). So in the best case scenario
           | - the one where the copyright laundering as a service tool is
           | not treated as just that, any code it produces is not
           | protectable by copyright, and anyone can just copy "your
           | work" without the license and (because you've said if you
           | weren't intending to violate copyright it's ok) they can say
           | they could not distinguish the non-copyright-protected work
           | from the protected work and assumed that therefore none of it
           | was subject to copyright. To be super sure though they
           | weren't violating any of your copyrights, they then ran an
           | "AI tool" to make the names better and better suit your
           | style.
           | 
           | I am so sick of these arguments where people spout nonsense
           | about "AI" systems magically "understanding" or "knowing"
           | anything - they are very expensive statistical models, the
           | produce statistically plausible strings of text, by a
           | combination of copying the text of others wholesale, and
           | filling the remaining space with bullshit that for basic
           | tasks is often correct enough, and for anything else is wrong
           | - because again they're just producing plausible sequences of
           | tokens and have no understanding of anything beyond that.
           | 
           | To be very very very clear: if an AI system "understood"
           | anything it was doing, it would not need to ingest
           | essentially all the text that anyone has ever written, just
           | to produce content that is at best only locally coherent, and
           | that is frequently incorrect in more or less every domain to
           | which it is applied. Take code completion (as in this case):
           | Developers can write code without essentially reading all the
           | code that has ever existed just so that they can write basic
           | code, because developers understand code. Developers don't
           | intermingle random unrelated and non-present variables or
           | functions in their code as they write, because they
           | understand what variables are and therefore they can't use
           | non existent ones. "AI" on the other hand required more power
           | than many countries to "learn" by reading as much as possible
           | all code ever written, and then produce nonsense output for
           | anything complex because they're still just generating a
           | string of tokens that is plausible according to their
           | statistical model - the result of these AIs is essentially
           | binary: it has been in effect asked to produce code that does
           | something that was in its training corpus and can be copied
           | essentially verbatim, with a transformation path to make it
           | fit, or it's not in the training corpus and you get random
           | and generally incorrect code - hopefully wrong enough it
           | fails to build, because they're also good at generating code
           | that looks plausible but only fails at runtime because
           | plausible sequence of tokens often overlaps with 'things a
           | compiler will accept'.
        
             | shkkmo wrote:
             | > Also the intent does not matter under law - not intending
             | to break the law is not a defense if you break the law
             | 
             | Intent frequently matters a great deal when applying laws.
             | 
             | In the specific area of copyright law, it doesn't itself
             | make the use non infringing, but it can absolutely impact
             | the damages or a fair use argument.
        
         | dmix wrote:
         | That's a significant over simplification of how it works though
         | to the point of almost not being a useful analogy.
         | 
         | If your analogy was you were a human who memorized every
         | variation of a problem (and every other known problem) and
         | there was a tiny perctange of a chance where you reproduced
         | that exact varation of one you memorized, but then added an
         | after the fact filter so you don't directly reproduce it...
         | 
         | It's more like musicians who basically copy a bunch of music
         | patterns or chord progressions before then notice their final
         | output sounds too similar to another song (which happens often
         | IRL) then changes it to be more original before releasing it to
         | the public.
        
           | ADeerAppeared wrote:
           | > If you analogy was you were a human who memorized every
           | variation of a problem (and every other known problem)
           | 
           | This is mere assumption. AI is _supposed to_ work like that,
           | but that 's a goal, and not the result of current
           | implementations. Research shows that they do memorize
           | _solutions_ as well, and quite regularly so. (This is an
           | unavoidable flaw in current LLMs; They must be capable of
           | memorizing input verbatim in order to learn specific facts.)
           | 
           | > and there was a tiny perctange of a chance where you
           | reproduced that exact varation of one you memorized
           | 
           | This is copyright infringement. _Actionable_ copyright
           | infringement. The big music publishers go after this kind of
           | accidental partial reproduction.
           | 
           | > but then added an after the fact filter so you don't
           | directly reproduce it...
           | 
           | "Legally distinct" is a gimmick that only works where the
           | copyright is on specific identifiable parts of a work.
           | 
           | Changing a variable name does not make a code snippet
           | "legally distinct", it's still copyright infringement.
        
             | dmix wrote:
             | Meh I still see that as a big oversimplification. Context
             | matters. Even if the copyright courts often ignore that for
             | wealthy entities. Someone reproducing a song using AI and
             | publishing it as their own copyright infringement, a person
             | specifically querying an AI engine, that sucked up billions
             | of lines of information and generates what you ask it do
             | with a sma probability it will reproduce a small subset of
             | a larger commercial project and sends it to someone in a
             | chatbox is not exactly the same IMO.
             | 
             | This is Github Copilot after all. I use it daily and it
             | autocompletes lines of code or generates functions you can
             | find on stackoverflow. It's not letting giving you the
             | source code to Twitter in full and letting you put it on
             | the internet as a business under another name.
        
           | belorn wrote:
           | We are currently seeing the music industry reacting to AI
           | learning a bunch of music patterns and chord progressions and
           | outputting works that sounds very similar to existing music
           | and artists. They are not liking it.
           | 
           | To just see how much they disliked it, youtube copyright
           | strikes is basically a trained AI to detect music patterns to
           | identify sound with slight variations or copyrighted songs
           | and take videos down. Generating slight variations was one of
           | the early method that videos used to bypass the take down
           | system.
        
         | archontes wrote:
         | You might not get your ass kicked. Copyright doesn't protect
         | function, to the point where the court will assess the degree
         | to which the style of the code can be separated from the
         | function. In the even that they aren't separable, the code is
         | not copyrightable.
         | 
         | https://www.wardandsmith.com/articles/supreme-court-announce...
         | 
         | https://easlerlaw.com/software-computer-code-copyrighted#:~:...
        
           | ADeerAppeared wrote:
           | The simple version is that code _is_ copyrightable as an
           | _expression_. And the underlaying algorithm is _patentable_.
           | 
           | The legal term you're looking for here is the "Abstraction-
           | Filtration-Comparison" test; What remains if you subtract all
           | the non-copyrightable elements from a given piece of code.
        
           | tomxor wrote:
           | US copyright does protect for "substantial similarity" [0].
           | And at the other end of the spectrum, this has been abused in
           | absurd ways to argue that substantially different code has
           | infringed.
           | 
           | In Zenimax vs Oculus they basically argued that a bunch of
           | really abstract yet entirely generic parts of the code were
           | shared, we are talking some nested for loops, certain
           | combinations of if statements, and due to a lack of a
           | qualitative understanding of code, syntax, common patterns,
           | and what might actually qualify for substantively novel code
           | in the courtroom, this was accepted as infringing. [1]
           | 
           | Point is, the legal system is highly selective when it comes
           | to corporate interests.
           | 
           | [0] https://en.wikipedia.org/wiki/Substantial_similarity
           | 
           | [1] https://arstechnica.com/gaming/2017/02/doom-co-creator-
           | defen...
        
             | talldayo wrote:
             | > Point is, the legal system is highly selective when it
             | comes to corporate interests.
             | 
             | I don't even think it's that. In recent cases like Oracle
             | v. Google and Corellium v. Apple, Fair Use prevailed with
             | all sorts of conflicting corporate interests at play. The
             | Zenimax v. Oculus case very much revolved around NDAs that
             | Carmack had signed and not the propagation of trade
             | secrets. Where IP is strictly the only thing being
             | concerned, the literal interpretation of Fair Use does
             | still seem to exist.
             | 
             | Or for a more literal example, Authors Guild. v. Google
             | where Google defended their indexing of thousands of
             | licensed books as Fair Use.
        
               | tpmoney wrote:
               | In fact, go to far as to argue your example of Authors
               | Guild v. Google is a good indication that most cases will
               | probably go an AI platform's way. It's a pretty parallel
               | case to a number of the arguments. Indexing required
               | ingesting whole works of copyright material verbatim. It
               | utilized that ingested data to produce a new commercial
               | work consisting of output derived from that data. If I
               | remember the case correctly, google even displayed
               | snippets when matching a search so the searcher could see
               | the match in context, reproducing the works verbatim for
               | those snippets and one could presume (though I don't
               | recall if it was coded against), that with sufficiently
               | clever search prompts, someone could get the index search
               | to reproduce a substantial portion of a work.
               | 
               | Arguably, the AI platforms have an even stronger case as
               | their nominal goal is not to have their systems reproduce
               | any part of the works verbatim.
        
         | ars wrote:
         | > I assume that I would get my ass kicked legally speaking.
         | 
         | Maybe, maybe not. It's not as simple as you made it out to be.
         | If you write a book with lots of stuff and you got inspiration
         | from other books, and even put in phrases wholesale, but
         | modified to use your own character names instead, I'm not
         | convinced you would lose.
         | 
         | The court would look at the work as a whole, not single pieces
         | of it.
         | 
         | They would also check if you are just copying things verbatim,
         | or if you memorize a pattern and emit the same pattern - for
         | example look at lawsuits about copying music, where they'll
         | claim this part of the music is the same as that part.
         | 
         | It's really not as cut and dry as you make it out to be.
        
         | williamcotton wrote:
         | Just to set the stage and not entirely specific to this
         | complaint... It really depends on what is and isn't subject to
         | copyright for software.
         | 
         | Broadly, there is the distinction between expressive and
         | functional code. [1]
         | 
         | And then there are the specific tests that have been developed
         | by the courts to separate the expressive and functional aspects
         | of software. [2] [3]
         | 
         | In practice it is very expensive for a plaintiff to do such
         | analysis. For the most part the damages related to copyright
         | are not worth the time and money. Plaintiffs tend to go for
         | trade secret related damages as they are not restricted by the
         | above tests.
         | 
         | There are also arguments to be made of de minimis infringements
         | that are not worth the time of the court.
         | 
         | Most importantly the plaintiff fundamentally has the burden of
         | proof and cannot just say that copying must have taken place.
         | They need concrete evidence.
         | 
         | [1] https://en.wikipedia.org/wiki/Idea-expression_distinction
         | 
         | [2]
         | https://en.wikipedia.org/wiki/Structure,_sequence_and_organi...
         | 
         | [3] https://en.wikipedia.org/wiki/Abstraction-Filtration-
         | Compari...
        
         | wvenable wrote:
         | You probably do this all the time. Forget memorizing but
         | undoubtedly you've read code, learned from it, and then likely
         | reproduced similar code. Probably nothing terribly important,
         | just a function here or there. Maybe even reproduced something
         | you did for a previous employer.
        
         | JoshTriplett wrote:
         | You have a much smaller lobbying budget than the AI industry,
         | and you didn't flagrantly rush to copy billions of copyrighted
         | works as quickly as possible and then push a narrative acting
         | like that's the immutable status quo that must continue to be
         | permitted lest the now-massive industry built atop copyright
         | violation be destroyed.
         | 
         | Violate one or two copyrights, get sued or DMCAed out of
         | existence. Violate billions, on the other hand, and you
         | magically become immune to the rules everyone else has to
         | follow.
        
           | nadermx wrote:
           | What about the copyrights purpose of furthering the arts and
           | sciences?
        
             | JoshTriplett wrote:
             | Copyright has utterly failed to serve that purpose for a
             | long time, and has been actively counterproductive.
             | 
             | But if you want to argue that copyright is
             | counterproductive, I completely agree. That's an argument
             | for reducing or eliminating it across the board, fairly,
             | for everyone; it's not an argument for giving a free pass
             | to AI training while still enforcing it on everyone _else_.
        
         | hyperpape wrote:
         | From the article:
         | 
         | > The most recently dismissed claims were fairly important,
         | with one pertaining to infringement under the Digital
         | Millennium Copyright Act (DMCA), section 1202(b), which
         | basically says you shouldn't remove without permission crucial
         | "copyright management" information, such as in this context who
         | wrote the code and the terms of use, as licenses tend to
         | dictate.
         | 
         | > It was argued in the class-action suit that Copilot was
         | stripping that info out when offering code snippets from
         | people's projects, which in their view would break 1202(b).
         | 
         | > The judge disagreed, however, on the grounds that the code
         | suggested by Copilot was not identical enough to the
         | developers' own copyright-protected work, and thus section
         | 1202(b) did not apply. Indeed, last year GitHub was said to
         | have tuned its programming assistant to generate slight
         | variations of ingested training code to prevent its output from
         | being accused of being an exact copy of licensed software.
         | 
         | So (not a lawyer!) this reads like the point about GitHub
         | tuning their model is not a generic defense against any and all
         | claims of copyright infringement, but a response to a specific
         | claim that this violates a provision of the DMCA.
         | 
         | I don't know whether this is a reasonable defense or not, but
         | your intuitions or mine about whether there is a general
         | copyright violation or what's fair are not necessarily relevant
         | to how the judge construes that very specific bit of legal
         | code.
        
         | skybrian wrote:
         | The machine alone doesn't do anything. The user and machine
         | together constitute a larger system, and with autocomplete, the
         | user is charge. What's the user's intent?
         | 
         | I suspect that a lot of copyright violations are enabled by
         | cut-and-paste and screenshot-taking functionality, and maybe we
         | need to be careful with autocomplete, too? It's the user's
         | responsibility to avoid this. We should be careful using our
         | tools. Do users take enough care in this case? Is it possible
         | to take enough care while still using CoPilot?
         | 
         | I've switched from CoPilot to Cody, but I use them the same
         | way, to write _my_ code. There 's no particular reason to use
         | CoPilot's output verbatim and lots of good reasons not to. By
         | the time I've adapted it to my code base and code style and
         | refactored it to hell and back, it's an expression of how _I_
         | want to solve a problem, and I 'm pretty confident claiming
         | ownership.
         | 
         | Is that confidence misplaced? Are other people more careless?
        
       | lnxg33k1 wrote:
       | Suddenly, you would steal a car
        
         | blooalien wrote:
         | | "Suddenly, you would steal a car"
         | 
         | Nah, but I would _download_ a _copy_ of one without
         | hesitation... ;)
        
       | hn_throwaway_99 wrote:
       | A slight aside, but this is the subtitle:
       | 
       | > A few devs versus the powerful forces of Redmond - who did you
       | think was going to win?
       | 
       | I hate that kind of obnoxious "journalism". Sometimes the little
       | guy is actually wrong. To clarify, I'm not commenting on the
       | specifics of this case, I just hate how fake our online discourse
       | has been by appealing to "big guy evil" before even bringing up
       | the specifics of the case.
        
         | epolanski wrote:
         | I think you're misinterpreting the sentence.
         | 
         | I think it merely implies MS has more resources to throw at the
         | legal case.
        
         | deciplex wrote:
         | > Sometimes the little guy is actually wrong.
         | 
         | He is, sometimes. Also sometimes, the moon passes exactly
         | between the sun and Earth, a new star appears in the sky, the
         | magnetic field of our planet reverses, a proton decays (jury is
         | still out on that one, actually). Etc.
         | 
         | Tools like Copilot are plagiarism machines. We know the data
         | they're being trained on, and a conclusion of "that's
         | plagiarism" is not - or anyway should not be - controversial.
         | I'm not terribly _against_ the notion of a plagiarism machine
         | but I am against the owners of such machines reaping profits
         | from them to the exclusion of the people who provide the source
         | material. This is theft.
         | 
         | More importantly, getting back to big guys and little guys: big
         | guys gang up on little guys all the time. It's usually how they
         | get to be big. They tend to be the ones who realize that
         | working together against the rest of us is to their benefit.
         | So, in the interest of pushing back on that a little, and
         | recognizing that I am after all a fellow "little guy"
         | (figuratively speaking anyway), I tend to support the "little
         | guy" unless I have overwhelming evidence confirming that they
         | are, in fact, both _wrong_ and that _supporting them anyway
         | would be against my best interest._ Neither is the case, here.
         | 
         | At any rate, the subtitle here references a pretty ubiquitous
         | and, I'm happy to report, increasingly well-known and
         | understood facet of our economic and social institutions, which
         | is that they absolutely positively do not work for us or
         | further our interests in any sense.
        
       | epolanski wrote:
       | I am not strongly opinionated on this, but the very fact
       | Microsoft used all the code it could find, bar their own has
       | always looked suspicious to me.
        
       ___________________________________________________________________
       (page generated 2024-07-09 23:00 UTC)