[HN Gopher] Judge dismisses DMCA copyright claim in GitHub Copil...
___________________________________________________________________
Judge dismisses DMCA copyright claim in GitHub Copilot suit
Author : samspenc
Score : 82 points
Date : 2024-07-09 18:25 UTC (4 hours ago)
(HTM) web link (www.theregister.com)
(TXT) w3m dump (www.theregister.com)
| rolph wrote:
| copilot was apparently snipping license bearing comments, and
| applying "semantic" variations of the remaining code.
|
| i would package the entire code as a series of comments, [ideally
| this would be snipped by the pliagarists] leaving a snippet of
| example code that no one of sound mind would allow to execute,
| being proffered by copilot.
| ChrisMarshallNY wrote:
| _> of sound mind_
|
| That's a reach, these days...
|
| I'm seeing some really ... _interesting_ ... behavior, being
| exhibited by folks that, at first blush, I think are kids, just
| out of bootcamp, but, on further inspection, turn out to be
| middle-aged professionals.
|
| I really think Teh Internets Tubes have been rather corrosive
| to collective mental health.
| klyrs wrote:
| The ability to think for oneself will diminish rapidly in an
| environment that rewards one for not doing so.
|
| Smart people still exist. They just aren't online.
| kirth_gersen wrote:
| Suicide by words, here?
| nyc_data_geek wrote:
| The Internet is still one of the easiest ways to find and
| participate in communities and conversations with other
| smart people, if you're invested in vetting and filtering
| who/what you're engaging with.
|
| That said, I expect the ease of such will continue to
| decline as we approach a largely dead Internet, primarily
| consisting of bots talking to bots trying to sell each
| other herbal brain force supplements or whatever
| satvikpendem wrote:
| From Plato's dialogue Phaedrus 14, 274c-275b:
|
| Socrates: I heard, then, that at Naucratis, in Egypt, was
| one of the ancient gods of that country, the one whose
| sacred bird is called the ibis, and the name of the god
| himself was Theuth. He it was who invented numbers and
| arithmetic and geometry and astronomy, also draughts and
| dice, and, most important of all, letters.
|
| Now the king of all Egypt at that time was the god Thamus,
| who lived in the great city of the upper region, which the
| Greeks call the Egyptian Thebes, and they call the god
| himself Ammon. To him came Theuth to show his inventions,
| saying that they ought to be imparted to the other
| Egyptians. But Thamus asked what use there was in each, and
| as Theuth enumerated their uses, expressed praise or blame,
| according as he approved or disapproved.
|
| "The story goes that Thamus said many things to Theuth in
| praise or blame of the various arts, which it would take
| too long to repeat; but when they came to the letters,
| "This invention, O king," said Theuth, "will make the
| Egyptians wiser and will improve their memories; for it is
| an elixir of memory and wisdom that I have discovered." But
| Thamus replied, "Most ingenious Theuth, one man has the
| ability to beget arts, but the ability to judge of their
| usefulness or harmfulness to their users belongs to
| another; and now you, who are the father of letters, have
| been led by your affection to ascribe to them a power the
| opposite of that which they really possess.
|
| "For this invention will produce forgetfulness in the minds
| of those who learn to use it, because they will not
| practice their memory. Their trust in writing, produced by
| external characters which are no part of themselves, will
| discourage the use of their own memory within them. You
| have invented an elixir not of memory, but of reminding;
| and you offer your pupils the appearance of wisdom, not
| true wisdom, for they will read many things without
| instruction and will therefore seem to know many things,
| when they are for the most part ignorant and hard to get
| along with, since they are not wise, but only appear wise."
| ChrisMarshallNY wrote:
| That's great!
|
| They nailed us, what, four thousand years ago?
| courseofaction wrote:
| Awesome. Serves as a counter-example - would HN consider
| literacy to be damaging to the mind, or are we similarly
| mistaken by thinking that LLMs necessarily degrade the
| abilities of their users?
|
| Pre-writing 'texts' (such as the Iliad) were memorized by
| poets, which is reflected in their forms which made more
| use of memory-friendly forms like rhyming, consistent
| meter, and close repetition.
|
| Writing allowed greater complexity and more
| complex/information dense literary forms.
|
| I feel that intelligent, critical LLM usage is just
| writing with less laboriousnes, which opens up the
| writer's ability to explore ideas more widely rather than
| spend their time on the technical aspects of knowledge
| production.
| klyrs wrote:
| Does it serve as a counterexample? Or did the predicted
| loss of memory function come to pass?
|
| Worth noting that people were smoking plain old opium
| back in those times; I'd be reluctant to apply their
| reasoning to fentanyl.
| klyrs wrote:
| Precisely the quote I was thinking of, thank you.
| rolph wrote:
| ..that suggests there is actually a chance that someone would
| go for such a boobytrap.
| bityard wrote:
| This is pretty interesting, and I have conflicted feelings about
| the (seemingly obvious) outcome of this trial.
|
| I wonder, if MS and OpenAI win, does that mean it will be legal
| for anyone to take the leaked source code for a proprietary
| product, train an LLM on it, and then ask the LLM to emit a
| version of it that is different enough to avoid copyright
| infringement?
|
| That would be quite the double-edged sword for proprietary
| software companies.
| ChrisMarshallNY wrote:
| I suspect that this is exactly what will happen; not just with
| code, but also prose and artwork.
|
| Someone is likely to design an LLM that is specifically trained
| to do exactly that.
|
| Lots of money to be made...
| devmor wrote:
| On the matter of artwork there's no need for suspicion - it
| is and has been happening for a while now. There are entire
| online databases dedicated to providing non-consenting
| artist's "styles" as downloadable model parameters by name.
| ChrisMarshallNY wrote:
| Try getting Mickey Mouse comics.
|
| That should be fun...
| satvikpendem wrote:
| Style is not copyrightable so I see nothing wrong with
| making essentially a robot that can paint in the style of
| someone else.
| falcolas wrote:
| In isolation, no. But the produced works can be too close
| for fair use (as demonstrated with the Prince pieces by
| Andy Warhol), and passing it off as a piece from the
| original artist can open you up to forgery/fraud charges.
|
| To put another way, the motivations to produce art in
| another artist's style can still land the artist/buyer in
| legal trouble regardless of fair use.
| satvikpendem wrote:
| Yes that is true, but I don't think the people who use
| style transfer are actually passing it off as the
| original, they just like it for the aesthetic value of
| their own images. In other words, no one using the Van
| Gogh LoRA is actually trying to forge the Starry Night.
| falcolas wrote:
| Given the value of an "authentic" painting of the Starry
| Night (or more realistically the value of something
| forged in, say, Samwise Didier's style) I can't agree
| with "no one".
|
| I have to imagine that it's likely quite popular to sell
| AI generated art that mimics or copies existing works.
| satvikpendem wrote:
| Do you use AI art generators? Flaws are extremely easily
| found out, it is only good for a rough snapshot (without
| much fiddling and even then, artifacts remain). I can
| guarantee you it is definitely _not_ popular to sell
| existing works made with AI, you are better off hiring an
| actual forger. In fact, your suggestion is even the first
| I 've even heard of such an idea.
| 8organicbits wrote:
| I guess there's always a greater fool, but forging an oil
| painting using AI digital images seems pretty far
| fetched.
| devmor wrote:
| The legality of using someone's copyrighted work to train
| a model to reproduce it without their consent is still
| under debate - but the morality of the act at least, is
| not related to its legality - be it positively or
| negatively; and I personally consider it abhorrent.
| satvikpendem wrote:
| Under what morals do you consider it "abhorrent?" I bet
| got a straight answer from those I've asked about this as
| the counter arguments seem too easy to make.
| CuriouslyC wrote:
| I sure wish I could non-consent to people observing me in
| the world, I'd like to move through society invisibly and
| only show myself when it benefitted me. Unfortunately, the
| only answer is to stay inside if I don't want people to see
| me.
| vkou wrote:
| > I sure wish I could non-consent to people observing me
| in the world,
|
| You aren't allowed to use photos _featuring_ a non-
| consenting person to, for instance promote a product.
|
| You are allowed to use photos _including_ a non-
| consenting person.
|
| There's a lot of complicated law, differing between
| different jurisdictions to cover this question, and to
| balance the needs of the public with commercial desires.
| It's not as simple as you make it sound, and there's no
| reason we should just default to bending over backwards
| for commercial interests.
|
| Laws exist to serve society, not the other way around.
| CuriouslyC wrote:
| I'm sure that the people who are being constantly
| victimized by paparazzi would like to know those rules
| that you just quoted, and have them be enforced.
| vkou wrote:
| If you had done a little research into this question,
| you'd realize that 1A use cases ('journalism') are
| treated by law quite differently than use of likeness for
| commercial intent.
|
| This is my whole point. There isn't a single, one-size-
| fits-all rule that a five year old can comprehend that
| describes any particular country's legal framework around
| the many, many different dimensions of tension between
| public and private interests on this incredibly broad
| question.
|
| And none of the existing frameworks fit the new use cases
| well, and we should probably have an open political
| debate about what we want to do going forward.
| CuriouslyC wrote:
| I'll happily take your picture against your will and put
| it on the internet with the tag "vkou mad at
| photographer, news at 11"
| vkou wrote:
| Okay? What will that prove? That you can be an ass?
|
| Being an ass is generally not illegal. _Particular_
| behaviours might be, but no legal or social system
| intends to censure you for every possible one, and most
| people who are experts in law or ethics don 't believe
| that they should.
|
| If you identify particular problems with the particular
| paparazzi laws in your country, that's an interesting
| conversation, and maybe, if framed well, an interesting
| data point for this discussion, but is not in itself the
| 'last word' on it. Just because you can torture an
| analogy, doesn't mean the analogy has a lot of power.
| sweeter wrote:
| > consent Careful... A lot of people online have
| selective understanding when it comes to this concept.
| It's selfishness and self-centredness taken to it's
| extreme, and not seeing other people as humans, but as
| tools for their consumption to be used and tossed aside
| for pleasure or for profit. It's one of the most
| disgusting things I've layed eyes on.
| devmor wrote:
| We are not discussing people observing people. We are
| discussing programs observing people.
| ADeerAppeared wrote:
| > Someone is likely to design an LLM that is specifically
| trained to do exactly that.
|
| Perplexity AI.
| chimeracoder wrote:
| > Perplexity AI.
|
| How does this describe Perplexity AI more than any other
| LLM?
| epolanski wrote:
| I feel like what really matters is who has more money to
| throw in tribunals.
|
| Somehow I feel if it was "Adobe vs dev that claims his code
| was spit by copilot" it would not end the same.
| devmor wrote:
| Following existing law and applying reasonable expectations, I
| would point to the old adage "intent is 9/10ths of the law".
|
| It would probably be legal to do this, as long as no one could
| reasonably show that you intentionally trained the LLM on said
| leaked source code with the intent to reproduce the product.
|
| Of course, civil suits could be another matter entirely. If you
| pick a product to rip off that's owned by a multi-billion
| dollar company, all that can save you is the ethical limits of
| their legal team's consciences.
| spencerflem wrote:
| Not unless a big company is the one doing it lol
| pennomi wrote:
| Or even those AI-powered decompilers people are working on...
| you could clone virtually any software with that. Surely there
| will be limitations.
| beeboobaa3 wrote:
| The limitation is the amount of money & political power the
| owner of the software you're cloning has.
| wongarsu wrote:
| The source code of Windows XP is widely available. Same with
| a ~2 year old version of Bing, Bing Maps, Cortana etc. Yet
| that doesn't seem to have had major negative effects on those
| products. If anything having the Windows source code
| available seems to be a net boon for Windows development.
| Sometimes looking at the source is just better if the
| documentation is unclear.
| Legend2440 wrote:
| I mean you can legally do this by hand right now. That's how
| they cloned the IBM bios back in the day. IBM sued and lost.
| marcosdumay wrote:
| No, that's not.
|
| They cloned the bios by observing how it behaved and writing
| code that behaved the same way. Nobody even looked at the
| bios code.
| wvenable wrote:
| That's not how they did it. They had one team read the BIOS
| source listings in the IBM PC Technical Reference Manual
| and create a technical specification and a second team take
| that specification and write a new BIOS [1]. The second
| team never saw the original code so therefore they could
| not have copied it.
|
| To do something similar with AI, you really need to train
| one AI on the source code and then have it explain that
| code to a second AI that never saw the original code.
|
| [1] https://en.wikipedia.org/wiki/Phoenix_Technologies
| axus wrote:
| I thought there was a "clean room", where the people reading
| it and the people writing it were different; and they made a
| written specification instead of a Vulcan mind meld.
| jeroenhd wrote:
| A Wine fork built using an LLM trained on leaked Windows code
| might be pretty useful.
| Spivak wrote:
| No, because judges aren't robots applying the law like code.
| Intent matters. If you do this it will be painfully obvious
| that your intent is to duplicate a large body copywritten code.
| yellowapple wrote:
| It's painfully obvious that the intent of GitHub Copilot is
| to duplicate a large body of copyrighted code.
| yellowapple wrote:
| That's exactly what I expect to happen with the source code to
| Microsoft's own software products, namely Windows.
|
| Hilarity will ensue :)
| daedrdev wrote:
| > The anonymous programmers have repeatedly insisted Copilot
| could, and would, generate code identical to what they had
| written themselves, which is a key pillar of their lawsuit since
| there is an identicality requirement for their DMCA claim.
| However, Judge Tigar earlier ruled the plaintiffs hadn't actually
| demonstrated instances of this happening, which prompted a
| dismissal of the claim with a chance to amend it.
|
| It sounds fair from how the article describes it
| whimsicalism wrote:
| Huh. There have definitely been well publicized examples of
| this happening, like the quake inverse square root
| polishTar wrote:
| Fast inverse square root is now part of the public domain.
|
| Also, even if this weren't the case you can't sue for damages
| to other people (they'd need to bring their own suit)
| anonymoushn wrote:
| Is the particular implementation that the model spits out
| 70+ years old?
| immibis wrote:
| Has it really already been 70 years since John Carmack
| died?
| polishTar wrote:
| Ah, you're right. I was wrong to say "public domain".
|
| It would be more correct to say Quake III Arena was
| released to the public as free software under the GPLv2
| license.
| KnightHawk3 wrote:
| There is a large gap between public domain and GPL. For
| starters if Copilot is emitting GPL code for closed
| source projects... that's copyright infringement.
| voxic11 wrote:
| You can't copyright a mathematical operation. Only a
| particular implementation of it, and even then it may not be
| copyrightable if its a straightforward and obvious
| implementation.
|
| That said the implementation doesn't appear to be totally
| trivial and copilot apparently even copies the comments which
| are almost certainly copyrightable in themselves.
|
| https://x.com/StefanKarpinski/status/1410971061181681674
| https://github.com/id-Software/Quake-III-
| Arena/blob/dbe4ddb1...
|
| However a twitter post on its own isn't evidence a court will
| accept. You would need the original poster to testify that
| what is seen in the post is actually what he got from copilot
| and not just a meme or joke that he made.
|
| Also the plaintiffs in this case don't include id-Software
| and there is some evidence that id-Software actually stole
| the fast inverse sqrt code from 3dfx so they might not want
| to bring a claim here anyways.
| beeboobaa3 wrote:
| https://en.wikipedia.org/wiki/Illegal_number
| whimsicalism wrote:
| Not sure where you thought I said you could copyright a
| mathematical operation, I was clearly referring to the
| implementation due to the mention of "quake".
|
| When it was reported, I was able to reproduce it myself.
| wongarsu wrote:
| It reads like the judge required them to show it happened to
| their code, not to any code in general. That's a much higher
| bar. There are thousands of instances of fast inverse square
| root in the training data but only one copy of your random
| github repositories. Getting to model to reproduce your code
| verbatim might be possible for all we know, but it isn't
| trivial.
| whimsicalism wrote:
| of course for standing. but it seems like with the right
| plaintiffs this could have gone forward
| daedrdev wrote:
| The article mentions that GitHub copilot has been trained to
| avoid directly copying specific cases it knows, and that
| although you can get it to spit out copyright code by
| prefixing the copyrighted code as a starting point, in normal
| us cases its quite rare.
| ADeerAppeared wrote:
| Where it gets ethnically dubious is that:
|
| 1. The copilot team rushed to slap a copyright filter on top to
| keep these verbatim examples from showing up, and now claims
| they never happen.
|
| 2. LLMs are prone to paraphrasing. Just because you filter out
| verbatim copies doesn't mean there isn't still copyright
| infringement/plagiarism/whatever you want to call it. The
| copyright filter is only a legal protection, not a practical
| protection against the issue of copyright infringement.
|
| Everyone who knows how these systems work understand this. The
| copilot FAQ to this day claims that you should run copyright
| scanning tools on your codebase because your developers might
| "copy code from an online source or library".
|
| Github has it's own research from 2021 showing that these tools
| do indeed copy their training data occasionally:
| https://github.blog/2021-06-30-github-copilot-research-recit...
|
| They clearly know the problem is real. Their own research
| agreed, their FAQs and legal documents are carefully phrased to
| avoid admitting it. But rather than owning up to the problem,
| it's "Ner ner ner ner ner, you can't prove it to a boomer
| judge".
| squarefoot wrote:
| > 1.
|
| Isn't that akin to destruction of evidence?
| ADeerAppeared wrote:
| Legally? No.
|
| In spirit? ... Probably?
|
| Unlike most LLMs, Github copilot can trivially solve their
| copyright problem by just using only code they have the
| right to reproduce.
|
| They have a giant corpus of code tagged with license,
| SELECT BY license MIT/Equivalent and you're done, problem
| solved because those licenses explicitly grant permission
| for this kind of reuse.
|
| (It's still not very cash money to take open source work
| for commercial gain without paying the original authors,
| and there's a humorous question if MIT-copilot would need
| to come with a multi-gigabyte attribution file, but
| everyone widely agrees it's legal and permitted.)
|
| The only reason you'd hack a filter on top rather than
| doing the above is if you'd want to hide the copyright
| problem. It's an objectively worse solution.
| spencerflem wrote:
| No, you misinterpret the MIT licence. It still requires
| attribution, it is not legal to copy MIT code without
| notice.
| abigail95 wrote:
| Not in any way I'm aware of - and would be required if they
| were served a DMCA notification/Cease and Desist against a
| specific prompt.
|
| The people that think Copilot is infringng their copyright
| would be happy with that I would think? Unless they take a
| much stricter definition of fair use than current courts
| do.
| mvdtnz wrote:
| What were the plaintiffs even thinking when they submitted a
| claim based on identicality without being able to produce a
| single instance of copilot generating a verbatim copy. Even the
| research they submitted was unable to make a claim any stronger
| than "it's possibly in theory but we've never seen it".
| loceng wrote:
| This kind of argument makes me feel like it also supports the
| abolition of patents: eventually multiple other people will come
| up with the same obvious solution, which becomes obvious once a
| person spends enough time looking at a problem.
| CodeWriter23 wrote:
| The Patent System is not intended to be a test of exclusive
| original thought.
|
| The function of the Patent System is to incentivize search for
| solutions by temporarily securing exclusive right to market
| novel devices and processes for the discoverer.
| loceng wrote:
| Of non-obvious inventions. My argument being all inventions
| are obvious once attention is applied to that area and scope.
| erik_seaberg wrote:
| Unfortunately USPTO takes "non-obvious" to mean that it wasn't
| already suggested by combining patents or other written work,
| so if you are the first to work a problem you can claim easy
| solutions that anyone with a clue would have quickly reached.
| Land rushes to fence off new fields seem inevitable.
| pledess wrote:
| I thought "the Copilot coding assistant was trained on open
| source software hosted on GitHub and as such would suggest
| snippets from those public projects to other programmers without
| care for licenses" was explicitly allowed by the GitHub Terms of
| Service: https://docs.github.com/en/site-policy/github-
| terms/github-t... "If you set your pages and repositories to be
| viewed publicly, you grant each User of GitHub a nonexclusive,
| worldwide license to use, display, and perform Your Content
| through the GitHub Service." In other words, in addition to
| what's allowed by the LICENSE file in your repo, you are also
| separately licensing your code "to use ... through the GitHub
| Service" and this would (in my interpretation) include use by
| Copilot for training, and use by Copilot to deliver snippets to
| any other GitHub user.
| dmitrygr wrote:
| Lots of my code is on github (eg
| https://github.com/syuu1228/uARM), uploaded by others. I gave
| no license for its use in training. What now?
| zdragnar wrote:
| If the person didn't have your permission or permission from
| the license to agree to github's terms, then you sue the
| person who uploaded it to GitHub.
|
| You don't get to go after GitHub because you have no
| contractual relationship with them. At best, you can get an
| injunction forcing them to take it down, though getting them
| to un-train copilot may not be feasible. At best you'd get a
| small cash offer, since you're unlikely to be able to justify
| any damages in a suit.
| dredmorbius wrote:
| 17 USC SS504 says otherwise:
|
| _... the copyright owner may elect, at any time before
| final judgment is rendered, to recover, instead of actual
| damages and profits, an award of statutory damages for all
| infringements ... in a sum of not less than $750 or more
| than $30,000. ... in a case where the copyright owner
| sustains the burden of proving, and the court finds, that
| infringement was committed willfully, the court in its
| discretion may increase the award of statutory damages to a
| sum of not more than $150,000._
|
| <https://www.law.cornell.edu/uscode/text/17/504>
|
| The issue isn't contract. It's copyright infringement.
| 201984 wrote:
| So hypothetically, if a developer publishes GPL software on
| Codeberg, and someone uploads it to GitHub, could the
| original developer file takedowns against the Github copy?
|
| I'm curious if Github's ToS make uploading GPL software you
| don't own a copyright violation.
| votepaunchy wrote:
| No, because the GPL is already more permissive than the
| GitHub TOS.
| pton_xd wrote:
| > then you sue the person who uploaded it to GitHub.
|
| > You don't get to go after GitHub because you have no
| contractual relationship with them
|
| What makes you say that? If someone eg uploads my
| copyrighted work to YouTube, I file a DMCA notice with
| YouTube to stop distributing my work. If YT ignores the
| notice then I can pursue them with a lawsuit.
|
| How is this situation different?
| singleshot_ wrote:
| DMCA explicitly gives you a cause of action against the
| party who does not properly comply with your request. GP
| asserts that you lack a cause of action against GitHub
| before they fail to comply with DMCA but I'm not certain
| I agree.
| stefan_ wrote:
| DMCA is a narrow protection for operators of public
| websites like GitHub. I don't see what it has to do with
| GitHub taking the data submitted to it with dubious
| sourcing and developing their CoPilot whatever based on
| it. That has nothing to do with the privileges in DMCA.
| singleshot_ wrote:
| That's right. You have lost the thread of what we are
| talking about: causes of action based on privity vs those
| created by statute.
| simion314 wrote:
| That will work if I upload only my code, but there are many
| open source projects where there are more then one author and
| GithHub did not acquired the rights from all the authors, the
| uploader to GitHub might not even be the author too.
| Brian_K_White wrote:
| That just means github can display the code, and you can see
| the code, but that does not mean you can then profit from or
| redistribute (profit or no) the code without attribution.
|
| Amazon has the rights to publish a book, and you have the right
| to receive a copy of the book, but neither of those gives you
| the right to re-publish the book under your own name.
| rurcliped wrote:
| "use, display, and perform Your Content through the GitHub
| Service" might allow a wide range of uses on GitHub Pages
| websites, even if https://example.github.io is monetized
| (monetization is permitted by
| https://docs.github.com/en/site-policy/github-
| terms/github-t... in a few cases)
| purpleblue wrote:
| Can you insist or put instructions that AIs do not train on your
| code? If they train on your code but don't produce the exact same
| output, is there any protection you can have from that?
| archontes wrote:
| When are people going to get that this isn't a right folks
| have?
|
| If your code is readable, the public can learn from it.
|
| Copyright doesn't extend to function.
| ADeerAppeared wrote:
| People aren't going to get it, because you don't get them.
|
| People have the right to learn _non-copyrightable elements_
| from your code.
|
| The claim is that AI learns _copyrightable elements_.
| archontes wrote:
| The comment chain you are replying to includes a request to
| not train an AI on one's code.
|
| I agree it's certainly possible for AI to produce
| infringing output.
|
| Nevertheless, people don't have the right to enforce a
| limitation on training.
| munificent wrote:
| _> Indeed, last year GitHub was said to have tuned its
| programming assistant to generate slight variations of ingested
| training code to prevent its output from being accused of being
| an exact copy of licensed software._
|
| If I, a human, were to:
|
| 1. Carefully read and memorize some copyrighted code.
|
| 2. Produce new code that is textually identical to that. But in
| the process of typing it up, I randomly mechanically tweak a few
| identifiers or something to produce code that has the exact same
| semantics but isn't character-wise identical.
|
| 3. Claim that as new original code without the original
| copyright.
|
| I assume that I would get my ass kicked legally speaking. That
| reads to me exactly like deliberate copyright infringement with
| willful obfuscation of my infringement.
|
| How is it any different when a machine does the same thing?
| singleshot_ wrote:
| The guy who owns the machine is really rich, while you are more
| or less (all due respect of course) not worth suing.
|
| That's why I think the opposite of what you claim is true: if
| you were to do this, absolutely nothing would happen. When they
| do it, they will get sued over and over until the law changes
| and they can't be sued, or they enter some mutually-beneficial
| relationship with the parties who keep suing.
| beeboobaa3 wrote:
| > if you were to do this, absolutely nothing would happen
|
| Read up on the DMCA and the impact it has on e.g. nintendo
| emulators and the developers thereof
| dmix wrote:
| Those emulators are very popular though to the point of
| potentially impacting another business's bottom line. Where
| an individual putting it out a small block of code isn't
| exactly going to attract expensive lawyers.
|
| I'm skeptical Github Copilot reproducing a couple functions
| potentially used by some random Github project is going to
| be a threat to another party's livelihood.
|
| When AI gets good enough to make full duplicates of apps
| I'd be more concerned about the source. Thousands of
| smaller pieces drawn from a million sources and being
| combined in novel ways is less worrying though.
| BadHumans wrote:
| There is no impact to a company's bottom line when you
| are emulating a product they do not sell.
| lcouturi wrote:
| Yuzu, the emulator that was sued by Nintendo, was
| emulating the Nintendo Switch, which is a product
| Nintendo does sell.
| BadHumans wrote:
| Yuzu is not the only emulator taken down by Nintendo and
| Nintendo is not the only company that has gone after
| emulators.
| beeboobaa3 wrote:
| Rules for thee but not for me (rich companies). Think of the
| shareholders!
| Analemma_ wrote:
| > How is it any different when a machine does the same thing?
|
| Because intent matters in the law. If you intended to reproduce
| copyrighted code verbatim but tried to hide your activity with
| a few tweaks, that's a very different thing from using a tool
| which _occasionally_ reproduces copyrighted code by accident
| but clearly was not designed for that purpose, and much more
| often than not outputs transformative works.
| archontes wrote:
| Not in copyright. The work speaks for itself, and the
| function of code is not a copyrightable aspect.
| bawolff wrote:
| The intent of the work can matter when determining if de
| minimis applies as well as fair use.
| olliej wrote:
| Um, the entire intent of these "AI" systems is explicitly to
| reproduce copyrighted work with mechanical changes to make it
| not appear to be a verbatim copy.
|
| That is the whole purpose and mechanism by which they
| operate.
|
| Also the intent does not matter under law - not intending to
| break the law is not a defense if you break the law. Not
| intending to take someone's property doesn't mean it becomes
| your property. You _might_ get less penalties and /or
| charges, due to intent (the obvious examples being murder vs
| manslaughter, etc).
|
| But here we have an entire ecosystem where the model is "scan
| copyrighted material" followed by "regurgitate that material
| with mechanical changes to fit the surrounding context and to
| appear to be 'new' content".
|
| Moreover given that this 'new' code is just a regurgitation
| of existing code with mutations to make it appear to fit the
| context and not directly identical to the existing code, then
| that 'new' code cannot be subject to copyright (you can't
| claim copyright to something you did not create, copyright
| does not protect output of mechanical or automatic
| transformations of other copyrighted content, and copyright
| does not protect the result of "natural processes", e.g 'I
| asked a statistical model to give me a statically plausible
| sequence of tokens and it did'). So in the best case scenario
| - the one where the copyright laundering as a service tool is
| not treated as just that, any code it produces is not
| protectable by copyright, and anyone can just copy "your
| work" without the license and (because you've said if you
| weren't intending to violate copyright it's ok) they can say
| they could not distinguish the non-copyright-protected work
| from the protected work and assumed that therefore none of it
| was subject to copyright. To be super sure though they
| weren't violating any of your copyrights, they then ran an
| "AI tool" to make the names better and better suit your
| style.
|
| I am so sick of these arguments where people spout nonsense
| about "AI" systems magically "understanding" or "knowing"
| anything - they are very expensive statistical models, the
| produce statistically plausible strings of text, by a
| combination of copying the text of others wholesale, and
| filling the remaining space with bullshit that for basic
| tasks is often correct enough, and for anything else is wrong
| - because again they're just producing plausible sequences of
| tokens and have no understanding of anything beyond that.
|
| To be very very very clear: if an AI system "understood"
| anything it was doing, it would not need to ingest
| essentially all the text that anyone has ever written, just
| to produce content that is at best only locally coherent, and
| that is frequently incorrect in more or less every domain to
| which it is applied. Take code completion (as in this case):
| Developers can write code without essentially reading all the
| code that has ever existed just so that they can write basic
| code, because developers understand code. Developers don't
| intermingle random unrelated and non-present variables or
| functions in their code as they write, because they
| understand what variables are and therefore they can't use
| non existent ones. "AI" on the other hand required more power
| than many countries to "learn" by reading as much as possible
| all code ever written, and then produce nonsense output for
| anything complex because they're still just generating a
| string of tokens that is plausible according to their
| statistical model - the result of these AIs is essentially
| binary: it has been in effect asked to produce code that does
| something that was in its training corpus and can be copied
| essentially verbatim, with a transformation path to make it
| fit, or it's not in the training corpus and you get random
| and generally incorrect code - hopefully wrong enough it
| fails to build, because they're also good at generating code
| that looks plausible but only fails at runtime because
| plausible sequence of tokens often overlaps with 'things a
| compiler will accept'.
| shkkmo wrote:
| > Also the intent does not matter under law - not intending
| to break the law is not a defense if you break the law
|
| Intent frequently matters a great deal when applying laws.
|
| In the specific area of copyright law, it doesn't itself
| make the use non infringing, but it can absolutely impact
| the damages or a fair use argument.
| dmix wrote:
| That's a significant over simplification of how it works though
| to the point of almost not being a useful analogy.
|
| If your analogy was you were a human who memorized every
| variation of a problem (and every other known problem) and
| there was a tiny perctange of a chance where you reproduced
| that exact varation of one you memorized, but then added an
| after the fact filter so you don't directly reproduce it...
|
| It's more like musicians who basically copy a bunch of music
| patterns or chord progressions before then notice their final
| output sounds too similar to another song (which happens often
| IRL) then changes it to be more original before releasing it to
| the public.
| ADeerAppeared wrote:
| > If you analogy was you were a human who memorized every
| variation of a problem (and every other known problem)
|
| This is mere assumption. AI is _supposed to_ work like that,
| but that 's a goal, and not the result of current
| implementations. Research shows that they do memorize
| _solutions_ as well, and quite regularly so. (This is an
| unavoidable flaw in current LLMs; They must be capable of
| memorizing input verbatim in order to learn specific facts.)
|
| > and there was a tiny perctange of a chance where you
| reproduced that exact varation of one you memorized
|
| This is copyright infringement. _Actionable_ copyright
| infringement. The big music publishers go after this kind of
| accidental partial reproduction.
|
| > but then added an after the fact filter so you don't
| directly reproduce it...
|
| "Legally distinct" is a gimmick that only works where the
| copyright is on specific identifiable parts of a work.
|
| Changing a variable name does not make a code snippet
| "legally distinct", it's still copyright infringement.
| dmix wrote:
| Meh I still see that as a big oversimplification. Context
| matters. Even if the copyright courts often ignore that for
| wealthy entities. Someone reproducing a song using AI and
| publishing it as their own copyright infringement, a person
| specifically querying an AI engine, that sucked up billions
| of lines of information and generates what you ask it do
| with a sma probability it will reproduce a small subset of
| a larger commercial project and sends it to someone in a
| chatbox is not exactly the same IMO.
|
| This is Github Copilot after all. I use it daily and it
| autocompletes lines of code or generates functions you can
| find on stackoverflow. It's not letting giving you the
| source code to Twitter in full and letting you put it on
| the internet as a business under another name.
| belorn wrote:
| We are currently seeing the music industry reacting to AI
| learning a bunch of music patterns and chord progressions and
| outputting works that sounds very similar to existing music
| and artists. They are not liking it.
|
| To just see how much they disliked it, youtube copyright
| strikes is basically a trained AI to detect music patterns to
| identify sound with slight variations or copyrighted songs
| and take videos down. Generating slight variations was one of
| the early method that videos used to bypass the take down
| system.
| archontes wrote:
| You might not get your ass kicked. Copyright doesn't protect
| function, to the point where the court will assess the degree
| to which the style of the code can be separated from the
| function. In the even that they aren't separable, the code is
| not copyrightable.
|
| https://www.wardandsmith.com/articles/supreme-court-announce...
|
| https://easlerlaw.com/software-computer-code-copyrighted#:~:...
| ADeerAppeared wrote:
| The simple version is that code _is_ copyrightable as an
| _expression_. And the underlaying algorithm is _patentable_.
|
| The legal term you're looking for here is the "Abstraction-
| Filtration-Comparison" test; What remains if you subtract all
| the non-copyrightable elements from a given piece of code.
| tomxor wrote:
| US copyright does protect for "substantial similarity" [0].
| And at the other end of the spectrum, this has been abused in
| absurd ways to argue that substantially different code has
| infringed.
|
| In Zenimax vs Oculus they basically argued that a bunch of
| really abstract yet entirely generic parts of the code were
| shared, we are talking some nested for loops, certain
| combinations of if statements, and due to a lack of a
| qualitative understanding of code, syntax, common patterns,
| and what might actually qualify for substantively novel code
| in the courtroom, this was accepted as infringing. [1]
|
| Point is, the legal system is highly selective when it comes
| to corporate interests.
|
| [0] https://en.wikipedia.org/wiki/Substantial_similarity
|
| [1] https://arstechnica.com/gaming/2017/02/doom-co-creator-
| defen...
| talldayo wrote:
| > Point is, the legal system is highly selective when it
| comes to corporate interests.
|
| I don't even think it's that. In recent cases like Oracle
| v. Google and Corellium v. Apple, Fair Use prevailed with
| all sorts of conflicting corporate interests at play. The
| Zenimax v. Oculus case very much revolved around NDAs that
| Carmack had signed and not the propagation of trade
| secrets. Where IP is strictly the only thing being
| concerned, the literal interpretation of Fair Use does
| still seem to exist.
|
| Or for a more literal example, Authors Guild. v. Google
| where Google defended their indexing of thousands of
| licensed books as Fair Use.
| tpmoney wrote:
| In fact, go to far as to argue your example of Authors
| Guild v. Google is a good indication that most cases will
| probably go an AI platform's way. It's a pretty parallel
| case to a number of the arguments. Indexing required
| ingesting whole works of copyright material verbatim. It
| utilized that ingested data to produce a new commercial
| work consisting of output derived from that data. If I
| remember the case correctly, google even displayed
| snippets when matching a search so the searcher could see
| the match in context, reproducing the works verbatim for
| those snippets and one could presume (though I don't
| recall if it was coded against), that with sufficiently
| clever search prompts, someone could get the index search
| to reproduce a substantial portion of a work.
|
| Arguably, the AI platforms have an even stronger case as
| their nominal goal is not to have their systems reproduce
| any part of the works verbatim.
| ars wrote:
| > I assume that I would get my ass kicked legally speaking.
|
| Maybe, maybe not. It's not as simple as you made it out to be.
| If you write a book with lots of stuff and you got inspiration
| from other books, and even put in phrases wholesale, but
| modified to use your own character names instead, I'm not
| convinced you would lose.
|
| The court would look at the work as a whole, not single pieces
| of it.
|
| They would also check if you are just copying things verbatim,
| or if you memorize a pattern and emit the same pattern - for
| example look at lawsuits about copying music, where they'll
| claim this part of the music is the same as that part.
|
| It's really not as cut and dry as you make it out to be.
| williamcotton wrote:
| Just to set the stage and not entirely specific to this
| complaint... It really depends on what is and isn't subject to
| copyright for software.
|
| Broadly, there is the distinction between expressive and
| functional code. [1]
|
| And then there are the specific tests that have been developed
| by the courts to separate the expressive and functional aspects
| of software. [2] [3]
|
| In practice it is very expensive for a plaintiff to do such
| analysis. For the most part the damages related to copyright
| are not worth the time and money. Plaintiffs tend to go for
| trade secret related damages as they are not restricted by the
| above tests.
|
| There are also arguments to be made of de minimis infringements
| that are not worth the time of the court.
|
| Most importantly the plaintiff fundamentally has the burden of
| proof and cannot just say that copying must have taken place.
| They need concrete evidence.
|
| [1] https://en.wikipedia.org/wiki/Idea-expression_distinction
|
| [2]
| https://en.wikipedia.org/wiki/Structure,_sequence_and_organi...
|
| [3] https://en.wikipedia.org/wiki/Abstraction-Filtration-
| Compari...
| wvenable wrote:
| You probably do this all the time. Forget memorizing but
| undoubtedly you've read code, learned from it, and then likely
| reproduced similar code. Probably nothing terribly important,
| just a function here or there. Maybe even reproduced something
| you did for a previous employer.
| JoshTriplett wrote:
| You have a much smaller lobbying budget than the AI industry,
| and you didn't flagrantly rush to copy billions of copyrighted
| works as quickly as possible and then push a narrative acting
| like that's the immutable status quo that must continue to be
| permitted lest the now-massive industry built atop copyright
| violation be destroyed.
|
| Violate one or two copyrights, get sued or DMCAed out of
| existence. Violate billions, on the other hand, and you
| magically become immune to the rules everyone else has to
| follow.
| nadermx wrote:
| What about the copyrights purpose of furthering the arts and
| sciences?
| JoshTriplett wrote:
| Copyright has utterly failed to serve that purpose for a
| long time, and has been actively counterproductive.
|
| But if you want to argue that copyright is
| counterproductive, I completely agree. That's an argument
| for reducing or eliminating it across the board, fairly,
| for everyone; it's not an argument for giving a free pass
| to AI training while still enforcing it on everyone _else_.
| hyperpape wrote:
| From the article:
|
| > The most recently dismissed claims were fairly important,
| with one pertaining to infringement under the Digital
| Millennium Copyright Act (DMCA), section 1202(b), which
| basically says you shouldn't remove without permission crucial
| "copyright management" information, such as in this context who
| wrote the code and the terms of use, as licenses tend to
| dictate.
|
| > It was argued in the class-action suit that Copilot was
| stripping that info out when offering code snippets from
| people's projects, which in their view would break 1202(b).
|
| > The judge disagreed, however, on the grounds that the code
| suggested by Copilot was not identical enough to the
| developers' own copyright-protected work, and thus section
| 1202(b) did not apply. Indeed, last year GitHub was said to
| have tuned its programming assistant to generate slight
| variations of ingested training code to prevent its output from
| being accused of being an exact copy of licensed software.
|
| So (not a lawyer!) this reads like the point about GitHub
| tuning their model is not a generic defense against any and all
| claims of copyright infringement, but a response to a specific
| claim that this violates a provision of the DMCA.
|
| I don't know whether this is a reasonable defense or not, but
| your intuitions or mine about whether there is a general
| copyright violation or what's fair are not necessarily relevant
| to how the judge construes that very specific bit of legal
| code.
| skybrian wrote:
| The machine alone doesn't do anything. The user and machine
| together constitute a larger system, and with autocomplete, the
| user is charge. What's the user's intent?
|
| I suspect that a lot of copyright violations are enabled by
| cut-and-paste and screenshot-taking functionality, and maybe we
| need to be careful with autocomplete, too? It's the user's
| responsibility to avoid this. We should be careful using our
| tools. Do users take enough care in this case? Is it possible
| to take enough care while still using CoPilot?
|
| I've switched from CoPilot to Cody, but I use them the same
| way, to write _my_ code. There 's no particular reason to use
| CoPilot's output verbatim and lots of good reasons not to. By
| the time I've adapted it to my code base and code style and
| refactored it to hell and back, it's an expression of how _I_
| want to solve a problem, and I 'm pretty confident claiming
| ownership.
|
| Is that confidence misplaced? Are other people more careless?
| lnxg33k1 wrote:
| Suddenly, you would steal a car
| blooalien wrote:
| | "Suddenly, you would steal a car"
|
| Nah, but I would _download_ a _copy_ of one without
| hesitation... ;)
| hn_throwaway_99 wrote:
| A slight aside, but this is the subtitle:
|
| > A few devs versus the powerful forces of Redmond - who did you
| think was going to win?
|
| I hate that kind of obnoxious "journalism". Sometimes the little
| guy is actually wrong. To clarify, I'm not commenting on the
| specifics of this case, I just hate how fake our online discourse
| has been by appealing to "big guy evil" before even bringing up
| the specifics of the case.
| epolanski wrote:
| I think you're misinterpreting the sentence.
|
| I think it merely implies MS has more resources to throw at the
| legal case.
| deciplex wrote:
| > Sometimes the little guy is actually wrong.
|
| He is, sometimes. Also sometimes, the moon passes exactly
| between the sun and Earth, a new star appears in the sky, the
| magnetic field of our planet reverses, a proton decays (jury is
| still out on that one, actually). Etc.
|
| Tools like Copilot are plagiarism machines. We know the data
| they're being trained on, and a conclusion of "that's
| plagiarism" is not - or anyway should not be - controversial.
| I'm not terribly _against_ the notion of a plagiarism machine
| but I am against the owners of such machines reaping profits
| from them to the exclusion of the people who provide the source
| material. This is theft.
|
| More importantly, getting back to big guys and little guys: big
| guys gang up on little guys all the time. It's usually how they
| get to be big. They tend to be the ones who realize that
| working together against the rest of us is to their benefit.
| So, in the interest of pushing back on that a little, and
| recognizing that I am after all a fellow "little guy"
| (figuratively speaking anyway), I tend to support the "little
| guy" unless I have overwhelming evidence confirming that they
| are, in fact, both _wrong_ and that _supporting them anyway
| would be against my best interest._ Neither is the case, here.
|
| At any rate, the subtitle here references a pretty ubiquitous
| and, I'm happy to report, increasingly well-known and
| understood facet of our economic and social institutions, which
| is that they absolutely positively do not work for us or
| further our interests in any sense.
| epolanski wrote:
| I am not strongly opinionated on this, but the very fact
| Microsoft used all the code it could find, bar their own has
| always looked suspicious to me.
___________________________________________________________________
(page generated 2024-07-09 23:00 UTC)