[HN Gopher] Microsofts AI boss thinks its perfectly OK to steal ...
___________________________________________________________________
Microsofts AI boss thinks its perfectly OK to steal content if its
on open web
Author : avivallssa
Score : 55 points
Date : 2024-06-29 21:00 UTC (2 hours ago)
(HTM) web link (www.theverge.com)
(TXT) w3m dump (www.theverge.com)
| avivallssa wrote:
| Will this make people who make indirect money through their
| content, less motivated from publishing their content on the Web
| ? This might be arguable.
|
| May be, there should be a similar amount of openness in
| publishing the content used for training commercial models.
|
| The copyright owner should have a privilege to ask for that
| content to be removed from training. This may also allow
| individual authors to gain their share with their Advanced RAG
| applications, that are specially focussed on the content they own
| and also published on the web.
| talldayo wrote:
| Steelmanning against this, when you publish any form of content
| online you have to be prepared for the consequences of it's
| digital proliferation. When Napster was all the rage people
| said the same thing about music, and before that they decried
| home taping as the death of the music industry. The music
| industry lived on, it just changed form as the shape of music
| did.
|
| If the online newsletter community that Substack and Medium
| built consequently dies from sheepish author syndrome, very
| little will change. The content they made will be replaced, and
| the internet will survive fine without them as it did for
| dozens of years before digital subscription services were a
| realistic revenue stream.
| LegitShady wrote:
| >Steelmanning against this, when you publish any form of
| content online you have to be prepared for the consequences
| of it's digital proliferation.
|
| when you walk down the street in skimpy dress you have you
| have to be prepared for the consequences too, right? Whatever
| you're trying to say it has nothing to do with what rights
| companies have to use content.
| talldayo wrote:
| The point is that you cannot claim damages to something
| that you give away for free. What did they take from you,
| notoriety? Content? Traffic?
|
| When the Author's Guild pressed Google for their indexing
| of plaintext copywritten books, they lost in court.
| Transforming freely-available content can't be gatekept
| because the intent isn't strictly what the author imagined.
| There is a degree of fair use that exists when you make
| anything public. Music, art, text, videos, all of it can be
| consumed in novel and unexpected ways. People haven't been
| concerned about the legal ramifications of abusing
| intellectual property since teenage Neil Cicierega made Mr
| Rogers fight Batman in 2005:
| https://www.youtube.com/watch?v=lrzKT-dFUjE
| Dalewyn wrote:
| >The copyright owner should have a privilege to ask for that
| content to be removed
|
| Just so you know, privileges can be (and probably will be in
| this case) denied.
|
| Rights on the other hand can't be denied.
| dialup_sounds wrote:
| Right, nobody on the internet has ever violated copyright and
| gotten away with it. /s
| tiahura wrote:
| The open web's ethos since its inception in the 1990s has been
| one of unrestricted access and fair use. Content published openly
| online inherently invites broad consumption, reproduction, and
| creative reuse by the public. This is not merely custom, but a
| fundamental aspect of fair use doctrine as applied to the digital
| realm.
|
| The four factors of fair use - purpose of use, nature of the
| copyrighted work, amount used, and effect on the market -
| overwhelmingly favor allowing free use of openly published web
| content. The transformative nature of most reuses, the public
| availability of the original works, the necessity of using entire
| works in many cases, and the lack of a traditional market for
| such content all support this interpretation.
|
| This longstanding practice has been the catalyst for
| unprecedented innovation and information dissemination. It
| represents a tacit social contract between content creators and
| users, establishing a de facto "freeware" model for open web
| content. Any attempt to retroactively impose strict copyright
| limitations would not only stifle innovation but also contradict
| decades of established legal precedent and digital norms.
|
| -As a side note, I'm not certain that training necessarily
| involves "copying."
|
| ---Lastly, if anyone really thinks the Robert's court is going to
| knee-cap AI, you're soft in the head.
| pessimizer wrote:
| > Content published openly online inherently invites broad
| consumption, reproduction, and creative reuse by the public.
|
| You're using "invites" as a weaselword here. Then you go into
| the law, as if the law is discussing "invitations."
|
| Counterpoint: content published online is usually extremely
| hostile to reproduction, and the people who produce it are
| deathly afraid of other people copying their work and
| outranking them with it.
|
| > ---Lastly, if anyone really thinks the Robert's court is
| going to knee-cap AI, you're soft in the head.
|
| If you think the Roberts court is going to be hostile to
| copyright, you're insane.
| amiantos wrote:
| I am optimistic but I feel sad when I remember when went
| through all this with sampling 30 years ago and now the music
| industry is more insular and controls every scrap of music they
| can find so no one can sample anything and even if your song
| has similar "vibes" to another, well, now it belongs to them as
| well.
|
| I'd wager that 90% of the people cheering on the RIAA (and
| copyright in general) now were singing a much different tune
| before they decided AI was a threat to their livelihoods.
| There's a lot of reasons to not be optimistic about the
| continued legal existence of freely available open source AI,
| because if the copyright holders have their way, they will be
| the only entities that control all the data needed to train it.
| And many people on the internet are all too happy to cheer this
| on without realizing that once the RIAA, MPAA, and publishers
| (when they finally figure out how to organize effectively like
| film and music has) hold the reigns, the AI will still exist,
| and it will still take all the same jobs, but all the open
| source and freely available AI that everyone could use, that
| could level the playing field for everyone... that's going to
| be illegal for people to use, without being able to pay the
| content rights holders. Just another way to keep it to have /
| have nots, same as it's always been.
| cjk2 wrote:
| Ah yes the implied social contract that it's ok because it
| happens all the time.
|
| That's how society falls.
| mewpmewp2 wrote:
| What exactly is wrong with the statement he has made?
| elicksaur wrote:
| The article does a pretty good job outlining the actual legal
| rights of published content and how his statement is not
| supported by laws and precedent.
| jprete wrote:
| The Verge quotes: "I think that with respect to content that's
| already on the open web, the social contract of that content
| since the '90s has been that it is fair use. Anyone can copy
| it, recreate with it, reproduce with it. That has been
| 'freeware,' if you like, that's been the understanding."
|
| This statement is quite wrong. People have complained about
| content farms and plagiarists ripping off their content for
| ages. The only kinds of Web content I can think of where his
| statement could apply are open source software (for the more
| permissive licenses) and Stack Overflow posts. Almost
| everything else was posted with the intent of copyright.
| trueismywork wrote:
| Content on web eithout any copyright notice is by default "all
| rights reserved" content. So you cannot use it to produce new
| content, unless you can prove that the old content was never
| reproduced ever along all the invocation of the LLM
| seanmcdirmid wrote:
| You can totally study it and learn from a bunch of content on
| the web and then create a new work having looked at other
| works before. My writing is influenced by lots of copyrighted
| work I've read already, there is no mental blockade going on
| when I write something. The only question is what does it
| mean when a machine does that instead?
|
| The old content can't be reproduced by an LLM unless the
| content is provided as part of its prompt. At least, that's
| what the AI companies claim and will have to show in court.
| fimdomeio wrote:
| So we've now learned that copyright is determined by
| communications protocol. If you're using torrents it's copyright
| infringement, if it's the web then it's public domain.
| echelon wrote:
| > if it's the web then it's public domain.
|
| Only if you're a big company. If you're an individual or small
| company, then it isn't.
| mordae wrote:
| Not sure why downvoted.
| jsyang00 wrote:
| No he doesn't.
|
| > I think that with respect to content that's already on the open
| web, the social contract of that content since the '90s has been
| that it is fair use. Anyone can copy it, recreate with it,
| reproduce with it. That has been "freeware," if you like, that's
| been the understanding.
|
| > There's a separate category where a website, or a publisher, or
| a news organization had explicitly said 'do not scrape or crawl
| me for any other reason than indexing me so that other people can
| find this content.' That's a grey area, and I think it's going to
| work its way through the courts.
| madeofpalk wrote:
| Yes, he does.
|
| > content that's already on the open web, the social contract
| of that content since the '90s has been that it is fair use.
| Anyone can copy it, recreate with it, reproduce with it. That
| has been "freeware," if you like, that's been the understanding
|
| Sure, it's his belief, but this statement is completely
| incorrect. The words that he's using mean specific things, and
| he's just got it wrong. You would expect he's smart enough to
| know this, but 'it is difficult to get a man to understand
| something, when his salary depends on his not understanding
| it'.
|
| > That's a grey area, and I think it's going to work its way
| through the courts.
|
| It's not a grey area. Putting up a robots.txt doesn't change
| copyright, and it certainly doesn't make it a 'grey area'.
| elicksaur wrote:
| Most people would read the first quote as totally aligned with
| the article's title.
| joe_the_user wrote:
| It's an incredibly fuzzy statement, really.
|
| You can copy copyrighted material you download for your own use
| and the use of your friends. You can't copy it to a different
| website and distribute it further on the "open web". You can
| modify material provided for download in your own home for any
| purpose you wish but unless that modification is in the
| category "fair use" (a small part, a parody, etc) you can't
| distribute it freely on the web either.
|
| His statement might mean this but it could mean a zillion other
| things too (what does "it is fair use" mean? etc)
|
| Whether using randomly obtained copyrighted material to "train"
| an LLM and then selling the output of that LLM is "fair use"
| seems like so far another "gray area" and the foggy statement
| seems oriented to reducing awareness of that situation.
| jprete wrote:
| Legally you cannot actually do the first thing without
| permission. The fact that it was technologically impossible
| to stop, and the damages would be impossible to prove, didn't
| make it a legal right.
| raincole wrote:
| > I think that with respect to content that's already on the
| open web, the social contract of that content since the '90s
| has been that it is fair use. Anyone can copy it, recreate with
| it, reproduce with it. That has been "freeware," if you like,
| that's been the understanding.
|
| is _literally_
|
| > I think it's perfectly OK to use content in arbitrary way if
| it's on open web
|
| The only difference between this and the title is he doesn't
| think this behavior is called "stealing".
| advael wrote:
| I mean, that seems to be exactly how he's defining "open web"
| here, actually. That which is - in the dichotomy presented by
| these two quotes - "the open web" is free game for any use, and
| he defines things that use language that explicitly disallows
| all uses except indexing as the complement of this category.
| Maybe he'd accept any site that effectively declares any
| "whitelist" of acceptable uses in this category too, though
| this isn't explicitly stated.
|
| His contention is an assumptive close, wrapping the assumption
| that anything not explicitly labeled otherwise must use a
| "blacklist" policy where any usage not specifically forbidden
| is permitted into "the social contract" that he claims to be so
| obvious as to not permit challenge
|
| He would like the "grey area" of legal debate on this matter,
| as he explained quite clearly, to be exclusively about whether
| AI models can be enforcably barred from training on content for
| which such a narrow whitelist of acceptable uses has been
| defined. Naturally this would mean both that the courts _could_
| decide such a blanket ban can 't bar msft (or anyone) from
| using this content to train AI models, but also that the court
| needn't or maybe even can't decide that failure to ban this use
| case explicitly (or adopt a similar "whitelist" style blanket
| ban) makes acceptance of it legally implied. Hell, he even
| leaves room for explicitly banning this use to be rendered
| legally unenforceable
|
| I can see why he would want that to be the overton window!
| FactKnower69 wrote:
| >I think that with respect to content that's already on the
| open web, the social contract of that content since the '90s
| has been that it is fair use. Anyone can copy it, recreate with
| it, reproduce with it.
|
| Good thing it doesn't matter what he "thinks" the "social
| contract" is, copyright is automatic.
| Forge36 wrote:
| That "separate category" isn't a explicit opt in. It's its two
| different mechanisms for indexing and copyright.
|
| Index is your out. You can use our content is opt in.
|
| Reproduction and recreation, especially when taken physically
| outside of the Internet or into products for sale has always
| been a against the rules. As mentioned by another post,
| torrents of music and movies solidified this stance legally.
|
| Unindexed connect can be open source no strings attached.
| krisoft wrote:
| > No he doesn't.
|
| Could you please help me see where you see the difference
| between the title and the quotes? Even after reading them it
| seems the title is substantially true?
|
| Or to be curt while mirroring your comment's style: "Yes he
| does."
| starik36 wrote:
| The more I read about this guy the more I get the feeling that he
| is an unscrupulous individual.
|
| robots.txt is a "grey idea" to him, instead of being a directive
| to keep moving? Wow.
| JonChesterfield wrote:
| I'll bet they don't consider the windows and office source code
| fair game for arbitrary reuse provided the other party found the
| copy on the web. Even if the person found the copy on GitHub.
| 29athrowaway wrote:
| One thing is a robots.txt policy, meant mostly for search
| crawlers.
|
| Another thing is the copyright of the content, terms of use
| policies, etc.
|
| Abiding by a robots.txt policy doesn't make you immune to
| copyright, terms of service, law in various jurisdictions, etc.
| If you think that you are probably a kleptomaniac.
|
| Just create a robots.txt with "User-Agent: one billion asterisks"
| so that the crawlers die when parsing it.
| KoolKat23 wrote:
| This is nothing but performative clickbait by the Verge.
|
| It is classified as fair use, the term is transformative use,
| where those using it are training models (their intention) if
| anyone wishes to Google it.
|
| The end.
| jimmaswell wrote:
| From the beginning, it's seemed completely intuitive to me that
| training a computer made of sand on publicly available content
| and then generating art later should be fair use, so long as it's
| fair use to train the meat computer in your head on the same
| content and then use it to generate art later. There's no
| meaningful difference to me as far as the ethics of the act are
| concerned.
| sircastor wrote:
| Part of the problem here is that the web has gone through lots of
| change as to what it is and how people understand it.
|
| Some people think of it as billboards posted on the highway. Some
| think it's a bulletin board. Some think it's a newspaper. A
| television, a "zine", a diary, graffiti. It has been all of these
| things, and is and isn't. And people who publish are really bad
| at explicitly stating which one they are. But they expect you to
| know.
___________________________________________________________________
(page generated 2024-06-29 23:02 UTC)