[HN Gopher] Microsofts AI boss thinks its perfectly OK to steal ...
       ___________________________________________________________________
        
       Microsofts AI boss thinks its perfectly OK to steal content if its
       on open web
        
       Author : avivallssa
       Score  : 55 points
       Date   : 2024-06-29 21:00 UTC (2 hours ago)
        
 (HTM) web link (www.theverge.com)
 (TXT) w3m dump (www.theverge.com)
        
       | avivallssa wrote:
       | Will this make people who make indirect money through their
       | content, less motivated from publishing their content on the Web
       | ? This might be arguable.
       | 
       | May be, there should be a similar amount of openness in
       | publishing the content used for training commercial models.
       | 
       | The copyright owner should have a privilege to ask for that
       | content to be removed from training. This may also allow
       | individual authors to gain their share with their Advanced RAG
       | applications, that are specially focussed on the content they own
       | and also published on the web.
        
         | talldayo wrote:
         | Steelmanning against this, when you publish any form of content
         | online you have to be prepared for the consequences of it's
         | digital proliferation. When Napster was all the rage people
         | said the same thing about music, and before that they decried
         | home taping as the death of the music industry. The music
         | industry lived on, it just changed form as the shape of music
         | did.
         | 
         | If the online newsletter community that Substack and Medium
         | built consequently dies from sheepish author syndrome, very
         | little will change. The content they made will be replaced, and
         | the internet will survive fine without them as it did for
         | dozens of years before digital subscription services were a
         | realistic revenue stream.
        
           | LegitShady wrote:
           | >Steelmanning against this, when you publish any form of
           | content online you have to be prepared for the consequences
           | of it's digital proliferation.
           | 
           | when you walk down the street in skimpy dress you have you
           | have to be prepared for the consequences too, right? Whatever
           | you're trying to say it has nothing to do with what rights
           | companies have to use content.
        
             | talldayo wrote:
             | The point is that you cannot claim damages to something
             | that you give away for free. What did they take from you,
             | notoriety? Content? Traffic?
             | 
             | When the Author's Guild pressed Google for their indexing
             | of plaintext copywritten books, they lost in court.
             | Transforming freely-available content can't be gatekept
             | because the intent isn't strictly what the author imagined.
             | There is a degree of fair use that exists when you make
             | anything public. Music, art, text, videos, all of it can be
             | consumed in novel and unexpected ways. People haven't been
             | concerned about the legal ramifications of abusing
             | intellectual property since teenage Neil Cicierega made Mr
             | Rogers fight Batman in 2005:
             | https://www.youtube.com/watch?v=lrzKT-dFUjE
        
         | Dalewyn wrote:
         | >The copyright owner should have a privilege to ask for that
         | content to be removed
         | 
         | Just so you know, privileges can be (and probably will be in
         | this case) denied.
         | 
         | Rights on the other hand can't be denied.
        
           | dialup_sounds wrote:
           | Right, nobody on the internet has ever violated copyright and
           | gotten away with it. /s
        
       | tiahura wrote:
       | The open web's ethos since its inception in the 1990s has been
       | one of unrestricted access and fair use. Content published openly
       | online inherently invites broad consumption, reproduction, and
       | creative reuse by the public. This is not merely custom, but a
       | fundamental aspect of fair use doctrine as applied to the digital
       | realm.
       | 
       | The four factors of fair use - purpose of use, nature of the
       | copyrighted work, amount used, and effect on the market -
       | overwhelmingly favor allowing free use of openly published web
       | content. The transformative nature of most reuses, the public
       | availability of the original works, the necessity of using entire
       | works in many cases, and the lack of a traditional market for
       | such content all support this interpretation.
       | 
       | This longstanding practice has been the catalyst for
       | unprecedented innovation and information dissemination. It
       | represents a tacit social contract between content creators and
       | users, establishing a de facto "freeware" model for open web
       | content. Any attempt to retroactively impose strict copyright
       | limitations would not only stifle innovation but also contradict
       | decades of established legal precedent and digital norms.
       | 
       | -As a side note, I'm not certain that training necessarily
       | involves "copying."
       | 
       | ---Lastly, if anyone really thinks the Robert's court is going to
       | knee-cap AI, you're soft in the head.
        
         | pessimizer wrote:
         | > Content published openly online inherently invites broad
         | consumption, reproduction, and creative reuse by the public.
         | 
         | You're using "invites" as a weaselword here. Then you go into
         | the law, as if the law is discussing "invitations."
         | 
         | Counterpoint: content published online is usually extremely
         | hostile to reproduction, and the people who produce it are
         | deathly afraid of other people copying their work and
         | outranking them with it.
         | 
         | > ---Lastly, if anyone really thinks the Robert's court is
         | going to knee-cap AI, you're soft in the head.
         | 
         | If you think the Roberts court is going to be hostile to
         | copyright, you're insane.
        
         | amiantos wrote:
         | I am optimistic but I feel sad when I remember when went
         | through all this with sampling 30 years ago and now the music
         | industry is more insular and controls every scrap of music they
         | can find so no one can sample anything and even if your song
         | has similar "vibes" to another, well, now it belongs to them as
         | well.
         | 
         | I'd wager that 90% of the people cheering on the RIAA (and
         | copyright in general) now were singing a much different tune
         | before they decided AI was a threat to their livelihoods.
         | There's a lot of reasons to not be optimistic about the
         | continued legal existence of freely available open source AI,
         | because if the copyright holders have their way, they will be
         | the only entities that control all the data needed to train it.
         | And many people on the internet are all too happy to cheer this
         | on without realizing that once the RIAA, MPAA, and publishers
         | (when they finally figure out how to organize effectively like
         | film and music has) hold the reigns, the AI will still exist,
         | and it will still take all the same jobs, but all the open
         | source and freely available AI that everyone could use, that
         | could level the playing field for everyone... that's going to
         | be illegal for people to use, without being able to pay the
         | content rights holders. Just another way to keep it to have /
         | have nots, same as it's always been.
        
       | cjk2 wrote:
       | Ah yes the implied social contract that it's ok because it
       | happens all the time.
       | 
       | That's how society falls.
        
       | mewpmewp2 wrote:
       | What exactly is wrong with the statement he has made?
        
         | elicksaur wrote:
         | The article does a pretty good job outlining the actual legal
         | rights of published content and how his statement is not
         | supported by laws and precedent.
        
         | jprete wrote:
         | The Verge quotes: "I think that with respect to content that's
         | already on the open web, the social contract of that content
         | since the '90s has been that it is fair use. Anyone can copy
         | it, recreate with it, reproduce with it. That has been
         | 'freeware,' if you like, that's been the understanding."
         | 
         | This statement is quite wrong. People have complained about
         | content farms and plagiarists ripping off their content for
         | ages. The only kinds of Web content I can think of where his
         | statement could apply are open source software (for the more
         | permissive licenses) and Stack Overflow posts. Almost
         | everything else was posted with the intent of copyright.
        
         | trueismywork wrote:
         | Content on web eithout any copyright notice is by default "all
         | rights reserved" content. So you cannot use it to produce new
         | content, unless you can prove that the old content was never
         | reproduced ever along all the invocation of the LLM
        
           | seanmcdirmid wrote:
           | You can totally study it and learn from a bunch of content on
           | the web and then create a new work having looked at other
           | works before. My writing is influenced by lots of copyrighted
           | work I've read already, there is no mental blockade going on
           | when I write something. The only question is what does it
           | mean when a machine does that instead?
           | 
           | The old content can't be reproduced by an LLM unless the
           | content is provided as part of its prompt. At least, that's
           | what the AI companies claim and will have to show in court.
        
       | fimdomeio wrote:
       | So we've now learned that copyright is determined by
       | communications protocol. If you're using torrents it's copyright
       | infringement, if it's the web then it's public domain.
        
         | echelon wrote:
         | > if it's the web then it's public domain.
         | 
         | Only if you're a big company. If you're an individual or small
         | company, then it isn't.
        
         | mordae wrote:
         | Not sure why downvoted.
        
       | jsyang00 wrote:
       | No he doesn't.
       | 
       | > I think that with respect to content that's already on the open
       | web, the social contract of that content since the '90s has been
       | that it is fair use. Anyone can copy it, recreate with it,
       | reproduce with it. That has been "freeware," if you like, that's
       | been the understanding.
       | 
       | > There's a separate category where a website, or a publisher, or
       | a news organization had explicitly said 'do not scrape or crawl
       | me for any other reason than indexing me so that other people can
       | find this content.' That's a grey area, and I think it's going to
       | work its way through the courts.
        
         | madeofpalk wrote:
         | Yes, he does.
         | 
         | > content that's already on the open web, the social contract
         | of that content since the '90s has been that it is fair use.
         | Anyone can copy it, recreate with it, reproduce with it. That
         | has been "freeware," if you like, that's been the understanding
         | 
         | Sure, it's his belief, but this statement is completely
         | incorrect. The words that he's using mean specific things, and
         | he's just got it wrong. You would expect he's smart enough to
         | know this, but 'it is difficult to get a man to understand
         | something, when his salary depends on his not understanding
         | it'.
         | 
         | > That's a grey area, and I think it's going to work its way
         | through the courts.
         | 
         | It's not a grey area. Putting up a robots.txt doesn't change
         | copyright, and it certainly doesn't make it a 'grey area'.
        
         | elicksaur wrote:
         | Most people would read the first quote as totally aligned with
         | the article's title.
        
         | joe_the_user wrote:
         | It's an incredibly fuzzy statement, really.
         | 
         | You can copy copyrighted material you download for your own use
         | and the use of your friends. You can't copy it to a different
         | website and distribute it further on the "open web". You can
         | modify material provided for download in your own home for any
         | purpose you wish but unless that modification is in the
         | category "fair use" (a small part, a parody, etc) you can't
         | distribute it freely on the web either.
         | 
         | His statement might mean this but it could mean a zillion other
         | things too (what does "it is fair use" mean? etc)
         | 
         | Whether using randomly obtained copyrighted material to "train"
         | an LLM and then selling the output of that LLM is "fair use"
         | seems like so far another "gray area" and the foggy statement
         | seems oriented to reducing awareness of that situation.
        
           | jprete wrote:
           | Legally you cannot actually do the first thing without
           | permission. The fact that it was technologically impossible
           | to stop, and the damages would be impossible to prove, didn't
           | make it a legal right.
        
         | raincole wrote:
         | > I think that with respect to content that's already on the
         | open web, the social contract of that content since the '90s
         | has been that it is fair use. Anyone can copy it, recreate with
         | it, reproduce with it. That has been "freeware," if you like,
         | that's been the understanding.
         | 
         | is _literally_
         | 
         | > I think it's perfectly OK to use content in arbitrary way if
         | it's on open web
         | 
         | The only difference between this and the title is he doesn't
         | think this behavior is called "stealing".
        
         | advael wrote:
         | I mean, that seems to be exactly how he's defining "open web"
         | here, actually. That which is - in the dichotomy presented by
         | these two quotes - "the open web" is free game for any use, and
         | he defines things that use language that explicitly disallows
         | all uses except indexing as the complement of this category.
         | Maybe he'd accept any site that effectively declares any
         | "whitelist" of acceptable uses in this category too, though
         | this isn't explicitly stated.
         | 
         | His contention is an assumptive close, wrapping the assumption
         | that anything not explicitly labeled otherwise must use a
         | "blacklist" policy where any usage not specifically forbidden
         | is permitted into "the social contract" that he claims to be so
         | obvious as to not permit challenge
         | 
         | He would like the "grey area" of legal debate on this matter,
         | as he explained quite clearly, to be exclusively about whether
         | AI models can be enforcably barred from training on content for
         | which such a narrow whitelist of acceptable uses has been
         | defined. Naturally this would mean both that the courts _could_
         | decide such a blanket ban can 't bar msft (or anyone) from
         | using this content to train AI models, but also that the court
         | needn't or maybe even can't decide that failure to ban this use
         | case explicitly (or adopt a similar "whitelist" style blanket
         | ban) makes acceptance of it legally implied. Hell, he even
         | leaves room for explicitly banning this use to be rendered
         | legally unenforceable
         | 
         | I can see why he would want that to be the overton window!
        
         | FactKnower69 wrote:
         | >I think that with respect to content that's already on the
         | open web, the social contract of that content since the '90s
         | has been that it is fair use. Anyone can copy it, recreate with
         | it, reproduce with it.
         | 
         | Good thing it doesn't matter what he "thinks" the "social
         | contract" is, copyright is automatic.
        
         | Forge36 wrote:
         | That "separate category" isn't a explicit opt in. It's its two
         | different mechanisms for indexing and copyright.
         | 
         | Index is your out. You can use our content is opt in.
         | 
         | Reproduction and recreation, especially when taken physically
         | outside of the Internet or into products for sale has always
         | been a against the rules. As mentioned by another post,
         | torrents of music and movies solidified this stance legally.
         | 
         | Unindexed connect can be open source no strings attached.
        
         | krisoft wrote:
         | > No he doesn't.
         | 
         | Could you please help me see where you see the difference
         | between the title and the quotes? Even after reading them it
         | seems the title is substantially true?
         | 
         | Or to be curt while mirroring your comment's style: "Yes he
         | does."
        
       | starik36 wrote:
       | The more I read about this guy the more I get the feeling that he
       | is an unscrupulous individual.
       | 
       | robots.txt is a "grey idea" to him, instead of being a directive
       | to keep moving? Wow.
        
       | JonChesterfield wrote:
       | I'll bet they don't consider the windows and office source code
       | fair game for arbitrary reuse provided the other party found the
       | copy on the web. Even if the person found the copy on GitHub.
        
       | 29athrowaway wrote:
       | One thing is a robots.txt policy, meant mostly for search
       | crawlers.
       | 
       | Another thing is the copyright of the content, terms of use
       | policies, etc.
       | 
       | Abiding by a robots.txt policy doesn't make you immune to
       | copyright, terms of service, law in various jurisdictions, etc.
       | If you think that you are probably a kleptomaniac.
       | 
       | Just create a robots.txt with "User-Agent: one billion asterisks"
       | so that the crawlers die when parsing it.
        
       | KoolKat23 wrote:
       | This is nothing but performative clickbait by the Verge.
       | 
       | It is classified as fair use, the term is transformative use,
       | where those using it are training models (their intention) if
       | anyone wishes to Google it.
       | 
       | The end.
        
       | jimmaswell wrote:
       | From the beginning, it's seemed completely intuitive to me that
       | training a computer made of sand on publicly available content
       | and then generating art later should be fair use, so long as it's
       | fair use to train the meat computer in your head on the same
       | content and then use it to generate art later. There's no
       | meaningful difference to me as far as the ethics of the act are
       | concerned.
        
       | sircastor wrote:
       | Part of the problem here is that the web has gone through lots of
       | change as to what it is and how people understand it.
       | 
       | Some people think of it as billboards posted on the highway. Some
       | think it's a bulletin board. Some think it's a newspaper. A
       | television, a "zine", a diary, graffiti. It has been all of these
       | things, and is and isn't. And people who publish are really bad
       | at explicitly stating which one they are. But they expect you to
       | know.
        
       ___________________________________________________________________
       (page generated 2024-06-29 23:02 UTC)