[HN Gopher] 9k authors say AI firms exploited books to train cha...
       ___________________________________________________________________
        
       9k authors say AI firms exploited books to train chatbots
        
       Author : webmaven
       Score  : 11 points
       Date   : 2023-07-20 21:59 UTC (1 hours ago)
        
 (HTM) web link (www.latimes.com)
 (TXT) w3m dump (www.latimes.com)
        
       | JoeAltmaier wrote:
       | Doesn't 'fair use' come into play at some point? I know this is
       | all new, but isn't this approximately what fair use is about?
       | Even a little modification of prior art is enough to qualify.
       | Embedding it into an AI corpus is quite a bit of modification.
       | 
       | Or am I way behind the curve on this one.
        
         | fxtentacle wrote:
         | The article itself says "Not only does the recent Supreme Court
         | decision in Warhol v. Goldsmith make clear that the high
         | commerciality of your use argues against fair use, but no court
         | would excuse copying illegally sourced works as fair use."
        
         | SllX wrote:
         | Fair use isn't straightforward to begin with, but basically
         | there's room for a Judge to narrow its applicability in this
         | instance.
         | 
         | Here's the four factor test courts use:
         | 
         | > In determining whether the use made of a work in any
         | particular case is a fair use the factors to be considered
         | shall include:
         | 
         | > 1. the purpose and character of the use, including whether
         | such use is of a commercial nature or is for nonprofit
         | educational purposes;
         | 
         | > 2. the nature of the copyrighted work;
         | 
         | > 3. the amount and substantiality of the portion used in
         | relation to the copyrighted work as a whole; and
         | 
         | > 4. the effect of the use upon the potential market for or
         | value of the copyrighted work.
         | 
         | I think points 1 and 3 are the least debatable: commercial LLMs
         | should easily fail 1 and all LLMs should pass 3. EDIT: actually
         | changed my mind on this. If they're using the whole work, then
         | LLMs should actually fail #3. It's not just how much of the
         | model it makes which is insignificant relative to all the rest
         | of the text in the model, but also how much of each individual
         | work is used. If that's 100% then that's actually an easy fail.
         | 
         | We could have a debate about #2 till the end of time and I'm
         | not really here for that, not today anyway, but it is probably
         | worth debating.
         | 
         | #4 is also absolutely debatable, but I think there is a large
         | potential for LLMs to harm the market for the original source
         | material that they were trained on.
        
         | webmaven wrote:
         | One of the claims is that the book corpora weren't even sourced
         | legally. It is harder to make a fair use claim if the copy the
         | model was trained on was pirated to begin with. Contemporary
         | in-copyright ebooks are a bit different in this regard than
         | images found on the public web.
        
       | EA-3167 wrote:
       | I don't think this is a reasonable argument, but I do think it
       | might still prevail in court, which from the POV of the authors
       | an their lawyers makes it a worthwhile filing. Personally I hope
       | it doesn't prevail, the implications for how people learn and
       | synthesize based on an existing corpus is too profound.
        
       | [deleted]
        
       ___________________________________________________________________
       (page generated 2023-07-20 23:03 UTC)