[HN Gopher] AI weights are not open "source"
       ___________________________________________________________________
        
       AI weights are not open "source"
        
       Author : subomi
       Score  : 196 points
       Date   : 2023-07-05 15:18 UTC (7 hours ago)
        
 (HTM) web link (opencoreventures.com)
 (TXT) w3m dump (opencoreventures.com)
        
       | TheRealPomax wrote:
       | So, Open Data. Got it. This is the same category as config files
       | that are kept up to date by a program as it runs.
       | 
       | - Is it "a program"? Very clearly not.
       | 
       | - Is it source code? You can argue either way. The program won't
       | work without it, but "this specific one" is not required for the
       | program to do _something_ , and that ambiguity means you probably
       | don't want to call it "source code" because it's too vague.
       | 
       | - Is it _data_ used by a program in order to perform its task?
       | Absolutely. It even uniquely defines the program behaviour, and
       | so is a thing onto itself within the context of the program it 's
       | used by.
        
       | Zetobal wrote:
       | If my own data is in the dataset even when I didn't give consent
       | is it a collaborator dataset?
        
         | tensor wrote:
         | If you posted your data into a service where the TOS allows
         | this use then yes.
        
       | iLoveOncall wrote:
       | [flagged]
        
       | worksonmine wrote:
       | > Unlike software licensing, AI isn't as simple as applying
       | current proprietary/open source software licenses. AI has
       | multiple components--the source code, weights, data, etc.--that
       | are licensed differently.
       | 
       | Software also has multiple components, often the same as the ones
       | listed by the author. But what do I know, to me AI is just
       | another example of software.
        
       | mensetmanusman wrote:
       | Weights are an information asset that require millions in capital
       | and burned-out GPUs to mine and refine.
        
         | daniel-cussen wrote:
         | [dead]
        
       | light_hue_1 wrote:
       | I think this is very shortsighted.
       | 
       | Weights are a program. CUDA is an interpreter for that program.
       | 
       | One day we will be able to decompile these programs into
       | something more human understandable.
        
       | bee_rider wrote:
       | > The ethical license category applies to licenses that allow
       | commercial use of the component but includes field of endeavor
       | and/or behavioral use restrictions set by the licensor.
       | 
       | I don't love the name, "ethical license" sounds like a
       | description of the license: this license is ethical. Really this
       | sort of license imposes a particular ethical framework on the
       | user.
       | 
       | Not to throw shade, though. It is actually hard to come up
       | neutral sounding name for this sort of license I think. I keep
       | thinking of things like "morality encumbered license," but that
       | sounds ridiculously euphemistic in a weird way.
        
         | version_five wrote:
         | Yes I was going to say the same thing. It's a branding that has
         | been applied by the license's proponents, and I personally
         | reject a lot of what they call "ethics" as well as the idea of
         | whatever monitoring and enforcement the restrictions entail -
         | maybe calling it a religious license would be better.
        
         | iandanforth wrote:
         | "Opinionated" is how I think about it.
        
           | bee_rider wrote:
           | That might be a good pick, IMO the word has negative
           | connotations elsewhere, but in tech circles seems basically
           | neutral.
        
         | 93po wrote:
         | I'd argue any licensing of IP is unethical. I'd use the word
         | "conditional"
        
       | cpcallen wrote:
       | I'm disappointed that the article is only making the (somewhat
       | pedantic) distinction between source code and weights. From the
       | quotation marks in the headline I hoped that it would instead be
       | making the distinction between human-readable source code and
       | machine-readable compiled form.
       | 
       | For example, IMHO (IANAL) an AI code-completion tool that had
       | been trained on GPL software is (or should be) only be legal to
       | distribute if it is accompanied by the training code _and all the
       | code ingested during training_ (or an offer to provide such code
       | upon request).
        
         | version_five wrote:
         | This is an interesting point. If you read the OSI open source
         | definition, specifically on source code (quoted below) I'm
         | inclined to treat the training data as part of the source code
         | for the purpose of determining whether to consider any model
         | open source.                 2. Source Code       The program
         | must include source code, and must allow distribution in source
         | code as well as compiled form. Where some form of a product is
         | not distributed with source code, there must be a well-
         | publicized means of obtaining the source code for no more than
         | a reasonable reproduction cost, preferably downloading via the
         | Internet without charge. The source code must be the preferred
         | form in which a programmer would modify the program.
         | Deliberately obfuscated source code is not allowed.
         | Intermediate forms such as the output of a preprocessor or
         | translator are not allowed.
         | 
         | https://opensource.org/osd/
        
       | TrackerFF wrote:
       | Weights are just matrices with values between a certain range. So
       | are digital images - just matrices with values. Images are
       | covered by copyright laws, so why shouldn't weights also be?
        
       | mellosouls wrote:
       | Hmm. Makes a few unsubstantiated claims, with hand-wavy appeals
       | to risks that our private corp overlords are presumably
       | protecting us humble users from, now that they've built their
       | product on open source and data by closing it down and changing
       | terminology to suit.
       | 
       | There's an intelligent discussion to be had, and I think this
       | otherwise-reasonable article could be part of it if it toned down
       | the presumption and condescension a little.
        
       | seydor wrote:
       | If it is extremely complex, then it can only be modeled by an AI
        
       | horsawlarway wrote:
       | If anything - this entire conversation just highlights (Over and
       | Over and Over and Over again) how absolutely bonkers abusive our
       | current copyright laws are.
       | 
       | The vast majority of small individuals are compelled by contract
       | to surrender their rights to large corporations. Those large
       | corporations then abuse the ever loving fuck out of those rights.
       | 
       | The express intent of copyright is now a sad joke.
       | 
       | Personally - I'm pretty over the entire show. This system is
       | generating an incredible amount of inequality. New and novel
       | content is absolutely NOT getting made, and these laws are
       | creating vicious infights that drain resources from well
       | intentioned companies & individuals and pass them along to
       | complete scam corporations.
       | 
       | We are told stories as children that we cannot retell in our own
       | voices decades later to our own children.
       | 
       | I am firmly ready to burn this copyright system to the fucking
       | ground. It's been 300 years since the Statute of Anne - I'm ready
       | for a different game.
        
         | PartiallyTyped wrote:
         | There is also the whole patent / copyright trolling issue too.
         | The fact that $BIG_CORP can hire armies of lawyers to freeze
         | competitors and beat them to market by filing frivolous
         | lawsuits is yet another example of insanity in the whole
         | system.
        
           | chongli wrote:
           | I recently watched the documentary _Fire in the Blood (2013)_
           | [1] about the use, by big pharma, of patents and WIPO to
           | obstruct access to affordable antiretrovirals (ARVs) in
           | Africa during the worst years of the AIDS epidemic, leading
           | to over ten million deaths. All of this when the African
           | market for these medications represented less than 1% of the
           | total market, in dollars. It's absolutely infuriating!
           | 
           | [1]
           | https://en.wikipedia.org/wiki/Fire_in_the_Blood_(2013_film)
        
           | drdaeman wrote:
           | It's a problem with legal system (not unique to any specific
           | country, mind you, the problem is global), not patent or
           | copyright system specifically. It grew incredible amounts of
           | complexity so _pro se_ became a sad joke in all but simplest
           | cases, and there 's no incentive to fix it - quite the
           | opposite, everyone in the system is all for keeping the
           | status quo, because it generates money.
        
             | TaylorAlexander wrote:
             | Personally I don't think patents do what people believe
             | they do (encourage innovation). It's a bigger discussion
             | but briefly, the only literal function of a patent is to
             | discourage innovation by legally barring anyone from using
             | a patented idea as part of a new innovation. The idea we
             | have is that the secondary effects of this will be
             | increased profits for inventors and therefore more
             | innovation. But actually there's loads of secondary effects
             | and often many of them outweigh the effect of increased
             | profit. For every one inventor that gets a patent there
             | might be 100 prevented from using that idea in a different
             | and innovative way.
             | 
             | A classic example is 3D printers. Stratasys spent 15 years
             | selling printers that cost tens of thousands of dollars. It
             | wasn't until the patent expired that people figured out how
             | to make them for $250. Those cheaper printers are enabling
             | mechanical engineers and designers to accelerate their
             | process and make other new innovations faster. Stratasys
             | had such a powerful patent they never bothered innovating
             | down in price, instead rested on their laurels selling $25k
             | printers to big customers.
             | 
             | So how many inventions were delayed or shelved because the
             | inventors couldn't afford a $25,000 3D printer, and $250
             | printers didn't exist yet? Both Stratasys and IBM held
             | patents related to 3D printing and they had to cross
             | license to go in to production, so how many others would
             | have come up with 3D printing in the 1990's if they had not
             | been patented? Would first mover advantage in a free market
             | have been enough to stimulate development of 3D printers?
             | Could we have had $2000 3D printers in the early 2000's
             | (Stratasys sold theirs for $30k) instead of ten years
             | later? How many engineers would have invented new gadgets
             | faster if they had a 3D printer ten years earlier?
        
               | PartiallyTyped wrote:
               | Another possible example is the x86 and x86-64 ISAs
               | locked between AMD and Intel. I don't think Intel would
               | have become complacent had there been more competitors...
               | 
               | ... or the whole "oracle vs google" over the java API.
        
             | jrumbut wrote:
             | But there are specific problems with copyright and patent
             | law that could be improved without a global systemic
             | overhaul that may never happen.
             | 
             | We have to take some small wins even in the presence of big
             | problems.
        
               | drdaeman wrote:
               | Of course. I'm just saying that the core problem is
               | larger than just the copyright and patent law.
        
         | loudmax wrote:
         | Fully agree that the existing copyright and intellectual
         | property systems are dire need of deep reform. But to get
         | people on board, you can't just propose burning it all down,
         | you need to point to a viable alternative. Say, limiting
         | copyrights to something sane like 15 or 30 years. Or making it
         | easier to invalidate obvious or trivial patents.
         | 
         | Or do you really want to do away with notions of intellectual
         | property altogether? You can make an argument for that, but
         | there would lead to deep economic changes, and you need to
         | anticipate what the end result would look like. You still need
         | some way to encourage the creation of new content.
         | 
         | Pointing out that our copyright/IP system is broken is easy.
         | And you're right, it's totally broken! Coming up with a fix is
         | hard work.
        
           | Frost1x wrote:
           | >Pointing out that our copyright/IP system is broken is easy.
           | And you're right, it's totally broken! Coming up with a fix
           | is hard work.
           | 
           | The problem I have with these arguments is they ultimately
           | tend to boil down to the devil you know or the devil you
           | don't know.
           | 
           | We keep claiming when something is broken we must provide a
           | "fix" and the assumption is that fix has to be better than
           | the current approach. There's pretty much no way to guarantee
           | this because the systems in place are the only systems with
           | evidence. So, because we have other ideas, we dare not try
           | them because they have to "fix" the problem. The amount of
           | inertia that keeps corruption in motion bothers me and at a
           | fundamental level most of the inertia comes down, ironically,
           | back to property ownership. If we abolish copyright or change
           | it we have to make sure things are fair/equitable. Well sure,
           | that's ideal, but what we have isn't even remotely fair and
           | equitable anymore, so even something broken is likely an
           | improvement.
           | 
           | We have no willingness as a society to try some modifications
           | and be willing to accept failure, then shift to the next
           | modification and iterate around until we get something sane
           | in place. As such, the systems in place remain in place and
           | more and more holes are found to exploit as time progress.
           | 
           | Our systems need to be more adaptable. Founders of the
           | country understood that which is why they made the legal
           | system a legal adaptable system. The question has always been
           | though, what is the threshold? We've played it safe so long
           | that much of the entire system designed to adapt to fix these
           | issues has itself been targeted and gummed up intentionally
           | to prevent that.
        
           | pessimizer wrote:
           | > You still need some way to encourage the creation of new
           | content.
           | 
           | Do you? What's the argument for this? Is there some sort of
           | extreme shortage of creative work that the state should find
           | it necessary to encourage it? How about we end copyright, and
           | if there's ever a problem, we offer copyrights for a short
           | period to fluff the commons up again. A copyright anti-
           | holiday, as it were.
           | 
           | Instead we do the opposite: automatically copyright
           | everything anyone produces, and make it very difficult to
           | surrender your copyright (unless Google or Microsoft want it,
           | then if you object you're literally a Luddite caveman who is
           | trying to turn back the clock on modernity because you're
           | old, stupid, and afraid of fire.)
        
           | Zaskoda wrote:
           | > But to get people on board, you can't just propose burning
           | it all down, you need to point to a viable alternative.
           | 
           | This may be true for most people, I don't know. However, I
           | personally am fully on board the "burn it all down" train and
           | have been for a while.
        
           | TaylorAlexander wrote:
           | I don't think we need to use the legal system to encourage
           | creation of new content! That's a natural thing people do. In
           | fact there's a lot of artistic remixing that is illegal or
           | ambiguously legal under the current copyright regime that can
           | be a powerful form of expression.
           | 
           | I really don't think we need government policy to encourage
           | artists to create art. (At least not of this sort - I am all
           | for art grants.)
        
           | capr wrote:
           | Pointing out that IP is broken is _not_ easy because most
           | people believe in the contradictory notion of intellectual
           | property, you included, not knowing the legal history of IP,
           | the legal and economic history of the concept of property,
           | and so on. If it were easy, it would be obvious to everybody
           | that 1) IP law is immoral and 2) nothing bad would happen if
           | it's abolished outright.
           | 
           | Here's a free ebook on the subject, written by a patent
           | lawyer no less: https://mises.org/library/against-
           | intellectual-property-0
        
             | mike_d wrote:
             | > 2) nothing bad would happen if it's abolished outright.
             | 
             | It is interesting that people living in the places with the
             | weakest IP laws will pay a premium to import baby formula
             | from the places with the strictest laws.
             | 
             | > Here's a free ebook on the subject
             | 
             | Of course it is some right-libertarian wonk piece.
        
               | immibis wrote:
               | This has nothing to do with IP laws, and everything to do
               | with baby formula laws. You don't seriously think that
               | without the ability to sell the recipe, nobody would
               | invent a safe and effective baby formula, right?
        
               | mike_d wrote:
               | I'm not sure you fully grasp all the dimensions of IP
               | law.
               | 
               | If you have two brands of baby formula, Death brand that
               | kills babies, and OK brand that is perfectly fine, and
               | you start putting Death brand in fake cans labeled OK
               | brand - that is absolutely an IP enforcement issue. The
               | desire of OK brand to protect their brand, and profits,
               | combined with reasonable IP laws allows them to lead
               | enforcement actions and protect consumers.
        
               | immibis wrote:
               | That is absolutely not the kind of intellectual property
               | that anyone hates.
               | 
               | The kind of intellectual property we are talking about is
               | the one where Death isn't allowed to make baby formula
               | that doesn't kill babies, because OK patented making baby
               | formula that doesn't kill babies and won't give them a
               | license.
        
               | xigoi wrote:
               | > you start putting Death brand in fake cans labeled OK
               | brand - that is absolutely an IP enforcement issue.
               | 
               | This is a matter of trademark, which is completely
               | orthogonal to copyright and nobody is protesting against
               | it here.
        
         | 111111IIIIIII wrote:
         | > _I am firmly ready to burn this copyright system to the
         | fucking ground._
         | 
         | Same, but the issue is not copyright, which is simply an effort
         | to wield the state to control intellectual property in the same
         | way the state is wielded to control physical property.
         | 
         | The compounding problem arises when property is _capital_ ,
         | defined as the means to convert labor into new value.
         | Capitalism is specifically a system in which one can wield
         | control of capital (intellectual or otherwise) to extract
         | profit from labor then trade that profit for more capital. As a
         | result, capital accumulates infinitely, independent of the
         | value produced by the labor which is provided to society.
         | 
         | Artists require capital to convert their labor into value just
         | as any other worker would, so where should that capital come
         | from if not from control of the value they produce? Society
         | must solve this problem or we will not have art to begin with.
         | Only looking at the demand side obfuscates such issues that
         | arise on the supply side, and the only reason we're talking
         | about them now is that digital technology has solved the
         | scarcity problem on the supply side. It has not solved the
         | scarcity problem on the demand side, however.
         | 
         | Finally, art, just like all technological progress, is always
         | the product of entire societies and the history of all mankind
         | that came before it. For this reason, all copyright and patents
         | have no rational basis and are merely bandaids for the ill side
         | effects of controlling capital to extract profit from labor to
         | begin with.
        
       | habitue wrote:
       | One thing I don't see discussed enough is that, ok let's say the
       | weights are unencumbered, and the source is under an OSI license:
       | the point of open source licenses and free software was to expose
       | the *human understandable* meaning of the final program.
       | 
       | That's why distributing binaries isn't allowed even though
       | technically all of the functionality is present in the machine
       | code. AI weights are basically binary blobs. We don't know what
       | they mean, there is really no source code for them. The best we
       | can do is various black box manipulations on them like LoRA, etc,
       | similar to what we can do to a binary blob.
        
         | phkahler wrote:
         | >> AI weights are basically binary blobs. We don't know what
         | they mean, there is really no source code for them.
         | 
         | No. You can do further training on them. If they are something
         | less than code I don't think it's going to warrant all this
         | talk about licensing. GPL, MIT, or some proprietary should
         | cover it.
        
           | habitue wrote:
           | You can do further training on them, just like you can patch
           | a binary blob. There are some surgeries you can do to the
           | weights, and there are analyses you can do to poke at them
           | and try to understand them, but ultimately they weren't
           | created from a human understandable spec, and without a ton
           | of reverse engineering work the weights by themselves aren't
           | human understandable: hence the "source" component is
           | missing.
           | 
           | The source code that generated the weights is one step
           | removed from the kind of source code we'd need to interpret a
           | bunch of AI weights. It's really meta-source code
        
         | [deleted]
        
       | Topfi wrote:
       | This post did cover many of the same ideas I have been ruminating
       | on concerning model weights and the nomenclature of current
       | efforts. That's also why I generally tend to stick with calling
       | these[0] "local/self hosted models" for the time being. A major
       | reason for my reluctance is that I see weights far closer to
       | binary than code, making a distinction important and current FOSS
       | concepts not really applicable.
       | 
       | Of course, this all hinges on the idea that weights by themselves
       | are inherently protected by current copyright, which still seems
       | to be an unsettled topic, hotly debated by both laypeople and
       | legal professionals. Authors generally are afforded copyright on
       | their work by default, and weights raises so many questions
       | concerning authorship that have never been considered.
       | 
       | This being such a contested issue, which will require new laws
       | and/or precedent (depending on the legal system), is very
       | problematic. Regardless of where you live, generally courts and
       | government entities are not famous for their speedy reaction to
       | new things, so clarity may take a while, at which point the
       | industry might have already settled on some agreement that then
       | may be adopted as a basis for actual legislation, which would
       | likely favor financially well baked entities already actively
       | lobbying for their interests, such as OpenAI.
       | 
       | Some have also pointed out that this is arguing semantics, and I
       | am tempted to agree in principle, but also want to emphasize that
       | I feel this is a situation where that can be valuable. Should
       | weights in some way be afforded copyright protection, clear
       | nomenclature will be needed. Putting some thought into this now
       | is definitely not the worst idea.
       | 
       | I very strongly feel that the specific word "ethical" as part of
       | defining licenses is not the best idea, though. "Ethical" can
       | carry vastly different connotations, depending on a myriad of
       | factors, many of which would go beyond the use-focused definition
       | laid out in the post. Due to this, I'd argue for "behavioral" or
       | "restricted use" over "ethical", as both more clearly state what
       | the intended effect is in cases such as Open RAIL-M[1].
       | 
       | Part of my strong feelings on the use of the word "ethical" come
       | from the fact that with weights and training data, there has been
       | a lot of discussion concerning both rights of and considerations
       | for creators whose published works have been used to create those
       | weights. Due to this, the use of "ethical" referring to a group
       | of licenses could give some the impression that this may indicate
       | that the training data used was "ethically sourced", i.e. in
       | agreement with the original creator. This is something that in my
       | eyes should also have clear labeling, though with weights being
       | very hard to reliably trace back to source data, it currently
       | seems impossible to verify, making this essentially just a good
       | faith effort.
       | 
       | [0] https://huggingface.co/tiiuae/falcon-40b-instruct
       | 
       | [1]
       | https://drive.google.com/file/d/16NqKiAkzyZ55NClubCIFup8pT2j...
        
       | ianbutler wrote:
       | I'm not sure OCV gets to decide any of this. Just like I don't
       | think OSI trying to be the sole dictator of the term "Open
       | Source" works out long term. My opinion is always received
       | controversially about things like this, but terms evolve to meet
       | the common usage by the people. If people are calling this "Open
       | Source", and there are more people who want to call this "Open
       | Source", than people who don't; unless you intend to legally bar
       | them from using the term, with actual action, like a lawsuit or
       | something then eventually this will will also be encompassed by
       | the term "Open Source" as people know it like it or not.
       | 
       | Yes I know this term is currently defined explicitly by OSI, no I
       | don't think language prescriptivism wins out regardless how hard
       | they try with it, and since I haven't seen any of the hundreds of
       | quasi Open Source, but not really, companies get dragged to court
       | over usage of the term, this is all toothless complaining in my
       | view.
       | 
       | As to their actual point, I might actually agree with them if it
       | were only the weights being shared. In most cases the
       | configuration is also shared which allows popular frameworks to
       | instantiate the model and then execute it for either inference or
       | further training making the release fully suitable for
       | modification and rerelease. I don't need the exact implementation
       | of FlashAttention they used if I can load the model into
       | Huggingface and use theirs, or mine or whatever.
       | 
       | Edit: This obviously doesn't apply to the models who have
       | restrictions placed on usage just in case people think I mean
       | every instance of sharing a model. Those are obviously restricted
       | use and I agree it muddies the term.
        
       | TZubiri wrote:
       | Agreed, output weighs are target code, and no one would argue the
       | contrary. Companies pretending to publish source code is nothing
       | new.
       | 
       | Stallman defines source code as "the preferred way in which
       | developers modify the program"
       | 
       | I wrote for wikipedia once that
       | 
       | "Stallman's definition thus contemplates JavaScript and HTML's
       | source-target ambivalence, as well as contemplating possible
       | future forms of software production, like visual programming
       | languages, or datasets in Machine Learning."
       | 
       | So the datasets could be a form or source code, but the most
       | appropriate source code would be the code that crawls or
       | downloads the dataset and modifies it.
       | 
       | Clear as water
        
       | dahart wrote:
       | > Some people have the perspective that if a license isn't open
       | source, it's proprietary. I think it's more nuanced than that and
       | believe there are three more license types worth naming: non-
       | commercial NDA, non-commercial public, and ethical.
       | 
       | It's very useful to remember the U.S. government definition of
       | commercial software: it is software that "Has been sold, leased,
       | _or licensed_ to the general public" [1]
       | 
       | This means that a "non-commercial license" is a bit of an
       | oxymoron to a lot of people. Their definition of commercial
       | includes all software with a license, and does not depend on
       | whether the software costs money. (Perhaps not entirely unlike
       | how FSF does not define "free software" based on whether it costs
       | money.)
       | 
       | [1] https://www.acquisition.gov/far/2.101
        
       | Traubenfuchs wrote:
       | [dead]
        
       | ndriscoll wrote:
       | The complexity described seems to be resting on the unestablished
       | idea that weights are copyrightable in the first place. If
       | they're not, then presumably "available weights", "ethical
       | weights", and "open weights" are all the same: open weights.
       | Either your weights are under NDA and presumably considered to be
       | a trade secret, or they are public, and the words in your
       | "license" mean absolutely nothing? That seems like a rather
       | important point to bring up when discussing the licensing
       | landscape for weights...
        
         | feoren wrote:
         | Some thought experiments:
         | 
         | What happens if we train a neural network on a single,
         | copyrighted work? Say it has one input node (or even zero, if
         | you like), and regardless of this input, its output is always
         | exactly the copyrighted work it was trained on. What do its
         | weights represent? Clearly, its weights represent a direct
         | encoding of the original work. Those weights _are_
         | copyrightable, but not by the person who trained the neural
         | network -- the copyright is held by the owner of the original
         | work.
         | 
         | What if we train the neural network on just two copyrighted
         | works? If its one input node is 0, it outputs the first, and if
         | it's 1, it outputs the 2nd. Almost certainly, its weights are a
         | complicated, tangled mix encoding both, like a compression
         | algorithm that completely rearranged its input. Who owns the
         | copyright to those weights? To whatever extent the weights can
         | be "factored out" into a set representing the first work and a
         | set representing the second, clearly the copyright holder of
         | the first work holds the copyright on the first "factored set",
         | and the 2nd on the 2nd. It seems obvious that we must be able
         | to do this "factoring out" _somehow_ (even if the topology of
         | the factored networks is different), because we know both works
         | are exactly represented by the weights, and the neural network
         | itself can use this information to reconstruct them both, so
         | they 're _in there_ ... somewhere. So is there a sort of
         | "joint copyright" on the combined weights, where nobody is
         | really allowed to do anything with it without approval of the
         | other? Regardless, it's still clear that whoever trained the
         | neural network has no claim on any copyright.
         | 
         | Where is the breaking point extending this from 2 works to a
         | billion? People make arguments like "drawing a car from memory
         | isn't infringing on copyright design of that car", which ...
         | are you _sure_? Reproducing a piece of music from memory (and
         | selling it) _is_ usually copyright infringement. You 're
         | allowed to _learn_ a Taylor Swift song as part of your musical
         | training, but you 're not usually allowed to then play it back
         | from memory and sell that recording (I'm not sure I morally
         | agree with this treatment of covers, nor if it's globally
         | applicable). So the argument that "surely neural networks are
         | allowed to _learn_ from copyrighted works " misses the point:
         | they can learn all they want, but as soon as they reproduce
         | verbatim (or close enough) a copyrighted work, they're
         | infringing. And if they're representing a complete copy of the
         | work within their weights (which they obviously are if they can
         | reproduce it), then the original copyright holder has a claim
         | on those weights. And never in this process has the trainer of
         | the NN acquired any copyright to anything. The real trainer is
         | a bunch of GPUs, after all.
         | 
         | If the neural network _cannot_ reproduce any of the copyrighted
         | works verbatim, then we 're getting closer to "fair use"
         | territory. Yes, it's permissible to write a summary of a
         | copyrighted work. That is so lossy as to not "compete" with the
         | original work in any meaningful way. If it could be
         | demonstrated that neural networks do not encode completed works
         | (no matter how hard the factorization would be), then one could
         | make this argument. Unfortunately, the evidence is that LLMs
         | are more than happy to completely regurgitate copyrighted works
         | verbatim. It seems to me the copyright holder of the original
         | work therefore must hold a share of the claim on the weights.
         | Still, the GPUs that trained the network do not magically
         | acquire copyright over anything.
         | 
         | I wonder if the real answer is that the weights are
         | copyrighted, and that copyright is held jointly by hundreds of
         | millions of people, and nobody can do anything with those
         | weights without the approval of all the others. I'm not saying
         | I like that universe, but I am saying it's the most internally
         | consistent answer I can think of, and seems to follow from the
         | above argument.
        
           | feoren wrote:
           | In fact, the "factoring out" process shouldn't even be that
           | hard: find the input vector that forces the ANN to output the
           | copyrighted work verbatim. There should be some simple method
           | of "baking in" the first step of the feedforward algorithm,
           | applying that vector to the first layer of weights, and then
           | considering the input layer as the first hidden layer of a
           | network with 0 input nodes. It is now equivalent to a neural
           | network that can only ever output a single copyrighted work,
           | and therefore its weights exactly encode (bloatedly!) that
           | work. The owner of the work holds copyright on those weights.
           | Importantly, if I'm thinking about this right, the weights of
           | this derived network are exactly the same as the original
           | except in the first layer.
           | 
           | On the other hand, we need the original input vector for this
           | to work, and one could argue that the network weights are
           | simply the algorithm for decoding the input vector into the
           | copyrighted work. So the originator holds copyright on the
           | _input vector_ , not the weights. Does it matter if the input
           | vector has smaller information content than the original
           | work? Clearly this argument relies on the input vector being
           | the "actual encoding", and therefore must have at least as
           | much information. If the input vector is an embedding of
           | "please show me the latest Tom Clancy novel in full", this
           | argument breaks down.
           | 
           | Okay, this is hard.
        
         | spullara wrote:
         | This has been my position from the beginning. It is very hard
         | for me to imagine that weights can be copyrighted at all.
         | IANAL.
        
         | nerdponx wrote:
         | Weights are equivalent to compiled object code IMO. All else
         | follows from there.
        
           | feoren wrote:
           | Compiled object code of a bunch of code _you didn 't write_.
           | I don't know why programmers are so eager to forget that
           | copyright is not at all about what something _is_ , and all
           | about where it _came from_. It 'd be hard to assert that you
           | hold copyright over object code compiled from code you didn't
           | write!
        
             | cubefox wrote:
             | Or is the fact that compiled code enjoys copyright
             | protection, even though it is not human generated, evidence
             | that being generated by a human is not overly important for
             | copyright protection?
        
               | feoren wrote:
               | An mp3 encoding of a wav file of a copyrighted song is
               | still copyrighted, despite those exact bits never having
               | existed before, and being created entirely by a computer.
        
             | xigency wrote:
             | > programmers are so eager to forget that copyright is not
             | at all about what something is, and all about where it came
             | from
             | 
             | See "What color are your bits?":
             | https://ansuz.sooke.bc.ca/entry/23
             | 
             | >> And very much of intellectual property law comes down to
             | rules regarding intangible attributes of bits - Who created
             | the bits? Where did they come from? Where are they going?
             | Are they copies of other bits?
        
         | Animats wrote:
         | > The complexity described seems to be resting on the
         | unestablished idea that weights are copyrightable in the first
         | place.
         | 
         | Yes. Weights probably aren't copyrightable in the US. See Feist
         | vs. Rural Telephone, in which the Supreme Court ruled that
         | telephone directories are not copyrightable. The copyright
         | clause in the Constitution ("To promote the Progress of Science
         | and useful Arts, by securing for limited Times to Authors and
         | Inventors the exclusive Right to their respective Writings and
         | Discoveries.") is understood to require human authorship. The
         | US does not have database copyright, or "sweat of the brow"
         | copyright. That it was expensive to produce some collection of
         | data does not make it copyrightable.
         | 
         | Outputs from LLMs, machine generated art, and machine generated
         | music probably are not copyrightable either. US Copyright
         | Office: "Based on the Office's understanding of the generative
         | AI technologies currently available, users do not exercise
         | ultimate creative control over how such systems interpret
         | prompts and generate material. Instead, these prompts function
         | more like instructions to a commissioned artist."[1]
         | 
         | [1] https://www.reuters.com/world/us/us-copyright-office-says-
         | so...
        
           | deepsun wrote:
           | > Outputs from LLMs, machine generated art, and machine
           | generated music probably are not copyrightable either.
           | 
           | Let me put a straw man, and try to find a middle point, when
           | the copyright argument stops being applicable:
           | 
           | 1. A painting was done by an artist.
           | 
           | 2. On a computer.
           | 
           | 3. With a help from an image processor software.
           | 
           | 4. Using some advanced filters, like super-resolution, that
           | utilize computer vision techniques. Like neural networks.
           | 
           | Many smartphones already automatically process your* photos
           | with some advanced CV algorithms. That can be called "machine
           | generated art".
           | 
           | I'd personally prefer to stop saying "neural network did X",
           | same way as we don't say "a bulldozer built a road, a crane
           | built a house".
        
             | comfypotato wrote:
             | The distinction, defining your straw man, is simply that
             | the image itself is generated by the "commissioned artist"
             | that is the AI.
             | 
             | Even non-generative-AI inside Photoshop only mutates
             | images. Generative AI is the _source_ of images.
        
               | drdeca wrote:
               | Is it though? What of the e.g. choice of prompt, guidance
               | scale, maybe a specification of a pose, etc.?
               | 
               | Or, is the distinction you are making based on there
               | being an image before the model is used?
        
           | radarsat1 wrote:
           | > Weights probably aren't copyrightable in the US. ... is
           | understood to require human authorship.
           | 
           | Are you arguing here that because the weights come from an
           | optimization program, they are not "human authored"? If so I
           | find that to be a strange assertion. If I'm working every day
           | on my model and training algorithm to ensure it produces the
           | best weights possible to solve my problem, I would be very
           | surprised for someone to tell me I have no ownership over
           | those weights because they are _generated_ from a program I
           | wrote and data that I own.
        
           | mirekrusin wrote:
           | Assuming that weights are not copyrightable, how much
           | restrictions can you put on output through API from those
           | networks/weights?
           | 
           | Ie. if ClosedAI says you can't use output of their API to
           | train competitive models - is that enforceable or not?
        
             | photonerd wrote:
             | That would be down to contract/terms of service. You'd be
             | in breach of that, not copyright
        
               | mirekrusin wrote:
               | But is it enforceable? Companies can put in contracts all
               | kind of nonsense, it doesn't mean all of it is
               | unconditionally enforceable, right?
               | 
               | Ie. if somebody creates company that sells milkshakes and
               | they say you can't use them to feed employees of
               | competing milkshakes companies - it wouldn't fly, would
               | it?
        
               | photonerd wrote:
               | Would strongly depend on the contract. Probably wouldn't
               | fly in a post sake terms of service agreement, but you'd
               | likely be in breach of a normal contract yeah.
        
           | dllthomas wrote:
           | > Outputs from LLMs, machine generated art, and machine
           | generated music probably are not copyrightable either.
           | 
           | I don't have a strong sense of whether this is reasonable (I
           | see arguments both ways) but I do think it's pretty strongly
           | at odds with how we treat photographs. There are a bunch of
           | photos on my phone where I unquestionably own the copyright,
           | despite putting in much less creativity than I did for some
           | AI images I've generated.
           | 
           | I don't think it's clear how to resolve this, but I do think
           | that _if_ we are going to protect photos and not prompted AI
           | images, the distinction needs to turn on something other than
           | whether  "sufficient creativity" was applied to the input of
           | the mechanical system.
           | 
           | Edited to add: It's probably also worth calling out that the
           | question of whether we protect the work produced by a
           | person's use of mechanical system is a separate one from
           | whether we protect the work of others when it is (in various
           | ways, to various degrees, with various likelihoods)
           | reproduced by use of those mechanical systems.
        
             | makeitdouble wrote:
             | On photography, the argument was condensed into "who pushed
             | the button". We saw it with the monkey auto-portrait
             | copyright fight where copyright was not granted to the
             | photographer, and other nature photography using photo
             | traps where the copyright stuck with the human basically
             | because they were the last operator of the camera.
             | 
             | The interesting part is, those controversial case are
             | pretty recent when the art of photography is century(ies?)
             | old now. I wouldn't expect super clear guidelines regarding
             | AI art before a few decades of weird cases fought tooth and
             | nails in court.
        
               | tokai wrote:
               | Eh? The copyright was the photographers and not the
               | monkeys.
               | 
               | https://petapixel.com/2018/04/24/photographer-wins-
               | monkey-se...
        
               | shagie wrote:
               | * * *
        
           | DannyBee wrote:
           | This is mostly right - It depends on what the weights
           | represent and how they were generated so I would not go as
           | far as the initial claim.
           | 
           | A collection of numbers is copyrightable if it's the encoded
           | result of a creative process. Just because it's represented
           | as a bunch of numbers does not make it non copyrightable.
           | That's why it says " original works of authorship fixed in
           | any tangible medium of expression, now known or later
           | developed, from which they can be perceived, reproduced, or
           | otherwise communicated, either directly or with the aid of a
           | machine or device. "
           | 
           | You can't just classify the weights as facts simply because
           | they are numbers. If they are creatively made by a human they
           | would be copyrightable. Mechanically computed from random
           | numbers, no. Somewhere in the middle? Harder
        
             | kevin42 wrote:
             | I'm not a lawyer, but it seems like you stood up a straw
             | man there.
             | 
             | >Just because it's represented as a bunch of numbers does
             | not make it non copyrightable.
             | 
             | Can you give an example of where the bunch of numbers is
             | copyrightable when it's not just a numeric encoding of
             | something that was already copyrightable? Taking music and
             | encoding it as a wav file is not a creative work, but it's
             | a representation of a copyrighted work.
             | 
             | Maybe you could create a long list of numbers and call it
             | an artistic impression, but that's clearly not what AI
             | weights are. I'm interested to hear an example of your
             | copyrightable numbers.
        
               | mlyle wrote:
               | The key factor of Feist v Rural is whether there was any
               | original or creative process in the way the facts were
               | arranged.
               | 
               | Here, there's a whole lot of creative decisions in
               | labelling and guiding of training that produces the
               | weights, so it's reasonable to think it might be
               | copyrightable.
               | 
               | That is, the numbers are a whole lot more original than
               | the issuance of phone numbers or part numbers.
        
               | Retric wrote:
               | The requirement for expertise doesn't necessarily imply
               | that that setting up perimeters for training AI is
               | necessarily copyrightable. A normal brick wall for
               | example needs skills to create but doesn't qualify as the
               | goal is not creative. If so the mechanical output of a
               | process that doesn't qualify for copyright is not going
               | to qualify.
               | 
               | Labeling training data may qualify for copyright, but if
               | the underlying training data doesn't taint the output as
               | a derivative work then labeling isn't going to qualify by
               | itself.
               | 
               | Thus without some new and very generous interpretation AI
               | companies are at best not going to benefit from copyright
               | and at worst may be forced to create all training data in
               | house. My suspicion is this generation of AI companies
               | are in a very difficult situation.
        
               | mlyle wrote:
               | > but if the underlying training data doesn't taint the
               | output as a derivative work then labeling isn't going to
               | qualify by itself.
               | 
               | It depends. If each individual training item has a small
               | impact on the output coefficients, then perhaps it's not
               | a derivative work of them. But if there's a large
               | creative process in determining model training procedure,
               | deciding labelling strategies, and applying those--
               | perhaps those numbers are strongly derived from _those_
               | things.
        
               | Retric wrote:
               | That sounds like wishful thinking, individual training
               | items have significant impact on the result.
               | 
               | Anyway, suppose you're building an AI to walk, there's
               | nothing creative about selecting 9.8m/s/s for gravity
               | that's simply the ideal value to achieve a desired goal.
               | Labeling an elephant as "Elephant" rather than "coat
               | hanger" is similarly a functional choice.
               | 
               | Just because a person is holding a camera and taking a
               | photo doesn't mean the result is copyrightable.
        
               | mlyle wrote:
               | > Anyway, suppose you're building an AI to walk, there's
               | nothing creative about selecting 9.8m/s/s for gravity
               | that's simply the ideal value to achieve a desired goal.
               | 
               | Suppose you're not building a strawman, but instead
               | building an AI to be an LLM. The exact sequence of what
               | you choose to do for instruction tuning, and the metrics
               | and labels that you choose, the prompt/response pairs you
               | write, and the loss functions you employ are quite
               | creative. They greatly affect the coefficients and are
               | not simple mechanical steps and are the result of a large
               | amount of creative choice.
               | 
               | We are nowhere near a point where they are an uncreative,
               | mechanical recipe to follow.
               | 
               | > Just because a person is holding a camera and taking a
               | photo doesn't mean the result is copyrightable.
               | 
               | No, but in the overwhelming majority of circumstances it
               | is. What it depends upon is whether the person holding
               | the camera is making a significant, original creative
               | choice.
               | 
               | I am not sure what courts will decide, but I am certain
               | that there is more creativity and originality employed
               | than you are giving OpenAI et al. credit for.
        
               | dragonwriter wrote:
               | > Here, there's a whole lot of creative decisions in
               | labelling and guiding of training that produces the
               | weights
               | 
               | Often, labelling is part of large public datasets that
               | are chosen for use for that exact reason, and/or is
               | otherwise not the work of the party claiming copyright in
               | the model.
        
               | [deleted]
        
               | mft_ wrote:
               | IANAL, but I'd wonder whether 'creativity' is really
               | present in labelling - and indeed, mightn't it be the
               | last thing you want? I'd argue labelling should be
               | strictly factual and reproducible, and ideally following
               | a logical structure... maybe akin to how addresses of
               | buildings might appear in a phone directory...
               | 
               | (Agree that the skill in knowing how to code and guide
               | the training of a model is probably very different
               | though. It's not just access to compute time that
               | separates me from OpenAI :) )
        
               | DannyBee wrote:
               | "Can you give an example of where the bunch of numbers is
               | copyrightable when it's not just a numeric encoding of
               | something that was already copyrightable?"
               | 
               | Sure, there are "poems" that consist of just a groups of
               | numbers that are copyrighted. They are not encodings,
               | it's just a string of numbers. It's indistinguishable
               | from a bunch of numbers. This is just one example, there
               | are lots.
               | 
               | They are enforceable to the degree it's creative, and to
               | the degree the infringing use is also creative.
               | 
               | So you would not be able to sue me for using those
               | numbers in a math equation. You would be able to sue me
               | for reproducing your poem in a book of poems :)
               | 
               | As feist says, the creativity required for copyright is
               | quite minimal. But it's still only as protectable as it
               | is creative.
               | 
               | Look - AI is not the first thing to have this "issue".
               | The answer remains the same as it always was - it's
               | mostly about the process not the output.
               | 
               | The output mostly matters is if the output is not
               | intended to be creative (or it's de minimis or ...).
               | 
               | Copyright as it currently exists is weird.
               | 
               | Like if you go to the copyright office and try to
               | register your ssh public key and say "this was generated
               | by ssh-keygen i had nothing to do with it" you _may_ get
               | a different result than if you said  "this is my new
               | visually stunning masterpiece, my ssh public key, which
               | was generated with computer help but I used 37 precisely
               | timed keyboard smashes to do it. Prints are available
               | from my gallery for $500"
        
               | mlyle wrote:
               | I fully agree with what you say, with one bit of nuance
               | to point out:
               | 
               | > Like if you go to the copyright office and try to
               | register your ssh public key and say "this was generated
               | by ssh-keygen i had nothing to do with it" you may get a
               | different result than if you said "this is my new
               | visually stunning masterpiece, my ssh public key, which
               | was generated with computer help but I used 37 precisely
               | timed keyboard smashes to do it. Prints are available
               | from my gallery for $500"
               | 
               | The important thing, of course, isn't whether the
               | copyright office denies to register your copyright, but
               | instead what courts will ultimately do when you attempt
               | to enforce your copyright.
               | 
               | We know the current administrative algorithms used by the
               | copyright offices. We have less clarity on what courts
               | will ultimately do.
        
             | dxbydt wrote:
             | > Mechanically computed from random numbers, no
             | 
             | Even random numbers are copyrightable.
             | 
             | Below is an implementation of Marsaglia's invention, from
             | p348, courtesy infamous NR[1]. Its a MWC (multiply with
             | carry) random number generator, with two parameters,
             | variable a and base b=2^32. --- For a, "The values below
             | are recommended with no particular ordering." ID a B1
             | 4294957665 B2 4294963023 B3 4162943475 B4 3947008974 B5
             | 3874257210 B6 2936881968 B7 2811536238 B8 2654432763 B9
             | 1640531364 --- as we all now know, the whole thing is
             | copyrighted - you can't redistribute that code and can't
             | use those specific numbers to generate random numbers
             | without purchasing a license, which only allows you to use
             | it once in your personal machine; that's why GSL[2]. The
             | pseudorandom numbers you would get from MWC if you use
             | above numbers are also copyrighted since they are work-
             | product.
             | 
             | [1]http://numerical.recipes/book/book.html
             | [2]https://www.gnu.org/software/gsl/design/gsl-design.html
        
             | saynay wrote:
             | I would say that is uncertain. Model weights are always
             | going to effectively be a huge collection of statistics
             | about the training corpus. Unless you are envisioning
             | artisanal, hand-crafted, free-range model weights where a
             | person used a non-mathematical method to purposely and
             | creatively choose each one?
        
         | WanderPanda wrote:
         | Of course weights are copyrightable. Otherwise nothing is
         | copyrightable
        
           | enlightens wrote:
           | Recipes, for example, are not copyrightable in the US.
           | Neither are some of the concepts behind creating a fillable
           | form. It's not an all-or-nothing system.
           | 
           | https://www.copyright.gov/circs/circ33.pdf
        
         | throwaway98721 wrote:
         | Why is it unestablished? Is a document not copyrightable based
         | on its contents? Weights are just a different kind of a
         | document.
        
           | Conscat wrote:
           | No, a document's contents aren't inherently copyrightable.
           | They have to be a creative work or a method of production,
           | and part of that basically means it has to be human generated
           | content (as opposed to computer or animal generated).
           | 
           | AI weights might be considered a method of production, but
           | that isn't clear yet.
        
             | throwaway98721 wrote:
             | [flagged]
        
               | ketzu wrote:
               | > Was there no work put into their creation by someone?
               | 
               | Putting work into something is not a sufficient cirteria
               | for copyright.
               | 
               | > All of it is just a stream of bytes that the computer
               | can interpret somehow
               | 
               | This is also not a sufficient or at all relevant cirteria
               | for assigning copyright.
               | 
               | Also, in the sense you presented, those files are not
               | fundamentally different from random noise. Which is not a
               | particularly useful reduction for this exercise.
        
               | dragonwriter wrote:
               | > Was there no work put into their creation by someone?
               | 
               | This is the "sweat of the brow" theory of
               | copyrightability, which courts have rejected (for good
               | reason based on the statute.)
               | 
               | "Someone did work to enable this thing to exist" is not
               | sufficient to make a thing copyright-protected.
               | 
               | > There's no fundamental difference between an image,
               | code, or weights.
               | 
               | And neither images, code, nor weights that are
               | mechanically produced with no creative input by a
               | particular author are subject to copyright in their own
               | right (depending on their relation to the source material
               | on which the mechanical process rests, they may be
               | covered by the copyright on the source material.)
               | 
               | The _best_ argument for weights being copyrightable (and
               | it probably applies better to some models than others) is
               | that the assembly of source material is a creative work
               | subject to a compilers copyright, and that the model
               | weights themselves are just a mechanical translation of
               | that compilation subject to its copyright.
        
           | jerf wrote:
           | Copyright is not for "documents", it is for works that have
           | creativity in them. The legal bar for that level of
           | creativity is low, so low that it is easy to come away
           | thinking that anything that can be cast as a "document" must
           | be copyrightable, but the bar is in fact not zero.
           | 
           | In particular, taking other documents and shoving them
           | through a process that generates a lot of other numbers with
           | no human or creative interaction is definitely something I'd
           | be concerned the courts would judge as not sufficiently
           | creative to be copyrightable. The process itself would
           | certainly consist of copyrightable code, but the output
           | doesn't necessarily. This would be somewhat similar to the
           | observation that there is no copyright to be had in a big
           | table of files and their MD5 hashes (or other hashes), such
           | as a Linux distro might use for integrity checking. Lots of
           | copyright in the original file contents, copyright available
           | on the process for producing these tables, but the _tables
           | themselves_ would likely be ruled not itself copyrightable as
           | there is no creativity in that output.
           | 
           | Note this also has absolutely nothing to do with the question
           | of whether AI output is copyrightable, this is about the huge
           | table of numbers that make up the neural net weights being
           | copyrightable. (Though it would be sort of an interesting
           | question for the legal system to grapple with as to how a
           | non-copyrightable set of numbers could then produce something
           | copyrightable. Call it a philosophical variation on the
           | "copyright washing" argument; can copyright spring from a
           | non-copyrightable source other than a human brain, thus
           | somehow "flowing uphill"? Would a human brain be
           | copyrightable? Stay tuned for those questions, I guess, or if
           | not you, your grandchildren.)
           | 
           | Per your other comments, "work" is not the bar, "creativity"
           | is. "Size" is not the bar either. Merely being a much larger
           | table of numbers than a list of hashes or a phone book is not
           | the question. No human is in that table of numbers creatively
           | saying "no, wait, this neural weight should be -1.5 instead
           | of 2.0 to produce this creative effect". No human is even
           | _capable_ of working in the medium of neural net weights in a
           | creative manner.
           | 
           | If you want to go the "novel legal theory" route, you could
           | play with claiming creativity in the selection of input
           | material and claim the resulting neural weights has a
           | copyright in compilation:
           | https://en.wikipedia.org/wiki/Copyright_in_compilation That's
           | a long way from a slam dunk though. Way out on a legal limb
           | there. It isn't entirely clear to me what exact rights would
           | result from such a claim either. It would be a landmark
           | copyright court case for sure.
        
             | AnimalMuppet wrote:
             | IANAL, but I suspect that the "novel legal theory" in your
             | last paragraph would fail. It might succeed if you gave GPT
             | a hand-curated list of materials; hoovering up the entire
             | internet is not that.
        
           | dragonwriter wrote:
           | Weights are the output of a mechanical process over the
           | training set with no element of human authorship, just as
           | much the output a model produces with a prompt is, which the
           | Copyright Office has already declared outside of copyright.
           | 
           | > Is a document not copyrightable based on its contents?
           | 
           | Creative process is the bigger issue.
           | 
           | > Weights are just a different kind of a document.
           | 
           | And who sits down and writes this document of weights?
        
           | raincole wrote:
           | > Is a document not copyrightable based on its contents?
           | 
           | Yes, exactly. It's copyright 101.
           | 
           | For example, if you write a random number generator, and
           | print 10000 randon numbers in a document, it's not
           | copyrightable.
           | 
           | Even if you invented a specific random number generation
           | algorithm, the document is still not copyrightable. Your code
           | is copyrightable.
           | 
           | Again it's just copyright 101. If any of above surprises you,
           | maybe you should read a few copyright case studies.
        
             | [deleted]
        
         | xg15 wrote:
         | Furthermore, _if_ weights are copyrightable, wouldn 't this
         | make the issue of training data licenses even more urgent?
         | 
         | IANAL, but if weights are IP, wouldn't they constitute a
         | "derived work" of the training data?
        
           | golemotron wrote:
           | In a sane legal system a new copyright law would be passed to
           | clarify all of this. In ours, the poor copyright office needs
           | to make things up on the fly.
           | 
           | Their recent decision that implies that anything that AI is
           | used to produce is non-copyrightable is silly, sad, and not
           | sustainable.
        
           | YetAnotherNick wrote:
           | No, weights are not just data fed but also the training
           | process itself. I think the whole argument hinges on how much
           | human thought and action is needed in training the model.
           | 
           | On the other end of the spectrum, AI generated content
           | couldn't be copyrighted if there is no human involvement. If
           | someone asks GPT to write 1000 poems, it couldn't be
           | copyrighted.
        
           | realusername wrote:
           | That's also my understanding, either the weights are
           | copyrightable and then all the models need explicit
           | agreements for any work they include in it because models
           | become derivatives or they are not copyrightable being just
           | machine data (the most likely scenario in my opinion), they
           | can't have it both ways.
        
             | AnimalMuppet wrote:
             | "Transformative use".
             | 
             | The inputs could be copyrighted _and_ the weights could be
             | copyrighted _if_ creating the weights from the inputs is
             | (legally) regarded as a transformative use. And I think it
             | could reasonably be considered to be transformative - the
             | weights don 't look anything like the input data.
             | 
             | Disclaimer: IANAL. So far as I know, no court has ruled on
             | whether this qualifies as a transformative use. I take no
             | position on how the courts will actually rule. I merely say
             | that they _could_ regard this as transformative use. (But
             | see jerf 's "creativity" argument for another hurdle that
             | weights must pass to be copyrightable.)
        
               | shagie wrote:
               | Transformative use doesn't necessarily mean
               | copyrightable.
               | 
               | Google's thumbnails are a purely mathematical
               | transformation on images (no copyright themselves), and
               | yet are considered a transformative use.
               | 
               | I believe that trained models are similarly a purely
               | mathematical transformation of {data}, but is
               | transformative in what that _can_ be used for going
               | forward.
               | 
               | "Can" bearing a lot of weight in that sentence.
               | 
               | It's how the human, with agency, uses the model that may
               | be a derivative or copyright infringing use - not the
               | model itself nor necessarily the output.
               | 
               | The output of a generative AI _may_ be similar enough to
               | an existing work that it is derivative of that work. It
               | is possible to construct a prompt that infringes on an
               | existing work _even if that work wasn 't part of the
               | training data_.
               | 
               | For that case, consider you drew a picture. That picture
               | that you just drew isn't part of any training data. I
               | could presumably look at it and describe it with
               | sufficient detail that something similar enough would be
               | generated... and that may be considered a derivative
               | work. The same test could be applied to me describing it
               | to someone on Fiverr with the same outcome.
               | 
               | If I were to publish that work by the generative AI or
               | Fiverr - who would be infringing on copyright? me? or the
               | black box that may be AI or Fiverr that created a picture
               | based on my prompts?
        
               | numpad0 wrote:
               | Another way to look at it is if a thing reproduced a data
               | subjectively resembling originals and then you used it
               | anyhow, then its non-transformative use, and methods used
               | is just extra details.
        
             | phantom784 wrote:
             | I think there could be an argument that it's copyrightable
             | but not a derivative work.
             | 
             | If I read a few books about a subject as research, and then
             | I write an article about the subject, it's my own
             | copyright. The fact that I did research doesn't make it
             | derivative of those books (correct me if I'm wrong, IANAL).
             | 
             | Perhaps a model created from copyrighted material be
             | treated in the same way?
        
               | floomk wrote:
               | That's because you are human and have rights that a
               | computer program doesn't
        
               | adamc wrote:
               | A fertile subject for sf stories.
        
               | bloak wrote:
               | > If I read a few books about a subject as research, and
               | then I write an article about the subject, it's my own
               | copyright.
               | 
               | Yes, because in that case you'd be the "author" doing
               | "creative work".
               | 
               | > Perhaps a model created from copyrighted material be
               | treated in the same way?
               | 
               | Who would be the author doing creative work in this case?
               | The people who decided what training material to use?
               | Perhaps, but it seems a stretch for the people who
               | selected the training material to be authors but not the
               | people who created the training material.
        
               | slaymaker1907 wrote:
               | The difference is you are person and have many more
               | rights than a machine.
        
             | OkayPhysicist wrote:
             | There is also a (IMO less likely, but still conceivable)
             | scenario where weights ARE copyrightable, but represent
             | fair use of the training data on grounds of being
             | "sufficiently transformative".
        
               | floomk wrote:
               | Sadly this seems to be the most likely considering how
               | the US is ran
        
               | 8note wrote:
               | I consider that one super likely, but then using the
               | model to make competing works with one the artists in
               | their own style is a non-fair use derivative work
        
               | Filligree wrote:
               | Style explicitly isn't copyrightable. It'll need to be
               | for some other reason.
        
               | OkayPhysicist wrote:
               | Your case wouldn't be about style, it would be about
               | specific elements that you posit were memorized and
               | regurgitated by the model. The fact that you're creating
               | art in the same style/medium as the author is what
               | negates the "sufficiently transformative" fair use
               | defense.
               | 
               | Basically, that world ignores the AI model completely. If
               | your resulting work wouldn't be fair use if you directly
               | were working with something from the training set, it
               | wouldn't be fair use if you fed it through an AI model
               | first.
        
               | shaky-carrousel wrote:
               | Also known as "having your cake and eating it".
        
         | jrm4 wrote:
         | Exactly. A _lot_ of the difficulty here is how they skip is the
         | hugely important issue:
         | 
         | An entirely reasonable, if not fully tested, statement is the
         | following:
         | 
         | Every single one of these AI weight things _itself_ is a result
         | of unencumbered, massive, law-breaking, right-violating
         | copyright infringement -- accordingly, it 's _extremely_
         | difficult to say anything morally justifiable or authoritative
         | about anyone elses  "rights" downstream, and to try to inject
         | the word "ethical" makes the whole thing even more ridiculous.
        
           | visarga wrote:
           | > is a result of unencumbered, massive, law-breaking, right-
           | violating copyright infringement
           | 
           | Why? Copyright covers expression not information, AIs can
           | learn information from any source regardless of copyright.
           | They should just not regurgitate copyrighted content, that's
           | all. And much of what organic content is online is common
           | knowledge, thus can't be copyright-controlled.
        
             | pfdietz wrote:
             | Copyright is for things that are the result of human
             | creativity. If the weights come from running an algorithm
             | on a training set (that one does not have a copyright to)
             | then how can the weights then be copyrightable? They might
             | be a derivative work, but that just means they infringe
             | copyright, not that they are copyrightable themselves.
        
               | shagie wrote:
               | Note that the requirements for copyright are not
               | consistent between nations.
               | 
               | The US has the "threshold of originality" as its
               | principle. Under that doctrine, it requires some _human_
               | (and this has been emphasized many times over the years)
               | originality in order for something to be copyrighted. It
               | 's a low bar for how original it needs to be, but it must
               | be human (monkeys taking selfies are not human).
               | 
               | https://en.wikipedia.org/wiki/Threshold_of_originality
               | 
               | In England, the doctrine is "sweat of the brow" instead.
               | 
               | https://en.wikipedia.org/wiki/Sweat_of_the_brow
               | 
               | > Under a "sweat of the brow" doctrine, the creator of a
               | work, even if it is completely unoriginal, is entitled to
               | have that effort and expense protected; no one else may
               | use such a work without permission, but must instead
               | recreate the work by independent research or effort.
               | 
               | The definitive case for this in the US that set the two
               | apart is Feist Publications, Inc., v. Rural Telephone
               | Service Co. ( https://en.wikipedia.org/wiki/Feist_Publica
               | tions,_Inc.,_v._R.... ) where it was deemed that a
               | telephone directory is not copyrightable in the US as
               | there is no originality in it... but under the sweat of
               | the brow doctrine it would have been.
               | 
               | So the "[c]opyright is for things that are the result of
               | human creativity" gets an "it depends" and it would be
               | curious to see if companies that are firmly in the
               | "models are valuable" camp go to the UK for what I
               | believe would be a more favorable copyright protection.
               | 
               | ... _However_ there are other IP laws around trade
               | secrets that may be better for it in the US (I 'm not as
               | familiar in that domain - I would be curious to find
               | out).
        
               | stale2002 wrote:
               | The answer would be if the weights are transformative
               | enough, and the copyright would come from the person who
               | decided what images to include in the training set.
               | 
               | The act of choosing to place images in a certain
               | arrangement, such as a collage, can be copyrightable. The
               | same could be said for the "act" of choosing what images
               | to include in a training set and which parameters to use
               | to train the model.
        
               | kevin42 wrote:
               | Does that mean if someone copies a phone book but leaves
               | out some numbers and adds some other numbers then it's a
               | creative work?
        
               | stale2002 wrote:
               | It would depend on how transformative the work is.
               | 
               | There is in fact a whole art form where people cut out
               | words from different newspapers and books, for example,
               | and re-arrange those words to form new and interesting
               | art.
               | 
               | So there are ways in which such a work would be a
               | creative work, and ways it which it would not, and it
               | would depend on the particular instance and example.
        
               | saynay wrote:
               | The legal system is not like a computer program. The line
               | between what is "creative" and what is not concrete, but
               | is instead up to the interpretation of the judge who
               | rules on it.
               | 
               | So your phonebook modifications may or may not be
               | considered "creative" depending on the judge and your
               | ability to convince them. The more your modify it, the
               | more likely you are to convince a judge it is a creative
               | work, though.
        
             | wheelie_boy wrote:
             | It seems very difficult to ensure that a model will never
             | output any of the copyrighted content that it was trained
             | on. I can only think of three ways, but perhaps there are
             | others
             | 
             | 1. Evaluate every output from the model to ensure that none
             | of the outputs are copyrighted
             | 
             | 2. Evaluate every input to a model to ensure that the
             | inputs are either not copyrighted or properly licensed
             | 
             | 3. Change the definition of copyright so that ML models can
             | do whatever they want
             | 
             | Nobody is doing #1, because that makes the business models
             | not work. Established brands (like Adobe) are doing #2. I
             | get the feeling that there are a lot of ML startups that
             | are hoping that #3 will happen, but it seems unlikely
        
               | og_kalu wrote:
               | Ensuring a model never outputs copyrighted content is
               | unimportant and tangential. It's irrelevant. You don't
               | look for a way to make humans output no copyrighted
               | content, you address each time they do case by case.
               | 
               | A model training being rendered fair use doesn't mean any
               | of its output can be used for whatever regardless.
        
               | wheelie_boy wrote:
               | > you address each time they do case by case.
               | 
               | That's what I listed as #1 - evaluate each individual
               | output of the model to see if it violates copyright.
        
             | wilde wrote:
             | Tell that to some illegal numbers:
             | https://en.wikipedia.org/wiki/Illegal_number
        
             | slaymaker1907 wrote:
             | My issue with this take is that machines are not people. We
             | only have lax rules for humans precisely because they are
             | humans, not on the basis that they can learn. Copyrighted
             | works are produced for people and given how human learning
             | works, applying the derivative works rule to humans would
             | be completely impractical and destroy the point of works
             | with copyright. The same cannot be said for AI companies
             | treating everything on the internet as fair use for
             | training AI.
        
             | version_five wrote:
             | A lot of people are just upset because their local
             | equilibrium has been disrupted and they think that means
             | they lost a natural right.
             | 
             | "You wouldn't look at a car and then remember what that
             | looked like when someone asks you to draw another"
        
               | jrm4 wrote:
               | These are not bad arguments, but I don't think they're
               | conclusive. I am a lawyer, and I could absolutely see
               | this going the other way. "You can't make these machine
               | things without literally feeding this copyrighted
               | information into them, therefore they do contain a copy.
               | You can see this by when they reproduce, e.g. the "getty
               | images" deal."
               | 
               | *this is not legal advice, dangit commenter person below
        
               | m4nu3l wrote:
               | > you can't make these machine things without literally
               | feeding this copyrighted information into them, therefore
               | they do contain a copy.
               | 
               | They don't necessarily do. Think about that. You can take
               | some copyrighted material and transform the information
               | contained in it (for instance a fictional book). You can
               | then write a summary. The summary contains information
               | that was present in the original but it has been
               | transformed and hence it's not a copy. The ML model
               | contains information that has been generalized by some
               | degree. So it's just a grey area IMO.
        
               | saynay wrote:
               | More over, you are clearly not in violation of copyright
               | if you are talking about statistics about the material.
               | In your example, printing out a "there were 7000
               | instances of the word 'the'" is certainly not a
               | violation. A ML model is just a huge pile of these
               | statistics.
               | 
               | However, saying "the first word of the book is 'The'"
               | would not be a violation, while repeating that for every
               | word in the book, as a whole, would be one.
        
               | version_five wrote:
               | I agree with you but I think it's important to have some
               | nuance. Imagine I build a statistical model for 10-word
               | sequences (10-grams) and then I trained it on a single
               | book. I probably could pick some starting words and get
               | most of the book back from the "statistics" I compiled.
               | If I trained the same model on a giant dataset, the one
               | book would just contribute to the stats.
               | 
               | All that to say, the models have potential to memorize,
               | but they don't, and if they do it's an undesirable
               | failure mode, not some deliberate copying.
        
               | jrm4 wrote:
               | I like this argument a lot; but again -- how does this
               | play out in the real world? It's pretty easy to refute
               | what will happen in real life. Think, e.g Batman. I could
               | write a very new and original "Batman" comic that doesn't
               | strongly resemble anything -- movie, toy, comic, whatever
               | -- that exists, but would be recognizable to fans.
               | 
               | Once it starts doing well, will DC come after me? You
               | bet.
        
               | ke88y wrote:
               | These models can definitely be used to intentionally
               | store and recall content that is copyrighted in a way
               | that's not subject to fair use. (eg: trivially, I can
               | very easily train a large model that has a small
               | subnetwork which encodes a compressed or even lossless
               | copy of a picture, and if I were to intentionally train a
               | model is that way then this would be no less a copyright
               | violation than distributing a JPEG of the same image
               | embedded in some large binary).
               | 
               | But also, an unintentional copy of a copyrighted image is
               | not a violation of copyright. (eg: an executable binary
               | which happens to contain the bits corresponding to a
               | picture of Batman -- but which are actually instruction
               | sequences and were provably not intended to encode the
               | picture -- clearly doesn't infringe.)
               | 
               | LLMs are somewhere in-between #1 and #2, and the intent
               | can happen both in the training and also the prompting.
               | 
               | Stack on top of this the fact that the models can also
               | definitely generate content that counts as fair use, or
               | which isn't copyrighted.
               | 
               | It's the multitude of possible outputs, across the
               | copyright spectrum, combined with the function of intent
               | in training and/or prompting, which make this such a
               | thorny legal issue for which existing copyright statute
               | and jurisprudence is ill-suited.
               | 
               | Taking your Batman example: DC would come after you for
               | trademark as well as copyright, and the copyright claims
               | would be very carefully evaluated with respect to your
               | very specific work. But here we are talking about a large
               | model that can generate tons of different work which
               | isn't subject to copyright or which is possibly fair use.
               | 
               | I don't think that existing jurisprudence (or even
               | statute?!) can handle this situation very well, at all,
               | without tons of arbitrary interpretative work on the
               | parts of juries/judges, because of the multitude and
               | vague intent issues described above.
               | 
               | (...Also presumably the merits of the DC case wouldn't
               | matter because your victory would be pyhrric unless you
               | are a mega-corp. Which from a legal theory perspective is
               | neither here nor there but from a legal practicality
               | perspective may inform how companies go about enforcing
               | copyright claims on model weights/outputs.)
               | 
               | Anyways. I think we have a right mess on our hands and
               | the legislature needs to do their damn jobs. Welcome to
               | America, I guess :)
               | 
               | Curious to hear your thoughts on these issues.
        
               | visarga wrote:
               | This is a great example. Summarizing or paraphrasing
               | copyrighted content, or simply using it as a seed to
               | generate input-output pairs - this kind of data
               | transformation prior to training could solve the issues
               | with copyright. It cleanly separates form from content.
        
               | pxoe wrote:
               | what is a 'copy'? byte accurate, or 'something with
               | general resemblance'? would a badly compressed "copy"
               | image of a copyrighted material still be 'a copy' or
               | would it be some other thing? would low quality image
               | compression be enough to skirt around copyright claims?
               | image formats and viewers just 'reproduce' an impression
               | of original data from derive compressed data. it is also
               | just 'information that's been generalized by some degree'
               | - for space saving purposes and so on. so, what if image
               | generators could be thought of as a 'very good multi-
               | image compression algorithm' that can output multiple
               | images as well, to a 'somewhat recognizable degree'.
        
               | hex4def6 wrote:
               | Badly compressed still counts. I think if the data allows
               | you to reconstruct a recognizable recreation of the
               | original work, you have a good chance of it being
               | considered a derivative copy.
               | 
               | A mono audio version of Star Wars, compressed down to
               | 320x240, filmed from the back of a theater on a VHS
               | camera, converted to Video CD, would under any reasonable
               | interpretation be just a copy of the original.
               | 
               | I assume it starts getting murky when there's some sort
               | of transformation done it it. What if I run motion
               | capture on it, and use that motion capture data to create
               | a cartoon version of Star Paws (my puppies in space
               | epic)? What if I do a scene for scene recreation as the
               | animated cartoon (removing any mentions to copyrighted
               | names -- Luke Skywalker is now Duke Dogwalker, for
               | example)? In this case, there's been no actual data
               | transfer -- all the sprites are hand drawn, backgrounds
               | etc.
               | 
               | What would be an interesting exercise would be to try and
               | create a series of artifacts that each on their own are
               | considered non-derivatives, but can be used together to
               | reconstitute the original. For example, create a
               | compression method that relies heavily on transforms /
               | macroblocks, but strip out any of the actual pixel data
               | from the film. That info might be supplied as palette
               | files which are themselves not really copyrighted data,
               | but together with the compressed transform stream can be
               | used to recreate the original video.
        
               | mike_d wrote:
               | > The summary contains information that was present in
               | the original but it has been transformed and hence it's
               | not a copy.
               | 
               | The summary also contains original thought, something is
               | added to it by a human to make it unique. AI models are
               | primarily deriviative.
               | 
               | A better example would be: if I take 1,000 different
               | copyrighted works and put them into a ZIP file, does that
               | resulting file violate copyright?
        
               | version_five wrote:
               | That example is awful, whatever side of the debate one is
               | on
        
               | kevin42 wrote:
               | Let's say you take the harry potter books and create a
               | spreadsheet with each word in it as a column, and the
               | number of times that word appears. Would that violate the
               | copyright? I'd be interested in the rationale if someone
               | thinks it would.
        
               | mike_d wrote:
               | If your table was the number of times a word was followed
               | by a chain of other words, that would be a closer
               | comparison to AI weights. In that case it would be
               | possible with reasonable accuracy to reconstruct passages
               | from the harry potter books (see GitHub Copilot).
               | 
               | The copyright aspect makes more sense when you start
               | thinking of AI training models as lossy compression for
               | the original works. Is a downsampled copy of the new Star
               | Wars movie still protected under copyright?
               | 
               | Just tabulating the word counts would not violate
               | copyright as it is considered facts and figures.
        
               | drdeca wrote:
               | It resembles lossy compression in some ways, but in other
               | important ways I think it doesn't?
               | 
               | Like, if one has access to such a model, and doesn't
               | count it towards the size cost of a
               | compression/decompression program nor as part of the
               | compressed size of the compressed images, then that
               | should allow for compressing images to have substantially
               | fewer bits than one would otherwise be able to achieve
               | (at least, assuming that one doesn't care about the
               | amount of time used to compress/decompress. Idk if this
               | is actually practical.)
               | 
               | But unlike say, a zip file, the model doesn't give you a
               | representation of like, a list of what images (or
               | image/caption pairs) it was trained on.
               | 
               | Or like, in your analogy with the lower resolution of the
               | movie, the lower resolution of it still tells you how
               | long the movie is (though maybe not as precisely due to
               | lower framerate, but that's just going to be off by less
               | than a second, unless you have an exceedingly low
               | framerate, but that's hardly a video at that point.)
               | 
               | There is a sense in which any model of some data yields a
               | way to compress data-points from it, where better models
               | generally give a smaller size. But, like, any (precisely
               | stated) description counts as a model?
               | 
               | So, whether it is "like lossy compression" in a way that
               | matters to copyright, I would think depends a lot on
               | things like,
               | 
               | Well, for one thing, isn't there some kind of "might
               | someone consume the allegedly infringing work as a
               | substitute for the original work, e.g. if cheaper?" test?
               | 
               | For a lower resolution version of Star Wars movie, people
               | clearly would.
               | 
               | But if one wanted to view some particular artwork that is
               | in the training set, I would think that one couldn't
               | really obtain such a direct substitute? (Well, without
               | using the work as an input to the trained model, asking
               | it to make a variation, but in that case one already has
               | the work separate from the model, so that's not really
               | relevant.)
               | 
               | If I wanted to know what happened in minute 33 of the
               | Star Wars movie, I could look at minute 33 of the
               | compressed version.
        
               | londons_explore wrote:
               | It's just a race for which test case gets to the supreme
               | court first really...
        
               | pmoriarty wrote:
               | ...and the Supreme Court could rule however it likes. It
               | doesn't matter what anyone else says, or what any law
               | says, what any lawyer or other judge says.
               | 
               | They could be completely biased, could completely ignore
               | everyone and everything else and rule however they want.
               | 
               | I'm almost surprised they still bother to write any kind
               | of "legal reasoning" in their ruling and don't simply
               | focus on what the ruling is rather than why they ruled
               | that way. But I guess such "reasoning" still serves a
               | propaganda purpose and still provides a fig leaf for
               | those who still believe in the quaint absurdity that "we
               | are a nation of laws, not men."
        
               | londons_explore wrote:
               | Supreme court precedent seems to impact a lot of
               | decisions...
               | 
               | Plenty of companies who have legal teams will keep an eye
               | on the legal landscape of court decisions, and use them
               | to decide if our T&C's or contracts need rewriting, or if
               | any precedent puts us at legal risk.
               | 
               | Sure - the supreme court could overthrow its precedent
               | anytime, but until it does, a lot of people will act as
               | if what they say is the law.
        
               | dragonwriter wrote:
               | > It's just a race for which test case gets to the
               | supreme court first really...
               | 
               | Not really for practical purposes. In the long term, the
               | Supreme Court can and does overrule its own precedent, so
               | the first case on the specific issue to get to the
               | Supreme Court doesn't end the discussion.
               | 
               | In the short-term, cases get resolved by lower courts and
               | parties either lack funds to do the maximum level of
               | appeals, or the Supreme Court chooses not to hear appeals
               | (they tend to prefer an issue to be well-developed with
               | circuit case law, often waiting till there is a conflict
               | between the Circuit Courts of Appeal, before taking it
               | up), so the state of the law _prior_ to any specific
               | ruling on the narrow topic by the Supreme Court matters
               | quite a bit.
        
               | staunton wrote:
               | This is the first time I ever saw a comment including the
               | text "I am a lawyer". Does that mean the comment
               | technically contains "legal advice"?
        
               | lcnPylGDnU4H9OF wrote:
               | As always, _a_ lawyer is not necessarily _your_ lawyer.
        
               | mike_d wrote:
               | If you choose to pay him, sure.
        
               | wahnfrieden wrote:
               | IP itself violates a natural right
               | 
               | (Yes the idea of rights is also unnatural and absent from
               | visions such as anarchy)
        
               | version_five wrote:
               | Yeah I didn't even think that was controversial. I'd
               | always been taught that copyright and patents exist to
               | explicitly restrict what people can do by granting a
               | monopoly to the owners in order to encourage invention
               | and creative work.
               | 
               | Edit to add I'm not saying I agree with the justification
               | or am trying to argue for it, only that the point above
               | is commonly raised as the justification, implying that
               | the intrusion on a person's rights is known and accepted.
        
               | wahnfrieden wrote:
               | [flagged]
        
               | dragonwriter wrote:
               | Natural rights are a fiction to pretend that someone's
               | moral code is a privileged aspect of physical reality in
               | a way every competing moral code is not.
        
               | ndriscoll wrote:
               | Even so, you can ask whether a given moral code is more
               | principled than another (e.g. in the sense of having some
               | algebraic structure), and use that to investigate what
               | might be considered "more natural". For example, one
               | might argue that if a "natural" right exists, then it
               | ought to be symmetric under exchange of humans (or
               | sentient beings or whatever). It's then "more natural" to
               | conclude that you have a right to perform actions that
               | have no interaction or consequences for other humans
               | (e.g. to sing a copyrighted song to yourself in an empty
               | room or downloading a song that you already have on CD
               | but don't feel like ripping yourself) than those that do
               | (e.g. taking food from someone because you'd otherwise
               | starve).
        
               | dragonwriter wrote:
               | > Even so, you can ask whether a given moral code is more
               | principled than another
               | 
               | What does "more principled" mean of a moral code? How
               | does one quantify "degree of principledness"?
               | 
               | > and use that to investigate what might be considered
               | "more natural".
               | 
               | What does the preceding (being "more principled") have to
               | do with being "more natural"? And what significance does
               | being "more natural" have?
               | 
               | And none of that has any relevance to what is usually
               | described as "natural rights"; its like taking existing
               | words and coming upnwith entirely novel meanings and then
               | a whole architecture around them, which is pretty
               | advanced equivocation.
        
               | __MatrixMan__ wrote:
               | That's going a bit far. They're just fictions that are
               | privileged over certain other fictions--it's like how you
               | can often cast magic missile in D&D but you can't usually
               | cast expelliarmus, it comes down to which fiction we
               | agree to inhabit.
        
             | svachalek wrote:
             | When the web was young, there was a lot of information
             | considered "public" like criminal record, marriage records,
             | birth certificates, property records, etc. But those were
             | still fairly veiled because of the amount of effort
             | required to see them. Suddenly these were getting blasted
             | all over the internet because now that was an easy thing to
             | do, and everyone had to rethink what "public" meant.
             | 
             | I suspect we're going to see the same kind of rethink about
             | intellectual property in the age of AI.
        
               | pulvinar wrote:
               | Not sure about criminal records, but the other records
               | are generally still public. Not blasted all over, but
               | there if you know where to look.
               | 
               | Not that we really had all that much privacy in the past,
               | as anyone who's browsed old newspapers knows.
        
           | ldoughty wrote:
           | > Every single one of these AI weight things itself is a
           | result of unencumbered, massive, law-breaking, right-
           | violating copyright infringement
           | 
           | Maybe the popular and free ones. Adobe has a product in beta
           | that uses "ethical training data" as a selling point.
        
             | jrm4 wrote:
             | Interesting. I wonder what they mean by "Ethical" --
             | instead of e.g. saying "definitely free and open." I'm
             | willing to bet "stuff they gathered from likely unwitting
             | Adobe users."
        
               | FooBarWidget wrote:
               | It means trained on data from stock photo sites they own,
               | for which all uploaders agreed to terms of service which
               | state that uploaded materials can be fed into AI
               | training.
        
               | jrm4 wrote:
               | So yes, exactly what I said. :)
        
               | themoonisachees wrote:
               | Adobe also happens to own Adobe stock, so maybe they
               | simply trained on their own corpus.
               | 
               | Who am I kidding this is Adobe of course they're fucking
               | over their users
        
               | Dr4kn wrote:
               | They at least say they did and other copyright free
               | artworks. They are a big company and know that they would
               | get sued, so it should be in their interest to do it this
               | way.
        
           | zitterbewegung wrote:
           | If I collect a set of copyright free data or public domain
           | data would we conclude that the weights are also public
           | domain?
        
             | jrm4 wrote:
             | That seems fair. I was under the impression that there
             | weren't too many out there like this.
        
             | jprete wrote:
             | No, that doesn't follow at all. The argument is that either
             | the training or the expression violated existing cooyrights
             | through the making of unlicensed copies. It's not based on
             | open source licensing. Although OSS viral licensing may
             | well apply if fair use is not a successful defense.
        
               | tedunangst wrote:
               | What copyright is violated by training on public domain
               | data?
        
               | numpad0 wrote:
               | DMCA works for free data too. It doesn't matter if your
               | gains are in the form of fiat or crypto or social
               | currency.
        
               | jprete wrote:
               | If it's public domain, then no copyright is violated. I'm
               | not talking about public-domain data; the G-G-GP
               | specifically mentioned the possible legal interpretation
               | that training on large amounts of publicly visible (but
               | not public domain) data is itself a copyright violation.
        
             | numpad0 wrote:
             | IANAL, I rather think the weight is not copyrightable
             | anyway, and, if I build a model on copyrighted data, I
             | would conclude that the inseparable but reproducible parts
             | of weight retains copyrights, despite the whole weight not
             | having its own.
        
             | 8note wrote:
             | Are they a work of art in and of themselves? I don't think
             | you could tell without litigation
        
           | [deleted]
        
           | hcks wrote:
           | > is a result of unencumbered, massive, law-breaking, right-
           | violating copyright infringement
           | 
           | Is there an official ruling? Or is it just a Reddit-style
           | over exaggeration?
        
             | Dr4kn wrote:
             | There is no official ruling... yet. We are very early in
             | this rapid public development. Laws and rulings take years
             | or decades.
             | 
             | They are trained on a lot of text. News sites, comments,
             | books etc. Most books and news sites fall under copyright.
             | Is this fair use? Who knows. Fair use is also an American
             | thing. ChatGPT can be used in the EU, which doesn't have
             | such a broad view of fair use.
             | 
             | If you make a game only out of a lot of copyrighted assets
             | without paying it isn't fair use. Are LLMs different?
             | 
             | What about image generation, which you can prompt the
             | models for specific styles of artists, which works are all
             | copyrighted, but still used for training?
        
       | barbariangrunge wrote:
       | completely off topic, but funny: I misread "opencoreventures" as
       | "opencorevultures"
        
         | Makhini wrote:
         | Funny
        
       | morpheuskafka wrote:
       | > AI also poses socio-ethical consequences that don't exist on
       | the same scale as computer software, necessitating more
       | restrictions like behavioral use restrictions
       | 
       | There's plenty of software that has, or could have, similar
       | restrictions. Consider software that allows you to plan vantage
       | points for a shooting or estimate the impact of using explosives
       | at various locations. And the government regulates all sorts of
       | software for export/download because it has military use--
       | everything from development tools to high performance chips that
       | could be used to crunch numbers for a nuclear program, CAD
       | software that can help you build (or destroy) a bridge, etc. The
       | CPUs and GPUs themselves are regulated at certain performance
       | levels, I think.
       | 
       | None of this is really new to AI.
        
       | ronsor wrote:
       | Model weights are not source code, but data. Arguably because of
       | how they are generated, they are not even copyrightable at all.
        
         | adamsmith143 wrote:
         | Corporate data is of course protect-able. Otherwise why don't
         | you just open up all your databases so anyone can access them?
        
           | WrongAssumption wrote:
           | Copyright protection is what gives protection when you put
           | something out into the public. The desire to not publish
           | something is evidence against having these protections,
           | because people know they are not copyright able, so for that
           | reason and others they keep it private. You just presented
           | evidence against your position.
        
           | DannyBee wrote:
           | They are only protectable by copyright you the degree they
           | are creative works of authorship. Copyright is not usually
           | how these are protected.
        
       | zarzavat wrote:
       | Weights might be copyrightable but in no universe are they
       | copyrightable by OpenAI, Google, etc just because they did the
       | training and spent money on GPUs.
       | 
       | The only people who can possibly own the copyright, if any such
       | copyright exists, are the authors of the training data.
       | 
       | I find this whole discussion about copyright of weights almost
       | absurd, the incredible amount of deference given to our corporate
       | lords is such that we are "hallucinating" new forms of IP
       | protection for NN weights that have never existed in any kind of
       | statue or case law and cut completely against the grain of all
       | the law that currently exists.
        
       | amelius wrote:
       | Just like you can't de-compile a binary without loss of
       | information, "source" means that you can reconstruct it, so the
       | training data should be available as well as the code that was
       | used to train it, and the build script that invoked it.
        
       | robomartin wrote:
       | Can someone give me a legal answer to this?
       | 
       | People, from early school, all the way up to university, use
       | copyrighted materials to learn various topics and obtain degrees.
       | This trains our brains using the work of others.
       | 
       | The same is true as we navigate life. We learn various skills and
       | subjects consuming the work of others.
       | 
       | And, yes, in the case of most people, we use that training to
       | pursue various careers, obtain work and get paid for it.
       | 
       | How can there be a claim of infringement on the part of LLM's and
       | not on every person who has ever used a book, website, article,
       | video or publication to learn something?
        
         | og_kalu wrote:
         | This is an argument yes. A model could certainly be considered
         | transformative enough to be fair use.
        
         | blharr wrote:
         | I am not a lawyer. But isn't this quite simple?
         | 
         | Copyrighted materials are either licensed specifically for a
         | human or it's implied that a human will use them to learn.
         | 
         | Naturally, human memory is going to distort and change that
         | information over time. But as soon as you use it in an AI,
         | which has superhuman capabilities of memory, that would go out
         | the window.
        
           | robomartin wrote:
           | > Copyrighted materials are either licensed specifically for
           | a human or it's implied that a human will use them to learn.
           | 
           | I don't think that's a part of copyright law at all. Maybe in
           | the future, not today. Which makes sense, since these laws
           | precede AI by a long time.
        
       | low_tech_punk wrote:
       | The lack of freedom to modification makes it not "open" either.
       | 
       | Comparing to traditional software, weights are actually worse
       | than binary. You can't "decompile" the weights into the training
       | source code so there is no way for the community to make useful
       | changes to them.
        
       | kmeisthax wrote:
       | >While the RAIL organization suggests adding the word "Open" to
       | RAIL licenses that include similar open-access and free-use as
       | open source (i.e. OpenRAIL-M), this is confusing since the
       | license is not open source so long as it includes usage
       | restrictions. A better name would be EthicalRAIL-M. Using the
       | term "ethical" to describe this category license clearly
       | indicates its functional difference from open source licenses.
       | 
       | I don't even think we should be using the word "ethical" because
       | it implies that anything more permissive is _un_ ethical. We
       | should call these morality clause licenses.
       | 
       | The question of whether or not we _should_ have morality clauses
       | involved is complicated. Most bad actors do not give a shit about
       | the licensing status of the code they are using. And these
       | licenses also cause headaches for people who want to follow the
       | rules[0] and avoid copyleft trolling[1]. On the other hand, the
       | morality clauses in OpenRAIL-M are relatively straightforward and
       | non-obnoxious.
       | 
       | [0] This also applies to "non-commercial" licensing, since that
       | is a concept entirely foreign to copyright law. As far as I'm
       | concerned the 'NC' clause in Creative Commons just means 'OK to
       | torrent'.
       | 
       | [1] A practice in which people abuse copyleft licenses to try and
       | extract licensing agreements for minor license violations. The
       | forgiveness periods added to GPLv3 and later versions of Creative
       | Commons are specifically to prevent this behavior.
        
       | c7b wrote:
       | Imho the weights are the real meat for most typical models, you
       | can run with them and continue training them with your own code.
       | It's not even guaranteed that the original code would be very
       | useful for that.
       | 
       | But if you are going to make that distinction, for which you can
       | make a case I think, shouldn't you include a third dimension,
       | 'data'? The code alone is hardly useful if you want to rebuild
       | the weights, but all it tells you is that they're loading their
       | proprietary data and then using PyTorch to set up and train the
       | model. You can't reproduce anything using just that. So the real
       | equivalent of open source would be imho either open weights, or
       | open data plus code plus weights (the latter are arguably
       | redundant, but still practical to include). Given that the size
       | of that repo will typically be gigantic, I think open weights is
       | the case we should really be focusing on. I'd rather have a paper
       | explaining the model together with the weights, rather than code
       | that I can't run anyway, if I'm designing an algorithm to
       | continue training the model.
        
       | meindnoch wrote:
       | According to whom?
       | 
       | Weights are a type of program, which are interpreted by the
       | neural network runtime. Same as Java bytecode interpreted by the
       | JVM runtime.
        
         | eigenket wrote:
         | x86 machine code is a type of program, which is interpreted by
         | the processor, but distributing the binary of my program
         | doesn't make it open source.
        
           | kfarr wrote:
           | Bingo, did a ctrl+f to find binary as that seems like the
           | closest analogy here.
        
         | slowmovintarget wrote:
         | Weights are data, not a type of program.
         | 
         | A computer program is a set of instructions that may be
         | executed. Weights are values that may be loaded by a program,
         | but are not a program in and of themselves.
        
           | earleybird wrote:
           | Weights are data in the same way that instruction codes in
           | memory is data.
        
             | slowmovintarget wrote:
             | Values for the variables do not the function make.
        
               | daniel-cussen wrote:
               | [dead]
        
           | rockinghigh wrote:
           | When people talk about weights, they talk about a network of
           | weights that takes an input and computes an output. There is
           | really not much difference between a saved model and a
           | program.
        
           | jstanley wrote:
           | It's a very difficult distinction to make.
           | 
           | Would you consider a Python program to be data rather than
           | program just because it is text input to the python
           | interpreter instead of machine code for the CPU?
        
             | slowmovintarget wrote:
             | It is not at all a difficult distinction.
             | 
             | Weights are literally numbers computed as output. They are
             | not instructions. The semantics of those numbers even when
             | emplaced (loaded) in an artificial neural net is such that
             | they do not execute. They are not instructions. LLM engines
             | and diffusers perform searches where the weights are used
             | to calculate additional output.
             | 
             | Is source code, like Python text, data? Yes. All code is
             | data. But not all data are source code.
             | 
             | If I gave you a web request log, you would not assert it is
             | a program. If I gave you a CSV file with time-series values
             | from a sensor, you would not assert it is a program. If I
             | hand you a database of contact information, you would not
             | assert it is a program. Weight files are the equivalent of
             | CSV files. They are are a dump of parameter values computed
             | from training.
             | 
             | They are not a program.
             | 
             | The definition of computer program is well worn. So is the
             | definition of source code, and the definition of
             | parameters. Weights are parameters.
        
               | xigoi wrote:
               | If a program has to consist of instructions, then source
               | code written in a declarative language is not a program.
        
               | jstanley wrote:
               | The difference between code and data only exists in our
               | minds. There is no distinction. Both code and data make
               | the computer do things (and, yes, both code and data
               | _only_ make the computer do things if other conditions
               | are permitting, for example if executed with the right
               | interpreter, or loaded with the right type of viewer).
               | Anything that can be expressed as code can be expressed
               | as data, and vice versa.
        
               | pravus wrote:
               | > Weights are literally numbers computed as output. They
               | are not instructions.
               | 
               | They are instructions if you consider the LLM system
               | itself to be a kind of weird, indirect virtual machine.
               | Each number can be mapped to a set of instructions that
               | are executed. Even your CPU uses numbers (machine codes)
               | to execute.
               | 
               | Join me in saying: ...code is data is code is data is
               | code is data...
        
           | graypegg wrote:
           | They're not data though, they're coefficients. They are the
           | only thing that significantly differentiates one model from
           | another.
           | 
           | If I told you the economy can be accurately modelled by
           | 
           | GDP(x) = Ax + B
           | 
           | But I don't define A And B for you because it's proprietary,
           | you haven't learned anything other than what you can glean
           | from the structure of the model itself (it's linear, there's
           | only a single input etc)
           | 
           | If most of these models are similarly structured, I'd say the
           | weights are the program.
        
             | slowmovintarget wrote:
             | The nature of the data as proprietary or not, important or
             | not, is not relevant.
             | 
             | Parameters, or actual arguments, are values; data. Not
             | instructions.
             | 
             | Valuable data is still data. It's significance doesn't
             | magically turn it into source code.
        
           | golemotron wrote:
           | No, declarative programs exist. They are not instructions.
           | 
           | There is no real line between code and data. This is an
           | observation that runs all the way from Turing Machines in
           | computability theory to the Von Neumann architecture and
           | homoiconicity in Lisp.
           | 
           | What we call 'data' is just code that needs a cleverer
           | interpreter.
        
             | Izkata wrote:
             | Less into theory and more into "wait wtf": Some of the
             | older projects I've worked on were written by people who
             | loved database-driven stuff, to the point they did things
             | like put perl code into one table column (with sentinel
             | values you had to find/replace before `eval`ing the code)
             | and sql into another table that retrieved values for those
             | find/replaces, both retrieved and executed by some really
             | generic code.
             | 
             | Code or data: Well... both.
        
           | killjoywashere wrote:
           | But not "raw" data. They are derived from other data and a
           | program. If this was a collaboration where one collaborator
           | did the processing and one sourced the data, they would
           | likely both claim some amount of ownership of the trained
           | weights.
           | 
           | At a minimum, it would be an active area of negotiation that
           | the attorneys would take notice of. Source: have negotiated
           | these agreements.
        
             | slowmovintarget wrote:
             | A curated data set is still a data set.
             | 
             | I imagine it is not settled law, but there's a clear
             | argument to be made that regardless of the difficulty in
             | curating the data set, it's still a data set.
             | 
             | Can it be licensed and sold. Yes, surely. Is it proper to
             | pretend an open source license is sufficient protection,
             | probably not.
        
           | programmarchy wrote:
           | This is a distinction without a difference. Code is data and
           | data is code.
        
             | mrguyorama wrote:
             | Maybe they've only worked with machines using a Harvard
             | Architecture
        
             | slowmovintarget wrote:
             | All source code is data. Not all data is source code. Data
             | may be encoded, but that doesn't make it source code
             | either.
        
         | tensor wrote:
         | I think the point here is that by being explicit you avoid the
         | need to have this argument.
        
         | cdelsolar wrote:
         | who wrote that program?
        
         | adamnemecek wrote:
         | Java bytecode is not "open source". At least for Java bytecode
         | there are decompilers.
        
       | thepangolino wrote:
       | I've always seen weights as akin to configuration files.
        
       | Makhini wrote:
       | What if you change the weights slightly? Kaboom, not breaking the
       | copyright anymore.
        
         | bskap wrote:
         | Then it's a derivative work and copyright law covers that too.
        
           | rpodraza wrote:
           | And you're basing this theory on what exactly?
        
             | sharcnick wrote:
             | Copyright law & the definition of a "derivative work." See
             | e.g. 17 USC SSSS 101 and 106. See also
             | https://www.copyright.gov/circs/circ14.pdf.
        
       | FrustratedMonky wrote:
       | Are the weights in our brain copyrightable?
       | 
       | Might want to get ahead of the curve on this one. How would this
       | work? Would I get a tattoo with a license spelling out covering
       | the contents of my body?
        
         | DannyBee wrote:
         | Has to be fixated (unchanging) and in a tangible medium.
        
           | RobotToaster wrote:
           | So I just need to cryogenically freeze my brain in order to
           | copyright it?
        
       | adamsmith143 wrote:
       | The question shouldn't be whether the weights are copyrightable
       | but whether they are protected under other electronic
       | communication/data privacy laws.
        
       | jkeisling wrote:
       | The article makes a good point: we should prevent "open-washing"
       | and draw a distinction between well-intentioned restrictive
       | licenses like "Open"RAIL and true open source. However, I worry
       | the name "ethical source" is itself a bit question-begging. While
       | outfits like Bloom may believe in good-faith ethical principles,
       | their definition of ethics isn't necessarily everyone's. If
       | restricted models are "ethical", is releasing open weights
       | "unethical"? Conversely, is releasing a model with PII or artist
       | styles in it "ethical" if a few known use cases are forbidden?
       | There's no one right answer. Labeling any one set of restrictions
       | as "ethical" off the bat makes discussion harder and puts open
       | source on the back foot to justify "not being ethical". Better to
       | just call them "restricted models" or "guarded models", and leave
       | it to individuals to decide if these restrictions are beneficial
       | or not.
        
         | A4ET8a8uTh0 wrote:
         | I think the more interesting aspect of all this is that the
         | confusion created by this new business model ( not sure to
         | classify it so business model had to do ) appears to be largely
         | intentional. The subject matter is complicated to begin with
         | experts being niche of a niche of a niche and the assumption
         | that the general public can even understand it ( and whether it
         | can even dumbed down to digestible sound bites ) is, in my
         | mind, very optimistic. Now, courts are not typically stacked
         | with dummies, but again how many are well versed in issues of
         | technology?
         | 
         | All in all, I don't disagree with the point you raised, but I
         | worry that all this will only further muddy the water for the
         | general population.
        
           | pmoriarty wrote:
           | _" Now, courts are not typically stacked with dummies, but
           | again how many are well versed in issues of technology?"_
           | 
           | Even if they are well versed in issues of technology that
           | does not mean they'll make what any given one of would
           | consider a good decision, as plenty of people well versed in
           | issues of technology disagree with each other on these
           | issues.
           | 
           | Nothing guarantees that on, on any issue, really, as you can
           | always find people who disagree.. and if they happen to be
           | judges, they get to decide unless another higher judge
           | overrule them.. and that judge has the same problem as the
           | first.
        
             | A4ET8a8uTh0 wrote:
             | Sure. My point is that I would so much rather have a
             | decision handed down that was considered on actual merits (
             | we might disagree, but at least I would be able to see some
             | sort of real consideration and not what amounts to talking
             | points from various lobbyists ). A judge that has zero
             | exposure in that area is at best 50/50 and regardless of
             | the ruling I will be annoyed that a person with zero
             | knowledge is declaring how something he knows little to no
             | about can be used ( just like I am more and more annoyed
             | about political class in Washington, but I am more inclined
             | to believe these days they know exactly what they are doing
             | -- serve their own interests ).
             | 
             | To your point, it is absolutely not panacea ( new blood is
             | inevitably ending in government and the result so far is in
             | line with what you said ), but it would at least be a
             | starting point.
        
       | tiffanyg wrote:
       | _AI licensing is extremely complex. Unlike software licensing, AI
       | isn't as simple as applying current proprietary /open source
       | software licenses. AI has multiple components--the source code,
       | weights, data, etc.--that are licensed differently._
       | 
       | Are you joking? This isn't _wrong_ , per se, but it's worded as
       | though written by someone with only the most casual / cursory
       | interaction and knowledge of this area of law / commerce (e.g.,
       | including licensing, copyright, trademark / service mark, patent,
       | etc.) ... until perhaps quite recently.
       | 
       | Yes, the AREA IS complicated. No, so-called "AI" is not
       | introducing all sorts of novel issues, structures, etc. "AI" has
       | some nuances distinct from much of what has come before (happens
       | basically every time more significant tech comes along) and some
       | possibly more unique questions related to economics, ethics,
       | philosophy, and the like, but the relevant areas of law and
       | practice have often been complicated and sort of "bleeding edge",
       | even going back before the industrial revolution.
       | 
       | Big money, powerful tech, large-scale economic forces, etc. =
       | lots of maneuvering, legislation, litigation, etc. = complicated
       | "rules of the game".
       | 
       | Drawing the distinction vs. software in general is reasonable -
       | but, the rather click-baity headline and "I just learned about
       | 'IP' law and bah gawd y'all are doin' it wrong" tone to the start
       | of this article suggest, to me, that this isn't likely to be the
       | best article to use as a reference to learn about these issues.
        
         | larodi wrote:
         | I was like going to write 'are u joking', but you make the same
         | point so well. This article is at best oversimplifying and
         | misleading.
         | 
         | Besides I doubt this 'my weights your weights' thing is a thing
         | at all.
        
       | cf141q5325 wrote:
       | A focus on licensing ignores that there are security incentive to
       | not run just any weights you find floating around the net.
       | Getting exploited through miss-aligned networks is a very real
       | threat and really hard to combat.
        
       ___________________________________________________________________
       (page generated 2023-07-05 23:01 UTC)