[HN Gopher] AI weights are not open "source"
___________________________________________________________________
AI weights are not open "source"
Author : subomi
Score : 196 points
Date : 2023-07-05 15:18 UTC (7 hours ago)
(HTM) web link (opencoreventures.com)
(TXT) w3m dump (opencoreventures.com)
| TheRealPomax wrote:
| So, Open Data. Got it. This is the same category as config files
| that are kept up to date by a program as it runs.
|
| - Is it "a program"? Very clearly not.
|
| - Is it source code? You can argue either way. The program won't
| work without it, but "this specific one" is not required for the
| program to do _something_ , and that ambiguity means you probably
| don't want to call it "source code" because it's too vague.
|
| - Is it _data_ used by a program in order to perform its task?
| Absolutely. It even uniquely defines the program behaviour, and
| so is a thing onto itself within the context of the program it 's
| used by.
| Zetobal wrote:
| If my own data is in the dataset even when I didn't give consent
| is it a collaborator dataset?
| tensor wrote:
| If you posted your data into a service where the TOS allows
| this use then yes.
| iLoveOncall wrote:
| [flagged]
| worksonmine wrote:
| > Unlike software licensing, AI isn't as simple as applying
| current proprietary/open source software licenses. AI has
| multiple components--the source code, weights, data, etc.--that
| are licensed differently.
|
| Software also has multiple components, often the same as the ones
| listed by the author. But what do I know, to me AI is just
| another example of software.
| mensetmanusman wrote:
| Weights are an information asset that require millions in capital
| and burned-out GPUs to mine and refine.
| daniel-cussen wrote:
| [dead]
| light_hue_1 wrote:
| I think this is very shortsighted.
|
| Weights are a program. CUDA is an interpreter for that program.
|
| One day we will be able to decompile these programs into
| something more human understandable.
| bee_rider wrote:
| > The ethical license category applies to licenses that allow
| commercial use of the component but includes field of endeavor
| and/or behavioral use restrictions set by the licensor.
|
| I don't love the name, "ethical license" sounds like a
| description of the license: this license is ethical. Really this
| sort of license imposes a particular ethical framework on the
| user.
|
| Not to throw shade, though. It is actually hard to come up
| neutral sounding name for this sort of license I think. I keep
| thinking of things like "morality encumbered license," but that
| sounds ridiculously euphemistic in a weird way.
| version_five wrote:
| Yes I was going to say the same thing. It's a branding that has
| been applied by the license's proponents, and I personally
| reject a lot of what they call "ethics" as well as the idea of
| whatever monitoring and enforcement the restrictions entail -
| maybe calling it a religious license would be better.
| iandanforth wrote:
| "Opinionated" is how I think about it.
| bee_rider wrote:
| That might be a good pick, IMO the word has negative
| connotations elsewhere, but in tech circles seems basically
| neutral.
| 93po wrote:
| I'd argue any licensing of IP is unethical. I'd use the word
| "conditional"
| cpcallen wrote:
| I'm disappointed that the article is only making the (somewhat
| pedantic) distinction between source code and weights. From the
| quotation marks in the headline I hoped that it would instead be
| making the distinction between human-readable source code and
| machine-readable compiled form.
|
| For example, IMHO (IANAL) an AI code-completion tool that had
| been trained on GPL software is (or should be) only be legal to
| distribute if it is accompanied by the training code _and all the
| code ingested during training_ (or an offer to provide such code
| upon request).
| version_five wrote:
| This is an interesting point. If you read the OSI open source
| definition, specifically on source code (quoted below) I'm
| inclined to treat the training data as part of the source code
| for the purpose of determining whether to consider any model
| open source. 2. Source Code The program
| must include source code, and must allow distribution in source
| code as well as compiled form. Where some form of a product is
| not distributed with source code, there must be a well-
| publicized means of obtaining the source code for no more than
| a reasonable reproduction cost, preferably downloading via the
| Internet without charge. The source code must be the preferred
| form in which a programmer would modify the program.
| Deliberately obfuscated source code is not allowed.
| Intermediate forms such as the output of a preprocessor or
| translator are not allowed.
|
| https://opensource.org/osd/
| TrackerFF wrote:
| Weights are just matrices with values between a certain range. So
| are digital images - just matrices with values. Images are
| covered by copyright laws, so why shouldn't weights also be?
| mellosouls wrote:
| Hmm. Makes a few unsubstantiated claims, with hand-wavy appeals
| to risks that our private corp overlords are presumably
| protecting us humble users from, now that they've built their
| product on open source and data by closing it down and changing
| terminology to suit.
|
| There's an intelligent discussion to be had, and I think this
| otherwise-reasonable article could be part of it if it toned down
| the presumption and condescension a little.
| seydor wrote:
| If it is extremely complex, then it can only be modeled by an AI
| horsawlarway wrote:
| If anything - this entire conversation just highlights (Over and
| Over and Over and Over again) how absolutely bonkers abusive our
| current copyright laws are.
|
| The vast majority of small individuals are compelled by contract
| to surrender their rights to large corporations. Those large
| corporations then abuse the ever loving fuck out of those rights.
|
| The express intent of copyright is now a sad joke.
|
| Personally - I'm pretty over the entire show. This system is
| generating an incredible amount of inequality. New and novel
| content is absolutely NOT getting made, and these laws are
| creating vicious infights that drain resources from well
| intentioned companies & individuals and pass them along to
| complete scam corporations.
|
| We are told stories as children that we cannot retell in our own
| voices decades later to our own children.
|
| I am firmly ready to burn this copyright system to the fucking
| ground. It's been 300 years since the Statute of Anne - I'm ready
| for a different game.
| PartiallyTyped wrote:
| There is also the whole patent / copyright trolling issue too.
| The fact that $BIG_CORP can hire armies of lawyers to freeze
| competitors and beat them to market by filing frivolous
| lawsuits is yet another example of insanity in the whole
| system.
| chongli wrote:
| I recently watched the documentary _Fire in the Blood (2013)_
| [1] about the use, by big pharma, of patents and WIPO to
| obstruct access to affordable antiretrovirals (ARVs) in
| Africa during the worst years of the AIDS epidemic, leading
| to over ten million deaths. All of this when the African
| market for these medications represented less than 1% of the
| total market, in dollars. It's absolutely infuriating!
|
| [1]
| https://en.wikipedia.org/wiki/Fire_in_the_Blood_(2013_film)
| drdaeman wrote:
| It's a problem with legal system (not unique to any specific
| country, mind you, the problem is global), not patent or
| copyright system specifically. It grew incredible amounts of
| complexity so _pro se_ became a sad joke in all but simplest
| cases, and there 's no incentive to fix it - quite the
| opposite, everyone in the system is all for keeping the
| status quo, because it generates money.
| TaylorAlexander wrote:
| Personally I don't think patents do what people believe
| they do (encourage innovation). It's a bigger discussion
| but briefly, the only literal function of a patent is to
| discourage innovation by legally barring anyone from using
| a patented idea as part of a new innovation. The idea we
| have is that the secondary effects of this will be
| increased profits for inventors and therefore more
| innovation. But actually there's loads of secondary effects
| and often many of them outweigh the effect of increased
| profit. For every one inventor that gets a patent there
| might be 100 prevented from using that idea in a different
| and innovative way.
|
| A classic example is 3D printers. Stratasys spent 15 years
| selling printers that cost tens of thousands of dollars. It
| wasn't until the patent expired that people figured out how
| to make them for $250. Those cheaper printers are enabling
| mechanical engineers and designers to accelerate their
| process and make other new innovations faster. Stratasys
| had such a powerful patent they never bothered innovating
| down in price, instead rested on their laurels selling $25k
| printers to big customers.
|
| So how many inventions were delayed or shelved because the
| inventors couldn't afford a $25,000 3D printer, and $250
| printers didn't exist yet? Both Stratasys and IBM held
| patents related to 3D printing and they had to cross
| license to go in to production, so how many others would
| have come up with 3D printing in the 1990's if they had not
| been patented? Would first mover advantage in a free market
| have been enough to stimulate development of 3D printers?
| Could we have had $2000 3D printers in the early 2000's
| (Stratasys sold theirs for $30k) instead of ten years
| later? How many engineers would have invented new gadgets
| faster if they had a 3D printer ten years earlier?
| PartiallyTyped wrote:
| Another possible example is the x86 and x86-64 ISAs
| locked between AMD and Intel. I don't think Intel would
| have become complacent had there been more competitors...
|
| ... or the whole "oracle vs google" over the java API.
| jrumbut wrote:
| But there are specific problems with copyright and patent
| law that could be improved without a global systemic
| overhaul that may never happen.
|
| We have to take some small wins even in the presence of big
| problems.
| drdaeman wrote:
| Of course. I'm just saying that the core problem is
| larger than just the copyright and patent law.
| loudmax wrote:
| Fully agree that the existing copyright and intellectual
| property systems are dire need of deep reform. But to get
| people on board, you can't just propose burning it all down,
| you need to point to a viable alternative. Say, limiting
| copyrights to something sane like 15 or 30 years. Or making it
| easier to invalidate obvious or trivial patents.
|
| Or do you really want to do away with notions of intellectual
| property altogether? You can make an argument for that, but
| there would lead to deep economic changes, and you need to
| anticipate what the end result would look like. You still need
| some way to encourage the creation of new content.
|
| Pointing out that our copyright/IP system is broken is easy.
| And you're right, it's totally broken! Coming up with a fix is
| hard work.
| Frost1x wrote:
| >Pointing out that our copyright/IP system is broken is easy.
| And you're right, it's totally broken! Coming up with a fix
| is hard work.
|
| The problem I have with these arguments is they ultimately
| tend to boil down to the devil you know or the devil you
| don't know.
|
| We keep claiming when something is broken we must provide a
| "fix" and the assumption is that fix has to be better than
| the current approach. There's pretty much no way to guarantee
| this because the systems in place are the only systems with
| evidence. So, because we have other ideas, we dare not try
| them because they have to "fix" the problem. The amount of
| inertia that keeps corruption in motion bothers me and at a
| fundamental level most of the inertia comes down, ironically,
| back to property ownership. If we abolish copyright or change
| it we have to make sure things are fair/equitable. Well sure,
| that's ideal, but what we have isn't even remotely fair and
| equitable anymore, so even something broken is likely an
| improvement.
|
| We have no willingness as a society to try some modifications
| and be willing to accept failure, then shift to the next
| modification and iterate around until we get something sane
| in place. As such, the systems in place remain in place and
| more and more holes are found to exploit as time progress.
|
| Our systems need to be more adaptable. Founders of the
| country understood that which is why they made the legal
| system a legal adaptable system. The question has always been
| though, what is the threshold? We've played it safe so long
| that much of the entire system designed to adapt to fix these
| issues has itself been targeted and gummed up intentionally
| to prevent that.
| pessimizer wrote:
| > You still need some way to encourage the creation of new
| content.
|
| Do you? What's the argument for this? Is there some sort of
| extreme shortage of creative work that the state should find
| it necessary to encourage it? How about we end copyright, and
| if there's ever a problem, we offer copyrights for a short
| period to fluff the commons up again. A copyright anti-
| holiday, as it were.
|
| Instead we do the opposite: automatically copyright
| everything anyone produces, and make it very difficult to
| surrender your copyright (unless Google or Microsoft want it,
| then if you object you're literally a Luddite caveman who is
| trying to turn back the clock on modernity because you're
| old, stupid, and afraid of fire.)
| Zaskoda wrote:
| > But to get people on board, you can't just propose burning
| it all down, you need to point to a viable alternative.
|
| This may be true for most people, I don't know. However, I
| personally am fully on board the "burn it all down" train and
| have been for a while.
| TaylorAlexander wrote:
| I don't think we need to use the legal system to encourage
| creation of new content! That's a natural thing people do. In
| fact there's a lot of artistic remixing that is illegal or
| ambiguously legal under the current copyright regime that can
| be a powerful form of expression.
|
| I really don't think we need government policy to encourage
| artists to create art. (At least not of this sort - I am all
| for art grants.)
| capr wrote:
| Pointing out that IP is broken is _not_ easy because most
| people believe in the contradictory notion of intellectual
| property, you included, not knowing the legal history of IP,
| the legal and economic history of the concept of property,
| and so on. If it were easy, it would be obvious to everybody
| that 1) IP law is immoral and 2) nothing bad would happen if
| it's abolished outright.
|
| Here's a free ebook on the subject, written by a patent
| lawyer no less: https://mises.org/library/against-
| intellectual-property-0
| mike_d wrote:
| > 2) nothing bad would happen if it's abolished outright.
|
| It is interesting that people living in the places with the
| weakest IP laws will pay a premium to import baby formula
| from the places with the strictest laws.
|
| > Here's a free ebook on the subject
|
| Of course it is some right-libertarian wonk piece.
| immibis wrote:
| This has nothing to do with IP laws, and everything to do
| with baby formula laws. You don't seriously think that
| without the ability to sell the recipe, nobody would
| invent a safe and effective baby formula, right?
| mike_d wrote:
| I'm not sure you fully grasp all the dimensions of IP
| law.
|
| If you have two brands of baby formula, Death brand that
| kills babies, and OK brand that is perfectly fine, and
| you start putting Death brand in fake cans labeled OK
| brand - that is absolutely an IP enforcement issue. The
| desire of OK brand to protect their brand, and profits,
| combined with reasonable IP laws allows them to lead
| enforcement actions and protect consumers.
| immibis wrote:
| That is absolutely not the kind of intellectual property
| that anyone hates.
|
| The kind of intellectual property we are talking about is
| the one where Death isn't allowed to make baby formula
| that doesn't kill babies, because OK patented making baby
| formula that doesn't kill babies and won't give them a
| license.
| xigoi wrote:
| > you start putting Death brand in fake cans labeled OK
| brand - that is absolutely an IP enforcement issue.
|
| This is a matter of trademark, which is completely
| orthogonal to copyright and nobody is protesting against
| it here.
| 111111IIIIIII wrote:
| > _I am firmly ready to burn this copyright system to the
| fucking ground._
|
| Same, but the issue is not copyright, which is simply an effort
| to wield the state to control intellectual property in the same
| way the state is wielded to control physical property.
|
| The compounding problem arises when property is _capital_ ,
| defined as the means to convert labor into new value.
| Capitalism is specifically a system in which one can wield
| control of capital (intellectual or otherwise) to extract
| profit from labor then trade that profit for more capital. As a
| result, capital accumulates infinitely, independent of the
| value produced by the labor which is provided to society.
|
| Artists require capital to convert their labor into value just
| as any other worker would, so where should that capital come
| from if not from control of the value they produce? Society
| must solve this problem or we will not have art to begin with.
| Only looking at the demand side obfuscates such issues that
| arise on the supply side, and the only reason we're talking
| about them now is that digital technology has solved the
| scarcity problem on the supply side. It has not solved the
| scarcity problem on the demand side, however.
|
| Finally, art, just like all technological progress, is always
| the product of entire societies and the history of all mankind
| that came before it. For this reason, all copyright and patents
| have no rational basis and are merely bandaids for the ill side
| effects of controlling capital to extract profit from labor to
| begin with.
| habitue wrote:
| One thing I don't see discussed enough is that, ok let's say the
| weights are unencumbered, and the source is under an OSI license:
| the point of open source licenses and free software was to expose
| the *human understandable* meaning of the final program.
|
| That's why distributing binaries isn't allowed even though
| technically all of the functionality is present in the machine
| code. AI weights are basically binary blobs. We don't know what
| they mean, there is really no source code for them. The best we
| can do is various black box manipulations on them like LoRA, etc,
| similar to what we can do to a binary blob.
| phkahler wrote:
| >> AI weights are basically binary blobs. We don't know what
| they mean, there is really no source code for them.
|
| No. You can do further training on them. If they are something
| less than code I don't think it's going to warrant all this
| talk about licensing. GPL, MIT, or some proprietary should
| cover it.
| habitue wrote:
| You can do further training on them, just like you can patch
| a binary blob. There are some surgeries you can do to the
| weights, and there are analyses you can do to poke at them
| and try to understand them, but ultimately they weren't
| created from a human understandable spec, and without a ton
| of reverse engineering work the weights by themselves aren't
| human understandable: hence the "source" component is
| missing.
|
| The source code that generated the weights is one step
| removed from the kind of source code we'd need to interpret a
| bunch of AI weights. It's really meta-source code
| [deleted]
| Topfi wrote:
| This post did cover many of the same ideas I have been ruminating
| on concerning model weights and the nomenclature of current
| efforts. That's also why I generally tend to stick with calling
| these[0] "local/self hosted models" for the time being. A major
| reason for my reluctance is that I see weights far closer to
| binary than code, making a distinction important and current FOSS
| concepts not really applicable.
|
| Of course, this all hinges on the idea that weights by themselves
| are inherently protected by current copyright, which still seems
| to be an unsettled topic, hotly debated by both laypeople and
| legal professionals. Authors generally are afforded copyright on
| their work by default, and weights raises so many questions
| concerning authorship that have never been considered.
|
| This being such a contested issue, which will require new laws
| and/or precedent (depending on the legal system), is very
| problematic. Regardless of where you live, generally courts and
| government entities are not famous for their speedy reaction to
| new things, so clarity may take a while, at which point the
| industry might have already settled on some agreement that then
| may be adopted as a basis for actual legislation, which would
| likely favor financially well baked entities already actively
| lobbying for their interests, such as OpenAI.
|
| Some have also pointed out that this is arguing semantics, and I
| am tempted to agree in principle, but also want to emphasize that
| I feel this is a situation where that can be valuable. Should
| weights in some way be afforded copyright protection, clear
| nomenclature will be needed. Putting some thought into this now
| is definitely not the worst idea.
|
| I very strongly feel that the specific word "ethical" as part of
| defining licenses is not the best idea, though. "Ethical" can
| carry vastly different connotations, depending on a myriad of
| factors, many of which would go beyond the use-focused definition
| laid out in the post. Due to this, I'd argue for "behavioral" or
| "restricted use" over "ethical", as both more clearly state what
| the intended effect is in cases such as Open RAIL-M[1].
|
| Part of my strong feelings on the use of the word "ethical" come
| from the fact that with weights and training data, there has been
| a lot of discussion concerning both rights of and considerations
| for creators whose published works have been used to create those
| weights. Due to this, the use of "ethical" referring to a group
| of licenses could give some the impression that this may indicate
| that the training data used was "ethically sourced", i.e. in
| agreement with the original creator. This is something that in my
| eyes should also have clear labeling, though with weights being
| very hard to reliably trace back to source data, it currently
| seems impossible to verify, making this essentially just a good
| faith effort.
|
| [0] https://huggingface.co/tiiuae/falcon-40b-instruct
|
| [1]
| https://drive.google.com/file/d/16NqKiAkzyZ55NClubCIFup8pT2j...
| ianbutler wrote:
| I'm not sure OCV gets to decide any of this. Just like I don't
| think OSI trying to be the sole dictator of the term "Open
| Source" works out long term. My opinion is always received
| controversially about things like this, but terms evolve to meet
| the common usage by the people. If people are calling this "Open
| Source", and there are more people who want to call this "Open
| Source", than people who don't; unless you intend to legally bar
| them from using the term, with actual action, like a lawsuit or
| something then eventually this will will also be encompassed by
| the term "Open Source" as people know it like it or not.
|
| Yes I know this term is currently defined explicitly by OSI, no I
| don't think language prescriptivism wins out regardless how hard
| they try with it, and since I haven't seen any of the hundreds of
| quasi Open Source, but not really, companies get dragged to court
| over usage of the term, this is all toothless complaining in my
| view.
|
| As to their actual point, I might actually agree with them if it
| were only the weights being shared. In most cases the
| configuration is also shared which allows popular frameworks to
| instantiate the model and then execute it for either inference or
| further training making the release fully suitable for
| modification and rerelease. I don't need the exact implementation
| of FlashAttention they used if I can load the model into
| Huggingface and use theirs, or mine or whatever.
|
| Edit: This obviously doesn't apply to the models who have
| restrictions placed on usage just in case people think I mean
| every instance of sharing a model. Those are obviously restricted
| use and I agree it muddies the term.
| TZubiri wrote:
| Agreed, output weighs are target code, and no one would argue the
| contrary. Companies pretending to publish source code is nothing
| new.
|
| Stallman defines source code as "the preferred way in which
| developers modify the program"
|
| I wrote for wikipedia once that
|
| "Stallman's definition thus contemplates JavaScript and HTML's
| source-target ambivalence, as well as contemplating possible
| future forms of software production, like visual programming
| languages, or datasets in Machine Learning."
|
| So the datasets could be a form or source code, but the most
| appropriate source code would be the code that crawls or
| downloads the dataset and modifies it.
|
| Clear as water
| dahart wrote:
| > Some people have the perspective that if a license isn't open
| source, it's proprietary. I think it's more nuanced than that and
| believe there are three more license types worth naming: non-
| commercial NDA, non-commercial public, and ethical.
|
| It's very useful to remember the U.S. government definition of
| commercial software: it is software that "Has been sold, leased,
| _or licensed_ to the general public" [1]
|
| This means that a "non-commercial license" is a bit of an
| oxymoron to a lot of people. Their definition of commercial
| includes all software with a license, and does not depend on
| whether the software costs money. (Perhaps not entirely unlike
| how FSF does not define "free software" based on whether it costs
| money.)
|
| [1] https://www.acquisition.gov/far/2.101
| Traubenfuchs wrote:
| [dead]
| ndriscoll wrote:
| The complexity described seems to be resting on the unestablished
| idea that weights are copyrightable in the first place. If
| they're not, then presumably "available weights", "ethical
| weights", and "open weights" are all the same: open weights.
| Either your weights are under NDA and presumably considered to be
| a trade secret, or they are public, and the words in your
| "license" mean absolutely nothing? That seems like a rather
| important point to bring up when discussing the licensing
| landscape for weights...
| feoren wrote:
| Some thought experiments:
|
| What happens if we train a neural network on a single,
| copyrighted work? Say it has one input node (or even zero, if
| you like), and regardless of this input, its output is always
| exactly the copyrighted work it was trained on. What do its
| weights represent? Clearly, its weights represent a direct
| encoding of the original work. Those weights _are_
| copyrightable, but not by the person who trained the neural
| network -- the copyright is held by the owner of the original
| work.
|
| What if we train the neural network on just two copyrighted
| works? If its one input node is 0, it outputs the first, and if
| it's 1, it outputs the 2nd. Almost certainly, its weights are a
| complicated, tangled mix encoding both, like a compression
| algorithm that completely rearranged its input. Who owns the
| copyright to those weights? To whatever extent the weights can
| be "factored out" into a set representing the first work and a
| set representing the second, clearly the copyright holder of
| the first work holds the copyright on the first "factored set",
| and the 2nd on the 2nd. It seems obvious that we must be able
| to do this "factoring out" _somehow_ (even if the topology of
| the factored networks is different), because we know both works
| are exactly represented by the weights, and the neural network
| itself can use this information to reconstruct them both, so
| they 're _in there_ ... somewhere. So is there a sort of
| "joint copyright" on the combined weights, where nobody is
| really allowed to do anything with it without approval of the
| other? Regardless, it's still clear that whoever trained the
| neural network has no claim on any copyright.
|
| Where is the breaking point extending this from 2 works to a
| billion? People make arguments like "drawing a car from memory
| isn't infringing on copyright design of that car", which ...
| are you _sure_? Reproducing a piece of music from memory (and
| selling it) _is_ usually copyright infringement. You 're
| allowed to _learn_ a Taylor Swift song as part of your musical
| training, but you 're not usually allowed to then play it back
| from memory and sell that recording (I'm not sure I morally
| agree with this treatment of covers, nor if it's globally
| applicable). So the argument that "surely neural networks are
| allowed to _learn_ from copyrighted works " misses the point:
| they can learn all they want, but as soon as they reproduce
| verbatim (or close enough) a copyrighted work, they're
| infringing. And if they're representing a complete copy of the
| work within their weights (which they obviously are if they can
| reproduce it), then the original copyright holder has a claim
| on those weights. And never in this process has the trainer of
| the NN acquired any copyright to anything. The real trainer is
| a bunch of GPUs, after all.
|
| If the neural network _cannot_ reproduce any of the copyrighted
| works verbatim, then we 're getting closer to "fair use"
| territory. Yes, it's permissible to write a summary of a
| copyrighted work. That is so lossy as to not "compete" with the
| original work in any meaningful way. If it could be
| demonstrated that neural networks do not encode completed works
| (no matter how hard the factorization would be), then one could
| make this argument. Unfortunately, the evidence is that LLMs
| are more than happy to completely regurgitate copyrighted works
| verbatim. It seems to me the copyright holder of the original
| work therefore must hold a share of the claim on the weights.
| Still, the GPUs that trained the network do not magically
| acquire copyright over anything.
|
| I wonder if the real answer is that the weights are
| copyrighted, and that copyright is held jointly by hundreds of
| millions of people, and nobody can do anything with those
| weights without the approval of all the others. I'm not saying
| I like that universe, but I am saying it's the most internally
| consistent answer I can think of, and seems to follow from the
| above argument.
| feoren wrote:
| In fact, the "factoring out" process shouldn't even be that
| hard: find the input vector that forces the ANN to output the
| copyrighted work verbatim. There should be some simple method
| of "baking in" the first step of the feedforward algorithm,
| applying that vector to the first layer of weights, and then
| considering the input layer as the first hidden layer of a
| network with 0 input nodes. It is now equivalent to a neural
| network that can only ever output a single copyrighted work,
| and therefore its weights exactly encode (bloatedly!) that
| work. The owner of the work holds copyright on those weights.
| Importantly, if I'm thinking about this right, the weights of
| this derived network are exactly the same as the original
| except in the first layer.
|
| On the other hand, we need the original input vector for this
| to work, and one could argue that the network weights are
| simply the algorithm for decoding the input vector into the
| copyrighted work. So the originator holds copyright on the
| _input vector_ , not the weights. Does it matter if the input
| vector has smaller information content than the original
| work? Clearly this argument relies on the input vector being
| the "actual encoding", and therefore must have at least as
| much information. If the input vector is an embedding of
| "please show me the latest Tom Clancy novel in full", this
| argument breaks down.
|
| Okay, this is hard.
| spullara wrote:
| This has been my position from the beginning. It is very hard
| for me to imagine that weights can be copyrighted at all.
| IANAL.
| nerdponx wrote:
| Weights are equivalent to compiled object code IMO. All else
| follows from there.
| feoren wrote:
| Compiled object code of a bunch of code _you didn 't write_.
| I don't know why programmers are so eager to forget that
| copyright is not at all about what something _is_ , and all
| about where it _came from_. It 'd be hard to assert that you
| hold copyright over object code compiled from code you didn't
| write!
| cubefox wrote:
| Or is the fact that compiled code enjoys copyright
| protection, even though it is not human generated, evidence
| that being generated by a human is not overly important for
| copyright protection?
| feoren wrote:
| An mp3 encoding of a wav file of a copyrighted song is
| still copyrighted, despite those exact bits never having
| existed before, and being created entirely by a computer.
| xigency wrote:
| > programmers are so eager to forget that copyright is not
| at all about what something is, and all about where it came
| from
|
| See "What color are your bits?":
| https://ansuz.sooke.bc.ca/entry/23
|
| >> And very much of intellectual property law comes down to
| rules regarding intangible attributes of bits - Who created
| the bits? Where did they come from? Where are they going?
| Are they copies of other bits?
| Animats wrote:
| > The complexity described seems to be resting on the
| unestablished idea that weights are copyrightable in the first
| place.
|
| Yes. Weights probably aren't copyrightable in the US. See Feist
| vs. Rural Telephone, in which the Supreme Court ruled that
| telephone directories are not copyrightable. The copyright
| clause in the Constitution ("To promote the Progress of Science
| and useful Arts, by securing for limited Times to Authors and
| Inventors the exclusive Right to their respective Writings and
| Discoveries.") is understood to require human authorship. The
| US does not have database copyright, or "sweat of the brow"
| copyright. That it was expensive to produce some collection of
| data does not make it copyrightable.
|
| Outputs from LLMs, machine generated art, and machine generated
| music probably are not copyrightable either. US Copyright
| Office: "Based on the Office's understanding of the generative
| AI technologies currently available, users do not exercise
| ultimate creative control over how such systems interpret
| prompts and generate material. Instead, these prompts function
| more like instructions to a commissioned artist."[1]
|
| [1] https://www.reuters.com/world/us/us-copyright-office-says-
| so...
| deepsun wrote:
| > Outputs from LLMs, machine generated art, and machine
| generated music probably are not copyrightable either.
|
| Let me put a straw man, and try to find a middle point, when
| the copyright argument stops being applicable:
|
| 1. A painting was done by an artist.
|
| 2. On a computer.
|
| 3. With a help from an image processor software.
|
| 4. Using some advanced filters, like super-resolution, that
| utilize computer vision techniques. Like neural networks.
|
| Many smartphones already automatically process your* photos
| with some advanced CV algorithms. That can be called "machine
| generated art".
|
| I'd personally prefer to stop saying "neural network did X",
| same way as we don't say "a bulldozer built a road, a crane
| built a house".
| comfypotato wrote:
| The distinction, defining your straw man, is simply that
| the image itself is generated by the "commissioned artist"
| that is the AI.
|
| Even non-generative-AI inside Photoshop only mutates
| images. Generative AI is the _source_ of images.
| drdeca wrote:
| Is it though? What of the e.g. choice of prompt, guidance
| scale, maybe a specification of a pose, etc.?
|
| Or, is the distinction you are making based on there
| being an image before the model is used?
| radarsat1 wrote:
| > Weights probably aren't copyrightable in the US. ... is
| understood to require human authorship.
|
| Are you arguing here that because the weights come from an
| optimization program, they are not "human authored"? If so I
| find that to be a strange assertion. If I'm working every day
| on my model and training algorithm to ensure it produces the
| best weights possible to solve my problem, I would be very
| surprised for someone to tell me I have no ownership over
| those weights because they are _generated_ from a program I
| wrote and data that I own.
| mirekrusin wrote:
| Assuming that weights are not copyrightable, how much
| restrictions can you put on output through API from those
| networks/weights?
|
| Ie. if ClosedAI says you can't use output of their API to
| train competitive models - is that enforceable or not?
| photonerd wrote:
| That would be down to contract/terms of service. You'd be
| in breach of that, not copyright
| mirekrusin wrote:
| But is it enforceable? Companies can put in contracts all
| kind of nonsense, it doesn't mean all of it is
| unconditionally enforceable, right?
|
| Ie. if somebody creates company that sells milkshakes and
| they say you can't use them to feed employees of
| competing milkshakes companies - it wouldn't fly, would
| it?
| photonerd wrote:
| Would strongly depend on the contract. Probably wouldn't
| fly in a post sake terms of service agreement, but you'd
| likely be in breach of a normal contract yeah.
| dllthomas wrote:
| > Outputs from LLMs, machine generated art, and machine
| generated music probably are not copyrightable either.
|
| I don't have a strong sense of whether this is reasonable (I
| see arguments both ways) but I do think it's pretty strongly
| at odds with how we treat photographs. There are a bunch of
| photos on my phone where I unquestionably own the copyright,
| despite putting in much less creativity than I did for some
| AI images I've generated.
|
| I don't think it's clear how to resolve this, but I do think
| that _if_ we are going to protect photos and not prompted AI
| images, the distinction needs to turn on something other than
| whether "sufficient creativity" was applied to the input of
| the mechanical system.
|
| Edited to add: It's probably also worth calling out that the
| question of whether we protect the work produced by a
| person's use of mechanical system is a separate one from
| whether we protect the work of others when it is (in various
| ways, to various degrees, with various likelihoods)
| reproduced by use of those mechanical systems.
| makeitdouble wrote:
| On photography, the argument was condensed into "who pushed
| the button". We saw it with the monkey auto-portrait
| copyright fight where copyright was not granted to the
| photographer, and other nature photography using photo
| traps where the copyright stuck with the human basically
| because they were the last operator of the camera.
|
| The interesting part is, those controversial case are
| pretty recent when the art of photography is century(ies?)
| old now. I wouldn't expect super clear guidelines regarding
| AI art before a few decades of weird cases fought tooth and
| nails in court.
| tokai wrote:
| Eh? The copyright was the photographers and not the
| monkeys.
|
| https://petapixel.com/2018/04/24/photographer-wins-
| monkey-se...
| shagie wrote:
| * * *
| DannyBee wrote:
| This is mostly right - It depends on what the weights
| represent and how they were generated so I would not go as
| far as the initial claim.
|
| A collection of numbers is copyrightable if it's the encoded
| result of a creative process. Just because it's represented
| as a bunch of numbers does not make it non copyrightable.
| That's why it says " original works of authorship fixed in
| any tangible medium of expression, now known or later
| developed, from which they can be perceived, reproduced, or
| otherwise communicated, either directly or with the aid of a
| machine or device. "
|
| You can't just classify the weights as facts simply because
| they are numbers. If they are creatively made by a human they
| would be copyrightable. Mechanically computed from random
| numbers, no. Somewhere in the middle? Harder
| kevin42 wrote:
| I'm not a lawyer, but it seems like you stood up a straw
| man there.
|
| >Just because it's represented as a bunch of numbers does
| not make it non copyrightable.
|
| Can you give an example of where the bunch of numbers is
| copyrightable when it's not just a numeric encoding of
| something that was already copyrightable? Taking music and
| encoding it as a wav file is not a creative work, but it's
| a representation of a copyrighted work.
|
| Maybe you could create a long list of numbers and call it
| an artistic impression, but that's clearly not what AI
| weights are. I'm interested to hear an example of your
| copyrightable numbers.
| mlyle wrote:
| The key factor of Feist v Rural is whether there was any
| original or creative process in the way the facts were
| arranged.
|
| Here, there's a whole lot of creative decisions in
| labelling and guiding of training that produces the
| weights, so it's reasonable to think it might be
| copyrightable.
|
| That is, the numbers are a whole lot more original than
| the issuance of phone numbers or part numbers.
| Retric wrote:
| The requirement for expertise doesn't necessarily imply
| that that setting up perimeters for training AI is
| necessarily copyrightable. A normal brick wall for
| example needs skills to create but doesn't qualify as the
| goal is not creative. If so the mechanical output of a
| process that doesn't qualify for copyright is not going
| to qualify.
|
| Labeling training data may qualify for copyright, but if
| the underlying training data doesn't taint the output as
| a derivative work then labeling isn't going to qualify by
| itself.
|
| Thus without some new and very generous interpretation AI
| companies are at best not going to benefit from copyright
| and at worst may be forced to create all training data in
| house. My suspicion is this generation of AI companies
| are in a very difficult situation.
| mlyle wrote:
| > but if the underlying training data doesn't taint the
| output as a derivative work then labeling isn't going to
| qualify by itself.
|
| It depends. If each individual training item has a small
| impact on the output coefficients, then perhaps it's not
| a derivative work of them. But if there's a large
| creative process in determining model training procedure,
| deciding labelling strategies, and applying those--
| perhaps those numbers are strongly derived from _those_
| things.
| Retric wrote:
| That sounds like wishful thinking, individual training
| items have significant impact on the result.
|
| Anyway, suppose you're building an AI to walk, there's
| nothing creative about selecting 9.8m/s/s for gravity
| that's simply the ideal value to achieve a desired goal.
| Labeling an elephant as "Elephant" rather than "coat
| hanger" is similarly a functional choice.
|
| Just because a person is holding a camera and taking a
| photo doesn't mean the result is copyrightable.
| mlyle wrote:
| > Anyway, suppose you're building an AI to walk, there's
| nothing creative about selecting 9.8m/s/s for gravity
| that's simply the ideal value to achieve a desired goal.
|
| Suppose you're not building a strawman, but instead
| building an AI to be an LLM. The exact sequence of what
| you choose to do for instruction tuning, and the metrics
| and labels that you choose, the prompt/response pairs you
| write, and the loss functions you employ are quite
| creative. They greatly affect the coefficients and are
| not simple mechanical steps and are the result of a large
| amount of creative choice.
|
| We are nowhere near a point where they are an uncreative,
| mechanical recipe to follow.
|
| > Just because a person is holding a camera and taking a
| photo doesn't mean the result is copyrightable.
|
| No, but in the overwhelming majority of circumstances it
| is. What it depends upon is whether the person holding
| the camera is making a significant, original creative
| choice.
|
| I am not sure what courts will decide, but I am certain
| that there is more creativity and originality employed
| than you are giving OpenAI et al. credit for.
| dragonwriter wrote:
| > Here, there's a whole lot of creative decisions in
| labelling and guiding of training that produces the
| weights
|
| Often, labelling is part of large public datasets that
| are chosen for use for that exact reason, and/or is
| otherwise not the work of the party claiming copyright in
| the model.
| [deleted]
| mft_ wrote:
| IANAL, but I'd wonder whether 'creativity' is really
| present in labelling - and indeed, mightn't it be the
| last thing you want? I'd argue labelling should be
| strictly factual and reproducible, and ideally following
| a logical structure... maybe akin to how addresses of
| buildings might appear in a phone directory...
|
| (Agree that the skill in knowing how to code and guide
| the training of a model is probably very different
| though. It's not just access to compute time that
| separates me from OpenAI :) )
| DannyBee wrote:
| "Can you give an example of where the bunch of numbers is
| copyrightable when it's not just a numeric encoding of
| something that was already copyrightable?"
|
| Sure, there are "poems" that consist of just a groups of
| numbers that are copyrighted. They are not encodings,
| it's just a string of numbers. It's indistinguishable
| from a bunch of numbers. This is just one example, there
| are lots.
|
| They are enforceable to the degree it's creative, and to
| the degree the infringing use is also creative.
|
| So you would not be able to sue me for using those
| numbers in a math equation. You would be able to sue me
| for reproducing your poem in a book of poems :)
|
| As feist says, the creativity required for copyright is
| quite minimal. But it's still only as protectable as it
| is creative.
|
| Look - AI is not the first thing to have this "issue".
| The answer remains the same as it always was - it's
| mostly about the process not the output.
|
| The output mostly matters is if the output is not
| intended to be creative (or it's de minimis or ...).
|
| Copyright as it currently exists is weird.
|
| Like if you go to the copyright office and try to
| register your ssh public key and say "this was generated
| by ssh-keygen i had nothing to do with it" you _may_ get
| a different result than if you said "this is my new
| visually stunning masterpiece, my ssh public key, which
| was generated with computer help but I used 37 precisely
| timed keyboard smashes to do it. Prints are available
| from my gallery for $500"
| mlyle wrote:
| I fully agree with what you say, with one bit of nuance
| to point out:
|
| > Like if you go to the copyright office and try to
| register your ssh public key and say "this was generated
| by ssh-keygen i had nothing to do with it" you may get a
| different result than if you said "this is my new
| visually stunning masterpiece, my ssh public key, which
| was generated with computer help but I used 37 precisely
| timed keyboard smashes to do it. Prints are available
| from my gallery for $500"
|
| The important thing, of course, isn't whether the
| copyright office denies to register your copyright, but
| instead what courts will ultimately do when you attempt
| to enforce your copyright.
|
| We know the current administrative algorithms used by the
| copyright offices. We have less clarity on what courts
| will ultimately do.
| dxbydt wrote:
| > Mechanically computed from random numbers, no
|
| Even random numbers are copyrightable.
|
| Below is an implementation of Marsaglia's invention, from
| p348, courtesy infamous NR[1]. Its a MWC (multiply with
| carry) random number generator, with two parameters,
| variable a and base b=2^32. --- For a, "The values below
| are recommended with no particular ordering." ID a B1
| 4294957665 B2 4294963023 B3 4162943475 B4 3947008974 B5
| 3874257210 B6 2936881968 B7 2811536238 B8 2654432763 B9
| 1640531364 --- as we all now know, the whole thing is
| copyrighted - you can't redistribute that code and can't
| use those specific numbers to generate random numbers
| without purchasing a license, which only allows you to use
| it once in your personal machine; that's why GSL[2]. The
| pseudorandom numbers you would get from MWC if you use
| above numbers are also copyrighted since they are work-
| product.
|
| [1]http://numerical.recipes/book/book.html
| [2]https://www.gnu.org/software/gsl/design/gsl-design.html
| saynay wrote:
| I would say that is uncertain. Model weights are always
| going to effectively be a huge collection of statistics
| about the training corpus. Unless you are envisioning
| artisanal, hand-crafted, free-range model weights where a
| person used a non-mathematical method to purposely and
| creatively choose each one?
| WanderPanda wrote:
| Of course weights are copyrightable. Otherwise nothing is
| copyrightable
| enlightens wrote:
| Recipes, for example, are not copyrightable in the US.
| Neither are some of the concepts behind creating a fillable
| form. It's not an all-or-nothing system.
|
| https://www.copyright.gov/circs/circ33.pdf
| throwaway98721 wrote:
| Why is it unestablished? Is a document not copyrightable based
| on its contents? Weights are just a different kind of a
| document.
| Conscat wrote:
| No, a document's contents aren't inherently copyrightable.
| They have to be a creative work or a method of production,
| and part of that basically means it has to be human generated
| content (as opposed to computer or animal generated).
|
| AI weights might be considered a method of production, but
| that isn't clear yet.
| throwaway98721 wrote:
| [flagged]
| ketzu wrote:
| > Was there no work put into their creation by someone?
|
| Putting work into something is not a sufficient cirteria
| for copyright.
|
| > All of it is just a stream of bytes that the computer
| can interpret somehow
|
| This is also not a sufficient or at all relevant cirteria
| for assigning copyright.
|
| Also, in the sense you presented, those files are not
| fundamentally different from random noise. Which is not a
| particularly useful reduction for this exercise.
| dragonwriter wrote:
| > Was there no work put into their creation by someone?
|
| This is the "sweat of the brow" theory of
| copyrightability, which courts have rejected (for good
| reason based on the statute.)
|
| "Someone did work to enable this thing to exist" is not
| sufficient to make a thing copyright-protected.
|
| > There's no fundamental difference between an image,
| code, or weights.
|
| And neither images, code, nor weights that are
| mechanically produced with no creative input by a
| particular author are subject to copyright in their own
| right (depending on their relation to the source material
| on which the mechanical process rests, they may be
| covered by the copyright on the source material.)
|
| The _best_ argument for weights being copyrightable (and
| it probably applies better to some models than others) is
| that the assembly of source material is a creative work
| subject to a compilers copyright, and that the model
| weights themselves are just a mechanical translation of
| that compilation subject to its copyright.
| jerf wrote:
| Copyright is not for "documents", it is for works that have
| creativity in them. The legal bar for that level of
| creativity is low, so low that it is easy to come away
| thinking that anything that can be cast as a "document" must
| be copyrightable, but the bar is in fact not zero.
|
| In particular, taking other documents and shoving them
| through a process that generates a lot of other numbers with
| no human or creative interaction is definitely something I'd
| be concerned the courts would judge as not sufficiently
| creative to be copyrightable. The process itself would
| certainly consist of copyrightable code, but the output
| doesn't necessarily. This would be somewhat similar to the
| observation that there is no copyright to be had in a big
| table of files and their MD5 hashes (or other hashes), such
| as a Linux distro might use for integrity checking. Lots of
| copyright in the original file contents, copyright available
| on the process for producing these tables, but the _tables
| themselves_ would likely be ruled not itself copyrightable as
| there is no creativity in that output.
|
| Note this also has absolutely nothing to do with the question
| of whether AI output is copyrightable, this is about the huge
| table of numbers that make up the neural net weights being
| copyrightable. (Though it would be sort of an interesting
| question for the legal system to grapple with as to how a
| non-copyrightable set of numbers could then produce something
| copyrightable. Call it a philosophical variation on the
| "copyright washing" argument; can copyright spring from a
| non-copyrightable source other than a human brain, thus
| somehow "flowing uphill"? Would a human brain be
| copyrightable? Stay tuned for those questions, I guess, or if
| not you, your grandchildren.)
|
| Per your other comments, "work" is not the bar, "creativity"
| is. "Size" is not the bar either. Merely being a much larger
| table of numbers than a list of hashes or a phone book is not
| the question. No human is in that table of numbers creatively
| saying "no, wait, this neural weight should be -1.5 instead
| of 2.0 to produce this creative effect". No human is even
| _capable_ of working in the medium of neural net weights in a
| creative manner.
|
| If you want to go the "novel legal theory" route, you could
| play with claiming creativity in the selection of input
| material and claim the resulting neural weights has a
| copyright in compilation:
| https://en.wikipedia.org/wiki/Copyright_in_compilation That's
| a long way from a slam dunk though. Way out on a legal limb
| there. It isn't entirely clear to me what exact rights would
| result from such a claim either. It would be a landmark
| copyright court case for sure.
| AnimalMuppet wrote:
| IANAL, but I suspect that the "novel legal theory" in your
| last paragraph would fail. It might succeed if you gave GPT
| a hand-curated list of materials; hoovering up the entire
| internet is not that.
| dragonwriter wrote:
| Weights are the output of a mechanical process over the
| training set with no element of human authorship, just as
| much the output a model produces with a prompt is, which the
| Copyright Office has already declared outside of copyright.
|
| > Is a document not copyrightable based on its contents?
|
| Creative process is the bigger issue.
|
| > Weights are just a different kind of a document.
|
| And who sits down and writes this document of weights?
| raincole wrote:
| > Is a document not copyrightable based on its contents?
|
| Yes, exactly. It's copyright 101.
|
| For example, if you write a random number generator, and
| print 10000 randon numbers in a document, it's not
| copyrightable.
|
| Even if you invented a specific random number generation
| algorithm, the document is still not copyrightable. Your code
| is copyrightable.
|
| Again it's just copyright 101. If any of above surprises you,
| maybe you should read a few copyright case studies.
| [deleted]
| xg15 wrote:
| Furthermore, _if_ weights are copyrightable, wouldn 't this
| make the issue of training data licenses even more urgent?
|
| IANAL, but if weights are IP, wouldn't they constitute a
| "derived work" of the training data?
| golemotron wrote:
| In a sane legal system a new copyright law would be passed to
| clarify all of this. In ours, the poor copyright office needs
| to make things up on the fly.
|
| Their recent decision that implies that anything that AI is
| used to produce is non-copyrightable is silly, sad, and not
| sustainable.
| YetAnotherNick wrote:
| No, weights are not just data fed but also the training
| process itself. I think the whole argument hinges on how much
| human thought and action is needed in training the model.
|
| On the other end of the spectrum, AI generated content
| couldn't be copyrighted if there is no human involvement. If
| someone asks GPT to write 1000 poems, it couldn't be
| copyrighted.
| realusername wrote:
| That's also my understanding, either the weights are
| copyrightable and then all the models need explicit
| agreements for any work they include in it because models
| become derivatives or they are not copyrightable being just
| machine data (the most likely scenario in my opinion), they
| can't have it both ways.
| AnimalMuppet wrote:
| "Transformative use".
|
| The inputs could be copyrighted _and_ the weights could be
| copyrighted _if_ creating the weights from the inputs is
| (legally) regarded as a transformative use. And I think it
| could reasonably be considered to be transformative - the
| weights don 't look anything like the input data.
|
| Disclaimer: IANAL. So far as I know, no court has ruled on
| whether this qualifies as a transformative use. I take no
| position on how the courts will actually rule. I merely say
| that they _could_ regard this as transformative use. (But
| see jerf 's "creativity" argument for another hurdle that
| weights must pass to be copyrightable.)
| shagie wrote:
| Transformative use doesn't necessarily mean
| copyrightable.
|
| Google's thumbnails are a purely mathematical
| transformation on images (no copyright themselves), and
| yet are considered a transformative use.
|
| I believe that trained models are similarly a purely
| mathematical transformation of {data}, but is
| transformative in what that _can_ be used for going
| forward.
|
| "Can" bearing a lot of weight in that sentence.
|
| It's how the human, with agency, uses the model that may
| be a derivative or copyright infringing use - not the
| model itself nor necessarily the output.
|
| The output of a generative AI _may_ be similar enough to
| an existing work that it is derivative of that work. It
| is possible to construct a prompt that infringes on an
| existing work _even if that work wasn 't part of the
| training data_.
|
| For that case, consider you drew a picture. That picture
| that you just drew isn't part of any training data. I
| could presumably look at it and describe it with
| sufficient detail that something similar enough would be
| generated... and that may be considered a derivative
| work. The same test could be applied to me describing it
| to someone on Fiverr with the same outcome.
|
| If I were to publish that work by the generative AI or
| Fiverr - who would be infringing on copyright? me? or the
| black box that may be AI or Fiverr that created a picture
| based on my prompts?
| numpad0 wrote:
| Another way to look at it is if a thing reproduced a data
| subjectively resembling originals and then you used it
| anyhow, then its non-transformative use, and methods used
| is just extra details.
| phantom784 wrote:
| I think there could be an argument that it's copyrightable
| but not a derivative work.
|
| If I read a few books about a subject as research, and then
| I write an article about the subject, it's my own
| copyright. The fact that I did research doesn't make it
| derivative of those books (correct me if I'm wrong, IANAL).
|
| Perhaps a model created from copyrighted material be
| treated in the same way?
| floomk wrote:
| That's because you are human and have rights that a
| computer program doesn't
| adamc wrote:
| A fertile subject for sf stories.
| bloak wrote:
| > If I read a few books about a subject as research, and
| then I write an article about the subject, it's my own
| copyright.
|
| Yes, because in that case you'd be the "author" doing
| "creative work".
|
| > Perhaps a model created from copyrighted material be
| treated in the same way?
|
| Who would be the author doing creative work in this case?
| The people who decided what training material to use?
| Perhaps, but it seems a stretch for the people who
| selected the training material to be authors but not the
| people who created the training material.
| slaymaker1907 wrote:
| The difference is you are person and have many more
| rights than a machine.
| OkayPhysicist wrote:
| There is also a (IMO less likely, but still conceivable)
| scenario where weights ARE copyrightable, but represent
| fair use of the training data on grounds of being
| "sufficiently transformative".
| floomk wrote:
| Sadly this seems to be the most likely considering how
| the US is ran
| 8note wrote:
| I consider that one super likely, but then using the
| model to make competing works with one the artists in
| their own style is a non-fair use derivative work
| Filligree wrote:
| Style explicitly isn't copyrightable. It'll need to be
| for some other reason.
| OkayPhysicist wrote:
| Your case wouldn't be about style, it would be about
| specific elements that you posit were memorized and
| regurgitated by the model. The fact that you're creating
| art in the same style/medium as the author is what
| negates the "sufficiently transformative" fair use
| defense.
|
| Basically, that world ignores the AI model completely. If
| your resulting work wouldn't be fair use if you directly
| were working with something from the training set, it
| wouldn't be fair use if you fed it through an AI model
| first.
| shaky-carrousel wrote:
| Also known as "having your cake and eating it".
| jrm4 wrote:
| Exactly. A _lot_ of the difficulty here is how they skip is the
| hugely important issue:
|
| An entirely reasonable, if not fully tested, statement is the
| following:
|
| Every single one of these AI weight things _itself_ is a result
| of unencumbered, massive, law-breaking, right-violating
| copyright infringement -- accordingly, it 's _extremely_
| difficult to say anything morally justifiable or authoritative
| about anyone elses "rights" downstream, and to try to inject
| the word "ethical" makes the whole thing even more ridiculous.
| visarga wrote:
| > is a result of unencumbered, massive, law-breaking, right-
| violating copyright infringement
|
| Why? Copyright covers expression not information, AIs can
| learn information from any source regardless of copyright.
| They should just not regurgitate copyrighted content, that's
| all. And much of what organic content is online is common
| knowledge, thus can't be copyright-controlled.
| pfdietz wrote:
| Copyright is for things that are the result of human
| creativity. If the weights come from running an algorithm
| on a training set (that one does not have a copyright to)
| then how can the weights then be copyrightable? They might
| be a derivative work, but that just means they infringe
| copyright, not that they are copyrightable themselves.
| shagie wrote:
| Note that the requirements for copyright are not
| consistent between nations.
|
| The US has the "threshold of originality" as its
| principle. Under that doctrine, it requires some _human_
| (and this has been emphasized many times over the years)
| originality in order for something to be copyrighted. It
| 's a low bar for how original it needs to be, but it must
| be human (monkeys taking selfies are not human).
|
| https://en.wikipedia.org/wiki/Threshold_of_originality
|
| In England, the doctrine is "sweat of the brow" instead.
|
| https://en.wikipedia.org/wiki/Sweat_of_the_brow
|
| > Under a "sweat of the brow" doctrine, the creator of a
| work, even if it is completely unoriginal, is entitled to
| have that effort and expense protected; no one else may
| use such a work without permission, but must instead
| recreate the work by independent research or effort.
|
| The definitive case for this in the US that set the two
| apart is Feist Publications, Inc., v. Rural Telephone
| Service Co. ( https://en.wikipedia.org/wiki/Feist_Publica
| tions,_Inc.,_v._R.... ) where it was deemed that a
| telephone directory is not copyrightable in the US as
| there is no originality in it... but under the sweat of
| the brow doctrine it would have been.
|
| So the "[c]opyright is for things that are the result of
| human creativity" gets an "it depends" and it would be
| curious to see if companies that are firmly in the
| "models are valuable" camp go to the UK for what I
| believe would be a more favorable copyright protection.
|
| ... _However_ there are other IP laws around trade
| secrets that may be better for it in the US (I 'm not as
| familiar in that domain - I would be curious to find
| out).
| stale2002 wrote:
| The answer would be if the weights are transformative
| enough, and the copyright would come from the person who
| decided what images to include in the training set.
|
| The act of choosing to place images in a certain
| arrangement, such as a collage, can be copyrightable. The
| same could be said for the "act" of choosing what images
| to include in a training set and which parameters to use
| to train the model.
| kevin42 wrote:
| Does that mean if someone copies a phone book but leaves
| out some numbers and adds some other numbers then it's a
| creative work?
| stale2002 wrote:
| It would depend on how transformative the work is.
|
| There is in fact a whole art form where people cut out
| words from different newspapers and books, for example,
| and re-arrange those words to form new and interesting
| art.
|
| So there are ways in which such a work would be a
| creative work, and ways it which it would not, and it
| would depend on the particular instance and example.
| saynay wrote:
| The legal system is not like a computer program. The line
| between what is "creative" and what is not concrete, but
| is instead up to the interpretation of the judge who
| rules on it.
|
| So your phonebook modifications may or may not be
| considered "creative" depending on the judge and your
| ability to convince them. The more your modify it, the
| more likely you are to convince a judge it is a creative
| work, though.
| wheelie_boy wrote:
| It seems very difficult to ensure that a model will never
| output any of the copyrighted content that it was trained
| on. I can only think of three ways, but perhaps there are
| others
|
| 1. Evaluate every output from the model to ensure that none
| of the outputs are copyrighted
|
| 2. Evaluate every input to a model to ensure that the
| inputs are either not copyrighted or properly licensed
|
| 3. Change the definition of copyright so that ML models can
| do whatever they want
|
| Nobody is doing #1, because that makes the business models
| not work. Established brands (like Adobe) are doing #2. I
| get the feeling that there are a lot of ML startups that
| are hoping that #3 will happen, but it seems unlikely
| og_kalu wrote:
| Ensuring a model never outputs copyrighted content is
| unimportant and tangential. It's irrelevant. You don't
| look for a way to make humans output no copyrighted
| content, you address each time they do case by case.
|
| A model training being rendered fair use doesn't mean any
| of its output can be used for whatever regardless.
| wheelie_boy wrote:
| > you address each time they do case by case.
|
| That's what I listed as #1 - evaluate each individual
| output of the model to see if it violates copyright.
| wilde wrote:
| Tell that to some illegal numbers:
| https://en.wikipedia.org/wiki/Illegal_number
| slaymaker1907 wrote:
| My issue with this take is that machines are not people. We
| only have lax rules for humans precisely because they are
| humans, not on the basis that they can learn. Copyrighted
| works are produced for people and given how human learning
| works, applying the derivative works rule to humans would
| be completely impractical and destroy the point of works
| with copyright. The same cannot be said for AI companies
| treating everything on the internet as fair use for
| training AI.
| version_five wrote:
| A lot of people are just upset because their local
| equilibrium has been disrupted and they think that means
| they lost a natural right.
|
| "You wouldn't look at a car and then remember what that
| looked like when someone asks you to draw another"
| jrm4 wrote:
| These are not bad arguments, but I don't think they're
| conclusive. I am a lawyer, and I could absolutely see
| this going the other way. "You can't make these machine
| things without literally feeding this copyrighted
| information into them, therefore they do contain a copy.
| You can see this by when they reproduce, e.g. the "getty
| images" deal."
|
| *this is not legal advice, dangit commenter person below
| m4nu3l wrote:
| > you can't make these machine things without literally
| feeding this copyrighted information into them, therefore
| they do contain a copy.
|
| They don't necessarily do. Think about that. You can take
| some copyrighted material and transform the information
| contained in it (for instance a fictional book). You can
| then write a summary. The summary contains information
| that was present in the original but it has been
| transformed and hence it's not a copy. The ML model
| contains information that has been generalized by some
| degree. So it's just a grey area IMO.
| saynay wrote:
| More over, you are clearly not in violation of copyright
| if you are talking about statistics about the material.
| In your example, printing out a "there were 7000
| instances of the word 'the'" is certainly not a
| violation. A ML model is just a huge pile of these
| statistics.
|
| However, saying "the first word of the book is 'The'"
| would not be a violation, while repeating that for every
| word in the book, as a whole, would be one.
| version_five wrote:
| I agree with you but I think it's important to have some
| nuance. Imagine I build a statistical model for 10-word
| sequences (10-grams) and then I trained it on a single
| book. I probably could pick some starting words and get
| most of the book back from the "statistics" I compiled.
| If I trained the same model on a giant dataset, the one
| book would just contribute to the stats.
|
| All that to say, the models have potential to memorize,
| but they don't, and if they do it's an undesirable
| failure mode, not some deliberate copying.
| jrm4 wrote:
| I like this argument a lot; but again -- how does this
| play out in the real world? It's pretty easy to refute
| what will happen in real life. Think, e.g Batman. I could
| write a very new and original "Batman" comic that doesn't
| strongly resemble anything -- movie, toy, comic, whatever
| -- that exists, but would be recognizable to fans.
|
| Once it starts doing well, will DC come after me? You
| bet.
| ke88y wrote:
| These models can definitely be used to intentionally
| store and recall content that is copyrighted in a way
| that's not subject to fair use. (eg: trivially, I can
| very easily train a large model that has a small
| subnetwork which encodes a compressed or even lossless
| copy of a picture, and if I were to intentionally train a
| model is that way then this would be no less a copyright
| violation than distributing a JPEG of the same image
| embedded in some large binary).
|
| But also, an unintentional copy of a copyrighted image is
| not a violation of copyright. (eg: an executable binary
| which happens to contain the bits corresponding to a
| picture of Batman -- but which are actually instruction
| sequences and were provably not intended to encode the
| picture -- clearly doesn't infringe.)
|
| LLMs are somewhere in-between #1 and #2, and the intent
| can happen both in the training and also the prompting.
|
| Stack on top of this the fact that the models can also
| definitely generate content that counts as fair use, or
| which isn't copyrighted.
|
| It's the multitude of possible outputs, across the
| copyright spectrum, combined with the function of intent
| in training and/or prompting, which make this such a
| thorny legal issue for which existing copyright statute
| and jurisprudence is ill-suited.
|
| Taking your Batman example: DC would come after you for
| trademark as well as copyright, and the copyright claims
| would be very carefully evaluated with respect to your
| very specific work. But here we are talking about a large
| model that can generate tons of different work which
| isn't subject to copyright or which is possibly fair use.
|
| I don't think that existing jurisprudence (or even
| statute?!) can handle this situation very well, at all,
| without tons of arbitrary interpretative work on the
| parts of juries/judges, because of the multitude and
| vague intent issues described above.
|
| (...Also presumably the merits of the DC case wouldn't
| matter because your victory would be pyhrric unless you
| are a mega-corp. Which from a legal theory perspective is
| neither here nor there but from a legal practicality
| perspective may inform how companies go about enforcing
| copyright claims on model weights/outputs.)
|
| Anyways. I think we have a right mess on our hands and
| the legislature needs to do their damn jobs. Welcome to
| America, I guess :)
|
| Curious to hear your thoughts on these issues.
| visarga wrote:
| This is a great example. Summarizing or paraphrasing
| copyrighted content, or simply using it as a seed to
| generate input-output pairs - this kind of data
| transformation prior to training could solve the issues
| with copyright. It cleanly separates form from content.
| pxoe wrote:
| what is a 'copy'? byte accurate, or 'something with
| general resemblance'? would a badly compressed "copy"
| image of a copyrighted material still be 'a copy' or
| would it be some other thing? would low quality image
| compression be enough to skirt around copyright claims?
| image formats and viewers just 'reproduce' an impression
| of original data from derive compressed data. it is also
| just 'information that's been generalized by some degree'
| - for space saving purposes and so on. so, what if image
| generators could be thought of as a 'very good multi-
| image compression algorithm' that can output multiple
| images as well, to a 'somewhat recognizable degree'.
| hex4def6 wrote:
| Badly compressed still counts. I think if the data allows
| you to reconstruct a recognizable recreation of the
| original work, you have a good chance of it being
| considered a derivative copy.
|
| A mono audio version of Star Wars, compressed down to
| 320x240, filmed from the back of a theater on a VHS
| camera, converted to Video CD, would under any reasonable
| interpretation be just a copy of the original.
|
| I assume it starts getting murky when there's some sort
| of transformation done it it. What if I run motion
| capture on it, and use that motion capture data to create
| a cartoon version of Star Paws (my puppies in space
| epic)? What if I do a scene for scene recreation as the
| animated cartoon (removing any mentions to copyrighted
| names -- Luke Skywalker is now Duke Dogwalker, for
| example)? In this case, there's been no actual data
| transfer -- all the sprites are hand drawn, backgrounds
| etc.
|
| What would be an interesting exercise would be to try and
| create a series of artifacts that each on their own are
| considered non-derivatives, but can be used together to
| reconstitute the original. For example, create a
| compression method that relies heavily on transforms /
| macroblocks, but strip out any of the actual pixel data
| from the film. That info might be supplied as palette
| files which are themselves not really copyrighted data,
| but together with the compressed transform stream can be
| used to recreate the original video.
| mike_d wrote:
| > The summary contains information that was present in
| the original but it has been transformed and hence it's
| not a copy.
|
| The summary also contains original thought, something is
| added to it by a human to make it unique. AI models are
| primarily deriviative.
|
| A better example would be: if I take 1,000 different
| copyrighted works and put them into a ZIP file, does that
| resulting file violate copyright?
| version_five wrote:
| That example is awful, whatever side of the debate one is
| on
| kevin42 wrote:
| Let's say you take the harry potter books and create a
| spreadsheet with each word in it as a column, and the
| number of times that word appears. Would that violate the
| copyright? I'd be interested in the rationale if someone
| thinks it would.
| mike_d wrote:
| If your table was the number of times a word was followed
| by a chain of other words, that would be a closer
| comparison to AI weights. In that case it would be
| possible with reasonable accuracy to reconstruct passages
| from the harry potter books (see GitHub Copilot).
|
| The copyright aspect makes more sense when you start
| thinking of AI training models as lossy compression for
| the original works. Is a downsampled copy of the new Star
| Wars movie still protected under copyright?
|
| Just tabulating the word counts would not violate
| copyright as it is considered facts and figures.
| drdeca wrote:
| It resembles lossy compression in some ways, but in other
| important ways I think it doesn't?
|
| Like, if one has access to such a model, and doesn't
| count it towards the size cost of a
| compression/decompression program nor as part of the
| compressed size of the compressed images, then that
| should allow for compressing images to have substantially
| fewer bits than one would otherwise be able to achieve
| (at least, assuming that one doesn't care about the
| amount of time used to compress/decompress. Idk if this
| is actually practical.)
|
| But unlike say, a zip file, the model doesn't give you a
| representation of like, a list of what images (or
| image/caption pairs) it was trained on.
|
| Or like, in your analogy with the lower resolution of the
| movie, the lower resolution of it still tells you how
| long the movie is (though maybe not as precisely due to
| lower framerate, but that's just going to be off by less
| than a second, unless you have an exceedingly low
| framerate, but that's hardly a video at that point.)
|
| There is a sense in which any model of some data yields a
| way to compress data-points from it, where better models
| generally give a smaller size. But, like, any (precisely
| stated) description counts as a model?
|
| So, whether it is "like lossy compression" in a way that
| matters to copyright, I would think depends a lot on
| things like,
|
| Well, for one thing, isn't there some kind of "might
| someone consume the allegedly infringing work as a
| substitute for the original work, e.g. if cheaper?" test?
|
| For a lower resolution version of Star Wars movie, people
| clearly would.
|
| But if one wanted to view some particular artwork that is
| in the training set, I would think that one couldn't
| really obtain such a direct substitute? (Well, without
| using the work as an input to the trained model, asking
| it to make a variation, but in that case one already has
| the work separate from the model, so that's not really
| relevant.)
|
| If I wanted to know what happened in minute 33 of the
| Star Wars movie, I could look at minute 33 of the
| compressed version.
| londons_explore wrote:
| It's just a race for which test case gets to the supreme
| court first really...
| pmoriarty wrote:
| ...and the Supreme Court could rule however it likes. It
| doesn't matter what anyone else says, or what any law
| says, what any lawyer or other judge says.
|
| They could be completely biased, could completely ignore
| everyone and everything else and rule however they want.
|
| I'm almost surprised they still bother to write any kind
| of "legal reasoning" in their ruling and don't simply
| focus on what the ruling is rather than why they ruled
| that way. But I guess such "reasoning" still serves a
| propaganda purpose and still provides a fig leaf for
| those who still believe in the quaint absurdity that "we
| are a nation of laws, not men."
| londons_explore wrote:
| Supreme court precedent seems to impact a lot of
| decisions...
|
| Plenty of companies who have legal teams will keep an eye
| on the legal landscape of court decisions, and use them
| to decide if our T&C's or contracts need rewriting, or if
| any precedent puts us at legal risk.
|
| Sure - the supreme court could overthrow its precedent
| anytime, but until it does, a lot of people will act as
| if what they say is the law.
| dragonwriter wrote:
| > It's just a race for which test case gets to the
| supreme court first really...
|
| Not really for practical purposes. In the long term, the
| Supreme Court can and does overrule its own precedent, so
| the first case on the specific issue to get to the
| Supreme Court doesn't end the discussion.
|
| In the short-term, cases get resolved by lower courts and
| parties either lack funds to do the maximum level of
| appeals, or the Supreme Court chooses not to hear appeals
| (they tend to prefer an issue to be well-developed with
| circuit case law, often waiting till there is a conflict
| between the Circuit Courts of Appeal, before taking it
| up), so the state of the law _prior_ to any specific
| ruling on the narrow topic by the Supreme Court matters
| quite a bit.
| staunton wrote:
| This is the first time I ever saw a comment including the
| text "I am a lawyer". Does that mean the comment
| technically contains "legal advice"?
| lcnPylGDnU4H9OF wrote:
| As always, _a_ lawyer is not necessarily _your_ lawyer.
| mike_d wrote:
| If you choose to pay him, sure.
| wahnfrieden wrote:
| IP itself violates a natural right
|
| (Yes the idea of rights is also unnatural and absent from
| visions such as anarchy)
| version_five wrote:
| Yeah I didn't even think that was controversial. I'd
| always been taught that copyright and patents exist to
| explicitly restrict what people can do by granting a
| monopoly to the owners in order to encourage invention
| and creative work.
|
| Edit to add I'm not saying I agree with the justification
| or am trying to argue for it, only that the point above
| is commonly raised as the justification, implying that
| the intrusion on a person's rights is known and accepted.
| wahnfrieden wrote:
| [flagged]
| dragonwriter wrote:
| Natural rights are a fiction to pretend that someone's
| moral code is a privileged aspect of physical reality in
| a way every competing moral code is not.
| ndriscoll wrote:
| Even so, you can ask whether a given moral code is more
| principled than another (e.g. in the sense of having some
| algebraic structure), and use that to investigate what
| might be considered "more natural". For example, one
| might argue that if a "natural" right exists, then it
| ought to be symmetric under exchange of humans (or
| sentient beings or whatever). It's then "more natural" to
| conclude that you have a right to perform actions that
| have no interaction or consequences for other humans
| (e.g. to sing a copyrighted song to yourself in an empty
| room or downloading a song that you already have on CD
| but don't feel like ripping yourself) than those that do
| (e.g. taking food from someone because you'd otherwise
| starve).
| dragonwriter wrote:
| > Even so, you can ask whether a given moral code is more
| principled than another
|
| What does "more principled" mean of a moral code? How
| does one quantify "degree of principledness"?
|
| > and use that to investigate what might be considered
| "more natural".
|
| What does the preceding (being "more principled") have to
| do with being "more natural"? And what significance does
| being "more natural" have?
|
| And none of that has any relevance to what is usually
| described as "natural rights"; its like taking existing
| words and coming upnwith entirely novel meanings and then
| a whole architecture around them, which is pretty
| advanced equivocation.
| __MatrixMan__ wrote:
| That's going a bit far. They're just fictions that are
| privileged over certain other fictions--it's like how you
| can often cast magic missile in D&D but you can't usually
| cast expelliarmus, it comes down to which fiction we
| agree to inhabit.
| svachalek wrote:
| When the web was young, there was a lot of information
| considered "public" like criminal record, marriage records,
| birth certificates, property records, etc. But those were
| still fairly veiled because of the amount of effort
| required to see them. Suddenly these were getting blasted
| all over the internet because now that was an easy thing to
| do, and everyone had to rethink what "public" meant.
|
| I suspect we're going to see the same kind of rethink about
| intellectual property in the age of AI.
| pulvinar wrote:
| Not sure about criminal records, but the other records
| are generally still public. Not blasted all over, but
| there if you know where to look.
|
| Not that we really had all that much privacy in the past,
| as anyone who's browsed old newspapers knows.
| ldoughty wrote:
| > Every single one of these AI weight things itself is a
| result of unencumbered, massive, law-breaking, right-
| violating copyright infringement
|
| Maybe the popular and free ones. Adobe has a product in beta
| that uses "ethical training data" as a selling point.
| jrm4 wrote:
| Interesting. I wonder what they mean by "Ethical" --
| instead of e.g. saying "definitely free and open." I'm
| willing to bet "stuff they gathered from likely unwitting
| Adobe users."
| FooBarWidget wrote:
| It means trained on data from stock photo sites they own,
| for which all uploaders agreed to terms of service which
| state that uploaded materials can be fed into AI
| training.
| jrm4 wrote:
| So yes, exactly what I said. :)
| themoonisachees wrote:
| Adobe also happens to own Adobe stock, so maybe they
| simply trained on their own corpus.
|
| Who am I kidding this is Adobe of course they're fucking
| over their users
| Dr4kn wrote:
| They at least say they did and other copyright free
| artworks. They are a big company and know that they would
| get sued, so it should be in their interest to do it this
| way.
| zitterbewegung wrote:
| If I collect a set of copyright free data or public domain
| data would we conclude that the weights are also public
| domain?
| jrm4 wrote:
| That seems fair. I was under the impression that there
| weren't too many out there like this.
| jprete wrote:
| No, that doesn't follow at all. The argument is that either
| the training or the expression violated existing cooyrights
| through the making of unlicensed copies. It's not based on
| open source licensing. Although OSS viral licensing may
| well apply if fair use is not a successful defense.
| tedunangst wrote:
| What copyright is violated by training on public domain
| data?
| numpad0 wrote:
| DMCA works for free data too. It doesn't matter if your
| gains are in the form of fiat or crypto or social
| currency.
| jprete wrote:
| If it's public domain, then no copyright is violated. I'm
| not talking about public-domain data; the G-G-GP
| specifically mentioned the possible legal interpretation
| that training on large amounts of publicly visible (but
| not public domain) data is itself a copyright violation.
| numpad0 wrote:
| IANAL, I rather think the weight is not copyrightable
| anyway, and, if I build a model on copyrighted data, I
| would conclude that the inseparable but reproducible parts
| of weight retains copyrights, despite the whole weight not
| having its own.
| 8note wrote:
| Are they a work of art in and of themselves? I don't think
| you could tell without litigation
| [deleted]
| hcks wrote:
| > is a result of unencumbered, massive, law-breaking, right-
| violating copyright infringement
|
| Is there an official ruling? Or is it just a Reddit-style
| over exaggeration?
| Dr4kn wrote:
| There is no official ruling... yet. We are very early in
| this rapid public development. Laws and rulings take years
| or decades.
|
| They are trained on a lot of text. News sites, comments,
| books etc. Most books and news sites fall under copyright.
| Is this fair use? Who knows. Fair use is also an American
| thing. ChatGPT can be used in the EU, which doesn't have
| such a broad view of fair use.
|
| If you make a game only out of a lot of copyrighted assets
| without paying it isn't fair use. Are LLMs different?
|
| What about image generation, which you can prompt the
| models for specific styles of artists, which works are all
| copyrighted, but still used for training?
| barbariangrunge wrote:
| completely off topic, but funny: I misread "opencoreventures" as
| "opencorevultures"
| Makhini wrote:
| Funny
| morpheuskafka wrote:
| > AI also poses socio-ethical consequences that don't exist on
| the same scale as computer software, necessitating more
| restrictions like behavioral use restrictions
|
| There's plenty of software that has, or could have, similar
| restrictions. Consider software that allows you to plan vantage
| points for a shooting or estimate the impact of using explosives
| at various locations. And the government regulates all sorts of
| software for export/download because it has military use--
| everything from development tools to high performance chips that
| could be used to crunch numbers for a nuclear program, CAD
| software that can help you build (or destroy) a bridge, etc. The
| CPUs and GPUs themselves are regulated at certain performance
| levels, I think.
|
| None of this is really new to AI.
| ronsor wrote:
| Model weights are not source code, but data. Arguably because of
| how they are generated, they are not even copyrightable at all.
| adamsmith143 wrote:
| Corporate data is of course protect-able. Otherwise why don't
| you just open up all your databases so anyone can access them?
| WrongAssumption wrote:
| Copyright protection is what gives protection when you put
| something out into the public. The desire to not publish
| something is evidence against having these protections,
| because people know they are not copyright able, so for that
| reason and others they keep it private. You just presented
| evidence against your position.
| DannyBee wrote:
| They are only protectable by copyright you the degree they
| are creative works of authorship. Copyright is not usually
| how these are protected.
| zarzavat wrote:
| Weights might be copyrightable but in no universe are they
| copyrightable by OpenAI, Google, etc just because they did the
| training and spent money on GPUs.
|
| The only people who can possibly own the copyright, if any such
| copyright exists, are the authors of the training data.
|
| I find this whole discussion about copyright of weights almost
| absurd, the incredible amount of deference given to our corporate
| lords is such that we are "hallucinating" new forms of IP
| protection for NN weights that have never existed in any kind of
| statue or case law and cut completely against the grain of all
| the law that currently exists.
| amelius wrote:
| Just like you can't de-compile a binary without loss of
| information, "source" means that you can reconstruct it, so the
| training data should be available as well as the code that was
| used to train it, and the build script that invoked it.
| robomartin wrote:
| Can someone give me a legal answer to this?
|
| People, from early school, all the way up to university, use
| copyrighted materials to learn various topics and obtain degrees.
| This trains our brains using the work of others.
|
| The same is true as we navigate life. We learn various skills and
| subjects consuming the work of others.
|
| And, yes, in the case of most people, we use that training to
| pursue various careers, obtain work and get paid for it.
|
| How can there be a claim of infringement on the part of LLM's and
| not on every person who has ever used a book, website, article,
| video or publication to learn something?
| og_kalu wrote:
| This is an argument yes. A model could certainly be considered
| transformative enough to be fair use.
| blharr wrote:
| I am not a lawyer. But isn't this quite simple?
|
| Copyrighted materials are either licensed specifically for a
| human or it's implied that a human will use them to learn.
|
| Naturally, human memory is going to distort and change that
| information over time. But as soon as you use it in an AI,
| which has superhuman capabilities of memory, that would go out
| the window.
| robomartin wrote:
| > Copyrighted materials are either licensed specifically for
| a human or it's implied that a human will use them to learn.
|
| I don't think that's a part of copyright law at all. Maybe in
| the future, not today. Which makes sense, since these laws
| precede AI by a long time.
| low_tech_punk wrote:
| The lack of freedom to modification makes it not "open" either.
|
| Comparing to traditional software, weights are actually worse
| than binary. You can't "decompile" the weights into the training
| source code so there is no way for the community to make useful
| changes to them.
| kmeisthax wrote:
| >While the RAIL organization suggests adding the word "Open" to
| RAIL licenses that include similar open-access and free-use as
| open source (i.e. OpenRAIL-M), this is confusing since the
| license is not open source so long as it includes usage
| restrictions. A better name would be EthicalRAIL-M. Using the
| term "ethical" to describe this category license clearly
| indicates its functional difference from open source licenses.
|
| I don't even think we should be using the word "ethical" because
| it implies that anything more permissive is _un_ ethical. We
| should call these morality clause licenses.
|
| The question of whether or not we _should_ have morality clauses
| involved is complicated. Most bad actors do not give a shit about
| the licensing status of the code they are using. And these
| licenses also cause headaches for people who want to follow the
| rules[0] and avoid copyleft trolling[1]. On the other hand, the
| morality clauses in OpenRAIL-M are relatively straightforward and
| non-obnoxious.
|
| [0] This also applies to "non-commercial" licensing, since that
| is a concept entirely foreign to copyright law. As far as I'm
| concerned the 'NC' clause in Creative Commons just means 'OK to
| torrent'.
|
| [1] A practice in which people abuse copyleft licenses to try and
| extract licensing agreements for minor license violations. The
| forgiveness periods added to GPLv3 and later versions of Creative
| Commons are specifically to prevent this behavior.
| c7b wrote:
| Imho the weights are the real meat for most typical models, you
| can run with them and continue training them with your own code.
| It's not even guaranteed that the original code would be very
| useful for that.
|
| But if you are going to make that distinction, for which you can
| make a case I think, shouldn't you include a third dimension,
| 'data'? The code alone is hardly useful if you want to rebuild
| the weights, but all it tells you is that they're loading their
| proprietary data and then using PyTorch to set up and train the
| model. You can't reproduce anything using just that. So the real
| equivalent of open source would be imho either open weights, or
| open data plus code plus weights (the latter are arguably
| redundant, but still practical to include). Given that the size
| of that repo will typically be gigantic, I think open weights is
| the case we should really be focusing on. I'd rather have a paper
| explaining the model together with the weights, rather than code
| that I can't run anyway, if I'm designing an algorithm to
| continue training the model.
| meindnoch wrote:
| According to whom?
|
| Weights are a type of program, which are interpreted by the
| neural network runtime. Same as Java bytecode interpreted by the
| JVM runtime.
| eigenket wrote:
| x86 machine code is a type of program, which is interpreted by
| the processor, but distributing the binary of my program
| doesn't make it open source.
| kfarr wrote:
| Bingo, did a ctrl+f to find binary as that seems like the
| closest analogy here.
| slowmovintarget wrote:
| Weights are data, not a type of program.
|
| A computer program is a set of instructions that may be
| executed. Weights are values that may be loaded by a program,
| but are not a program in and of themselves.
| earleybird wrote:
| Weights are data in the same way that instruction codes in
| memory is data.
| slowmovintarget wrote:
| Values for the variables do not the function make.
| daniel-cussen wrote:
| [dead]
| rockinghigh wrote:
| When people talk about weights, they talk about a network of
| weights that takes an input and computes an output. There is
| really not much difference between a saved model and a
| program.
| jstanley wrote:
| It's a very difficult distinction to make.
|
| Would you consider a Python program to be data rather than
| program just because it is text input to the python
| interpreter instead of machine code for the CPU?
| slowmovintarget wrote:
| It is not at all a difficult distinction.
|
| Weights are literally numbers computed as output. They are
| not instructions. The semantics of those numbers even when
| emplaced (loaded) in an artificial neural net is such that
| they do not execute. They are not instructions. LLM engines
| and diffusers perform searches where the weights are used
| to calculate additional output.
|
| Is source code, like Python text, data? Yes. All code is
| data. But not all data are source code.
|
| If I gave you a web request log, you would not assert it is
| a program. If I gave you a CSV file with time-series values
| from a sensor, you would not assert it is a program. If I
| hand you a database of contact information, you would not
| assert it is a program. Weight files are the equivalent of
| CSV files. They are are a dump of parameter values computed
| from training.
|
| They are not a program.
|
| The definition of computer program is well worn. So is the
| definition of source code, and the definition of
| parameters. Weights are parameters.
| xigoi wrote:
| If a program has to consist of instructions, then source
| code written in a declarative language is not a program.
| jstanley wrote:
| The difference between code and data only exists in our
| minds. There is no distinction. Both code and data make
| the computer do things (and, yes, both code and data
| _only_ make the computer do things if other conditions
| are permitting, for example if executed with the right
| interpreter, or loaded with the right type of viewer).
| Anything that can be expressed as code can be expressed
| as data, and vice versa.
| pravus wrote:
| > Weights are literally numbers computed as output. They
| are not instructions.
|
| They are instructions if you consider the LLM system
| itself to be a kind of weird, indirect virtual machine.
| Each number can be mapped to a set of instructions that
| are executed. Even your CPU uses numbers (machine codes)
| to execute.
|
| Join me in saying: ...code is data is code is data is
| code is data...
| graypegg wrote:
| They're not data though, they're coefficients. They are the
| only thing that significantly differentiates one model from
| another.
|
| If I told you the economy can be accurately modelled by
|
| GDP(x) = Ax + B
|
| But I don't define A And B for you because it's proprietary,
| you haven't learned anything other than what you can glean
| from the structure of the model itself (it's linear, there's
| only a single input etc)
|
| If most of these models are similarly structured, I'd say the
| weights are the program.
| slowmovintarget wrote:
| The nature of the data as proprietary or not, important or
| not, is not relevant.
|
| Parameters, or actual arguments, are values; data. Not
| instructions.
|
| Valuable data is still data. It's significance doesn't
| magically turn it into source code.
| golemotron wrote:
| No, declarative programs exist. They are not instructions.
|
| There is no real line between code and data. This is an
| observation that runs all the way from Turing Machines in
| computability theory to the Von Neumann architecture and
| homoiconicity in Lisp.
|
| What we call 'data' is just code that needs a cleverer
| interpreter.
| Izkata wrote:
| Less into theory and more into "wait wtf": Some of the
| older projects I've worked on were written by people who
| loved database-driven stuff, to the point they did things
| like put perl code into one table column (with sentinel
| values you had to find/replace before `eval`ing the code)
| and sql into another table that retrieved values for those
| find/replaces, both retrieved and executed by some really
| generic code.
|
| Code or data: Well... both.
| killjoywashere wrote:
| But not "raw" data. They are derived from other data and a
| program. If this was a collaboration where one collaborator
| did the processing and one sourced the data, they would
| likely both claim some amount of ownership of the trained
| weights.
|
| At a minimum, it would be an active area of negotiation that
| the attorneys would take notice of. Source: have negotiated
| these agreements.
| slowmovintarget wrote:
| A curated data set is still a data set.
|
| I imagine it is not settled law, but there's a clear
| argument to be made that regardless of the difficulty in
| curating the data set, it's still a data set.
|
| Can it be licensed and sold. Yes, surely. Is it proper to
| pretend an open source license is sufficient protection,
| probably not.
| programmarchy wrote:
| This is a distinction without a difference. Code is data and
| data is code.
| mrguyorama wrote:
| Maybe they've only worked with machines using a Harvard
| Architecture
| slowmovintarget wrote:
| All source code is data. Not all data is source code. Data
| may be encoded, but that doesn't make it source code
| either.
| tensor wrote:
| I think the point here is that by being explicit you avoid the
| need to have this argument.
| cdelsolar wrote:
| who wrote that program?
| adamnemecek wrote:
| Java bytecode is not "open source". At least for Java bytecode
| there are decompilers.
| thepangolino wrote:
| I've always seen weights as akin to configuration files.
| Makhini wrote:
| What if you change the weights slightly? Kaboom, not breaking the
| copyright anymore.
| bskap wrote:
| Then it's a derivative work and copyright law covers that too.
| rpodraza wrote:
| And you're basing this theory on what exactly?
| sharcnick wrote:
| Copyright law & the definition of a "derivative work." See
| e.g. 17 USC SSSS 101 and 106. See also
| https://www.copyright.gov/circs/circ14.pdf.
| FrustratedMonky wrote:
| Are the weights in our brain copyrightable?
|
| Might want to get ahead of the curve on this one. How would this
| work? Would I get a tattoo with a license spelling out covering
| the contents of my body?
| DannyBee wrote:
| Has to be fixated (unchanging) and in a tangible medium.
| RobotToaster wrote:
| So I just need to cryogenically freeze my brain in order to
| copyright it?
| adamsmith143 wrote:
| The question shouldn't be whether the weights are copyrightable
| but whether they are protected under other electronic
| communication/data privacy laws.
| jkeisling wrote:
| The article makes a good point: we should prevent "open-washing"
| and draw a distinction between well-intentioned restrictive
| licenses like "Open"RAIL and true open source. However, I worry
| the name "ethical source" is itself a bit question-begging. While
| outfits like Bloom may believe in good-faith ethical principles,
| their definition of ethics isn't necessarily everyone's. If
| restricted models are "ethical", is releasing open weights
| "unethical"? Conversely, is releasing a model with PII or artist
| styles in it "ethical" if a few known use cases are forbidden?
| There's no one right answer. Labeling any one set of restrictions
| as "ethical" off the bat makes discussion harder and puts open
| source on the back foot to justify "not being ethical". Better to
| just call them "restricted models" or "guarded models", and leave
| it to individuals to decide if these restrictions are beneficial
| or not.
| A4ET8a8uTh0 wrote:
| I think the more interesting aspect of all this is that the
| confusion created by this new business model ( not sure to
| classify it so business model had to do ) appears to be largely
| intentional. The subject matter is complicated to begin with
| experts being niche of a niche of a niche and the assumption
| that the general public can even understand it ( and whether it
| can even dumbed down to digestible sound bites ) is, in my
| mind, very optimistic. Now, courts are not typically stacked
| with dummies, but again how many are well versed in issues of
| technology?
|
| All in all, I don't disagree with the point you raised, but I
| worry that all this will only further muddy the water for the
| general population.
| pmoriarty wrote:
| _" Now, courts are not typically stacked with dummies, but
| again how many are well versed in issues of technology?"_
|
| Even if they are well versed in issues of technology that
| does not mean they'll make what any given one of would
| consider a good decision, as plenty of people well versed in
| issues of technology disagree with each other on these
| issues.
|
| Nothing guarantees that on, on any issue, really, as you can
| always find people who disagree.. and if they happen to be
| judges, they get to decide unless another higher judge
| overrule them.. and that judge has the same problem as the
| first.
| A4ET8a8uTh0 wrote:
| Sure. My point is that I would so much rather have a
| decision handed down that was considered on actual merits (
| we might disagree, but at least I would be able to see some
| sort of real consideration and not what amounts to talking
| points from various lobbyists ). A judge that has zero
| exposure in that area is at best 50/50 and regardless of
| the ruling I will be annoyed that a person with zero
| knowledge is declaring how something he knows little to no
| about can be used ( just like I am more and more annoyed
| about political class in Washington, but I am more inclined
| to believe these days they know exactly what they are doing
| -- serve their own interests ).
|
| To your point, it is absolutely not panacea ( new blood is
| inevitably ending in government and the result so far is in
| line with what you said ), but it would at least be a
| starting point.
| tiffanyg wrote:
| _AI licensing is extremely complex. Unlike software licensing, AI
| isn't as simple as applying current proprietary /open source
| software licenses. AI has multiple components--the source code,
| weights, data, etc.--that are licensed differently._
|
| Are you joking? This isn't _wrong_ , per se, but it's worded as
| though written by someone with only the most casual / cursory
| interaction and knowledge of this area of law / commerce (e.g.,
| including licensing, copyright, trademark / service mark, patent,
| etc.) ... until perhaps quite recently.
|
| Yes, the AREA IS complicated. No, so-called "AI" is not
| introducing all sorts of novel issues, structures, etc. "AI" has
| some nuances distinct from much of what has come before (happens
| basically every time more significant tech comes along) and some
| possibly more unique questions related to economics, ethics,
| philosophy, and the like, but the relevant areas of law and
| practice have often been complicated and sort of "bleeding edge",
| even going back before the industrial revolution.
|
| Big money, powerful tech, large-scale economic forces, etc. =
| lots of maneuvering, legislation, litigation, etc. = complicated
| "rules of the game".
|
| Drawing the distinction vs. software in general is reasonable -
| but, the rather click-baity headline and "I just learned about
| 'IP' law and bah gawd y'all are doin' it wrong" tone to the start
| of this article suggest, to me, that this isn't likely to be the
| best article to use as a reference to learn about these issues.
| larodi wrote:
| I was like going to write 'are u joking', but you make the same
| point so well. This article is at best oversimplifying and
| misleading.
|
| Besides I doubt this 'my weights your weights' thing is a thing
| at all.
| cf141q5325 wrote:
| A focus on licensing ignores that there are security incentive to
| not run just any weights you find floating around the net.
| Getting exploited through miss-aligned networks is a very real
| threat and really hard to combat.
___________________________________________________________________
(page generated 2023-07-05 23:01 UTC)