[HN Gopher] Not all tokens are meant to be forgotten
___________________________________________________________________
Not all tokens are meant to be forgotten
Author : MarcoDewey
Score : 48 points
Date : 2025-06-04 23:15 UTC (23 hours ago)
(HTM) web link (arxiv.org)
(TXT) w3m dump (arxiv.org)
| pixl97 wrote:
| > However, they tend to memorize unwanted information, such as
| private or copyrighted content,
|
| I mean humans don't forget copyrighted information. We just
| typically adjust it enough (some of the time) to avoid getting a
| copyright strike while modifying it in some way useful.
|
| We don't forget 'private' information either. We might not tell
| other people that information, but it still influences our
| thoughts.
|
| The idea of a world where we have AI minds forget vast amounts of
| information that humans have to deal with every day is concerning
| and dystopian to me.
| johnjreiser wrote:
| I'd counter with an anecdote; I had a colleague that boasted
| how he memorized a classmate's SSN in college and would greet
| him by SSN when seeing him years later. Is the goal of AI to
| replicate the entirety of the human experience (including
| social pressures, norms, and shame) or a tool to complement
| human decision making?
|
| While, yes, you can argue the slippery slope, it may be
| advantageous to flag certain training material as exempt. We as
| humans often make decisions without perfect knowledge, and
| "knowing more" isn't a guarantee that it produces better
| outcomes, given the types of information consumed.
| lmm wrote:
| Knowing more might not improve your accuracy but it's not
| going to harm it. Forcibly forgetting true parts of your
| knowledge seems far more likely to have unintended
| consequences.
| conception wrote:
| Counterpoint: There are plenty examples of breakthroughs
| from folks who are ignorant of the "right" way to go about
| it. A fresh take isn't always bad.
| Dylan16807 wrote:
| I disagree. Actively fighting against your memory will slow
| you down in any context where some memorized idea is
| similar to what you're doing but you shouldn't be using the
| memorized idea.
| lou1306 wrote:
| One obvious consequence: the model might still produce
| copyright infringement because it thinks its creative ideas
| are novel.
| genewitch wrote:
| If the copyrighted content is not in the training data,
| and I mean explicitly, and the AI produces a copyrighted
| output, I'd argue it's a clean room re-implementation,
| and also it ought devalue the original work, moreso if
| the work is more recent. Maybe.
|
| I get that "first to publish" matters to a lot of people,
| but, say 5 unrelated people are writing unique
| screenplays about a series of events that seems important
| to them or culture or whatever; if they all come up with
| very similar plots and locations and scenes, it just
| means that the idea is more obvious than non-obvious.
|
| Please, argue. I haven't fully reconciled a lot of this
| to myself, but off the cuff this'll do.
|
| The logic being - if an AI without taint produces some
| other work, that work drew on the same information the
| model did, and came to the same "conclusion" - which
| means with a time machine, you could wipe the LLM, go
| back to the period of the original work, train the LLM,
| and produce the work contemporaneous to the original.
| Hope that made sense.
| lmm wrote:
| > If the copyrighted content is not in the training data,
| and I mean explicitly, and the AI produces a copyrighted
| output, I'd argue it's a clean room re-implementation
|
| You can't claim it's a clean room without actually doing
| the legwork of making a clean room. Not including the
| copyrighted work verbatim isn't enough, you would need to
| show that the AI hadn't seen anything derived from that
| copyrighted work, or that it had seen only non-
| copyrightable pieces.
| lou1306 wrote:
| > The logic being - if an AI without taint produces some
| other work, that work drew on the same information the
| model did, and came to the same "conclusion" - which
| means with a time machine, you could wipe the LLM, go
| back to the period of the original work, train the LLM,
| and produce the work contemporaneous to the original.
| Hope that made sense.
|
| This logic would immediately get shot down by an
| "Objection, speculation" in an actual litigation.
| Besides, the technicalities of how the work was produced
| don't really play a role in assessing infringement. PK
| Dick wrote "The man in the high castle" by extensively
| using the I Ching, but if I use it and recreate the novel
| by complete accident I would still be infringing.
|
| By the way, I highly suggest Borges's "Pierre Menard,
| Author of the Quixote" as a great story on the topic of
| authorship :)
| genewitch wrote:
| > PK Dick wrote "The man in the high castle" by
| extensively using the I Ching, but if I use it and
| recreate the novel by complete accident I would still be
| infringing.
|
| I touched on this, with the comment that we love "first
| to market." That multiple people coming up with the same
| output may mean that the idea isn't that novel. whether
| that matters or not isn't really relevant to me.
|
| The part you quoted was just a thought experiment to
| explain why i compared it to a "clean room
| implementation" - note it also avoids this argument from
| a sibling comment:
|
| >need to show that the AI hadn't seen anything derived
| from that copyrighted work
|
| since there could not possibly be any derived work prior
| to the "original" work being published. For the sake of
| argument.
| kgwgk wrote:
| I highly suggest Borges's "Pierre Menard, Author of the
| Quixote" as a great story on the topic of authorship :)
| JadeNB wrote:
| The repetition of the end of lou1306's comment
| (https://news.ycombinator.com/item?id=44190054) "By the
| way, I highly suggest Borges's 'Pierre Menard, Author of
| the Quixote' as a great story on the topic of authorship
| :)" has to be a joke ... right?
| kgwgk wrote:
| Good question! Is Pierre Menard's Quixote a repetition of
| Cervantes' or is it a completely different work that just
| happens to contain the same words?
| JadeNB wrote:
| > Is Pierre Menard's Quixote a repetition of Cervantes'
| or is it a completely different work that just happens to
| contain the same words?
|
| I think that that is not the right question. It is a
| repetition of Cervantes's work by design, at least if one
| takes, as I do, 'repetition' to mean saying or writing
| the same words in the same order. I think the question is
| whether it is therefore the same work, or a different
| work that contains the same words.
| lynx97 wrote:
| The goal of AI is to make money. All the moralisation is very
| human, but also extremely naive.
|
| BTW, I don't really understand what "social pressure" and
| "shame" has to do with your story? In my book, the person
| with a good memory isn't to blame. They're just demonstrating
| a security issue, which is a good thing.
| falcor84 wrote:
| In that example, the mnemonist should be demonstrating the
| security issue to the government, and not to their friend.
| We have social taboos for this reason. As an extreme
| example, I wouldn't greet a person by their penis size
| after noticing it in the locker room - some information
| should still be considered private, regardless of how we
| came to obtain it.
|
| Same with an LLM, when it got sensitive information in its
| weights, regardless of how it obtained it, I think we
| should apply pressure/shame/deletion/censorship (whatever
| you call it) to stop it from using that information in any
| future interactions.
| lynx97 wrote:
| I am probably too autistic to recognize remembering a
| personal datum as a taboo.
|
| However, I am totally on your side regarding LLMs
| learning data they shouldn't have seen in the first
| place. IMO, we as a society are too much chicken to act
| on the current situation. Its plain insane that everyone
| and their dog knows that libgen has been used to train
| models, and the companies who did this experiencing NO
| consequences at all. After that, we shouldn't be surpised
| if things go downhill from here on.
| squidbeak wrote:
| I agree. As far as copyrighted and artistic works go, I've
| never fully understood what the objection is. If the work is
| being remixed not copied then it surely falls under fair use?
| Meanwhile, if it creates something new in an artist's style,
| it's only doing what talented imitators routinely do. There's
| the economic argument. But if that's accepted, then for
| fairness it would have to be extended to every other profession
| which stands to be wiped out by AI, which would be daft.
|
| New works in familiar styles are something I can't wait for.
| The idea that the best Beethoven symphony hasn't been composed
| yet, or that the best Basquiat hasn't been painted yet, or that
| if the tech ever gets far enough, Game of Thrones might
| actually be done properly with the same actors, is a pretty
| mouthwatering prospect. Also styles we haven't discovered, that
| AI can anticipate. How's it to do that without a full
| understanding of culture? Hobbling the delight it could bring
| generally for the sake of protected classes will just make the
| tech less human and a lot less exciting.
| wizardforhire wrote:
| Mind if I ask a few questions? Whats your current address,
| dob, ssn or NINO or equivalent, your full legal name, mothers
| maiden name, fathers place of birth, mothers place of birth,
| country of origin, do you drive? Whats your license number?
| How about a bank? Could I have your account and routing
| number as well as the answers to any security questions? How
| about investments I'm gonna need your accounts and passwords
| for these as well...
|
| > As far as copyrighted and artistic works go, I've never
| fully understood what the objection is ... > But if that's
| accepted, then for fairness it would have to be extended to
| every other profession which stands to be wiped out by AI,
| which would be daft. ... > Hobbling the delight it could
| bring generally for the sake of protected classes will just
| make the tech less human and a lot less exciting.
|
| So let me get this straight, you want to ruin the livelihoods
| of everyone so you can have a fancier toy to play with?
|
| When your life is ruined and can't make a living you'll have
| the answers you desire and understand the objections to why
| you can't have fancier toys.
|
| But heres the thing, and with the way the world is going atm,
| not being able to make a living is going to be the least of
| your and everyone else's worries that feel the way you do if
| ya'll get your way.
|
| People don't like having their livelihoods taken away, and
| when you threaten the livelihoods of their children... people
| tend towards violence.
|
| I really wish there was a more polite way to put this. Alas
| what you're proposing is all out war for what? A better game
| of thrones?
| squidbeak wrote:
| Violent artists with pitchforks, eh? Aside from their
| supposed predisposition to vengeful bloodlust, is there any
| other reason these protected classes should enjoy a
| different status to any other worker?
| wat10000 wrote:
| If it's remixed then it would be a derivative work and you'd
| need permission from the original copyright holder, just like
| if you literally remixed a song, or made a movie based on a
| novel.
|
| IMO the only reason there's even a question about whether
| LLMs can legally be trained on copyrighted works without
| permission is that the training is being done by (agents
| working on behalf of) rich people. If you or I scraped up
| every copyrighted work we could get our hands on without ever
| asking permission, trained an LLM on it, and then tried to
| sell access to the result? Just ask Aaron Swartz how that
| sort of thing goes, and his actions were orders of magnitude
| less.
|
| Humans don't forget copyrighted material but we also don't
| normally memorize it. It takes substantial time and effort to
| be able to reproduce copyrighted material with just your
| brain.
| JimDabell wrote:
| There's a related paper that Meta published a couple of days ago
| that is worth looking at:
|
| > How much do language models memorize?
|
| -- https://arxiv.org/abs/2505.24832
|
| -- https://news.ycombinator.com/item?id=44171363
|
| It shows that models are limited in how much they can memorise
| (~3.6 bits per parameter), and once that threshold is reached,
| the model starts to generalise instead of memorise.
___________________________________________________________________
(page generated 2025-06-05 23:02 UTC)