[HN Gopher] Identifying and eliminating CSAM in generative ML tr...
___________________________________________________________________
Identifying and eliminating CSAM in generative ML training data and
models
Author : pulisse
Score : 27 points
Date : 2023-12-20 17:29 UTC (5 hours ago)
(HTM) web link (purl.stanford.edu)
(TXT) w3m dump (purl.stanford.edu)
| causality0 wrote:
| _This methodology detected many hundreds of instances of known
| CSAM in the training set_
|
| Damn, I knew training data was kind of wild-westy, but I didn't
| know it was _that_ sloppy.
| goles wrote:
| I think to be fair to the study, it's building off a previous
| study that demonstrated Diffusion models can produce CG-CSAM.
| Which seems like it could have some pretty serious
| consequences, for victims, law enforcement, AI Ethics, and the
| Justice system.
|
| They reached that conclusion then naturally asked, if the model
| can produce that, how much of it does it need to be trained on
| to produce it? And they discovered the answer to that is very
| little.
|
| Which could be for any number of reasons that are far beyond
| how I understand ML.
| elpocko wrote:
| It doesn't need to be trained on CSAM at all. You can make a
| 2-pass workflow using Stable Diffusion: start generating
| normal porn for a few steps, and for the rest generate
| children. I don't think this can be stopped.
| goles wrote:
| I agree it has no place in the training data. They discuss
| identifying, removing, preventing, and changes to training,
| hosting section 5 page 11.
|
| https://stacks.stanford.edu/file/druid:kh752sm9123/ml_train
| i...
| akira2501 wrote:
| > Which seems like it could have some pretty serious
| consequences, for victims, law enforcement, AI Ethics, and
| the Justice system.
|
| What actual difference is there between a machine generating
| this material or a human being drawing it by hand? It seems
| like we're living with the "pretty serious consequences"
| already, and yet, it seems like we've found a way to manage
| those effectively already.
|
| > if the model can produce that, how much of it does it need
| to be trained on to produce it? And they discovered the
| answer to that is very little.
|
| The answer may very well be "none at all," particularly if
| these systems can create image fragments by inference and a
| different system or even a human can assemble them.
|
| > are far beyond how I understand ML.
|
| If you have a crappy product, get it regulated, it grants it
| the imprimatur of credibility and it hamstrings your
| competitors and startups in the space. An understanding of ML
| may not be required at all to understand this regulatory
| situation.
| Eisenstein wrote:
| 825 out of 5,850,000,000.
| akomtu wrote:
| Are they going to "lobotomize" ML models to make csam unthinkable
| for it, or are they going to teach it that csam is bad?
| GaggiX wrote:
| They are going to report the content to the CDNs and try to
| remove the entry URL from the LAION dataset and that's mostly
| it.
| tedivm wrote:
| The two paragraph abstract states that the goal is to identify
| CSAM in the training data itself and remove it. They ran a
| study and found "many hundreds of instances of known CSAM in
| the [LAION-5B] training set".
| darkwraithcov wrote:
| What happens when all the fake diffusion model generated csam
| that is flooding the darker areas of the clearnet gets inevitably
| trained, setting aside the notion of model collapse?
|
| I don't see this issue going away anytime soon, especially since
| all that fake csam is still legal and basically unstoppable.
| Footnote7341 wrote:
| Such a vanishingly small percentage of the images its not even
| worth calculating. Of course search engines also contain these
| links, LAION-5B only contains links not images as well....
|
| protect the kids!!! or something. Can we do a scandal about how
| many 'extremist' images are in the data-set next too? anti-vaxer,
| climate denier, nazi, religious extremist propaganda, scientific
| misinformation. Maybe we're all safer off using corporate models
| with closed data sets so no one gets any of the wrong ideas.
| WheatMillington wrote:
| For the children who have been victimised, I doubt the small
| percentage provides any comfort.
___________________________________________________________________
(page generated 2023-12-20 23:00 UTC)