[HN Gopher] Identifying and eliminating CSAM in generative ML tr...
       ___________________________________________________________________
        
       Identifying and eliminating CSAM in generative ML training data and
       models
        
       Author : pulisse
       Score  : 27 points
       Date   : 2023-12-20 17:29 UTC (5 hours ago)
        
 (HTM) web link (purl.stanford.edu)
 (TXT) w3m dump (purl.stanford.edu)
        
       | causality0 wrote:
       | _This methodology detected many hundreds of instances of known
       | CSAM in the training set_
       | 
       | Damn, I knew training data was kind of wild-westy, but I didn't
       | know it was _that_ sloppy.
        
         | goles wrote:
         | I think to be fair to the study, it's building off a previous
         | study that demonstrated Diffusion models can produce CG-CSAM.
         | Which seems like it could have some pretty serious
         | consequences, for victims, law enforcement, AI Ethics, and the
         | Justice system.
         | 
         | They reached that conclusion then naturally asked, if the model
         | can produce that, how much of it does it need to be trained on
         | to produce it? And they discovered the answer to that is very
         | little.
         | 
         | Which could be for any number of reasons that are far beyond
         | how I understand ML.
        
           | elpocko wrote:
           | It doesn't need to be trained on CSAM at all. You can make a
           | 2-pass workflow using Stable Diffusion: start generating
           | normal porn for a few steps, and for the rest generate
           | children. I don't think this can be stopped.
        
             | goles wrote:
             | I agree it has no place in the training data. They discuss
             | identifying, removing, preventing, and changes to training,
             | hosting section 5 page 11.
             | 
             | https://stacks.stanford.edu/file/druid:kh752sm9123/ml_train
             | i...
        
           | akira2501 wrote:
           | > Which seems like it could have some pretty serious
           | consequences, for victims, law enforcement, AI Ethics, and
           | the Justice system.
           | 
           | What actual difference is there between a machine generating
           | this material or a human being drawing it by hand? It seems
           | like we're living with the "pretty serious consequences"
           | already, and yet, it seems like we've found a way to manage
           | those effectively already.
           | 
           | > if the model can produce that, how much of it does it need
           | to be trained on to produce it? And they discovered the
           | answer to that is very little.
           | 
           | The answer may very well be "none at all," particularly if
           | these systems can create image fragments by inference and a
           | different system or even a human can assemble them.
           | 
           | > are far beyond how I understand ML.
           | 
           | If you have a crappy product, get it regulated, it grants it
           | the imprimatur of credibility and it hamstrings your
           | competitors and startups in the space. An understanding of ML
           | may not be required at all to understand this regulatory
           | situation.
        
       | Eisenstein wrote:
       | 825 out of 5,850,000,000.
        
       | akomtu wrote:
       | Are they going to "lobotomize" ML models to make csam unthinkable
       | for it, or are they going to teach it that csam is bad?
        
         | GaggiX wrote:
         | They are going to report the content to the CDNs and try to
         | remove the entry URL from the LAION dataset and that's mostly
         | it.
        
         | tedivm wrote:
         | The two paragraph abstract states that the goal is to identify
         | CSAM in the training data itself and remove it. They ran a
         | study and found "many hundreds of instances of known CSAM in
         | the [LAION-5B] training set".
        
       | darkwraithcov wrote:
       | What happens when all the fake diffusion model generated csam
       | that is flooding the darker areas of the clearnet gets inevitably
       | trained, setting aside the notion of model collapse?
       | 
       | I don't see this issue going away anytime soon, especially since
       | all that fake csam is still legal and basically unstoppable.
        
       | Footnote7341 wrote:
       | Such a vanishingly small percentage of the images its not even
       | worth calculating. Of course search engines also contain these
       | links, LAION-5B only contains links not images as well....
       | 
       | protect the kids!!! or something. Can we do a scandal about how
       | many 'extremist' images are in the data-set next too? anti-vaxer,
       | climate denier, nazi, religious extremist propaganda, scientific
       | misinformation. Maybe we're all safer off using corporate models
       | with closed data sets so no one gets any of the wrong ideas.
        
         | WheatMillington wrote:
         | For the children who have been victimised, I doubt the small
         | percentage provides any comfort.
        
       ___________________________________________________________________
       (page generated 2023-12-20 23:00 UTC)