[HN Gopher] DRuGS: Deep Random Micro-Glitch Sampling
       ___________________________________________________________________
        
       DRuGS: Deep Random Micro-Glitch Sampling
        
       Author : tosh
       Score  : 155 points
       Date   : 2023-12-30 16:25 UTC (6 hours ago)
        
 (HTM) web link (github.com)
 (TXT) w3m dump (github.com)
        
       | cranberryturkey wrote:
       | cute :P
        
         | revskill wrote:
         | What does cute mean ?
        
           | pastacacioepepe wrote:
           | kawaii
        
       | kstrauser wrote:
       | I appreciate the commitment to the metaphor. Bravo.
        
         | mylons wrote:
         | it is excellent copy, at the very least
        
       | Log_out_ wrote:
       | Nothing like good old monk humor about sex.. Unless druggy humor
       | by the notoriously drug free..
        
         | golergka wrote:
         | Yes, because HN crowd or AI/software crowd in general is famous
         | for being drug free and never indulging in various psychedelics
         | out in the desert.
        
           | mettamage wrote:
           | If they would, they would be burning man. Nothing to ignite
           | some passion than a deep dive in the desert ;)
        
           | BossingAround wrote:
           | It's kind of funny, I don't really see psychedelics as
           | "drugs" in the same way that you wouldn't view coffee as a
           | drug. I don't think SW engineers are, for example, known for
           | taking meth or heroin, so I'd definitely call the group
           | fairly "drug free".
           | 
           | I'd think (based on nothing but my experience) that your
           | average nurse might have higher likelihood of taking hard
           | drugs than an average SWE :)
        
             | golergka wrote:
             | > I don't think SW engineers are, for example, known for
             | taking meth
             | 
             | Some people think meth means "amphetamine" (although it
             | usually methamphetamine), and there's a lot of software
             | engineers who take different types of prescribed
             | amphetamines like Adderal and Ritalin.
             | 
             | And don't even get me started on weed and millenials.
        
               | danielbln wrote:
               | Methylphenidate is not an amphetamine but your points
               | stand.
        
             | knutzui wrote:
             | From Wikipedia: > A drug is any chemical substance that
             | when consumed causes a change in an organism's physiology,
             | including its psychology, if applicable.
             | 
             | Coffee, psychedelics and alcohol are drugs just like
             | heroin. Whether you believe they are useful to consume is a
             | different matter.
        
           | baq wrote:
           | Out in the desert is more extraordinary than microdosing
           | while WFH I guess
        
       | stavros wrote:
       | Ahh yes, the old drmgs.
        
       | nyolfen wrote:
       | i was just wondering yesterday what happened to this guy
        
       | madamelic wrote:
       | I appreciate the README etc, it's very funny and cute.
       | 
       | As a bit of a layperson / just an AI integrator, I want to get
       | some clarification of how this interfaces with a generative AI. I
       | haven't dug too deep in and I have a mostly elementary
       | understanding of neural nets. No need to _entirely_ dumb it down
       | though, just needing a smarter person to confirm or correct my
       | understanding. :)
       | 
       | Is this influencing the probabilities of token output? My
       | understanding is that currently it is a static value that is
       | essentially "how wild do you want it to be" where each generation
       | has a different consistent static value of "wildness".
       | 
       | So rather than using a static value for an entire generation,
       | this dynamically injects randomness in so rather than the output
       | being 'monotone', it is more dynamic rather than by-the-book that
       | AI tends to be?
        
         | andy99 wrote:
         | Typically the output of an LLM is a classifier that gives
         | scores reflecting the "best" next token, for some measure of
         | best. You can deterministically just take the highest score or
         | you can randomly sample in a way that makes it more likely to
         | get a token with a higher score, according to different
         | schemes. The point is that this all rests in the performance of
         | the output classifier layer.
         | 
         | From my skim, this proposes to inject randomness throughout the
         | model (as opposed to just during sampling the classifier
         | output) which according to the author (I'm paraphrasing) lets
         | the full power of the model be used to make the most of that
         | randomness, instead of just jittering the output. So the output
         | is still random but it's a function of random noise being added
         | through the model layers instead of just at the output,
         | supposedly demonstrating better properties. I have no opinion
         | about it as I haven't studied it.
        
           | madamelic wrote:
           | > You can deterministically just take the highest score or
           | you can randomly sample in a way that makes it more likely to
           | get a token with a higher score, according to different
           | schemes.
           | 
           | Ah ok, so it's like the difference between greedily always
           | taking the shortest / highest value vs using A* and exploring
           | alternatives in hopes of finding something better? That
           | probably doesn't line up as a metaphor as it isn't
           | "exploring" or running multiple parallel generations.
           | 
           | And it employs the method in the layers so that the benefit
           | can be amplified rather than only employing the method in the
           | final layer / output when it selects the output to take
        
             | doctorpangloss wrote:
             | The difference is between something that makes sense (the
             | way the model is intended to be used) and something that
             | makes almost no sense (this repository).
             | 
             | > benefit
             | 
             | In an intellectually honest way, it is not conclusive if
             | there is a benefit to the approach in this repository.
             | However, as a betting man: there is no benefit.
             | 
             | There are a lot of surprises! Riffusion and LCMs for me
             | were super surprising. With no examples or meaningful
             | investigation, I don't think this is going to yield
             | anything.
        
       | Philpax wrote:
       | Isn't this the GamerGate guy?
        
         | IshKebab wrote:
         | Apparently so:
         | 
         | https://ggwiki.deepfreeze.it/index.php/Eron_Gjoni
         | 
         | Unclear how "the gamergate guy" he was Vs just "accidentally
         | triggered gamergate". I'm not going down that rabbit hole
         | though - where's Internet Historian when you need him?
        
           | sleepybrett wrote:
           | Probably plagiarizing something.
        
           | sureglymop wrote:
           | Internet Historian recently got cancelled for being a Nazi
           | and hiding "Nazi symbolism" in his videos is what a quick
           | google told me.
        
             | zoklet-enjoyer wrote:
             | I'm still not really sure what gamer gate was or who the
             | Internet historian is
        
             | GaggiX wrote:
             | Where did you find that Internet Historian is a nazi? I
             | only found something when searching "Internet Historian
             | nazi", a reddit post that also claims that Jontron is a
             | white supremacist. This post: https://www.reddit.com/r/196/
             | comments/18erlt0/internet_histo...
             | 
             | If you actually googled "Internet Historian nazi" for a guy
             | you don't even know who he is, that would be pretty funny.
        
               | sureglymop wrote:
               | I have seen an Internet Historian video before. I just
               | googled "Internet Historian" and a cancellation style
               | reddit post was the third result. I even mentioned that I
               | only did a quick google; I encourage you to do your own
               | research here and form your own opinion on this. I have
               | heard of Jontron but I am not familiar nor have I read
               | anything about that.
        
               | tadfisher wrote:
               | JonTron actually said white supremacist stuff on video,
               | for which he later apologized (why he's not part of Game
               | Grumps anymore).
        
               | GaggiX wrote:
               | What do you mean? He left Game Grumps in 2013, way before
               | he said anything controversial.
        
             | IshKebab wrote:
             | I think to be "cancelled" you have to actually get
             | cancelled from somewhere... His Youtube channel seems to be
             | going fine.
             | 
             | It's very obvious from his videos that he is a 4chan troll.
             | Didn't know he was full on Nazi though.
        
           | bongripper wrote:
           | "Despite GamerGate's positive changes on the industry, it has
           | been met with heavy criticism from the journalists and
           | publications under fire and their supporters, as well as
           | third-wave feminists involved who have attributed the harsh
           | criticism towards female offenders to misogyny and sexism."
           | 
           | That summary seems a bit tendencious!
        
       | Der_Einzige wrote:
       | It's so funny how little love that work on the decoder/sampler
       | side with LLMs gets. You can take bad LLM and make them decent,
       | or good LLMs and make them garbage with intelligent sampling
       | settings. It's cool to see further innovation in this space.
        
       | RestlessAPI wrote:
       | I wonder if this works for Stable Diffusion as well...
        
         | samus wrote:
         | There are transformer-based approaches to vision models. It
         | should work for those.
        
       | junon wrote:
       | Great Readme. Strikes the perfect balance between quickly
       | informative and enjoyably cute.
        
       | bjornsing wrote:
       | So instead of the mathematically sound approach of sampling from
       | the categorical output distribution (that the model parameters
       | were fitted to with maximum likelihood) you just inject some
       | noise in the activations, and then pick the highest probability
       | output?
        
         | gwern wrote:
         | Injecting noise has many mathematically-sound interpretations,
         | like the Bayesian interpretations of dropout for ensembling or
         | posterior sampling.
        
       ___________________________________________________________________
       (page generated 2023-12-30 23:01 UTC)