[HN Gopher] Dummy's Guide to Modern LLM Sampling
       ___________________________________________________________________
        
       Dummy's Guide to Modern LLM Sampling
        
       Author : nkko
       Score  : 136 points
       Date   : 2025-05-04 16:26 UTC (6 hours ago)
        
 (HTM) web link (rentry.co)
 (TXT) w3m dump (rentry.co)
        
       | antonvs wrote:
       | This is great! "Sampling" covers much more than I expected.
        
       | blt wrote:
       | This is pretty interesting. I didn't realize so much manipulation
       | was happening after the initial softmax temperature choice.
        
         | minimaxir wrote:
         | It's worth noting that only _some_ of these techniques are
         | configurable in modern LLM API outputs. (usually only
         | temperature /top-p/top-k since other penalties require
         | overhead)
        
           | Der_Einzige wrote:
           | Most other penalties don't require much overhead (min_p is
           | basically free).
           | 
           | Most techniques are not made available by API providers
           | because they enable alignment breaking. It's the only
           | explanation for why we are still stuck with only top_p,
           | top_k, and temp of 0-2.
           | 
           | If you want proper sampler settings to be available, your
           | options are oobabooga, sillytavern (dependent on your
           | backend, so vllm backend for example doesn't have top-n sigma
           | yet), or directly running huggingface code. There might be
           | some marginal options here too but in general, sampling
           | innovation is firmly in the hands of open source coomers
           | right now and not in the hands of academics.
        
       | Der_Einzige wrote:
       | Related to this, our min_p paper was ranked #18 out of 12000
       | submission at ICLR and got an oral:
       | 
       | https://iclr.cc/virtual/2025/oral/31888
       | 
       | Our poster was popular:
       | 
       | poster:
       | https://iclr.cc/media/PosterPDFs/ICLR%202025/30358.png?t=174...
       | 
       | oral presentation (watch me roast yoshua bengio on this topic and
       | then have him be the first questioner, 2nd speaker starting
       | around 19:30 min mark. My slides for the presentation are there
       | too and really funny.):
       | https://iclr.cc/virtual/2025/session/31936
       | 
       | paper: https://arxiv.org/abs/2407.01082
       | 
       | As one of the min_p authors, I can confirm that Top N sigma is
       | currently the best general purpose sampler by far. Also,
       | temperature can and should be scaled far higher than it is today.
       | Temps of 100 are totally fine with techniques like min_p and top
       | N sigma.
       | 
       | Also, the special case of top_k = 2 with ultra high temperature
       | (one thing authors recommend against near the end) is very
       | interesting in its own right. Doing it leads to spelling errors
       | every ~10th word - but also seems to have a certain creativity to
       | it that's quite interesting.
        
         | toxik wrote:
         | Are there any samplers that aren't basically greedy? I.e.
         | actually searches the tree. I realize it's an absolutely insane
         | branching factor and quite expensive to expand nodes at that,
         | but it always seemed odd to me that we don't actually search.
        
           | Kubuxu wrote:
           | Beam Search sampling is sometimes getting used
        
           | Der_Einzige wrote:
           | Besides beam search and it's variants? (there are many
           | including the little known but awesomely powerful constrained
           | beam search: https://huggingface.co/blog/constrained-beam-
           | search)
           | 
           | Does MBR (minimal bayes risk) sampling count?
           | 
           | Also there was this paper at ICLR which is relevant to this
           | question: https://arxiv.org/abs/2410.03968
           | 
           | This paper basically claims that non-heuristic methods (like
           | beam search) are harmful compared to the heuristic ones.
        
       | orbital-decay wrote:
       | One thing not said here is that samplers have no access to
       | model's internal state. It's basic math applied to the output
       | distribution, which technically carries some semantics but you
       | can't decode it without being as smart as the model itself.
       | 
       | Certain samplers described here like repetition penalty or DRY
       | are just like this - the model could repeat itself in a myriad of
       | ways, the only way to prevent all of them is better training, not
       | n-gram search or other classic NLP methods. This is basically
       | trying to plug every hole with a finger. How many fingers do you
       | have?
       | 
       | Hacking the autoregressive process has some some low-hanging
       | fruits like Min-P that can make some improvement and certain
       | nifty tricks possible, but if you're doing it to turn a bad model
       | into a good one, you're doing it wrong.
        
         | Der_Einzige wrote:
         | No, it's done to turn an uncreative model into a creative
         | model. This idea that sampling isn't that important or is some
         | violation of the bitter lesson is exactly why I had to call out
         | the whole academic field as having a giant blindspot for this
         | kind of research in our oral presentation at ICLR!
         | 
         | Top n sigma has been around since mid 2024, min_p around since
         | 2023 and we are still waiting for these innovations to be
         | integrated outside of open source stuff (i.e. outside of
         | HF/vllm). It's being done slowly on purpose by API providers
         | because they don't want to deal with the risk of models being
         | "too creative" (also high temp likely breaks their
         | watermarking)
         | 
         | One other thing - making models aware of their own sampling
         | settings is super easy if you just feed it back to the model
         | every token or generation (say, using structured generation).
         | Models can control their own sampling settings and thus "have
         | access to its internal states" with just a tiny bit of extra
         | programming (the model can write that code for you now lol)
        
           | orbital-decay wrote:
           | I guess variance is a better word for this. Creativity is a
           | pretty loose term, for example most people will describe R1
           | as creative in RP/stories for its tendency to derail
           | everything in an unhinged way, but it still lacks variance
           | like every other modern model (kill the reasoning chain and
           | look at logprobs to get what I mean). The bitter lesson is
           | not some threshold and can't be violated, it describes a
           | curve of diminishing returns. As long as you're on the easy
           | part, it's fine.
           | 
           | But the bigger problem is that the concepts are expressed
           | before they're decoded into the output distribution. You can
           | steer them to a degree by hacking the autoregressive
           | transport, but if the model itself learned that this concept
           | corresponds to that particular concept, not a set of concepts
           | (and RL tends to do exactly that), fixing it with sampling is
           | usually hard to impossible, you'll just lose accuracy/make it
           | dumber as you basically force out-of-distribution outputs.
        
         | neuroelectron wrote:
         | The primary concern here (in this guide) seems to be efficiency
         | and preventing complexity explosions.
        
       | mdp2021 wrote:
       | When the attempt is though to have the LLM output an "idea", not
       | just a "next token", the selection over the logits vector should
       | break that original idea... If the idea is complete, there should
       | be no need to use sampling over the logits.
       | 
       | The sampling, in this framework, should not happen near the
       | output level ("what will the next spoke word be").
        
         | minimaxir wrote:
         | LLMs are trained to maximize the probability of correct guesses
         | for the next token, not "ideas". You cannot define an idea as a
         | training loss objective.
        
           | mdp2021 wrote:
           | That is an architectural problem. If you want the post to be
           | rephrased: it is paradoxical to have changes made near the
           | output level, "changing words before it says them", given
           | that the expected is to work with ideas. (And even then,
           | selection would not be at the output level - it would be
           | during the definition of the structure.)
           | 
           | So, articles like this submission - while interesting from
           | many points of view - make the elephant in the room more
           | evident.
           | 
           | > _You cannot define an idea as a training loss objective_
           | 
           | What tells you so? If you see a technical limit, note e.g.
           | that sentences and paragraphs can have their own position in
           | an embedding space.
        
           | orbital-decay wrote:
           | Interpretability studies offer several orthogonal ways to
           | look at this, it's like Newtonian vs Lagrangian mechanics.
           | Autoregressive token prediction, pattern matching, idea
           | conceptualization, pathfinding in the extremely
           | multidimensional space...
        
       | neuroelectron wrote:
       | Love this and the way everything is mapped out and explained
       | simply really opens up the opportunity for trying new things, and
       | where you can do that effectively.
       | 
       | For instance, _why not_ use whole words as tokens? Make a
       | "robot" with a limited "robot dialect." Yes, no capacity for new
       | words or rare words, but you could modify the training data and
       | input data to translate those words into the existing vocabulary.
       | Now you have a much smaller mapping that's literally robot-like
       | and kind of gives the user an expectation of what kind of answers
       | the robot can answer well, like C-3PO.
        
         | minimaxir wrote:
         | > For instance, why not use whole words as tokens?
         | 
         | Word-only tokenizers what people did in the RNN/LSTM days.
         | There's no functional improvement over tokenization schemes
         | like BPE or even WordPiece/SentencePiece, and it results in
         | worse quality since you can't use meaningful semantic hints
         | such as punctuation.
        
           | neuroelectron wrote:
           | You can encode semantic hints in the layers instead.
           | Admittedly, this is more expensive which is kind to counter
           | of the word-as-tokens idea.
        
       | simonw wrote:
       | This is a really useful document - the explanations are very
       | clear and it covers a lot of ground.
       | 
       | Anyone know who wrote it? It's not credited and it's pubished on
       | a free Markdown pastebin.
       | 
       | The section on DRY - "repetition penalties" - was interesting to
       | me. I often want LLMs to deliberately output exact copies of
       | their input. When summarizing a long conversation for example I
       | tend to ask for _exact quotes_ that are most illustrative of the
       | points being made. These are easy to fact check later by
       | searching for them in the source material.
       | 
       | The DRY penalty seems to me that it would run counter to my goal
       | there.
        
         | nkko wrote:
         | I didn't realize that it wasn't attributed; it was written by
         | @AlpinDale.
        
       | smcleod wrote:
       | I had a go at writing a bit of a sampling guide for
       | Ollama/llama.cpp as well recently, open to any feedback /
       | corrections - https://smcleod.net/2025/04/comprehensive-guide-to-
       | llm-sampl...
        
       | ltbarcly3 wrote:
       | Calling things modern that are updates to techniques to use
       | technologies only invented a few years ago is borderline
       | illiterate. Modern vs what, classical LLM sampling?
        
         | eddyzh wrote:
         | LLM are way older. The Nobel prize for it shows how they made
         | many of the breakthroughs decades ago ChatGTP was the popular
         | breakthrough. Even then your Smartphone keyboard has been using
         | an LLM for a decade.
        
         | Der_Einzige wrote:
         | Many of these algorithms were invented in like 2019 (i.e. TFS)
         | or even earlier (temperature)
        
       ___________________________________________________________________
       (page generated 2025-05-04 23:00 UTC)