[HN Gopher] Show HN: Tiny Diffusion - A character-level text dif...
       ___________________________________________________________________
        
       Show HN: Tiny Diffusion - A character-level text diffusion model
       from scratch
        
       This is a character-level language diffusion model for text
       generation.  The model is a modified version of Nanochat's GPT
       implementation and is trained on Tiny Shakespeare!  It is only 10.7
       million parameters, so you can try it out locally.
        
       Author : nathan-barry
       Score  : 65 points
       Date   : 2025-11-10 15:13 UTC (4 days ago)
        
 (HTM) web link (github.com)
 (TXT) w3m dump (github.com)
        
       | yugretcx wrote:
       | Why do these text diffusion demos always look like the number of
       | allowed tokens is fixed for a specific unfilled region?
       | 
       | Is this the case?
       | 
       | Ie. if the region only has four tokens(here characters) but
       | calculates the best word is "forget" does it just abandon the
       | best fit or truncate it to fit?
       | 
       | Are there text diffusion models with lax infill directives?
        
         | rand0mwalk wrote:
         | Tokens start as a special [MASK] token. Then as the diffusion
         | process runs they are "unmasked" i.e. sampled.
         | 
         | So yes, you define a sequence of [MASK] tokens with some length
         | ahead of time.
         | 
         | In practice, if a model wants to write a shorter sequence,
         | it'll just fill the remaining tokens with empty content. If it
         | wants to write a longer sequence, you'll have to identify this
         | and extend the sequence with more [MASK] tokens. This is
         | typically obvious since there's no "end of sequence" token
         | present if the model wants to generate more.
        
         | nathan-barry wrote:
         | Yes, this is the case. During training, the model will get a
         | sequence of text (ex, 512 tokens long) with a percentage of
         | them masked out (with a special <MASK> token). It learns how to
         | unmask those tokens to construct the original text.
         | 
         | In the case that you mentioned, if we had 4 <MASK> tokens in a
         | row, all we are doing for decoding is predicting what those 4
         | tokens should be.
         | 
         | Generally, this does not seem to be a significant problem, as
         | there are usually multiple ways to express an idea in varying
         | lengths. Also, with confidence-aware parallel decoding, it can
         | usually avoid the scenario you mentioned, as focusing on
         | decoding the highest confident tokens will generally avoid such
         | scenarios with a well trained model.
        
         | throwaway314155 wrote:
         | This comment was originally downvoted and flagged.
         | 
         | To those who did this - what the fuck? Why? It's legitimate
         | comment. Fuck you for abusing that system.
        
       | simonw wrote:
       | This is really neat.
       | 
       | I noticed the diffusion-process.py demo was using matplotlib in a
       | window, but I figured it would be cute if it used a terminal UI
       | instead - so I had Claude Code convert it to use curses. Code and
       | demo GIF here:
       | https://gist.github.com/simonw/9033ebd8dd17b4c0ad101ddda7a54...
        
       | Majromax wrote:
       | The basic MLP block in this model uses a ReLU^2 activation
       | function (x <- ReLU(x)^2). That seems to be copied from the
       | nanochat project, and it's not present in nanoGPT. Is there some
       | documentation on the choice of this activation function?
        
       | mlmonkey wrote:
       | I'm curious: has there been any work done on generating embedding
       | vectors instead of discrete tokens via diffusion? What would that
       | look like? Please point me to some references. Thanks!
        
       | volodia wrote:
       | There is also this one that was released in October:
       | https://github.com/kuleshov/char-mdlm
        
       | embedding-shape wrote:
       | Fun project, easy to understand and nice looking results,
       | everything one could ask for! I played around with it locally,
       | did some optimizations of low hanging fruits without making it
       | much more complicated, and was gonna send over a PR. But then I
       | noticed there is no license attached to the project. What are
       | your plans regarding the licensing for this?
        
         | nathan-barry wrote:
         | Hey, I'll add the MIT licenses later today!
        
       ___________________________________________________________________
       (page generated 2025-11-14 23:00 UTC)