[HN Gopher] Can you save on LLM tokens using images instead of t...
       ___________________________________________________________________
        
       Can you save on LLM tokens using images instead of text?
        
       Author : lpellis
       Score  : 44 points
       Date   : 2025-11-01 22:34 UTC (7 days ago)
        
 (HTM) web link (pagewatch.ai)
 (TXT) w3m dump (pagewatch.ai)
        
       | bikeshaving wrote:
       | Does this mean we'll finally get empirical proof for the aphorism
       | "a picture is worth a thousand words"?
       | 
       | https://en.wikipedia.org/wiki/A_picture_is_worth_a_thousand_...
        
         | heltale wrote:
         | I suppose it's only worth 256 words at a time right now. ;)
         | 
         | https://arxiv.org/abs/2010.11929
        
           | estebarb wrote:
           | The CALM paper https://shaochenze.github.io/blog/2025/CALM/
           | says it is possible to compress 4 tokens in a single
           | embedding, so... image = 4x256=1024 words > 1000 words. QED
        
             | bikeshaving wrote:
             | 2.4% relative error is not bad.
        
               | pastor_williams wrote:
               | Reminds me of Babbage making allowance for meter.
               | 
               | """                   ... it is said that he [Babbage]
               | sent the following letter to Alfred, Lord Tennyson about
               | a couplet in "The Vision of Sin":                   Every
               | minute dies a man,              Every minute one is born
               | I need hardly point out to you that this calculation
               | would tend to keep the sum total of the world's
               | population in a state of perpetual equipoise, whereas it
               | is a well-known fact that the said sum total is
               | constantly on the increase. I would therefore take the
               | liberty of suggesting that in the next edition of your
               | excellent poem the erroneous calculation to which I refer
               | should be corrected as follows:                   Every
               | minute dies a man,              And one and a sixteenth
               | is born              I may add that the exact figures are
               | 1.167, but something must, of course, be conceded to the
               | laws of metre.
               | 
               | """                   Charles Babbage and his Calculating
               | Engines
        
               | zahlman wrote:
               | Wouldn't "one and a sixth" be more accurate in both
               | respects?
        
             | behnamoh wrote:
             | how do you decompress all those 4 words from one token?
        
               | estebarb wrote:
               | Not from one token, from one embedding. Text contains a
               | low amount of information: it is possible to compress a
               | few token embeddings into a single tiken embedding.
               | 
               | The how is variable. The calm paper seems to have used a
               | MLP to compress from and ND input (N embeddings of size
               | D) into a single D embedding and other for decompress
               | them back
        
               | HarHarVeryFunny wrote:
               | The mechanism would be prediction (learnt during
               | training), not decompression.
               | 
               | It's the same as LLMs being able to "decode" Base64, or
               | work with sub-word tokens for that matter, it just learns
               | to predict that:
               | 
               | <compressed representation> will be followed by (or
               | preceded by) <decompressed representation>, or vice
               | versa.
        
       | floodfx wrote:
       | Why are completion tokens more with image prompts yet the text
       | output was about the same?
        
         | Garlef wrote:
         | "Thinking" Mode
        
           | nunodonato wrote:
           | it doesn't say that anywhere.
        
         | cma wrote:
         | Some multimodal models may have a hidden captioning step that
         | may take completion tokens, others work on a fully native
         | representation, and some do both I think.
        
       | ashed96 wrote:
       | In my experience, LLMs tend to take noticeably longer to process
       | images than text.
        
         | psadri wrote:
         | I wonder if these stay in the prefix cache?
        
         | weird-eye-issue wrote:
         | It has to get the image data first, basically just IO time
         | before processing it
        
       ___________________________________________________________________
       (page generated 2025-11-08 23:02 UTC)