[HN Gopher] Can you save on LLM tokens using images instead of t...
___________________________________________________________________
Can you save on LLM tokens using images instead of text?
Author : lpellis
Score : 44 points
Date : 2025-11-01 22:34 UTC (7 days ago)
(HTM) web link (pagewatch.ai)
(TXT) w3m dump (pagewatch.ai)
| bikeshaving wrote:
| Does this mean we'll finally get empirical proof for the aphorism
| "a picture is worth a thousand words"?
|
| https://en.wikipedia.org/wiki/A_picture_is_worth_a_thousand_...
| heltale wrote:
| I suppose it's only worth 256 words at a time right now. ;)
|
| https://arxiv.org/abs/2010.11929
| estebarb wrote:
| The CALM paper https://shaochenze.github.io/blog/2025/CALM/
| says it is possible to compress 4 tokens in a single
| embedding, so... image = 4x256=1024 words > 1000 words. QED
| bikeshaving wrote:
| 2.4% relative error is not bad.
| pastor_williams wrote:
| Reminds me of Babbage making allowance for meter.
|
| """ ... it is said that he [Babbage]
| sent the following letter to Alfred, Lord Tennyson about
| a couplet in "The Vision of Sin": Every
| minute dies a man, Every minute one is born
| I need hardly point out to you that this calculation
| would tend to keep the sum total of the world's
| population in a state of perpetual equipoise, whereas it
| is a well-known fact that the said sum total is
| constantly on the increase. I would therefore take the
| liberty of suggesting that in the next edition of your
| excellent poem the erroneous calculation to which I refer
| should be corrected as follows: Every
| minute dies a man, And one and a sixteenth
| is born I may add that the exact figures are
| 1.167, but something must, of course, be conceded to the
| laws of metre.
|
| """ Charles Babbage and his Calculating
| Engines
| zahlman wrote:
| Wouldn't "one and a sixth" be more accurate in both
| respects?
| behnamoh wrote:
| how do you decompress all those 4 words from one token?
| estebarb wrote:
| Not from one token, from one embedding. Text contains a
| low amount of information: it is possible to compress a
| few token embeddings into a single tiken embedding.
|
| The how is variable. The calm paper seems to have used a
| MLP to compress from and ND input (N embeddings of size
| D) into a single D embedding and other for decompress
| them back
| HarHarVeryFunny wrote:
| The mechanism would be prediction (learnt during
| training), not decompression.
|
| It's the same as LLMs being able to "decode" Base64, or
| work with sub-word tokens for that matter, it just learns
| to predict that:
|
| <compressed representation> will be followed by (or
| preceded by) <decompressed representation>, or vice
| versa.
| floodfx wrote:
| Why are completion tokens more with image prompts yet the text
| output was about the same?
| Garlef wrote:
| "Thinking" Mode
| nunodonato wrote:
| it doesn't say that anywhere.
| cma wrote:
| Some multimodal models may have a hidden captioning step that
| may take completion tokens, others work on a fully native
| representation, and some do both I think.
| ashed96 wrote:
| In my experience, LLMs tend to take noticeably longer to process
| images than text.
| psadri wrote:
| I wonder if these stay in the prefix cache?
| weird-eye-issue wrote:
| It has to get the image data first, basically just IO time
| before processing it
___________________________________________________________________
(page generated 2025-11-08 23:02 UTC)