[HN Gopher] From 300KB to 69KB per Token: How LLM Architectures ...
___________________________________________________________________
From 300KB to 69KB per Token: How LLM Architectures Solve the KV
Cache Problem
Author : future-shock-ai
Score : 69 points
Date : 2026-03-28 22:42 UTC (3 days ago)
(HTM) web link (news.future-shock.ai)
(TXT) w3m dump (news.future-shock.ai)
| az09mugen wrote:
| Unrelated, but 69KB is how much RAM Voyager 1 has.
| gregman1 wrote:
| Voyager as a token of curiosity
| LuxBennu wrote:
| good overview of the architecture side but worth mentioning
| there's another axis that stacks on top of all of this: you can
| quantize the kv cache itself at inference time. in llama.cpp you
| can run q8 for keys and q4 for values and it cuts cache memory
| roughly in half again on top of whatever gqa or mla already saves
| you. i run qwen 70b 4-bit on m2 max 96gb and the kv quant is what
| actually made longer contexts fit without running out of unified
| memory. keys need more precision because they drive attention
| scores but values are way more tolerant of lossy compression, so
| the asymmetry works out.
| suprjami wrote:
| Some models really suffer badly from KV quantisation. You can
| also take a speed hit using dissimilar K and V types.
|
| TurboQuant seems to be the next big thing in context memory
| usage. Polar coordinates achieving ~5x reduction in memory
| usage with minimal/no quality loss, and even a slight speedup
| in some cases.
| coppsilgold wrote:
| There are also interesting approaches to more directly compress a
| large document or an entire codebase into a smaller set of tokens
| without getting the LLM to wing it. For example, Cartridges:
| <https://hazyresearch.stanford.edu/blog/2025-06-08-cartridges>
|
| They basically get gradient descent to optimize the KV cache
| while freezing the network.
___________________________________________________________________
(page generated 2026-03-31 23:00 UTC)