[HN Gopher] From 300KB to 69KB per Token: How LLM Architectures ...
       ___________________________________________________________________
        
       From 300KB to 69KB per Token: How LLM Architectures Solve the KV
       Cache Problem
        
       Author : future-shock-ai
       Score  : 69 points
       Date   : 2026-03-28 22:42 UTC (3 days ago)
        
 (HTM) web link (news.future-shock.ai)
 (TXT) w3m dump (news.future-shock.ai)
        
       | az09mugen wrote:
       | Unrelated, but 69KB is how much RAM Voyager 1 has.
        
         | gregman1 wrote:
         | Voyager as a token of curiosity
        
       | LuxBennu wrote:
       | good overview of the architecture side but worth mentioning
       | there's another axis that stacks on top of all of this: you can
       | quantize the kv cache itself at inference time. in llama.cpp you
       | can run q8 for keys and q4 for values and it cuts cache memory
       | roughly in half again on top of whatever gqa or mla already saves
       | you. i run qwen 70b 4-bit on m2 max 96gb and the kv quant is what
       | actually made longer contexts fit without running out of unified
       | memory. keys need more precision because they drive attention
       | scores but values are way more tolerant of lossy compression, so
       | the asymmetry works out.
        
         | suprjami wrote:
         | Some models really suffer badly from KV quantisation. You can
         | also take a speed hit using dissimilar K and V types.
         | 
         | TurboQuant seems to be the next big thing in context memory
         | usage. Polar coordinates achieving ~5x reduction in memory
         | usage with minimal/no quality loss, and even a slight speedup
         | in some cases.
        
       | coppsilgold wrote:
       | There are also interesting approaches to more directly compress a
       | large document or an entire codebase into a smaller set of tokens
       | without getting the LLM to wing it. For example, Cartridges:
       | <https://hazyresearch.stanford.edu/blog/2025-06-08-cartridges>
       | 
       | They basically get gradient descent to optimize the KV cache
       | while freezing the network.
        
       ___________________________________________________________________
       (page generated 2026-03-31 23:00 UTC)