[HN Gopher] Lossless LLM 3x Throughput Increase by LMCache
___________________________________________________________________
Lossless LLM 3x Throughput Increase by LMCache
Author : lihanc111
Score : 130 points
Date : 2025-06-24 16:18 UTC (4 days ago)
(HTM) web link (github.com)
(TXT) w3m dump (github.com)
| lihanc111 wrote:
| Our team has built this open source project, LMCache, to reduce
| repetitive computation in LLM inference and make systems serve
| more people (3x more throughput in chat applications) and it has
| been used in IBM's open source LLM inference stack.
|
| In LLM serving, the input is computed into intermediate states
| called KV cache to further provide answers. These data are
| relatively large (~1-2GB for long context) and are often evicted
| when GPU memory is not enough. In these cases, when users ask a
| follow up question, the software needs to recompute for the same
| KV Cache. LMCache is designed to combat that by efficiently
| offloading and loading these KV cache to and from DRAM and disk.
|
| Ask us anything!
| dist-epoch wrote:
| How is it possible to do non-prefix KV cache? I was under the
| impression that the V for one token potentially depends on the
| V of all previous ones.
| da-x wrote:
| Yes, there's KV cache 'Blending' see [1].
|
| Future versions of LMCache are aiming to support this.
|
| [1] CacheBlend: Fast Large Language Model Serving for RAG
| with Cached Knowledge Fusion-
| https://arxiv.org/abs/2405.16444
| pama wrote:
| Is your aim targetting the inference at scale or
| specialized/new/simpler inference pipelines? Sglang and vllm
| have disaggregated prefix and decoding serving (eg
| https://docs.vllm.ai/examples/online_serving/disaggregated_s...
| or https://github.com/sgl-project/sglang/issues/3554 and
| https://github.com/sgl-project/sglang/issues/4655) -- could
| your solution enable a model-agnostic cache store/server or is
| that orthogonal to what you are trying to achieve?
| nativeit wrote:
| Has it been used _in_ IBM 's inference stack, or used _with_
| IBM 's inference stack? In other words, has this been merged
| into IBM's own repositories, or has someone just tested it
| using them?
| behnamoh wrote:
| > Our team
|
| So this is something that might in the future turning to a
| commercial product? something like Langchain and thousands of
| open source projects that started as "open source" but then
| ended up implementing proprietary features for a cost.
| Tokumei-no-hito wrote:
| i don't see anything wrong with that approach, do you?
| behnamoh wrote:
| Give it time and you'll come to my conclusion.
| 0xjunhao wrote:
| Hi, I had a quick question. Would it be correct to say the
| following?
|
| 1. For long inputs and short outputs, the inference can be
| arbitrarily number of times faster, as it avoids repeated KV
| computation.
|
| 2. Conversely, for short inputs and long outputs, it might be
| slightly slower, since loading and storing the KV cache are on
| the critical path of the execution.
| iLoveOncall wrote:
| Is this any different than prompt caching?
| smcleod wrote:
| Have you considered integrating it with the likes of llama.cpp?
| m3kw9 wrote:
| How would it work if a user wants to do 1 of n tries?
| kcorbitt wrote:
| Looks cool! With vLLM v1, prefix caching is enabled by default
| and seems quite performant. Is the advantage of LMCache the fact
| that you can offload to CPU and disk as well? How much is
| throughput/latency affected if you need to pull a large KV cache
| from disk/cpu instead of GPU RAM?
|
| Also, how realistic would it be to share the KV cache across vllm
| nodes within a data center? It would be really nice to be able to
| freely distribute requests to a pool of vLLM workers without
| worrying about prefix-aware routing, but maybe that isn't the
| right approach because moving the KV cache around would be too
| slow?
| guywhocodes wrote:
| This is exactly what llm-d is
| ekianjo wrote:
| wasn't this already implemented in llama.cpp?
| sgammon wrote:
| Hey LMCache team! Saw you guys at OSS N.A. but wasn't able to set
| aside time to say hello. We'd love to chat about collaborating.
| Is there an email we can reach out to?
| refulgentis wrote:
| Word to the wise:
|
| "Lossless 3x Throughput Increase" == "Cache _all inputs and
| output_ across _everyone_ , in RAM and _on disk_ , and if you
| assume the next request is covered by cache, its 3x faster!"
|
| I'm more surprised it's only advertised as 3x under those
| conditions: my llama.cpp wrapper does the same -- caching in RAM
| while running locally seems fine to me -- and when input is
| cached, TTFT is ~instantaneous, modulo any add'l prompt you add.
|
| I supposed it creates a little more distance, in that, instead of
| infinity times faster for latency, we measure _throughput_ , and
| then our speedup can be adjusted as desired by adjusting output
| length, and thus we can pick a more reasonable-sounding metric
| like 3x. (though, the GitHub README still frames it in terms of
| latency / TTFT)
| varispeed wrote:
| Sometimes I think the entire engineering profession collectively
| underwent a lobotomy. Techniques like caching partial computation
| results to avoid repeating expensive work were so basic a few
| decades ago that no one would have bothered to dignify them with
| a paper, let alone brand them with a fancy acronym and announce
| them like the second coming of Turing. Now we get breathless blog
| posts and community calls over the mind-blowing discovery that
| storing KV caches of repeated text speeds things up. Next we'll
| get a paper on using hash tables to look things up faster.
| Meanwhile, actual difficult problems in large-scale distributed
| inference and model interpretability get hand-waved so we can
| posture about reinventing memoisation. Tech never fails to take
| the obvious, put a bow on it, and sell it back to us as
| groundbreaking.
| vlovich123 wrote:
| Partial caching as a concept doesn't matter. The hard part is
| figuring out how to make it work for cross attention which sets
| up a data dependency for every entry on every preceding entry.
| So prefix caching of KV cache is brain dead easy. Computing a
| KV cache for random bits of text and then combining unrelated
| text in a way that makes the LLM still work coherently and
| correctly? That to me seems much harder.
|
| It seems to me like you're easily hand waving away a hard
| problem in a different part of the stack you're less familiar
| with.
| varispeed wrote:
| Let's be honest: it's fundamentally about analysing memory
| access patterns, spotting reuse opportunities, and
| orchestrating data flows. That's classic systems engineering.
| Useful, yes. Rocket science, no. The real joke is how the
| profession has sunk so low that anything beyond a trivial
| for-loop becomes a grounds for whitepapers, corporate
| branding, and breathless conference talks. In the past, we'd
| have quietly shipped this and moved on. Frankly, I'm
| surprised they haven't patented it yet.
| vlovich123 wrote:
| Caching and reuse broadly yes. Getting cross attention to
| work mathematically correctly by stitching the pre computed
| KV cache for snippets of text is not that unless you've
| redefined what classical systems engineering is.
|
| Again, the novelty is in getting cross attention to work
| correctly despite the fact that you're stitching together
| arbitrary caches together. It's akin to taking snippets of
| compressed portions of random compressed files and
| reconstructing a new correct plain text. That's obviously
| not possible but clearly this has been accomplished with
| the KV cache for _arbitrary_ models (ie not trained for it)
| despite the KV cache working like decompression where all
| the preceding bytes have to be computed correctly for the
| subsequent token to be correct.
| varispeed wrote:
| I get the argument, but let's be blunt: every serious
| cache system deals with consistency, partial reuse, and
| correctness. That's standard engineering - regardless of
| how much intimidating jargon you layer over it. Useful,
| sure. But watching the industry throw a circus around
| basic cache management, complete with papers and
| corporate branding, is exactly why so much of modern tech
| feels like a hype-driven clown show rather than a
| disciplined craft.
| notjoemama wrote:
| I've noticed this too. I wonder if it is the difference in
| experience levels. It feels odd seeing excitement at
| rediscovering a (what you and I think of as well-known)
| solution. To be fair, I was that kid at one time too. Still, it
| feels a bit like these more simple things ought to be taught at
| university so new grads can focus more on solving domain
| problems.
|
| I suppose, combine this with pressure from public or private
| investment, and the way to get ahead is to package anything
| into a prospect of revenue generation. I'm sure that's part of
| it too. Everything has to monetize because some business school
| graduate hasn't "made it" until they have a yacht like their
| ivy league friends.
|
| Eh, probably comes across as curmudgeonly or "who moved my
| cheese". But if there is an area that can improve this
| longstanding problem in tech, my guess is teaching the right
| skills and concepts at the collegiate level. And that's not a
| simple thing either.
|
| Edit > reading a bit more, this focuses on chat applications
| and seems to be a decent caching implementation tailored to
| that domain, of which, I'm guessing will allow AT&T and Verizon
| to save money on their gobsmackingly horrible AI chat bot in
| their mobile app. As an individual, it's unclear how this
| benefits me though. I don't think it does. ME: asks chat bot
| question about insurance coverage, CHATBOT: immediately serves
| canned response in no time about how that's covered in my
| individual insurance plan which I read more about on their
| website (pro-tip: no, I can't, those details are actually never
| on the website)
| nativeit wrote:
| It seems odd to me that so many of these projects are being
| launched by people who have only just discovered and/or joined
| HN. I'm worried this is just becoming LinkedIn for AI
| opportunists.
| parpfish wrote:
| I've got a side project that I may (someday) do a show HN with.
| However, I'd probably make a new account for that because the
| project is connected to my real name/portfolio and I don't want
| that connected with my pseudonymous comments here
| nativeit wrote:
| I considered that, but then why would anyone obfuscate this
| really very reasonable scenario by choosing another
| ostensibly pseudonymous username?
| fsmv wrote:
| [deleted]
| parpfish wrote:
| I imagine that this is a common problem and it could be
| another cool "unlockable" on HN, like the downvotes at 500
| karma.
|
| Once you get X karma or account age >Y years, you can make
| one anonymous submissions each quarter that comes from an
| non-user but still get some sort of "verified" badge that
| proves it comes from a legit user.
| refulgentis wrote:
| You nailed it IMHO.
|
| I quit my job at Google 2 years ago to do LLM stuff, was
| looking forward to having HN around, but discussions re: LLMs
| here are a minefield.
|
| Why?
|
| _Everyone_ knows at least a _little_ , and everyone has a
| strong opinion on it given the impact of it. People sharing
| stuff sell it _way_ high, and as with any new thing where
| people are selling, there 's a lot of skeptics. Then, throw in
| human bias towards disliking what seems like snark /
| complaining, so stuff with substance gets downvotes.
|
| SNR ratio is continually decreasing.
|
| Let's dig into why this one is weird:
|
| My work inferences using either 3P provider, which do caching,
| or llama.cpp, in which I do caching. (basically, picture it as
| there's a super expensive step that you can skip by keeping
| Map<input string, gpu state>)
|
| So I log into HN and see this and say to myself: 3x! throughput
| increase? This is either really clever or salesmanship, no way
| an optimization like that has been sitting around on the groud.
|
| So I read the GitHub, see it's just "write everyones inputs and
| outputs to disk, you can then use them to cobble together what
| the GPU state would be for an incoming request!", and write a
| mostly-polite comment below flagging "hey, this means writing
| everything to disk"
|
| Then I start replying to you...but then I throw away the
| comment, because I'm inviting drive-by downvotes. I.e. the
| minefield describe up top, and if you look like you're being
| mean, you'll eat downvotes, especially on a weekend.
|
| And to your average reader, maybe I just don't understand vLLM,
| and am taking it out in good hackers just pushing code.
|
| Then, when I go back, I immediately see a comment from someone
| who _does_ use vLLM noting it already does caching.
|
| Sigh.
| nativeit wrote:
| Thanks for sharing. You certainly aren't alone in your
| sentiments. I am seeing similar trends in arXiv submissions,
| as it seems it has become something of a means to inflate the
| value of one's own product(s) with a veneer of academic
| rigor. There seems to be a S.O.P. emerging for AI tools that
| follows many of the same trends as the less-than-reputable
| blockchain/crypto projects.
| pama wrote:
| I had related questions and checked out the project a bit
| deeper though I havent tested it seriously yet. The project
| did start work over a year ago based on relevant papers,
| before vllm or sglang had decent solutions; it might still be
| adding performance in some workflows though I havent tested
| it and some of the published measurements in the project are
| now stale. Caching LLM kv-cache to disk or external memory
| servers can be very helpful at scale. Cache management and
| figuring out cache invalidation is hard anyways and I am not
| sure at what level a tight integration with inference servers
| or specialized inference popelines can help vs a lose
| coupling that could advance each component separately. It
| would be nice if there were decent protocols used by all
| inference engines to help this decoupling.
| nativeit wrote:
| I'll just be unambiguous about this:
|
| > Please don't use HN primarily for promotion. It's ok to post
| your own stuff part of the time, but the primary use of the
| site should be for curiosity.
|
| https://news.ycombinator.com/newsguidelines.html
| Aurornis wrote:
| A couple months ago another project claimed to have sped up
| llama.cpp (IIRC) on the front page of HN, from another green
| name account.
|
| It gathered hundreds of GitHub stars and was on the front page
| all day. When some of us finally had time to look at the code
| we discovered they didn't invent anything new at all. They took
| some existing command line options for llama.cpp and then
| changed the wording slightly to make them appear novel.
|
| The strangest part was that everyone who pointed it out was
| downvoted at first. The first comment to catch it was even
| flagged away! You couldn't see it unless you had showdead
| turned on.
|
| At first glance I don't see this repo as being in the same
| category, though the "3X throughput increase" claim is very
| clearly dependent on the level of caching for subsequent
| responses and the "lossless" claim doesn't hold up as analyzed
| by another top-level comment.
|
| I think AI self-promoters have realized how easy it is to game
| Hacker News and GitHub stars if you use the right wording. You
| can make some big claims that are hard to examine in the quick
| turnaround times of a Hacker News front page cycle.
| wg0 wrote:
| Seems like snake oil to me. I mean lacks clear explanation of how
| exactly it works if at all.
| ahmedhawas123 wrote:
| Like this a lot and thanks for making it open source. Does this
| support ollama today? I only saw vLLM
| jbentley1 wrote:
| Is this the same as the prompt caching that other API's
| (Anthropc, OpenAI, etc) have had, just open source and for vLLM?
| alyxya wrote:
| I skimmed over a couple of the papers referenced to get an idea
| of what optimizations LMCache is doing.
|
| * KV cache compression - compressing the bytes of the KV cache,
| taking advantage of patterns in the KV cache and with dynamic
| levels of compression
|
| * KV cache blending - concatenating the KV caches of multiple
| reused prompts with minimal KV cache recomputation for use cases
| like RAG, where it's more performant than the standard lossless
| KV cache prefix optimization, and gives better results than
| naively concatenating the KV caches for the reused prompts
|
| These optimizations are pretty cool and different than the
| standard KV cache optimizations. The title saying lossless seems
| misleading though.
| tucnak wrote:
| "Blending," or translating arbitrary substrings to prefixes, is
| a real curious one, & likely become a prerequisite for running
| dataset-scale LLM inferences at scale.
|
| See https://arxiv.org/abs/2405.16444v3
|
| > To speed up the prefill of the long LLM inputs, one can pre-
| compute the KV cache of a text and re-use the KV cache when the
| context is reused as the prefix of another LLM input. However,
| the reused text chunks are not always the input prefix, which
| makes precomputed KV caches not directly usable since they
| ignore the text's cross-attention with the preceding texts.
| Thus, the benefits of reusing KV caches remain largely
| unrealized.
|
| > This paper tackles just one challenge: when an LLM input
| contains multiple text chunks, how to quickly combine their
| precomputed KV caches in order to achieve the same generation
| quality as the expensive full prefill (i.e., without reusing KV
| cache)? [..] We present a scheme that reuses the pre-computed
| KV caches, regardless prefix or not, and selectively recomputes
| the KV values of a small subset of tokens to partially update
| each reused KV cache.
|
| I had recently touched on benefits of compute-in-network for KV
| cache management https://news.ycombinator.com/item?id=44371227
| largely making arguments contra Bluefield. The CacheBlend
| authors note that the delay from recomputing some tokens can be
| hidden by pipelining it with KV loads. Note that the various
| systolic array/NoC architectures are well-suited for
| accelerating string matching tasks. A compute-in-network FPGA
| could therefore manage the entire process: identify viable
| chunks by indexing and matching of the hot substrings, prefetch
| the corresponding KV caches from network storage, and stitch up
| a new prefix before passing it to the primary inference
| hardware. It may as well be one of those weird cases where
| hard-coding the algorithm is possible in theory, but
| intractable in practice--because the optimal paths would be
| highly-dependent on topology.
|
| Nobody wants one-trick hardware.
|
| In view of Xilinx acquisition, AMD's death in the AI space
| appears to be greatly exaggerated!
___________________________________________________________________
(page generated 2025-06-28 23:01 UTC)