[HN Gopher] DeepSeek Open Source FlashMLA - MLA Decoding Kernel ...
___________________________________________________________________
DeepSeek Open Source FlashMLA - MLA Decoding Kernel for Hopper GPUs
Author : helloericsf
Score : 403 points
Date : 2025-02-24 01:37 UTC (21 hours ago)
(HTM) web link (github.com)
(TXT) w3m dump (github.com)
| helloericsf wrote:
| X:https://x.com/deepseek_ai/status/1893836827574030466 BF16
| support Paged KV cache (block size 64) 3000 GB/s memory-bound &
| 580 TFLOPS compute-bound on H800
| WithinReason wrote:
| That's 90% bandwidth efficiency and 60% compute efficiency
|
| https://www.nvidia.com/en-us/data-center/h100/
| helloericsf wrote:
| They don't have h100. wink,wink.
| rfoo wrote:
| They have H800s which have exactly same memory bandwidth
| and max FLOPS.
| pk-protect-ai wrote:
| What about NVLink? Does it plays a role here?
| rfoo wrote:
| For FlashMLA? No. The code here runs on one GPU only and
| do not have a builtin communication part.
| deyiao wrote:
| I heard their inferencing framework is way lower than typical
| deployment methods. Can this be verified from that open-source
| project? How does it stack up against vllm or llama.cpp
| helloericsf wrote:
| What do you mean by "lower"? To my understanding, they will
| open 5 infra related repos this week. Let's revisit your
| comparison question on Friday.
| find0x90 wrote:
| I don't see any use of PTX, might be in one of the other repos
| they plan to release.
| DesiLurker wrote:
| right, I think PTX use is a bigger deal than its getting
| coverage for. this opens an opening for other vendors to get
| their foot in with PTX to LLVM-ir translation for existing
| cuda kernels.
| reissbaker wrote:
| By "lower" you mean cheaper/better?
|
| I suspect it's much higher throughput than vLLM, which in turn
| is much higher throughput than llama.cpp. The MLA kernel they
| just open-sourced seems to indicate that, although we'll see
| how it does in third party benchmarks on non-hobbled GPUs vs
| FlashAttention. They only released the BF16 version -- whereas
| most people, including DeepSeek themselves, serve in FP8 -- so
| it might not be immediately useful to most companies quite yet,
| although I imagine there'll be FP8 ports soon enough.
| nialv7 wrote:
| i think they meant lower level.
| bee_rider wrote:
| It seems hard to guess. Could be lower level, lower
| performance, or lower compute cost.
| feverzsj wrote:
| Maybe. Apple ditched them in China, because their infra can't
| handle large scale users.
| helloericsf wrote:
| Don't think the decision is based on infra, or any technical
| reasons. It's more on the service support side. How a
| 200-person company supports 44M iPhone users in China?
| chvid wrote:
| Is that true? I thought Apple was going to use their own
| infrastructure.
| tw1984 wrote:
| deepseek doesn't have any experience on support a 50 million
| user base. that was the reason cited by apple a few weeks
| ago.
| mohsen1 wrote:
| I'm confused. Wasn't there sanctions against Chinese companies
| about Hopper GPUs? Are they just admitting that they had access
| to H100 against the US sanctions?!
| thot_experiment wrote:
| Just the H100, the H800 is a region-specific version of the
| card for china with shitty nvlink bandwidth which makes it
| rougher for making big clusters, but deepseek was able to
| mitigate the impact of that by being clever (rumored to have
| made significant use of PTX assembly instead of just using
| CUDA, we'll probably find out in the releases this week)
| Tiberium wrote:
| H800 is the export variant that they had access to. They
| directly reference it in the repo:
|
| >Achieving up to 3000 GB/s in memory-bound configuration and
| 580 TFLOPS in computation-bound configuration on H800 SXM5,
| using CUDA 12.6.
| ahofmann wrote:
| It isn't illegal for chinese companies to buy H100 cards. It is
| illegal for USA companies to sell them to China. So the "admit"
| part wouldn't be on Chinas side.
| jofzar wrote:
| It's also totally legal to sell h100 cards to a country that
| is very close to China.
|
| Unrelated, it's always impressed me how Singapore buys 15% of
| the world's h100's. Really is the AI development capital of
| the world.
| xbmcuser wrote:
| Not really Singapore is a trading hub a lot of multi
| national companies have regional offices or head offices in
| Singapore so if the head office buys anything for any where
| the purchase will show up as Singapore. Despite Nvidia
| showing such a large revenue from Singapore actual number
| of gpu shipped to Singapore is not that high. Not that some
| of the gpus are not going China but their is a valid reason
| for the Nvidia Singapore revenue numbers.
|
| https://www.tomshardware.com/tech-industry/deepseek-gpu-
| smug...
| samvher wrote:
| I can't tell if you're insinuating that Singapore is a
| pass-through for H100's heading towards China or whether
| there is some significant development taking place in
| Singapore that I'm unaware of?
| dist-epoch wrote:
| > Singapore plays a vital role in Nvidia's global
| business, accounting for 22% of its revenue as of Q3
| FY2025, up from 9% in Q3 FY2023 when the first
| significant restrictions on AI GPU sales to Chinese were
| introduced
| eagleislandsong wrote:
| I suggest that you read this comment, which explains why
| your quote is misleading:
| https://news.ycombinator.com/item?id=43159362
| janalsncm wrote:
| It's funny how this claim is able to make the rounds. I
| originally heard it here:
| https://m.youtube.com/watch?v=_1f-o0nqpEI
|
| Singapore is the billing location, not the shipping
| location, which makes sense because they're the HQ of a lot
| of companies in the region.
| amelius wrote:
| Also breaking the law to growth-hack happens all the time,
| see Uber.
| kridsdale1 wrote:
| And the British East India Company
| WiSaGaN wrote:
| H20 is a Hopper GPU, and they are allowed to be sold in China.
| feverzsj wrote:
| The secret ingredient is smuggling.
| tasuki wrote:
| I'd be very careful when using that word in this situation.
| If China wants X, and another country has X, _who are you_ to
| say they shouldn 't trade with each other?
| randomNumber7 wrote:
| Donald Trump?
| blackeyeblitzar wrote:
| Why does anyone need to be careful using that word? What a
| bizarre way to try to intimidate someone over speech.
|
| Another country has X because they were expected (in the
| terms of their purchase) to not sell it to an adversary. So
| yes they're supposed to honor that agreement and are not
| supposed to trade that particular thing X with each other.
| Not doing so invites sanctions and other consequences. Is
| it worth the risk just to do business with a dictatorship?
| Probably not.
| defrost wrote:
| If free citizens in the USofA have {X} and China has
| sanctioned Germany from having {X} should the free
| citizens of the USofA honor that agreement they made with
| China to not sell to Germany when they acquired {X} from
| China?
|
| How about if they got {X} from Mexico ( _who got it from
| Agnes_ .. ) ?
| Keyframe wrote:
| Some purchases come with strict protocols coded into
| contracts. Try buying F-35 and selling it to China, for
| example; See what happens. Other risk you not being able
| to purchase for yourself anymore and possible sanctions.
| H100 and others are under export control, I'm just not
| sure if it's an explicit export control or automatic,
| like what famously made PowerMac G4 a weapon export. I
| found a source there was an executive order for hardware
| exceeding 1e26 floating point operations or 1e23 integer
| operations. In any case, if an item is under export
| control that means paperwork and, if you're eligible to
| purchase, paperwork includes you also signing what you
| can and cannot do with the item purchased.
| yieldcrv wrote:
| I feel the same about capital controls and crypto
|
| People say "it's used for money laundering" as if we're
| supposed to be on China's side about restricting people's
| ability to move money out of the country over certain
| amounts
|
| Like, oh you're against freedom from a repressive regime?
| Or oh you're only against it when it's the American
| government restricting US citizens flow of capital? like
| I'm confused, pick a lane
|
| Capital controls are obsoleted under any context
| sangnoir wrote:
| It's not intimidation, its merely correcting and an
| inappropriate usage of a word. What exactly do you think
| smuggling is?
| 55555 wrote:
| Smuggling is normally thought of as hiding something when
| crossing a border/checkpoint. In this case, it would
| simply be nvidia violating US sanctions. The goods would
| have never entered or exited the USA so it's a strange or
| incorrect use of the word smuggling.
| quantum_state wrote:
| We should forget about the sanction BS ... it damages US
| industry when it has money to make while motivating others
| to be more self reliant and build the product to compete
| ...
| 7952 wrote:
| Do you think that would be morally wrong? Honest question.
| amelius wrote:
| No, especially considering that they open sourced
| everything. (not OP)
|
| Also, they could have outsourced the computation to a
| subsidiary company in the US, I suppose.
| jonplackett wrote:
| Can everyone stop downvoting people just for asking questions -
| this isn't Stack Overflow!
| behnamoh wrote:
| Open AI is back!
| echelon wrote:
| The real "Open" AI.
| fsndz wrote:
| DeepSeek is just the gift that keeps on giving. I now agree
| with people who say open source AI will win:
| https://open.substack.com/pub/transitions/p/deepseek-is-
| comi...
| baq wrote:
| Open sourcing is the runner-up's way to ensure the current
| best player doesn't steal the whole market. The elephant in
| the room is obviously the cluster size required, it hardly
| matters for normal people that the weights are free. We
| needed more efficiency breakthroughs.
| PeterStuer wrote:
| It matters a lot, even if you never intend to run it
| yourself or look at the code.
|
| It means that people can and will provide this service,
| and 1000's will build on this and make offers that you
| can use in either a commodity base market, or with a
| specific niche target.
|
| It means regulatory capture and control will be much,
| much harder to execute.
|
| It means AI might continue to be a benefit also to you
| rather than just a way to control, propagandize and
| exploit you.
| fsndz wrote:
| absolutely on point!
| helsinkiandrew wrote:
| > .. it hardly matters for normal people that the weights
| are free. We needed more efficiency breakthroughs.
|
| That atleast allows other companies/research labs to
| develop competing cutting edge LLM technology and come up
| with efficiency breakthroughs. The alternative is for the
| tech to be hidden inside OpenAI and FANGs or released as
| old versions.
| echelon wrote:
| Today's H100 cluster models are tomorrow's computing at
| the edge models.
|
| With the next wave of investment targeting local on-
| device robotics, I'm way more bullish about local AI than
| vertical SaaS AI.
| rvz wrote:
| This is the minimum bar that I expect very elite programmers
| should be striving for in the age of AI and DeepSeek should be
| studied as an example and this is the only just the first of many
| projects from them.
|
| There is an extremely high chance (in fact a 99.9% chance) that
| an AI did not build this and the ones who are able to build or
| adapt projects like this which are deep into hardware systems
| will be the most sort after.
|
| Not the horrendous JS or even TS slop across GitHub that is
| extremely easy for an AI to generate correctly.
|
| You've got until 2030 to decide. And my advice is to study the
| codebases of pytorch (backends), DeepSeek, tinygrad and ggml.
| beernet wrote:
| LLM generated comments are so 2024
| BoorishBears wrote:
| Nothing about that comment implies it's LLM generated, and
| it's bizzare how it's being received since it's a pretty
| reasonable take.
| rnewme wrote:
| I don't find it a reasonable take, it's like saying
| stackoverflow.com is taking developer jobs by making it
| easy to code, we better develop new stackoverflow.com
| jbm wrote:
| It's an interesting opinion, but I read the exact same opinions
| about JS developers in 2008 too.
|
| I do agree that if you are "only" a developer, you will have to
| be in some sort of tightly defined niche, and how long those
| niches survive is anyone's guess.
| KeplerBoy wrote:
| What do you mean with "only" developer? Someone who just
| knows how to code when given a spec but lacking domain
| knowledge (in this case ai math and hardware optimization)
| and larger context?
| jbm wrote:
| Personally, I don't think having deep domain knowledge is
| as important. However, being able to write the spec based
| on interactions with the client / customer / stakeholders
| is. (The AI Math and hardware optimization "never" being
| doable by an AI seems like an arbitrary distinction to
| justify one's choices.)
|
| Incidentally, I put the word "only" in quotes because I
| morally and aesthetically appreciate the strength of
| someone who can write to spec. I have no interest in
| demeaning the effort it takes to do so. I have worked with
| supposedly senior developers who ignore specs completely,
| even when the specs are done by a technical person and
| include details / unit tests.
| WithinReason wrote:
| AI is already writing optimized GPU code:
|
| https://sakana.ai/ai-cuda-engineer/
| mirekrusin wrote:
| Comments around that page suggest it's more of a facepalm
| than anything else.
| CamperBob2 wrote:
| x2 speed increase for ggml by optimizing SIMD:
| https://github.com/ggml-org/llama.cpp/pull/11453
|
| "99% written by DeepSeek-R1" according to the author.
| rfoo wrote:
| Speaks more about how many low hanging fruits remaining
| in "NOOOOO I DON'T WANT TO DOWNLOAD 200MiB PYTORCH I'D
| BETTER REINVENT THE WHEEL"-gang inference stacks.
|
| To be fair torch didn't try very hard optimizing on CPU
| either.
| badsectoracula wrote:
| FWIW as someone who "NOOO DOESN'T WANT TO DOWNLOAD
| 200MB[0] PYTORCH"s i'm glad for those who make
| alternative minimal/no-dependency stacks that are based
| on C/C++, like ggml.
|
| [0] 200MB is actually a very generous number, i tried to
| download some AI thing via pip3 the other day and it
| wanted 600MB or so of CUDA stuff. Meanwhile i do not even
| have an Nvidia GPU.
| rfoo wrote:
| The wheel of CPU-only PyTorch 2.6.0 for Python 3.12 is
| ~170MiB in size.
|
| It is indeed pretty silly that's not the default and you
| have to go to https://pytorch.org/get-started/locally/,
| copy the argument `--index-url
| https://download.pytorch.org/whl/cpu` to install CPU-only
| torch. But the alternative would be having the worlds
| scientists wondering why they can't use their GPUs after
| `pip install torch` so /shrug
| wrsh07 wrote:
| But as a response to the parent saying "LLMs will be
| great at ts/js slop but not for infra" it's quite
| reasonable to say: here's an example of someone applying
| it to backend optimizations today.
|
| Fwiw, there are always many attempts at optimizing code
| (assembly etc). This is good! Great to try new
| techniques. However, you get what you constrain. So I've
| seen optimized code that drops checks that the compiler
| authors say are required in the standard. So, if you
| don't explicitly tell your optimizer "this is a case I
| care about, this is the desired output" it will ignore
| that case.
|
| Did we find a faster implementation than the compiler
| creates? Well, I mean, sure, if you don't know why the
| compiler is doing what is doing
| menaerus wrote:
| I agree that DeepSeek continues to prove themselves as a great
| example of engineering but the number of job positions
| requiring this type of knowledge IME is typically very very low
| so I am not sure if this would be the right advice to follow.
| Though I wish it was different.
| PeterStuer wrote:
| Honest question:
|
| Do you feel GenAI coding is substantially different from the
| lineage of 4GL to 'low code' approaches?
|
| Reason I'm asking is because despite all promises al suffered
| from what Spolsky coined the 'leaky abstraction' problem.
|
| Once something goes wrong, the user is left without recourse in
| a sea of additional complexity created by the tooling that was
| meant to not have to deal with it in the first place.
|
| My own opinion is that GenAI _is_ different because of (a) its
| recursive reflexive potential (you can use the tool itself to
| help you past the failure) and (b) it shifts the input out of
| the necessity for algorithmic /systemic thinking (which may
| come as a surprise to the audience here but my experience has
| taught me is alien to dare I say the majority of people).
|
| Now don't get me wrong. We have not reached the point where
| (a)+(b) make it to where you don't need application layer devs,
| but we are definitely seeing some progress.
|
| As for going deeper into the stack to "escape" AI, I would
| venture that is probably a non starter as the deeper you go the
| more constrained the domain is, so your escape strategy relies
| on AI reasoning making little progress, where AI reasoning has
| always been more successful in smaller well defined spaces.
| rob_c wrote:
| Yeah, you're hitting the nail on the head. Low tier coding work
| can be reduced and the high end developers can now avoid boiler
| plate type coding problems and get back to high level work at
| reengineering complex frameworks.
|
| Yes, this unfortunately does mean a reduction in the less
| skilled workforce, but frankly that's an on the whole good
| thing. Does anyone really enjoy writing and testing boilerplate
| day in day out for low pay, it's the same as the old white
| collar pushing paper around until retirement...
| m3kw9 wrote:
| MHGA making hopper great again
| eigenvalue wrote:
| Nice, probably saved a bunch of FANG devs a lot of hours of work
| trying to knock this off.
| nicce wrote:
| There were likely some startups that tried to sell the same
| thing...
| anon389r58r58 wrote:
| You mean like Modular?
| nicce wrote:
| Or Silo AI (as an example of why) :
| https://www.silo.ai/blog/amd-to-acquire-silo-ai-to-expand-
| en...
| refibrillator wrote:
| vLLM supports MLA for Deepseek models as of 3 weeks ago. 3x
| higher generation throughput and 10x token memory capacity.
|
| https://github.com/vllm-project/vllm/releases/tag/v0.7.1
|
| MHA is still faster in low QPS regime apparently.
|
| https://neuralmagic.com/blog/enhancing-deepseek-models-with-...
|
| Also published this month was theoretical proof showing that for
| the same KV Cache overhead, MLA consistently offers greater
| expressive power than GQA. Furthermore, widely used GQA-based
| pre-trained models (e.g. LLaMA, Qwen, Mixtral) can be converted
| into MLA-based models.
|
| https://arxiv.org/pdf/2502.07864
| albertzeyer wrote:
| I also just read that paper. But I wonder, even though MLA is
| strictly more powerful, do you really gain by that in
| experiments? This paper doesn't really do too much experimental
| comparisons. GQA on the other side should still be faster (no
| need to an extra linear transformation).
| menaerus wrote:
| Pretty significant improvements. However, my back on the napkin
| math suggests that MLA, FlashAttention and similar
| optimizations will provide the benefits only when memory access
| time dominates the compute in attention implementation? Those
| would be the prefill-phase (or TTFT) and training (when
| batch_size >> 1) but not the decode phase (inference)?
| rfoo wrote:
| You've got it backwards. After FlashAttention, it's the
| decoding part being bound mainly by memory access. With FA as
| long as you have enough batch size you can push
| training/prefill to be compute-bound.
| menaerus wrote:
| I don't think I got it backwards, I believe what I said is
| correct - FA does not improve inference time.
|
| From the authors of FlashAttention:
|
| > This [decoding] operation has been optimized with
| FlashAttention (v1 and v2 recently) in the training case,
| where the bottleneck is the memory bandwidth to read and
| write the intermediate results
|
| And then they continue with:
|
| > However, these optimizations don't apply directly to the
| inference case, because the bottlenecks are different. For
| training, FlashAttention parallelizes across the batch size
| and query length dimensions. During inference, the query
| length is typically 1 ... With a batch size of 1,
| FlashAttention will use less than 1% of the GPU!
|
| And then they come up with a different proposal,
| FlashDecoding, that optimizes for inference time:
|
| > Our new approach Flash-Decoding is based on
| FlashAttention, and adds a new parallelization dimension:
| the keys/values sequence length. It combines the benefits
| of the 2 approaches from above. Like FlashAttention, it
| stores very little extra data to global memory, however it
| fully utilizes the GPU even when the batch size is small,
| as long as the context length is large enough.
|
| Link:
| https://crfm.stanford.edu/2023/10/12/flashdecoding.html
| rfoo wrote:
| That's correct, because FA can't turn inference time from
| memory-access bound into compute-bound. But your claim on
| that decoding is compute-bound is plainly wrong.
|
| FA, compared to naive implementation, made training /
| prefill (i.e. when you can have multiple tokens in the
| same sequence visible) compute-bound instead of memory-
| access bound.
|
| So, currently, on MHA/GQA, with Flash Attention,
| training/prefill is compute-bound, whereas decoding is
| memory-access-bound.
|
| Before FA, both prefill / decode are bound by memory-
| access. FA solved the problem of training/prefill. But
| because kvcache is large, decoding is inherently bound by
| memory-access.
|
| Our goal is always to make everything compute-bound.
| rfoo wrote:
| ... and batching does not help, you batch more requests
| and get more kvcache to load, still memory-access bound.
|
| MLA made it possible to cache a smaller form of k/v,
| mitigating (but not completely solve, on shorter context
| & smaller batches it's still memory-access bound) the
| problem.
| menaerus wrote:
| > But your claim on that decoding is compute-bound is
| plainly wrong.
|
| I did not say anything like that? What I said is that
| FlashAttention and arguably MLA will not make any
| significant gains in the inference time. And this is
| true.
|
| Also, FWIW there are certainly model shapes that are
| compute-bound in the decode phase so saying that decoding
| is universally inherently bound by memory access is what
| is plain wrong, if I were to use your dictionary.
| rfoo wrote:
| Apologize if I got it wrong, but:
|
| > MLA, FlashAttention and similar optimizations will
| provide the benefits only when memory access time
| dominates
|
| > Those would be [...] not the decode phase
|
| This does sound like you are saying that memory access
| time does NOT dominate during the decode phase. But it
| does.
|
| Reading your quotes, it looks like maybe you are talking
| about GPU utilization issues? (i.e. not launching enough
| threads). Due to the parallelization strategy of the
| original FA it indeed does not even keep the GPU busy if
| q*bs is too small. But this is not an inherent limitation
| of FA-style kernels and can be solved and people did
| solve it. Or you simply batch more. Now you can keep the
| GPUs busy at 100% waiting for memory access, but memory
| access time still dominates, hence "memory-access-bound".
| And here comes MLA.
|
| > FWIW there are certainly model shapes that are compute-
| bound in the decode phase
|
| Yeah. But so far all I read don't really work ("work"
| means being at least just slightly worse than
| alternatives) under same wall-clock time compute budget.
| Do you have any pointer to a working example, even on
| smaller 3B-ish models?
| menaerus wrote:
| > This does sound like you are saying that memory access
| time does NOT dominate during the decode phase. But it
| does.
|
| Let's take llama3-8B for an example. GFLOPS needed for
| self-attention per-layer per-token is roughly 0.15
| GFLOPS. For simplicity reasons let's assume that we store
| all our weights in FP8 precision, then our load memory-
| bandwidth required for the same is 0.05 GB. Store memory-
| bandwidth is negligible. If we expand this further to a
| 1k tokens context, this becomes ~180 GFLOPS and ~0.35 GB
| per-layer per-1k-ctx.
|
| Assuming that our HW is H100, is this compute-bound or
| memory-bound?
| rfoo wrote:
| You need to load cached k/v tensor, in addition to
| weights. It's going to take me some minutes to find out
| what's wrong in this napkin math. Will edit or reply this
| comment later.
| menaerus wrote:
| Re-computing everything every time is the worst-case
| scenario and which is why I included it in the example
| (1k tokens). In that case, KV-cache is obviously set to 0
| but it is also obvious that it is a much worse
| alternative than using the KV-cache. Which is pretty much
| the reason why we have the KV-cache. Therefore the
| argument about loading the cached tensors doesn't make a
| difference at all.
|
| > It's going to take me some minutes to find out what's
| wrong in this napkin math.
|
| I am sure you will. Please don't be so entitled.
| rfoo wrote:
| > Therefore the argument about loading the cached tensors
| doesn't make a difference at all.
|
| Sorry, what? Who the fuck in this world runs decode
| without k/v cache??! If you run without k/v cache you are
| basically doing prefill for every token you generate and
| that's not what we called "decode". That's what we called
| "prefill".
|
| k/v cache, while named "cache", is a lot more important
| than what people would perceive as a "cache". It's the
| essential part of the algorithm. If you lose your k/v
| cache you must run prefill again. If you run prefill for
| every token you generate it's not O(n^2), it's going to
| be O(n^3).
|
| And yeah, you can run prefill 1000 times to generate a
| 1000 tokens output. Or you can run prefill once and with
| the persisted k/v cache run decode 1000 times. Tradeoff
| has to be made here but it simply makes no sense to drop
| a k/v cache in the middle of generating a response, as
| your number shows, recomputing is guaranteed to be slower
| than loading k/v cache.
|
| > Please don't be so entitled.
|
| When someone came up with a wrong number, I try to be
| nice and run the numbers myself and figure out why
| someone would end up with such a number and point out the
| specific mistake, instead of dumping a page of my own
| calculation. It's usually just a missing factor
| somewhere. Guess I shouldn't be so nice to retards who
| keep insisting that you can be fine without k/v cache
| during decoding. Also in this case I admit I failed to
| have a theory on why your number is so off because giving
| out prefill numbers and claiming it's decode isn't in my
| book.
|
| Yeah, I know this sounds extremely mean, feel free to
| downvote, but I hope readers can feel my frustration now.
| menaerus wrote:
| I am not gonna downvote you but you will need to find
| your manners. People around and most certainly your
| colleagues will be grateful for that. Perhaps also learn
| to deal with the arguments and different opinions without
| coloring them "plain wrong" in advance. Give people
| around you the benefit of a doubt.
|
| > Also in this case I admit I failed to have a theory on
| why your number is so off because giving out prefill
| numbers and claiming it's decode isn't in my book.
|
| Maybe it's because it is not off? It's not terribly
| difficult to sum up all the matmul calculcations and
| number of bytes one needs to load and store per each
| layer in self-attention. My number could be off for a bit
| but it is certainly not terribly off.
| rfoo wrote:
| Thanks for your advice.
|
| > different opinions
|
| I won't argue with you so hard if it's your "opinions".
| What you described is not an opinion. And facts could be
| wrong. Plainly wrong.
|
| > Maybe it's because it is not off?
|
| Yeah, as I said earlier your number might be correct as
| an estimation for prefilling 1000 tokens on Llama 3 8B.
| That's not what everybody here called "decode". Your
| number shows that prefill is compute-bound. So what?
| imtringued wrote:
| You're confusing two things.
|
| Classic softmax attention aka Softmax(Q K^T/sqrt(d_k))V
| consists of two matrix multiplications.
|
| This means QK^T=O and then softmax(O/sqrt(d_k)V.
|
| The matrix O is quadratic with respect to the number of
| input tokens. Writing the O matrix to main memory is
| bound by the maximum bandwidth of your memory.
|
| Then it has to be read out again to be multiplied against
| V.
|
| What flash attention does is change the algorithm. Flash
| attention is numerically similar to softmax attention,
| but not equivalent. The changed algorithm allows you to
| fuse the independent kernels.
|
| Instead of writing out the O matrix to main memory, its
| softmax is calculated against V immediately. The double
| memory roundtrip is now gone. This in itself does not
| change the fact that both softmax attention and flash
| attention are quadratic with respect to the input, but it
| sure as hell improves the speed of "prefill".
|
| If you tile the Q, K, V matrices into n blocks each, you
| will still have to load O(n^2) blocks.
|
| But here is the thing. Matrix multiplication is an
| operation with a significant amount of shared data. This
| means the multipliers calculating the dot products are
| being fed from the same flip flops, or the data is
| shifted around via a systolic array. You end up in a
| situation with an insignificant memory load, but a
| massive amount of arithmetic.
|
| In addition to that, you have all the tokens already, so
| the MLPs at the end of the layer can be processed as GEMM
| instead of GEMV.
|
| This is why "prefill" is compute intensive instead of
| memory intensive.
|
| During token generation, you need to perform attention
| for the next token, with all the tokens already in the KV
| cache. You load n entries from the KV cache, then do GEMV
| on the MLP and you have to do this over and over again in
| a sequential fashion. This means that memory bandwidth is
| the deciding factor for token generation.
|
| Now here is a caveat: if SRAM is limited Vs your TOPS,
| then it is possible that even flash attention is memory
| bound, but for a different reason. It's memory bound,
| because the maximum tile size that can be held in SRAM
| can be processed faster than it takes to load it from
| system memory or VRAM and you are performing a quadratic
| amount of tile loading operations. This is only
| noticeable near the extreme top end of context lengths
| between 32k and 128k tokens.
| menaerus wrote:
| Let's just summarize the FlashAttention into the
| following: Att(i) computation without FA runs in
| O(seq_len*dk + seq_len^2)
|
| whereas Att(i) computation with FA runs in
| O(seq_len^2*dk^2/SRAM_size)
|
| Q, K, V computation remains the same. And ATTN(0,n)*Wo
| also remains the same.
|
| In a smaller model, with N=12, D=768, dk=64, seq_len=1k,
| SRAM=32KB, ..., FA optimization would roughly translate
| to 0.5M vs 4.5M per-head(att(i)). So ~10x improvement but
| in the grand scheme of things, in per-attention-layer it
| becomes ~91M vs ~45M so ~2x of net improvement.
|
| > This is why "prefill" is compute intensive instead of
| memory intensive.
|
| Yes, I think I agree and I have corrected myself
| elsewhere in the thread. The original thought that I
| actually wanted to convey in my initial comment which was
| somehow lost throughout the discussion is that -
| prefill/training will benefit from the FlashAttention/MLA
| but the inference will not. I can agree that the
| formulation "only when memory access time dominates the
| compute in attention implementation" was wrong.
|
| > During token generation ... memory bandwidth is the
| deciding factor for token generation.
|
| LLama3-70B MLP layer roughly takes 1 TFLOPS and 0.6 GB of
| bandwidth for 1024 tokens. Assuming that 1023 entries are
| taken from a KV-cache, attention layer computation for a
| single token will take ~0.6 GFLOPS and ~0.2 GB of
| bandwidth. To load the rest of the values from KV-cache
| at FP16 precision, it will take us 1023*0.1MB or ~1 GB.
|
| So, ~1 TFLOPS and ~1 GB of bandwidth per each
| Transformers layer. On hardware such as H100, this still
| looks like a compute-bound problem to me. OTOH on the CPU
| with 15 TFLOPS of compute but only <1TB/s of memory
| bandwidth, it becomes memory-bound problem. Or no?
| rfoo wrote:
| For Llama 3 70B, batch size = 1, each MLP layer roughly
| takes 1x8192x26872x2 + 1x8192x26872x2 + 1x26872x8192x2
| FLOPS ~= 1.31 GFLOPS, instead of ~1 TFLOPS.
|
| Since the number differs by roughly 1024x, maybe you
| forgot that you just need to work on the last decoded
| token for MLP, too? Because you don't need hidden state
| for previous tokens in Attn now.
| FL33TW00D wrote:
| You have it backwards.
|
| Training and prefill are compute bound. Decode is memory
| bound. FlashAttention massively increases the arithmetic
| intensity of naive MHA, such that you can remain compute
| bound at lower batch sizes during decode.
| menaerus wrote:
| > Decode is memory bound.
|
| > FlashAttention ... such that you can remain compute bound
| at lower batch sizes during decode.
|
| So, which one is it then?
| FL33TW00D wrote:
| It depends on the batch size and the accelerator you're
| running on! Decode is *typically* memory bound unless you
| can hit high batch sizes (in the hundreds), which is hard
| during serving due to the contention between batch size
| and low TTFT.
|
| https://jax-ml.github.io/scaling-book/inference/ - good
| read!
| shihab wrote:
| For future readers, note that those 3x and 10x figures are
| compared to vLLM's own previous release, and NOT compared to
| Deepseek's implementation.
|
| I am very curious to see how well-optimized Deepseek's code is
| compared to leading LLM serving softwares like vLLM or SGLang.
| lhl wrote:
| It's great to see vLLM getting faster/better for DeepSeek. I
| tested vLLM vs SGLang a couple weeks ago and SGLang's DeepSeek
| support was much better/faster (on 2 x p5 H100 nodes). It's
| great that no one's standing still, I saw this recent AMD
| article that reported SGLang perf on MI300X has increased by 4X
| over the past couple weeks:
| https://rocm.blogs.amd.com/artificial-intelligence/DeepSeekR...
|
| (w/ the extra memory V3/R1 fits on a single MI300X or H200
| node)
|
| It'll be interesting to see if either project can take
| advantage/get any benefits from this FlashMLA implementation.
| FL33TW00D wrote:
| It seems to me that MLA will become the standard from here on
| out.
|
| If Deepseek R1 had used standard MHA, they would need 1749KB per
| token for KV cache storage. This means that once the conversation
| reaches ~46,000 tokens, the KV cache will have exceeded the
| entire storage capacity of a single H100.
|
| Using MLA, each token now consumes 125KB. This means you can hit
| ~640,000 tokens (2x Ulysses) before overflowing.
| ur-whale wrote:
| For those who wonder ... it's somewhat likely that MLA mean
| Multi-head latent attention
|
| https://verticalserve.medium.com/group-query-attention-58283...
|
| https://paperswithcode.com/method/multi-head-attention
| rob_c wrote:
| Great work any plans to integrate with pyT or TF I wonder?
|
| (Showing my lack of breadth of knowledge in the ecosystem (s))
| mclau156 wrote:
| Was really hoping we could get flash games back with AI
| kridsdale1 wrote:
| Ask an LLM to write you some ActionScript3
| imranq wrote:
| Dang only forward passes. The real secret was in the backward
| pass! I was also curious to learn how they implemented the
| dualpipe scheduler
| rfoo wrote:
| Do they even have an optimized backward? It looks like
| optimizations like this aren't needed during training. Their V2
| paper also suggests so.
| syntex wrote:
| What i can do with that?
| rfoo wrote:
| Probably nothing.
|
| Inference providers like Fireworks, or major clouds, can use
| this to reduce their cost, if they don't already have a
| replication with similar perf.
|
| vLLM and SGLang may integrate this to be faster at serving
| DeepSeek-V2/V2.5/V3/R1 on H100/H800s.
|
| I believe that's why they didn't release this back then, this
| is part of their "moat" (pretty weak tho) and it only benefits
| competitors.
|
| Open sourcing this after being very popular may indicate that
| they don't want all the users to use their API/Chat and now
| want the world to serve it instead? Idk.
___________________________________________________________________
(page generated 2025-02-24 23:01 UTC)