[HN Gopher] AMD unveils Ryzen Pro 8000-series processors
___________________________________________________________________
AMD unveils Ryzen Pro 8000-series processors
Author : marban
Score : 150 points
Date : 2024-04-16 13:38 UTC (9 hours ago)
(HTM) web link (www.tomshardware.com)
(TXT) w3m dump (www.tomshardware.com)
| InTheArena wrote:
| While everyone has focused on Apple's power-efficiency on the M
| series chips, one thing that has been very interesting is how
| powerful the unified memory model (by having the memory on-
| package with CPU) with large bandwidth to the memory actually is.
| Hence a lot of people in the local LLMA community are really
| going after high-memory Macs.
|
| It's great to see NPUs here with the new Ryzen cores - but I
| wonder how effective they will be with off-die memory versus the
| Apple approach.
|
| That said, it's nothing but great to see these capabilities in
| something other then a expensive NVIDIA card. Local NPUs may
| really help with edge deploying more conferencing capabilities.
|
| Edited - sorry, ,meant on-package.
| vbezhenar wrote:
| Apple does not make on-die RAM.
| chaostheory wrote:
| What Apple has is theoretically great on paper, but it fails to
| live up to expectations. Whats the point of having the RAM for
| running an LLM locally when the performance is abysmal compared
| to running it on even a consumer Nvidia GPU. It's a missed
| opportunity that I hope either the M4 or M5 addresses
| InTheArena wrote:
| The performance of oolama on my M1 MAX is pretty solid - and
| does things that my 2070 GPU can't do because of memory.
| dangus wrote:
| Not that I don't believe you but the 2070 is two
| generations and 5 years old. Maybe a comparison to a 4000
| series would be more appropriate?
| Kirby64 wrote:
| The M1 Max is also 2 generations old, and ~3 years old at
| this point. Seems like a fair comparison to me.
| dangus wrote:
| The 4000 series still has a bigger gap in how much of a
| generational leap that product was.
|
| The M3 Max has something like 33% faster overall graphics
| performance than the M1 Max (average benchmark) while the
| 4090 is something like 138% faster than the 2080Ti.
|
| Depending on which 2070 and 4070 models you compare the
| difference is similar, close to or exceeding 100% uplift.
| whizzter wrote:
| Googling power draw the 4090 goes up to 450w whilst the
| 2080ti was at 250w, adjusting for power consumption the
| increase is somewhere around 32%. Some architectural
| gains and probably optimizations in chipset workings but
| we're not seeing as many amazing generational leaps
| anymore regardless of manufacturer/designer.
| talldayo wrote:
| Maybe it's controversial, but I don't think comparing 5nm
| mobile hardware from 2021 is a fair fight against 12nm
| desktop hardware from 2018.
|
| And still, performance-wise, the 2070 still wins out by a
| ~33% margin: https://browser.geekbench.com/opencl-
| benchmarks
| chessgecko wrote:
| For this comparison the generation of chip doesn't really
| matter because the llm decode (which is the costly step)
| barely uses any of the perf and just needs the model
| weights to fit in memory
| JudasGoat wrote:
| I found it interesting that the Apple M3 scored nearly
| identical to the Radeon 780M. I know the memory bandwidth
| is slower but you can add 2 32gb sodimms to the AMD APU
| for short money.
| Teever wrote:
| Well, you know that it would still be able to do more
| than a 4000 series GPU from Nvidia because you can have
| more system memory in a mac than you can have video ram
| in a 4000 series GPU.
| dangus wrote:
| Yes, obviously I'm aware that you can throw more RAM at
| an M-series GPU.
|
| But of course that's only helpful for specific workflows.
| bearjaws wrote:
| It's a 25w processor. How will it ever live up to a 400w GPU?
| Also you can't even run large models on a single 4090, but
| you can on M series laptops with enough RAM.
|
| The fact a laptop can run 70B+ parameter models is a miracle,
| it's not what the chip was built to do at all.
| wongarsu wrote:
| It's a valid comparison in the very limited sense of "I
| have $2000 to spend on a way to run LLMs, should I get an
| RTX4090 for the computer I have or should I get a 24GB
| MacBook", or "I have $5000, should I get an RTX A6000 48GB
| or a 96GB MacBook".
|
| Those comparisons are unreasonable in a sense, but they are
| implied by statements like GPs "Hence a lot of people in
| the local LLMA community are really going after high-memory
| Macs".
| fckgw wrote:
| No, it is not a valid comparison to make between an
| entire laptop and a single PC part. "The computer I have"
| is doing a ton of heavy lifting here.
| wongarsu wrote:
| "Should I upgrade what I have or buy something new" is a
| completely normal everyday decision. Of course it doesn't
| apply to everyone since it presumes you have something
| compatible to upgrade, but it is a real decision lots of
| people are making
| fckgw wrote:
| But it's also assuming everyone has a desktop PC capable
| of this stuff that can be upgraded.
| michaelt wrote:
| Sorta yes, sorta no.
|
| You're certainly right that with a macbook you get a
| whole computer, so you're getting more for your money.
| And it's a luxury high-end computer too!
|
| But personally, I've never seen anyone step directly from
| not-even-having-a-PC to buying a 4090 for $1800. Folks
| that aren't technically inclined by and large stick with
| hosted models like ChatGPT.
|
| More common in my experience is for technical folks with,
| say, an 8GB GPU to experiment with local ML, decide
| they're interested in it, then step up to a 4090 or
| something.
| oceanplexian wrote:
| The answer depends on what you plan to do with it.
|
| Do you need to do fine tuning on a smaller model and need
| the highest inference performance with smaller models?
| Are you planning to use it as a lab to learn how to work
| with tools that are used in Big Tech (i.e. CUDA)? Or do
| you just want to do slow inference on super huge models
| (e.g Grok)?
|
| Personally, I chose the Nvidia route because as a backend
| engineer, Macs aren't seriously used in datacenters. The
| same frameworks I use to develop on a 3090 are
| transferable to massive, infiniband-connected clusters
| with TB of VRAM.
| 0x457 wrote:
| I think there are two communities:
|
| - the "hobbyists" with $5k GPUs
|
| - People that work in the industry that never used "not
| mac" or even if they did - explaining to IT that you need
| a PC with RTX A6000 48GB instead of a mac like literally
| everyone else in the company is a loosing battle.
| wongarsu wrote:
| There is also an important third group:
|
| - people that work outside Silicon Valley, where the
| entire company uses Windows centrally managed through
| Active Directory, and explaining IT that you need an Mac
| is an uphill battle. So you just submit your request for
| an RTX A6000 48GB to be added to your existing
| workstation
|
| Those people are the intended target customer of the
| A6000, and there are a lot of them.
| chaostheory wrote:
| The problem is that it extends to both Mac Studio and Mac
| Pro.
| zitterbewegung wrote:
| Buying a m3 max with 128gb of RAM while will underperform any
| consumer NVIDIA GPU it will be able to load larger models in
| practice but slowly.
|
| I think a way for the m series chips to aggressively target
| GPU inference or training would need a strategy that
| increases the speed of the RAM to start to match GDDR6 or
| HBM3 or use it directly.
| chaostheory wrote:
| You summed up my point better than I did
| evilduck wrote:
| That completely depends on your expectations and uses.
|
| I have a gaming rig with a 4080 with 16GB of RAM and it can't
| even run Mixtral (kind of the minimum bar of a useful generic
| LLM in my opinion) without being heavily quantized. Yeah it's
| fast when something fits on it, but I don't see much point in
| very fast generation of bad output. A refurbished M1 Max with
| 32GB of RAM will enable you to generate better quality LLM
| output than even a 4090 with 24GB of VRAM and for ~$300 less,
| and it's a whole computer instead of a single part that still
| needs a computer around it. Compared to my 4080, that GPU and
| the surrounding computer get you half the VRAM capacity for
| greater cost than the Mac.
|
| If you're building a rig with multiple GPUs to support many
| users or for internal private services and are willing to
| drop more than $3k then I think the equation swings back in
| favor of Nvidia, but not until then.
| soupbowl wrote:
| Just buy another 16gb of ram for 80$....
| elzbardico wrote:
| GPU ram?
| evilduck wrote:
| Running your larger-than-your-GPU-VRAM LLM model on
| regular DDR ram will completely slaughter your token/s
| speed to the point that the Mac comes out ahead again.
| programd wrote:
| Depends on what you're doing. Just chatting with the AI?
|
| I'm getting about 7 tokens per sec for Mistral with the
| Q6_K on a bog standard Intel i5-11400 desktop with 32G of
| memory and no discrete GPU (the CPU has Intel UHD
| Graphics 730 built in). 2 year old low end CPU that goes
| for, what $150? these days. As far as I'm concerned
| that's conversational speed. Pop in some 8 core modern
| CPU and I'm betting you can double that, without even
| involving any GPU.
|
| People way overestimate what they need in order to play
| around with models these days. Use llama.cpp and buy that
| extra $80 worth of RAM and pay about half the price of a
| comparable Mac all in. Bigger models? Buy more RAM, which
| is very cheap these days.
|
| There's a $487 special on Newegg today with an
| i7-12700KF, motherboard and 32G of ram. Add another $300
| worth of case, power supply, SSD and more RAM and you're
| under the price of a Macbook Air. There's your LLM
| inference machine (not for training obviously) which can
| run even the 70B models at home at acceptable
| conversational speed.
| cjk2 wrote:
| Yeah this.
|
| Also as an anecdote, my daily driver machine is a bottom
| end M2 Mac mini because I am a cheap ass. I paid less than
| it for the 4070 card in my desktop PC. The M2 Mac does a
| dehaze from RAW in lightroom in 18 seconds. My 4070 takes 9
| seconds. So the GPU is twice as fast but the mac has a
| whole free computer stuck to it.
| instagib wrote:
| One thing I would consider is usage throttling on a MacBook
| Pro. Would repeated LLM usage run into throttling?
|
| No idea what specifically everyone is pulling their
| performance data from or what task(s).
|
| Here is a video to help visualize the differences with a
| maxed out m3 max vs 16gbm1 pro vs 4090 on llm 7B/13b/70b
| llama 2. https://youtu.be/jaM02mb6JFM
|
| Here's a Reddit comparison of 4090 vs M2 Ultra 96gb with
| tokens/s
|
| https://old.reddit.com/r/LocalLLaMA/comments/14319ra/rtx_409.
| ..
|
| M3 pro memory BW 150 gb/s M3 max 10/30 300 gb/s M3 max 12/40
| 400 gb/s
|
| "Llama models are mostly limited by memory bandwidth. rtx
| 3090 has 935.8 gb/s rtx 4090 has 1008 gb/s m2 ultra has 800
| gb/s m2 max has 400 gb/s so 4090 is 10% faster for llama
| inference than 3090 and more than 2x faster than apple m2 max
| https://github.com/turboderp/exllama using exllama you can
| get 160 tokens/s in 7b model and 97 tokens/s in 13b model
| while m2 max has only 40 tokens/s in 7b model and 24 tokens/s
| in 13b apple 40/s Memory bandwidth cap is also the reason why
| llamas work so well on cpu (...)
|
| buying second gpu will increase memory capacity to 48gb but
| has no effect on bandwidth so 2x 4090 will have 48gb vram and
| 1008 gb/s bandwidth and 50% utilization"
| john_alan wrote:
| what are you talking about, it's literally the fastest single
| core retail CPU globally, and multicore is close too -
| https://browser.geekbench.com/processor-benchmarks
| jwr wrote:
| Hmm. I'm running decent LLMs locally (deepseek-
| coder:33b-instruct-q8_0, mistral:7b-instruct-v0.2-q8_0,
| mixtral:8x7b-instruct-v0.1-q4_0) on my MacBook Pro and they
| respond pretty quickly. At least for interactive use they are
| fine and comparable to Anthropic Opus in speed.
|
| That MacBook has an M3 Max and 64GB RAM.
|
| I'd say it does live up to my expectations, perhaps even
| slightly exceeds them.
| atty wrote:
| Apples memory is on package, not on die.
| thsksbd wrote:
| Old becomes new, the SGI O2 had (off chip) a unified memory
| model for performance reasons.
|
| Not a CS guy, but it seems to me that NUMA like architecture
| has to come back. Large RAM on chip (balancing a thermal budget
| between #ofcores vs ram), a much larger RAM off chip and even
| more RAM through a fast interconnect on a single kernel image.
| Like the Origin 300 had.
| Rinzler89 wrote:
| UMA in the SGI machines (and gaming consoles) made sense
| because all the memory chips at that time were equally slow,
| or fast, depending how you wanna look at it.
|
| PC HW split the video memory from system memory once GDDRAM
| become so much faster than system RAM, but GDDRAM has too
| high latency for CPUs and DDR has too low bandwidth for GPUs,
| so the separation made sense for each's strengths and still
| does to this day. Unifying it again, like with AMD's APUs,
| means either compromises for the CPU or for the GPU. There's
| no free lunch.
|
| Currently AMD APUs on the PC use unified DDRAM so CPU
| performance is top but GPU/NPU perforce is bottlenecked. If
| they were to use unified GDDRAM like in the PS5/Xbox then
| GPU/NPU performance would be top and CPU performance would be
| bottlenecked.
| Dalewyn wrote:
| >Unifying it again, like with AMD's APUs, means either
| compromises for the CPU or for the GPU. There's no free
| lunch.
|
| I think the lunch here (it still ain't free) is that RAM
| speed means nothing if you don't have enough RAM in the
| first place, and this is a compromise solution to that
| practical problem.
| Rinzler89 wrote:
| _> if you don't have enough RAM in the first place_
|
| Enough RAM for what task exactly? System RAM is plentiful
| and cheap nowadays(unless you buy Apple). I got new a
| laptop with 32GB RAM for about 750 Euros. But the speeds
| are too low for high-end gamming or LLM training for the
| poor APU.
| numpad0 wrote:
| Enough RAM for LLM. There are GPUs faster than M2 Ultra
| but can't run LLMs normally, which make that speed a moot
| point for LLM use-cases.
| Dylan16807 wrote:
| The real problem is a lack of competition in GPU
| production.
|
| GDDR is not very expensive. You should be able to get a
| GPU with a mid-level chip and tons of memory, but it's
| just not offered. Instead, please pay triple or quadruple
| the price of a high end gaming GPU to get a model with
| double the memory and basically the same core.
|
| The level of markup is astonishing. I can go from 8GB to
| 16GB on AMD for $60, but going from 24GB to 48GB costs
| $3000. And nvidia isn't better.
| zozbot234 wrote:
| > You should be able to get a GPU with a mid-level chip
| and tons of memory, but it's just not offered.
|
| Apple unified memory is the easiest way to get exactly
| that. There is a markup on memory upgrades but it's quite
| reasonable, not at all extreme.
| Dylan16807 wrote:
| But I can't put that GPU into my existing machine, so I'm
| still paying $3000 extra if I don't want that Apple
| machine to _be_ my computer.
| nsteel wrote:
| This isn't my area but won't it be quite expensive to use
| that GDDR? The PHY and the controller are complicated and
| you've got max 2GB devices, so if you want more memory
| you need a wider bus. That requires more beachfront and
| therefore a bigger die. That must make it expensive once
| you go beyond what you can fit on your small, cheap chip.
| Do their 24GB+ cards really use GDDR?*
|
| And you need to ensure you don't shoot yourself in the
| foot by making anything (relatively) cheap that could be
| useful for AI...
|
| *Edit: wow yeh, they do! A 384-bit interface on some!
| Sounds hot.
| numpad0 wrote:
| I suspect there are difficulties with DRAM latency and/or
| signal integrity with APUs and RAM-expandable GPUs. Wasted
| ALUs are wasted if you'd be stalling deep SIMD pipelines
| all the time.
| fulafel wrote:
| The fancier and more expensive SGIs had higher bandwisth
| non UMA memory systems on the GPUs.
|
| The thing about bandwidth is that you can just make wider
| buses with more of the same memory chips in parallel.
| aidenn0 wrote:
| Several PC graphics standards attempted to offer high-speed
| access to system memory (though I think only VLB offered
| direct access to the memory controller at the same speed as
| the CPU). Not having to have dedicated GPU memory has obvious
| advantages, but it's hard to get right.
| Dylan16807 wrote:
| > Old becomes new
|
| I disagree. This is like pointing at a 2024 model hatchback
| and saying "old becomes new" because you can cite a hatchback
| from 50 years ago.
|
| There's a bevy of ways to isolate or combine memory pools,
| and mainstream hardware has consistently used _many_ of these
| methods _the entire time_.
| v1sea wrote:
| edit: I was wrong.
| oflordal wrote:
| You can do that on both HIP and cuda through e.g.
| hipHostMalloc and the cuda equivalent (Not officially
| supported on the AMD APUs but works in practice). With a
| discrete GPU the GPU will access memory across PCIe but on an
| APU it will go full speed to RAM as far as I can tell.
| smallmancontrov wrote:
| > low latency results each frame
|
| What does that do to your utilization?
|
| I've been out of this space for a while, but in game dev any
| backwards information flow (GPU->CPU) completely murdered
| performance. "Whatever you do, don't stall the pipeline."
| Instant 50%-90% performance hit. Even if you had to awkwardly
| duplicate calculations on the CPU, it was almost always worth
| it, and not by a small amount. The caveat to "almost" was
| that if you were willing to wait 2-4 frames to get data back,
| you could do that without stalling the pipeline.
|
| I didn't think this was a memory architecture thing, I
| thought it was a data dependency thing. If you have to finish
| all calculations before readback and if you have to readback
| before starting new calculations, the time for all cores to
| empty out and fill back up is guaranteed to be dead time,
| regardless of whether the job was rendering polygons or
| finite element calculations or neural nets.
|
| Does shared memory actually change this somehow? Or does it
| just make it more convenient to shoot yourself in the foot?
|
| EDIT: or is the difference that HPC operates in a regime
| where long "frame time" dwarfs the pipeline empty/refill
| "dead time"?
| v1sea wrote:
| It was probably from my workloads being relatively small
| that I could get away with 90Hz read on the cpu side. I'll
| need to dig deeper into it. The metrics I was seeing were
| showing 200-300 microseconds of GPU time for physics
| calculations and within the same frame the cpu reading from
| that buffer. Maybe I'm wrong, need to test more.
| smallmancontrov wrote:
| > edit: I was wrong.
|
| If the only reason you were "wrong" was because you
| intuitively understood that it wasn't worth a large
| amount of valuable human time to save a small amount of
| cheap machine time, you were right in the way that
| matters (time allocation) and should keep it up :)
| numpad0 wrote:
| Note that while UMA is great in the sense that they allow LLM
| models to be run at all, M-series chips aren't faster[1] when
| the model fits in VRAM. 1: screenshot from[2]:
| https://www.igorslab.de/wp-
| content/uploads/2023/06/Apple-M2-ULtra-SoC-Geekbench-5-OpenCL-
| Compute.jpg 2: https://wccftech.com/apple-m2-ultra-soc-
| isnt-faster-than-amd-intel-last-year-desktop-cpus-50-slower-
| than-nvidia-rtx-4080/
| cstejerean wrote:
| The problem is you're limited to 24 GB of VRAM unless you pay
| through the nose for datacenter GPUs, whereas you can get an
| M-series chip with 128 GB or 192 GB of unified memory.
| numpad0 wrote:
| Surely! The point is that they're not million times faster
| magic chips that makes NVIDIA bankrupt tomorrow. That's
| all. A laptop with up to 128GB "VRAM" is a great option,
| absolutely no doubt about that.
| john_alan wrote:
| They are powerful, but I agree with you, it's nice to be
| able to run Goliath locally, but it's a lot slower than
| my 4070.
| paulmd wrote:
| that's openCL compute, LLM models ideally should be hitting
| the neural accelerator, not running on generalized gpu
| compute shaders.
| spamizbad wrote:
| My understanding is the unified RAM on the M-series die does
| not contribute significantly to their performance. You get a
| little bit better latency but not much. The real value to Apple
| is likely it greatly simplifies your design since you don't
| have to route out tons of DRAM signaling and power management
| on your logic board. Might make DRAM training easier too but
| that's way beyond my expertise.
| john_alan wrote:
| also provides the GPU with serious RAM allocation, a 64GB M3
| chip comes with more ram for the GPU than a 4090
| AceJohnny2 wrote:
| > _unified memory model (by having the memory on-package with
| CPU)_
|
| That's not what "unified memory model" means.
|
| It means that the CPU and GPU (and ANE!) have access to the
| same banks of memory, unlike PC GPUs that have their own
| memory, separated from the CPU's by the PCIe bottleneck (as
| fast as that is, it's still smaller than direct shared DRAM
| access).
|
| It allows the hardware more flexibility in how the single pool
| of memory is allocated across devices, and faster sharing of
| data across devices. (throughput/latency depends on the
| internal system bus ports and how many each device have access
| to)
|
| The Apple M-Series chips _also_ has the memory on-package with
| the CPU (technically SoC, "System-on-Chip"), but that provides
| different benefits.
| cmovq wrote:
| Having separated GPU memory also has its benefits. Once the
| data makes it through the PCIe bus, graphics memory typically
| has much higher bandwidth which also doesn't need to split
| with the CPU.
| crawshaw wrote:
| An M2 Ultra has 800GB/s of memory bandwidth, an Nvidia 4090
| has 1008GB/s. Apple have chosen to use relatively little
| system memory at unusually high bandwidth.
| fulafel wrote:
| Most x86 machines have integrated GPUs and hardware-wise are
| UMA.
| alacritas0 wrote:
| integrated GPUs are not powerful in comparison to dedicated
| GPUs
| Havoc wrote:
| Are these NPUs addressable with a standard PyTorch LLM stack?
|
| NPU seems to mean very different things depending on
| device/vendor
| bearjaws wrote:
| The focus on TOPS seems a bit out of line with reality for LLMs.
| TOPs doesn't matter for LLMs if your memory bandwidth can't keep
| up. Since it doesn't have quad channel memory mentioned I guess
| it's still dual channel?
|
| Even top of the line DDR5 is around 128GB/s vs a M1 at 400GB/s.
|
| At the end of the day, it still seems like AI in consumer chips
| is chasing a buzzword, what is the killer feature?
|
| On mobile there are image processing benefits and voice to text,
| translation... but on desktop those are no where near common use
| cases.
| postalrat wrote:
| https://www.neatvideo.com/blog/post/m3
|
| That says M1 is 68.25 GB/s
| givinguflac wrote:
| The op were obviously talking about M1 Max.
| postalrat wrote:
| How is it obvious? Anyone reading that could assume that
| any M1 gets that bandwidth.
| bearjaws wrote:
| M1 Max sorry, I don't mean to compare a 4 year old tablet
| processor to the latest generation of laptop CPUs.
| VHRanger wrote:
| The killer feature is presumably inference at the edge, but I
| don't see that being used on desktop much at all right now.
|
| Especially since most desktop applications people use are web
| apps. Of the native apps people use that leverage this sort of
| stuff, almost all are GPU accelerated already (eg. image and
| video editing AI tools)
| jzig wrote:
| What does "at the edge" mean here?
| georgeecollins wrote:
| Not using AI on the cloud. So if your connection is
| uncertain or you want use your bandwidth for something
| else-- like video conferencing or gaming. Probably the
| killer app is something that wants to use AI but doesn't
| involve paying a cloud provider. I was talking to a vendor
| about their chat bot built to put into MMOs or mobile
| games. It woudl be killer to have a character have life
| like conversation in those kinds of experiences. But the
| last thing you want to do is increase your server costs the
| way this AI would. Edge computing could solve that.
| PeterSmit wrote:
| Not in the cloud.
| Zach_the_Lizard wrote:
| I'm guessing "the edge" is doing inference work in the
| browser, etc. as opposed to somewhere in the backend of the
| web app.
|
| Maybe your local machine can run, I don't know, a model to
| make suggestions as you're editing a Google Doc, which
| frees up the Big Machine in the Sky to do other things.
|
| As this becomes more technically feasible, it reduces the
| effective cost of inference for a new service provider,
| since you, the client, are now running their code.
|
| The Jevons paradox might kick in, causing more and more
| uses of LLMs for use cases that were too expensive before.
| VHRanger wrote:
| Edge is doing computing on the client (eg. browser, phone,
| laptop, etc.) instead of the server
| Dylan16807 wrote:
| Half the definitions I see of edge include client
| devices, and half of them don't include client devices.
|
| I like the latter. Why even use a new word if it's just
| going to be the same as "client"?
| futureshock wrote:
| Upscaling for gaming or video.
|
| Local context aware search
|
| Offline Speech to text and TTS
|
| Offline generation of clip art or stock images for document
| editing
|
| Offline LLM that can work with your documents as context and
| access application and OS APIs
|
| Improved enemy AI in gaming
|
| Webcam effects like background removal or filters.
|
| Audio upscaling and interpolation like for bad video call
| connections.
| bearjaws wrote:
| > Upscaling for gaming or video.
|
| Already exists on all three major GPU manufacturers, and it
| definitely makes sense as a GPU workload.
|
| > Local context aware search
|
| You don't need an AI processor to do this, Windows search
| used to work better and had even less compute resources to
| work with.
|
| > Offline Speech to text and TTS
|
| See my point about not a very common use case for desktops &
| laptops vs cell phones.
|
| > Offline LLM that can work with your documents as context
| and access application and OS APIs
|
| Maybe for some sort of background task or only using really
| small models <13B parameters. Anything real time is going to
| run at 1-2t/s with a large model.
|
| Small models are pretty terrible though, I doubt people want
| even more incorrect information and hallucinations.
|
| > Improved enemy AI in gaming
|
| See Ageia PhysX
|
| > Webcam effects like background removal or filters.
|
| We already have this without NPUs.
|
| > Audio upscaling and interpolation like for bad video call
| connections.
|
| I could see this, or noise cancellation.
| bayindirh wrote:
| It's about power management, and doing more things with
| less power. These specialized IP blocks on CPUs allow these
| things to be done with less power and less latency.
|
| Intel's bottom of the barrel N95 & N100 CPUs have Gaussian
| & Neural accelerators for simple image processing and
| object detection tasks, plus a voice processor for low
| power voice based activation and command capture and
| process.
|
| You can always add more power hungry, general purpose
| components to add capabilities. Heck, video post processing
| entered hardware era with ATI Radeon 8500. But doing these
| things with negligible power costs is the new front.
|
| Apple is not adding coprocessors to their iPhones because
| it looks nice. All of these coprocessors reduce CPU wake-up
| cycles tremendously and allows the device to monitor tons
| of things out of bands with negligible power costs.
| pdpi wrote:
| >> Upscaling for gaming or video.
|
| > Already exists on all three major GPU manufacturers, and
| it definitely makes sense as a GPU workload.
|
| "makes sense as a GPU workload" is underselling it a bit.
| Doing it on the CPU is basically insane. Games typically
| upscale only the world view (the expensive part to render)
| while rendering the UI at full res. So to do CPU-side
| upscaling we're talking about a game rendering a surface on
| the GPU, sending it to the CPU, upscaling it there, sending
| it back to the GPU, then compositing with the UI. It's just
| needlessly complicated.
| futureshock wrote:
| > Upscaling for gaming or video. > Already exists on all
| three major GPU manufacturers, and it definitely makes
| sense as a GPU workload. These AMD chips are APUs that are
| often the only GPU, not every user will have a dedicated
| GPU.
|
| > Local context aware search > You don't need an AI
| processor to do this, Windows search used to work better
| and had even less compute resources to work with. You could
| still improve it with increased natural language
| understanding instead of simple keyword. "Give me all
| documents about dogs" instead of searching for each breed
| as a keyword.
|
| > Offline Speech to text and TTS > See my point about not a
| very common use case for desktops & laptops vs cell phones.
| Maybe not for you but accessibility is a key feature for
| many users. You think blind users should suffer through bad
| TTS?
|
| > Offline LLM that can work with your documents as context
| and access application and OS APIs > Maybe for some sort of
| background task or only using really small models <13B
| parameters. Anything real time is going to run at 1-2t/s
| with a large model. > Small models are pretty terrible
| though, I doubt people want even more incorrect information
| and hallucinations. Small model have been improving and
| better capabilities in consumer chips will allow larger
| models to run faster.
|
| > Improved enemy AI in gaming > See Ageia PhysX Surely
| you're not suggesting that enemy AI is solved problem in
| gaming?
|
| > Webcam effects like background removal or filters. > We
| already have this without NPUs. Sure but it could go from
| obvious and distracting to seamless and convincing.
|
| > Audio upscaling and interpolation like for bad video call
| connections. I could see this, or noise cancellation.
| kanbankaren wrote:
| All of this(except upscaling) is possible with iGPU/CPU
| without breaking a sweat?
| bayindirh wrote:
| The things which doesn't make GPU to break a sweat has its
| own specialized (or semi-specialized) processing blocks on
| the GPU, too.
| kanbankaren wrote:
| I meant the current generation of GPUs that don't have
| any AI acceleration blocks.
| bayindirh wrote:
| They are MATMUL machines by design already. They do not
| need to "accelerate" AI to begin with.
|
| Their cores/shaders can be programmed to do that.
|
| Also, name a current gen GPU which doesn't have video
| encoding/decoding capabilities/facilities in silicon,
| even ones which do not allow shaders to be used in this
| process for post-processing. It's impossible (to not to
| have these capabilities) at this point in time.
| kanbankaren wrote:
| I was talking about AI blocks and you moved the goal post
| to video codec blocks.
| bayindirh wrote:
| No. I didn't move anything.
|
| I said that the core (3D rendering hardware) of a GPU
| with shaders is _the_ AI block already, and said that
| other tasks like video encoders have their own blocks,
| but still pull capabilities from the "core" to improve
| things.
| phkahler wrote:
| Am I missing something? These look just like the APUs with the
| addition of "management" and "security" features and without the
| iGPU. Is that right?
| c0l0 wrote:
| They also support ECC UDIMMs with ECC enabled, which has been
| the "PRO" series APU killer feature on AM4 for me. The
| non-"PRO" APUs will run fine with ECC UDIMMs, but cannot make
| use of the extra parity information (maybe for reasons of
| market segmentation - I don't know if anyone outside of AMD
| knows). This is probably less of a concern with DDR5 platforms
| and their "on-DIE ECC" (which you cannot monitor for
| Correctable Errors at least, afaik), but it's still gonna
| matter for me.
| nwah1 wrote:
| Yes. And it actually does have an iGPU, depending on which
| model.
| c0l0 wrote:
| In case you don't need an integrated GPU (that's somewhat
| powerful/potent), you can go with any other Ryzen AM5 CPU
| to receive proper ECC-enabled ECC UDIMM support, afaik :)
| pedrocr wrote:
| Since in AM5 all CPUs have a basic iGPU, for a home server
| all the normal CPUs already work fine. The advantage of the
| APUs is they're on a monolithic die so should have quite a
| bit lower idle power usage which is important if you have a
| NAS or other homelab server running 24/7.
| kokonoko wrote:
| I hope (but I doubt) that this will be more than a marketing
| stand to include something "AI" in their product line. Every
| vendor has their own hardware that is badly supported by tools,
| and even if it is supported, only a fraction of the available
| software uses it. In the meantime it takes precious die area and
| resources.
| Aissen wrote:
| A quick search into it shows that this Ryzen AI NPU's support
| isn't integrated into upstream inference frameworks yet -- so
| right now it's just useless silicon surface you pay for :-/
| Rinzler89 wrote:
| Some AMD laptops haven't even yet enabled the NPU in firmware
| even on the 7000 series wich are about a year old. Meaning it's
| still useless.
|
| I was kinda bummed out they released the 8000 series after I
| just bought a laptop with 7000 series, but I think I actually
| dodged a bullet here since it doesn't look like much of an
| upgrade and the AI silicone screams of very early first gen
| product to me, as if they rushed it out the door because
| everyone else was doing "AI" and they needed to also cash in on
| the hype, kinda like the first gen RTX cards.
|
| I think by the time I'll actually upgrade, the AI/NPU tech
| would have matured considerably and actually be useful.
| robocat wrote:
| Does anyone have any mental heuristics for judging how
| "useless" a feature is?
|
| Over decades I have a growing antipathy towards products with
| too many features. Especially new versions/models where the
| vaunted features of the previous version/model seem to never
| have been used by anyone.
| Rinzler89 wrote:
| _> Does anyone have any mental heuristics for judging how
| "useless" a feature is? _
|
| My favorite example is the story I got to live through of
| the first generations of consumer 64 bit CPUs.
|
| When the first AMD Athlon 64 came out, everyone I knew was
| buying them because they though they were getting something
| totally future proof by jumping early on the 64 bit
| bandwagon, in 2003, when nobody yet had 4GB+ of RAM and
| neither Windows nor any software would see 64bit releases
| till several years later when Vista came out which everyone
| avoided and staid on Windows XP 32bit waiting for Windows
| 7.
|
| And by the time RAM sizes over 4GB and 64 bit software
| became even remotely mainstream, we already had dual- and
| quad-core CPUs miles ahead of those original 64 bit CPUs
| which were now obsolete (tech progress back then was wild).
|
| So just like how 64bit silicone was a useless feature on
| consumer CPUs, and like the first GPUs with raytracing, I
| feel like now we're in the same boat with AI silicone in
| PCs, no much SW support for them and when it does come,
| these early chips will be obsolete. It's the price of being
| an early adopter.
| sva_ wrote:
| > Some AMD laptops haven't even yet enabled the NPU in
| firmware
|
| This is entirely the fault of the OEMs though, not AMD. It is
| activated on mine for example. But pretty much unusable under
| Linux at the moment (unless you're willing to run a custom
| kernel for it[0].)
|
| 0. https://github.com/amd/xdna-driver
| Rinzler89 wrote:
| _> This is entirely the fault of the OEMs though, not AMD._
|
| Not true. AMD can demand how OEMs integrate and use their
| chips in their products as part of the sales agreement,
| same how Nvidia does.
|
| AMD could have said to every system integrator buying 7000
| series chips and up, that the NPU must be active in the
| final product.
|
| So if the end products suck, AMD bares most of the blame
| for not ensuring a minimum level of QA with its integrators
| who release half-assed stuff since it all reflects poorly
| on them in the end. It's one of the reason why Nvidia keeps
| such a tight grip over its integrators on how their chips
| are to used.
| dhruvdh wrote:
| There is a VitisAI execution provider for ONNX, and you can use
| ONNX backends for inference frameworks that support it. More
| info here - https://ryzenai.docs.amd.com/en/latest/
|
| But regardless, 16 TOPs is no good for LLMs. Though there is a
| Ryzen AI demo that shows Llama 7B running on these at 8
| tokens/sec. A sub-par experience for a sub-par LLM.
| Aissen wrote:
| Thanks, I was looking for information on this, it seems to be
| lower speed than pure-CPU inference on M2, and probably much
| worse than a ROCm GPU-based solution?
| p_l wrote:
| Because the NPU isn't for high-end inferencing. It's a
| relatively small coprocessor that is supposed to do bunch
| of tasks with high TOPS/watt without engaging the way more
| power hungry GPU.
|
| At release time, the windows driver for example included
| few video processing offloads used by Windows Frameworks
| used for example by MS Teams for background removal - so
| that such tasks use less battery on laptops and free up
| CPU/GPU for other tasks on desktop.
|
| For higher end processing you can use the same AIE-ML
| coprocessors various chips available previously from Xilinx
| and now under AMD brand.
| fpgamlirfanboy wrote:
| > the same AIE-ML coprocessors
|
| they're not the same - versal acaps (whatever you want to
| call them) have AIE1 arch while phoenix has AIE2 arch.
| there are significant differences between the two arches
| (local memory, bfloat16, etc.)
| p_l wrote:
| Phoenix has AIE-ML (what you call AIE2), Versal has
| choice of AIE (AIE1) and AIE-ML (AIE2) depending on chip
| you buy.
|
| Essentially, AMD is making two tile designs optimized for
| slightly different computations and claims that they are
| going to offer both in Versal, but NPUs use exclusively
| the ML-optimized ones.
| markdog12 wrote:
| Wow, that's simply embarrassing.
| myself248 wrote:
| What does "commercial market" mean here? The article says these
| features were already available in the "consumer" versions of
| these chips -- were those given away for free or something? What
| about them was not "commercial"?
|
| I'm sure there's some market segmentation thing at work here, but
| this just sounds like a Hallmark-holiday excuse to rehash an old
| press-release and pretend it's another revolution all over again.
| dhruvdh wrote:
| You can't buy these pro variants from Microcenter for example,
| but you can buy them from pre-built OEM desktops. Mostly meant
| for enterprise customers who buy in bulk, I think.
| transpute wrote:
| _> the article says these features were already available_
|
| Actually, the article says they were _not_ available:
| the Pro series is based on AMD's existing consumer-oriented
| processor models but comes with additional features
|
| Pro (enterprise) CPUs include remote management, memory
| encryption and other security features:
| https://www.amd.com/en/ryzen-pro
| baarsh wrote:
| How come new series comes with options only up to 8 cores, while
| 5900X and 5950X were already 12 and 16 cores few years ago?
| Arrath wrote:
| Better sustained performance with the thermal envelope afforded
| to only 8 cores vs 12+?
| wmf wrote:
| Desktop Ryzen = up to 16 cores
|
| Laptop Ryzen = up to 8 cores (the 8000G are laptop CPUs in a
| desktop socket)
| protastus wrote:
| Dragon Range (e.g., 7945HX) is a laptop/mobile workstation
| Ryzen with up to 16 cores, but power inefficient compared to
| the 8000-series due to the chiplet design. Already in market,
| mostly in gaming laptops.
| JonChesterfield wrote:
| Zen4 is _very_ pretty, see
| https://news.ycombinator.com/item?id=32983406.
|
| I like their APUs a lot. Using a 4800U in the cable tray under
| the desk to drive screens on which I'm writing this. One in a
| laptop for whenever I'm away from the desk.
|
| If you're sufficiently determined the compute units on these
| things are totally usable for running arbitrary code. As in a
| program that spawns a bunch of threads to work stuff out could
| have some of those "threads" running on the GPU cores.
|
| I wouldn't say the software stack is totally there for out of the
| box convenience. As in you'll be writing in freestanding ~C and
| maybe a bit of assembly. I got partway through implementing that
| and got sidetracked. The GPU libc in LLVM is roughly the
| production version of some of that hacking. Between these
| machines coming out and the MI300A landing I really should put
| something up on github which looks like a pthread_create that
| executes on the GPU instead.
___________________________________________________________________
(page generated 2024-04-16 23:02 UTC)