[HN Gopher] AMD unveils Ryzen Pro 8000-series processors
       ___________________________________________________________________
        
       AMD unveils Ryzen Pro 8000-series processors
        
       Author : marban
       Score  : 150 points
       Date   : 2024-04-16 13:38 UTC (9 hours ago)
        
 (HTM) web link (www.tomshardware.com)
 (TXT) w3m dump (www.tomshardware.com)
        
       | InTheArena wrote:
       | While everyone has focused on Apple's power-efficiency on the M
       | series chips, one thing that has been very interesting is how
       | powerful the unified memory model (by having the memory on-
       | package with CPU) with large bandwidth to the memory actually is.
       | Hence a lot of people in the local LLMA community are really
       | going after high-memory Macs.
       | 
       | It's great to see NPUs here with the new Ryzen cores - but I
       | wonder how effective they will be with off-die memory versus the
       | Apple approach.
       | 
       | That said, it's nothing but great to see these capabilities in
       | something other then a expensive NVIDIA card. Local NPUs may
       | really help with edge deploying more conferencing capabilities.
       | 
       | Edited - sorry, ,meant on-package.
        
         | vbezhenar wrote:
         | Apple does not make on-die RAM.
        
         | chaostheory wrote:
         | What Apple has is theoretically great on paper, but it fails to
         | live up to expectations. Whats the point of having the RAM for
         | running an LLM locally when the performance is abysmal compared
         | to running it on even a consumer Nvidia GPU. It's a missed
         | opportunity that I hope either the M4 or M5 addresses
        
           | InTheArena wrote:
           | The performance of oolama on my M1 MAX is pretty solid - and
           | does things that my 2070 GPU can't do because of memory.
        
             | dangus wrote:
             | Not that I don't believe you but the 2070 is two
             | generations and 5 years old. Maybe a comparison to a 4000
             | series would be more appropriate?
        
               | Kirby64 wrote:
               | The M1 Max is also 2 generations old, and ~3 years old at
               | this point. Seems like a fair comparison to me.
        
               | dangus wrote:
               | The 4000 series still has a bigger gap in how much of a
               | generational leap that product was.
               | 
               | The M3 Max has something like 33% faster overall graphics
               | performance than the M1 Max (average benchmark) while the
               | 4090 is something like 138% faster than the 2080Ti.
               | 
               | Depending on which 2070 and 4070 models you compare the
               | difference is similar, close to or exceeding 100% uplift.
        
               | whizzter wrote:
               | Googling power draw the 4090 goes up to 450w whilst the
               | 2080ti was at 250w, adjusting for power consumption the
               | increase is somewhere around 32%. Some architectural
               | gains and probably optimizations in chipset workings but
               | we're not seeing as many amazing generational leaps
               | anymore regardless of manufacturer/designer.
        
               | talldayo wrote:
               | Maybe it's controversial, but I don't think comparing 5nm
               | mobile hardware from 2021 is a fair fight against 12nm
               | desktop hardware from 2018.
               | 
               | And still, performance-wise, the 2070 still wins out by a
               | ~33% margin: https://browser.geekbench.com/opencl-
               | benchmarks
        
               | chessgecko wrote:
               | For this comparison the generation of chip doesn't really
               | matter because the llm decode (which is the costly step)
               | barely uses any of the perf and just needs the model
               | weights to fit in memory
        
               | JudasGoat wrote:
               | I found it interesting that the Apple M3 scored nearly
               | identical to the Radeon 780M. I know the memory bandwidth
               | is slower but you can add 2 32gb sodimms to the AMD APU
               | for short money.
        
               | Teever wrote:
               | Well, you know that it would still be able to do more
               | than a 4000 series GPU from Nvidia because you can have
               | more system memory in a mac than you can have video ram
               | in a 4000 series GPU.
        
               | dangus wrote:
               | Yes, obviously I'm aware that you can throw more RAM at
               | an M-series GPU.
               | 
               | But of course that's only helpful for specific workflows.
        
           | bearjaws wrote:
           | It's a 25w processor. How will it ever live up to a 400w GPU?
           | Also you can't even run large models on a single 4090, but
           | you can on M series laptops with enough RAM.
           | 
           | The fact a laptop can run 70B+ parameter models is a miracle,
           | it's not what the chip was built to do at all.
        
             | wongarsu wrote:
             | It's a valid comparison in the very limited sense of "I
             | have $2000 to spend on a way to run LLMs, should I get an
             | RTX4090 for the computer I have or should I get a 24GB
             | MacBook", or "I have $5000, should I get an RTX A6000 48GB
             | or a 96GB MacBook".
             | 
             | Those comparisons are unreasonable in a sense, but they are
             | implied by statements like GPs "Hence a lot of people in
             | the local LLMA community are really going after high-memory
             | Macs".
        
               | fckgw wrote:
               | No, it is not a valid comparison to make between an
               | entire laptop and a single PC part. "The computer I have"
               | is doing a ton of heavy lifting here.
        
               | wongarsu wrote:
               | "Should I upgrade what I have or buy something new" is a
               | completely normal everyday decision. Of course it doesn't
               | apply to everyone since it presumes you have something
               | compatible to upgrade, but it is a real decision lots of
               | people are making
        
               | fckgw wrote:
               | But it's also assuming everyone has a desktop PC capable
               | of this stuff that can be upgraded.
        
               | michaelt wrote:
               | Sorta yes, sorta no.
               | 
               | You're certainly right that with a macbook you get a
               | whole computer, so you're getting more for your money.
               | And it's a luxury high-end computer too!
               | 
               | But personally, I've never seen anyone step directly from
               | not-even-having-a-PC to buying a 4090 for $1800. Folks
               | that aren't technically inclined by and large stick with
               | hosted models like ChatGPT.
               | 
               | More common in my experience is for technical folks with,
               | say, an 8GB GPU to experiment with local ML, decide
               | they're interested in it, then step up to a 4090 or
               | something.
        
               | oceanplexian wrote:
               | The answer depends on what you plan to do with it.
               | 
               | Do you need to do fine tuning on a smaller model and need
               | the highest inference performance with smaller models?
               | Are you planning to use it as a lab to learn how to work
               | with tools that are used in Big Tech (i.e. CUDA)? Or do
               | you just want to do slow inference on super huge models
               | (e.g Grok)?
               | 
               | Personally, I chose the Nvidia route because as a backend
               | engineer, Macs aren't seriously used in datacenters. The
               | same frameworks I use to develop on a 3090 are
               | transferable to massive, infiniband-connected clusters
               | with TB of VRAM.
        
               | 0x457 wrote:
               | I think there are two communities:
               | 
               | - the "hobbyists" with $5k GPUs
               | 
               | - People that work in the industry that never used "not
               | mac" or even if they did - explaining to IT that you need
               | a PC with RTX A6000 48GB instead of a mac like literally
               | everyone else in the company is a loosing battle.
        
               | wongarsu wrote:
               | There is also an important third group:
               | 
               | - people that work outside Silicon Valley, where the
               | entire company uses Windows centrally managed through
               | Active Directory, and explaining IT that you need an Mac
               | is an uphill battle. So you just submit your request for
               | an RTX A6000 48GB to be added to your existing
               | workstation
               | 
               | Those people are the intended target customer of the
               | A6000, and there are a lot of them.
        
             | chaostheory wrote:
             | The problem is that it extends to both Mac Studio and Mac
             | Pro.
        
           | zitterbewegung wrote:
           | Buying a m3 max with 128gb of RAM while will underperform any
           | consumer NVIDIA GPU it will be able to load larger models in
           | practice but slowly.
           | 
           | I think a way for the m series chips to aggressively target
           | GPU inference or training would need a strategy that
           | increases the speed of the RAM to start to match GDDR6 or
           | HBM3 or use it directly.
        
             | chaostheory wrote:
             | You summed up my point better than I did
        
           | evilduck wrote:
           | That completely depends on your expectations and uses.
           | 
           | I have a gaming rig with a 4080 with 16GB of RAM and it can't
           | even run Mixtral (kind of the minimum bar of a useful generic
           | LLM in my opinion) without being heavily quantized. Yeah it's
           | fast when something fits on it, but I don't see much point in
           | very fast generation of bad output. A refurbished M1 Max with
           | 32GB of RAM will enable you to generate better quality LLM
           | output than even a 4090 with 24GB of VRAM and for ~$300 less,
           | and it's a whole computer instead of a single part that still
           | needs a computer around it. Compared to my 4080, that GPU and
           | the surrounding computer get you half the VRAM capacity for
           | greater cost than the Mac.
           | 
           | If you're building a rig with multiple GPUs to support many
           | users or for internal private services and are willing to
           | drop more than $3k then I think the equation swings back in
           | favor of Nvidia, but not until then.
        
             | soupbowl wrote:
             | Just buy another 16gb of ram for 80$....
        
               | elzbardico wrote:
               | GPU ram?
        
               | evilduck wrote:
               | Running your larger-than-your-GPU-VRAM LLM model on
               | regular DDR ram will completely slaughter your token/s
               | speed to the point that the Mac comes out ahead again.
        
               | programd wrote:
               | Depends on what you're doing. Just chatting with the AI?
               | 
               | I'm getting about 7 tokens per sec for Mistral with the
               | Q6_K on a bog standard Intel i5-11400 desktop with 32G of
               | memory and no discrete GPU (the CPU has Intel UHD
               | Graphics 730 built in). 2 year old low end CPU that goes
               | for, what $150? these days. As far as I'm concerned
               | that's conversational speed. Pop in some 8 core modern
               | CPU and I'm betting you can double that, without even
               | involving any GPU.
               | 
               | People way overestimate what they need in order to play
               | around with models these days. Use llama.cpp and buy that
               | extra $80 worth of RAM and pay about half the price of a
               | comparable Mac all in. Bigger models? Buy more RAM, which
               | is very cheap these days.
               | 
               | There's a $487 special on Newegg today with an
               | i7-12700KF, motherboard and 32G of ram. Add another $300
               | worth of case, power supply, SSD and more RAM and you're
               | under the price of a Macbook Air. There's your LLM
               | inference machine (not for training obviously) which can
               | run even the 70B models at home at acceptable
               | conversational speed.
        
             | cjk2 wrote:
             | Yeah this.
             | 
             | Also as an anecdote, my daily driver machine is a bottom
             | end M2 Mac mini because I am a cheap ass. I paid less than
             | it for the 4070 card in my desktop PC. The M2 Mac does a
             | dehaze from RAW in lightroom in 18 seconds. My 4070 takes 9
             | seconds. So the GPU is twice as fast but the mac has a
             | whole free computer stuck to it.
        
           | instagib wrote:
           | One thing I would consider is usage throttling on a MacBook
           | Pro. Would repeated LLM usage run into throttling?
           | 
           | No idea what specifically everyone is pulling their
           | performance data from or what task(s).
           | 
           | Here is a video to help visualize the differences with a
           | maxed out m3 max vs 16gbm1 pro vs 4090 on llm 7B/13b/70b
           | llama 2. https://youtu.be/jaM02mb6JFM
           | 
           | Here's a Reddit comparison of 4090 vs M2 Ultra 96gb with
           | tokens/s
           | 
           | https://old.reddit.com/r/LocalLLaMA/comments/14319ra/rtx_409.
           | ..
           | 
           | M3 pro memory BW 150 gb/s M3 max 10/30 300 gb/s M3 max 12/40
           | 400 gb/s
           | 
           | "Llama models are mostly limited by memory bandwidth. rtx
           | 3090 has 935.8 gb/s rtx 4090 has 1008 gb/s m2 ultra has 800
           | gb/s m2 max has 400 gb/s so 4090 is 10% faster for llama
           | inference than 3090 and more than 2x faster than apple m2 max
           | https://github.com/turboderp/exllama using exllama you can
           | get 160 tokens/s in 7b model and 97 tokens/s in 13b model
           | while m2 max has only 40 tokens/s in 7b model and 24 tokens/s
           | in 13b apple 40/s Memory bandwidth cap is also the reason why
           | llamas work so well on cpu (...)
           | 
           | buying second gpu will increase memory capacity to 48gb but
           | has no effect on bandwidth so 2x 4090 will have 48gb vram and
           | 1008 gb/s bandwidth and 50% utilization"
        
           | john_alan wrote:
           | what are you talking about, it's literally the fastest single
           | core retail CPU globally, and multicore is close too -
           | https://browser.geekbench.com/processor-benchmarks
        
           | jwr wrote:
           | Hmm. I'm running decent LLMs locally (deepseek-
           | coder:33b-instruct-q8_0, mistral:7b-instruct-v0.2-q8_0,
           | mixtral:8x7b-instruct-v0.1-q4_0) on my MacBook Pro and they
           | respond pretty quickly. At least for interactive use they are
           | fine and comparable to Anthropic Opus in speed.
           | 
           | That MacBook has an M3 Max and 64GB RAM.
           | 
           | I'd say it does live up to my expectations, perhaps even
           | slightly exceeds them.
        
         | atty wrote:
         | Apples memory is on package, not on die.
        
         | thsksbd wrote:
         | Old becomes new, the SGI O2 had (off chip) a unified memory
         | model for performance reasons.
         | 
         | Not a CS guy, but it seems to me that NUMA like architecture
         | has to come back. Large RAM on chip (balancing a thermal budget
         | between #ofcores vs ram), a much larger RAM off chip and even
         | more RAM through a fast interconnect on a single kernel image.
         | Like the Origin 300 had.
        
           | Rinzler89 wrote:
           | UMA in the SGI machines (and gaming consoles) made sense
           | because all the memory chips at that time were equally slow,
           | or fast, depending how you wanna look at it.
           | 
           | PC HW split the video memory from system memory once GDDRAM
           | become so much faster than system RAM, but GDDRAM has too
           | high latency for CPUs and DDR has too low bandwidth for GPUs,
           | so the separation made sense for each's strengths and still
           | does to this day. Unifying it again, like with AMD's APUs,
           | means either compromises for the CPU or for the GPU. There's
           | no free lunch.
           | 
           | Currently AMD APUs on the PC use unified DDRAM so CPU
           | performance is top but GPU/NPU perforce is bottlenecked. If
           | they were to use unified GDDRAM like in the PS5/Xbox then
           | GPU/NPU performance would be top and CPU performance would be
           | bottlenecked.
        
             | Dalewyn wrote:
             | >Unifying it again, like with AMD's APUs, means either
             | compromises for the CPU or for the GPU. There's no free
             | lunch.
             | 
             | I think the lunch here (it still ain't free) is that RAM
             | speed means nothing if you don't have enough RAM in the
             | first place, and this is a compromise solution to that
             | practical problem.
        
               | Rinzler89 wrote:
               | _> if you don't have enough RAM in the first place_
               | 
               | Enough RAM for what task exactly? System RAM is plentiful
               | and cheap nowadays(unless you buy Apple). I got new a
               | laptop with 32GB RAM for about 750 Euros. But the speeds
               | are too low for high-end gamming or LLM training for the
               | poor APU.
        
               | numpad0 wrote:
               | Enough RAM for LLM. There are GPUs faster than M2 Ultra
               | but can't run LLMs normally, which make that speed a moot
               | point for LLM use-cases.
        
               | Dylan16807 wrote:
               | The real problem is a lack of competition in GPU
               | production.
               | 
               | GDDR is not very expensive. You should be able to get a
               | GPU with a mid-level chip and tons of memory, but it's
               | just not offered. Instead, please pay triple or quadruple
               | the price of a high end gaming GPU to get a model with
               | double the memory and basically the same core.
               | 
               | The level of markup is astonishing. I can go from 8GB to
               | 16GB on AMD for $60, but going from 24GB to 48GB costs
               | $3000. And nvidia isn't better.
        
               | zozbot234 wrote:
               | > You should be able to get a GPU with a mid-level chip
               | and tons of memory, but it's just not offered.
               | 
               | Apple unified memory is the easiest way to get exactly
               | that. There is a markup on memory upgrades but it's quite
               | reasonable, not at all extreme.
        
               | Dylan16807 wrote:
               | But I can't put that GPU into my existing machine, so I'm
               | still paying $3000 extra if I don't want that Apple
               | machine to _be_ my computer.
        
               | nsteel wrote:
               | This isn't my area but won't it be quite expensive to use
               | that GDDR? The PHY and the controller are complicated and
               | you've got max 2GB devices, so if you want more memory
               | you need a wider bus. That requires more beachfront and
               | therefore a bigger die. That must make it expensive once
               | you go beyond what you can fit on your small, cheap chip.
               | Do their 24GB+ cards really use GDDR?*
               | 
               | And you need to ensure you don't shoot yourself in the
               | foot by making anything (relatively) cheap that could be
               | useful for AI...
               | 
               | *Edit: wow yeh, they do! A 384-bit interface on some!
               | Sounds hot.
        
             | numpad0 wrote:
             | I suspect there are difficulties with DRAM latency and/or
             | signal integrity with APUs and RAM-expandable GPUs. Wasted
             | ALUs are wasted if you'd be stalling deep SIMD pipelines
             | all the time.
        
             | fulafel wrote:
             | The fancier and more expensive SGIs had higher bandwisth
             | non UMA memory systems on the GPUs.
             | 
             | The thing about bandwidth is that you can just make wider
             | buses with more of the same memory chips in parallel.
        
           | aidenn0 wrote:
           | Several PC graphics standards attempted to offer high-speed
           | access to system memory (though I think only VLB offered
           | direct access to the memory controller at the same speed as
           | the CPU). Not having to have dedicated GPU memory has obvious
           | advantages, but it's hard to get right.
        
           | Dylan16807 wrote:
           | > Old becomes new
           | 
           | I disagree. This is like pointing at a 2024 model hatchback
           | and saying "old becomes new" because you can cite a hatchback
           | from 50 years ago.
           | 
           | There's a bevy of ways to isolate or combine memory pools,
           | and mainstream hardware has consistently used _many_ of these
           | methods _the entire time_.
        
         | v1sea wrote:
         | edit: I was wrong.
        
           | oflordal wrote:
           | You can do that on both HIP and cuda through e.g.
           | hipHostMalloc and the cuda equivalent (Not officially
           | supported on the AMD APUs but works in practice). With a
           | discrete GPU the GPU will access memory across PCIe but on an
           | APU it will go full speed to RAM as far as I can tell.
        
           | smallmancontrov wrote:
           | > low latency results each frame
           | 
           | What does that do to your utilization?
           | 
           | I've been out of this space for a while, but in game dev any
           | backwards information flow (GPU->CPU) completely murdered
           | performance. "Whatever you do, don't stall the pipeline."
           | Instant 50%-90% performance hit. Even if you had to awkwardly
           | duplicate calculations on the CPU, it was almost always worth
           | it, and not by a small amount. The caveat to "almost" was
           | that if you were willing to wait 2-4 frames to get data back,
           | you could do that without stalling the pipeline.
           | 
           | I didn't think this was a memory architecture thing, I
           | thought it was a data dependency thing. If you have to finish
           | all calculations before readback and if you have to readback
           | before starting new calculations, the time for all cores to
           | empty out and fill back up is guaranteed to be dead time,
           | regardless of whether the job was rendering polygons or
           | finite element calculations or neural nets.
           | 
           | Does shared memory actually change this somehow? Or does it
           | just make it more convenient to shoot yourself in the foot?
           | 
           | EDIT: or is the difference that HPC operates in a regime
           | where long "frame time" dwarfs the pipeline empty/refill
           | "dead time"?
        
             | v1sea wrote:
             | It was probably from my workloads being relatively small
             | that I could get away with 90Hz read on the cpu side. I'll
             | need to dig deeper into it. The metrics I was seeing were
             | showing 200-300 microseconds of GPU time for physics
             | calculations and within the same frame the cpu reading from
             | that buffer. Maybe I'm wrong, need to test more.
        
               | smallmancontrov wrote:
               | > edit: I was wrong.
               | 
               | If the only reason you were "wrong" was because you
               | intuitively understood that it wasn't worth a large
               | amount of valuable human time to save a small amount of
               | cheap machine time, you were right in the way that
               | matters (time allocation) and should keep it up :)
        
         | numpad0 wrote:
         | Note that while UMA is great in the sense that they allow LLM
         | models to be run at all, M-series chips aren't faster[1] when
         | the model fits in VRAM.                 1: screenshot from[2]:
         | https://www.igorslab.de/wp-
         | content/uploads/2023/06/Apple-M2-ULtra-SoC-Geekbench-5-OpenCL-
         | Compute.jpg       2: https://wccftech.com/apple-m2-ultra-soc-
         | isnt-faster-than-amd-intel-last-year-desktop-cpus-50-slower-
         | than-nvidia-rtx-4080/
        
           | cstejerean wrote:
           | The problem is you're limited to 24 GB of VRAM unless you pay
           | through the nose for datacenter GPUs, whereas you can get an
           | M-series chip with 128 GB or 192 GB of unified memory.
        
             | numpad0 wrote:
             | Surely! The point is that they're not million times faster
             | magic chips that makes NVIDIA bankrupt tomorrow. That's
             | all. A laptop with up to 128GB "VRAM" is a great option,
             | absolutely no doubt about that.
        
               | john_alan wrote:
               | They are powerful, but I agree with you, it's nice to be
               | able to run Goliath locally, but it's a lot slower than
               | my 4070.
        
           | paulmd wrote:
           | that's openCL compute, LLM models ideally should be hitting
           | the neural accelerator, not running on generalized gpu
           | compute shaders.
        
         | spamizbad wrote:
         | My understanding is the unified RAM on the M-series die does
         | not contribute significantly to their performance. You get a
         | little bit better latency but not much. The real value to Apple
         | is likely it greatly simplifies your design since you don't
         | have to route out tons of DRAM signaling and power management
         | on your logic board. Might make DRAM training easier too but
         | that's way beyond my expertise.
        
           | john_alan wrote:
           | also provides the GPU with serious RAM allocation, a 64GB M3
           | chip comes with more ram for the GPU than a 4090
        
         | AceJohnny2 wrote:
         | > _unified memory model (by having the memory on-package with
         | CPU)_
         | 
         | That's not what "unified memory model" means.
         | 
         | It means that the CPU and GPU (and ANE!) have access to the
         | same banks of memory, unlike PC GPUs that have their own
         | memory, separated from the CPU's by the PCIe bottleneck (as
         | fast as that is, it's still smaller than direct shared DRAM
         | access).
         | 
         | It allows the hardware more flexibility in how the single pool
         | of memory is allocated across devices, and faster sharing of
         | data across devices. (throughput/latency depends on the
         | internal system bus ports and how many each device have access
         | to)
         | 
         | The Apple M-Series chips _also_ has the memory on-package with
         | the CPU (technically SoC,  "System-on-Chip"), but that provides
         | different benefits.
        
           | cmovq wrote:
           | Having separated GPU memory also has its benefits. Once the
           | data makes it through the PCIe bus, graphics memory typically
           | has much higher bandwidth which also doesn't need to split
           | with the CPU.
        
             | crawshaw wrote:
             | An M2 Ultra has 800GB/s of memory bandwidth, an Nvidia 4090
             | has 1008GB/s. Apple have chosen to use relatively little
             | system memory at unusually high bandwidth.
        
           | fulafel wrote:
           | Most x86 machines have integrated GPUs and hardware-wise are
           | UMA.
        
             | alacritas0 wrote:
             | integrated GPUs are not powerful in comparison to dedicated
             | GPUs
        
       | Havoc wrote:
       | Are these NPUs addressable with a standard PyTorch LLM stack?
       | 
       | NPU seems to mean very different things depending on
       | device/vendor
        
       | bearjaws wrote:
       | The focus on TOPS seems a bit out of line with reality for LLMs.
       | TOPs doesn't matter for LLMs if your memory bandwidth can't keep
       | up. Since it doesn't have quad channel memory mentioned I guess
       | it's still dual channel?
       | 
       | Even top of the line DDR5 is around 128GB/s vs a M1 at 400GB/s.
       | 
       | At the end of the day, it still seems like AI in consumer chips
       | is chasing a buzzword, what is the killer feature?
       | 
       | On mobile there are image processing benefits and voice to text,
       | translation... but on desktop those are no where near common use
       | cases.
        
         | postalrat wrote:
         | https://www.neatvideo.com/blog/post/m3
         | 
         | That says M1 is 68.25 GB/s
        
           | givinguflac wrote:
           | The op were obviously talking about M1 Max.
        
             | postalrat wrote:
             | How is it obvious? Anyone reading that could assume that
             | any M1 gets that bandwidth.
        
           | bearjaws wrote:
           | M1 Max sorry, I don't mean to compare a 4 year old tablet
           | processor to the latest generation of laptop CPUs.
        
         | VHRanger wrote:
         | The killer feature is presumably inference at the edge, but I
         | don't see that being used on desktop much at all right now.
         | 
         | Especially since most desktop applications people use are web
         | apps. Of the native apps people use that leverage this sort of
         | stuff, almost all are GPU accelerated already (eg. image and
         | video editing AI tools)
        
           | jzig wrote:
           | What does "at the edge" mean here?
        
             | georgeecollins wrote:
             | Not using AI on the cloud. So if your connection is
             | uncertain or you want use your bandwidth for something
             | else-- like video conferencing or gaming. Probably the
             | killer app is something that wants to use AI but doesn't
             | involve paying a cloud provider. I was talking to a vendor
             | about their chat bot built to put into MMOs or mobile
             | games. It woudl be killer to have a character have life
             | like conversation in those kinds of experiences. But the
             | last thing you want to do is increase your server costs the
             | way this AI would. Edge computing could solve that.
        
             | PeterSmit wrote:
             | Not in the cloud.
        
             | Zach_the_Lizard wrote:
             | I'm guessing "the edge" is doing inference work in the
             | browser, etc. as opposed to somewhere in the backend of the
             | web app.
             | 
             | Maybe your local machine can run, I don't know, a model to
             | make suggestions as you're editing a Google Doc, which
             | frees up the Big Machine in the Sky to do other things.
             | 
             | As this becomes more technically feasible, it reduces the
             | effective cost of inference for a new service provider,
             | since you, the client, are now running their code.
             | 
             | The Jevons paradox might kick in, causing more and more
             | uses of LLMs for use cases that were too expensive before.
        
             | VHRanger wrote:
             | Edge is doing computing on the client (eg. browser, phone,
             | laptop, etc.) instead of the server
        
               | Dylan16807 wrote:
               | Half the definitions I see of edge include client
               | devices, and half of them don't include client devices.
               | 
               | I like the latter. Why even use a new word if it's just
               | going to be the same as "client"?
        
         | futureshock wrote:
         | Upscaling for gaming or video.
         | 
         | Local context aware search
         | 
         | Offline Speech to text and TTS
         | 
         | Offline generation of clip art or stock images for document
         | editing
         | 
         | Offline LLM that can work with your documents as context and
         | access application and OS APIs
         | 
         | Improved enemy AI in gaming
         | 
         | Webcam effects like background removal or filters.
         | 
         | Audio upscaling and interpolation like for bad video call
         | connections.
        
           | bearjaws wrote:
           | > Upscaling for gaming or video.
           | 
           | Already exists on all three major GPU manufacturers, and it
           | definitely makes sense as a GPU workload.
           | 
           | > Local context aware search
           | 
           | You don't need an AI processor to do this, Windows search
           | used to work better and had even less compute resources to
           | work with.
           | 
           | > Offline Speech to text and TTS
           | 
           | See my point about not a very common use case for desktops &
           | laptops vs cell phones.
           | 
           | > Offline LLM that can work with your documents as context
           | and access application and OS APIs
           | 
           | Maybe for some sort of background task or only using really
           | small models <13B parameters. Anything real time is going to
           | run at 1-2t/s with a large model.
           | 
           | Small models are pretty terrible though, I doubt people want
           | even more incorrect information and hallucinations.
           | 
           | > Improved enemy AI in gaming
           | 
           | See Ageia PhysX
           | 
           | > Webcam effects like background removal or filters.
           | 
           | We already have this without NPUs.
           | 
           | > Audio upscaling and interpolation like for bad video call
           | connections.
           | 
           | I could see this, or noise cancellation.
        
             | bayindirh wrote:
             | It's about power management, and doing more things with
             | less power. These specialized IP blocks on CPUs allow these
             | things to be done with less power and less latency.
             | 
             | Intel's bottom of the barrel N95 & N100 CPUs have Gaussian
             | & Neural accelerators for simple image processing and
             | object detection tasks, plus a voice processor for low
             | power voice based activation and command capture and
             | process.
             | 
             | You can always add more power hungry, general purpose
             | components to add capabilities. Heck, video post processing
             | entered hardware era with ATI Radeon 8500. But doing these
             | things with negligible power costs is the new front.
             | 
             | Apple is not adding coprocessors to their iPhones because
             | it looks nice. All of these coprocessors reduce CPU wake-up
             | cycles tremendously and allows the device to monitor tons
             | of things out of bands with negligible power costs.
        
             | pdpi wrote:
             | >> Upscaling for gaming or video.
             | 
             | > Already exists on all three major GPU manufacturers, and
             | it definitely makes sense as a GPU workload.
             | 
             | "makes sense as a GPU workload" is underselling it a bit.
             | Doing it on the CPU is basically insane. Games typically
             | upscale only the world view (the expensive part to render)
             | while rendering the UI at full res. So to do CPU-side
             | upscaling we're talking about a game rendering a surface on
             | the GPU, sending it to the CPU, upscaling it there, sending
             | it back to the GPU, then compositing with the UI. It's just
             | needlessly complicated.
        
             | futureshock wrote:
             | > Upscaling for gaming or video. > Already exists on all
             | three major GPU manufacturers, and it definitely makes
             | sense as a GPU workload. These AMD chips are APUs that are
             | often the only GPU, not every user will have a dedicated
             | GPU.
             | 
             | > Local context aware search > You don't need an AI
             | processor to do this, Windows search used to work better
             | and had even less compute resources to work with. You could
             | still improve it with increased natural language
             | understanding instead of simple keyword. "Give me all
             | documents about dogs" instead of searching for each breed
             | as a keyword.
             | 
             | > Offline Speech to text and TTS > See my point about not a
             | very common use case for desktops & laptops vs cell phones.
             | Maybe not for you but accessibility is a key feature for
             | many users. You think blind users should suffer through bad
             | TTS?
             | 
             | > Offline LLM that can work with your documents as context
             | and access application and OS APIs > Maybe for some sort of
             | background task or only using really small models <13B
             | parameters. Anything real time is going to run at 1-2t/s
             | with a large model. > Small models are pretty terrible
             | though, I doubt people want even more incorrect information
             | and hallucinations. Small model have been improving and
             | better capabilities in consumer chips will allow larger
             | models to run faster.
             | 
             | > Improved enemy AI in gaming > See Ageia PhysX Surely
             | you're not suggesting that enemy AI is solved problem in
             | gaming?
             | 
             | > Webcam effects like background removal or filters. > We
             | already have this without NPUs. Sure but it could go from
             | obvious and distracting to seamless and convincing.
             | 
             | > Audio upscaling and interpolation like for bad video call
             | connections. I could see this, or noise cancellation.
        
           | kanbankaren wrote:
           | All of this(except upscaling) is possible with iGPU/CPU
           | without breaking a sweat?
        
             | bayindirh wrote:
             | The things which doesn't make GPU to break a sweat has its
             | own specialized (or semi-specialized) processing blocks on
             | the GPU, too.
        
               | kanbankaren wrote:
               | I meant the current generation of GPUs that don't have
               | any AI acceleration blocks.
        
               | bayindirh wrote:
               | They are MATMUL machines by design already. They do not
               | need to "accelerate" AI to begin with.
               | 
               | Their cores/shaders can be programmed to do that.
               | 
               | Also, name a current gen GPU which doesn't have video
               | encoding/decoding capabilities/facilities in silicon,
               | even ones which do not allow shaders to be used in this
               | process for post-processing. It's impossible (to not to
               | have these capabilities) at this point in time.
        
               | kanbankaren wrote:
               | I was talking about AI blocks and you moved the goal post
               | to video codec blocks.
        
               | bayindirh wrote:
               | No. I didn't move anything.
               | 
               | I said that the core (3D rendering hardware) of a GPU
               | with shaders is _the_ AI block already, and said that
               | other tasks like video encoders have their own blocks,
               | but still pull capabilities from the  "core" to improve
               | things.
        
       | phkahler wrote:
       | Am I missing something? These look just like the APUs with the
       | addition of "management" and "security" features and without the
       | iGPU. Is that right?
        
         | c0l0 wrote:
         | They also support ECC UDIMMs with ECC enabled, which has been
         | the "PRO" series APU killer feature on AM4 for me. The
         | non-"PRO" APUs will run fine with ECC UDIMMs, but cannot make
         | use of the extra parity information (maybe for reasons of
         | market segmentation - I don't know if anyone outside of AMD
         | knows). This is probably less of a concern with DDR5 platforms
         | and their "on-DIE ECC" (which you cannot monitor for
         | Correctable Errors at least, afaik), but it's still gonna
         | matter for me.
        
           | nwah1 wrote:
           | Yes. And it actually does have an iGPU, depending on which
           | model.
        
             | c0l0 wrote:
             | In case you don't need an integrated GPU (that's somewhat
             | powerful/potent), you can go with any other Ryzen AM5 CPU
             | to receive proper ECC-enabled ECC UDIMM support, afaik :)
        
           | pedrocr wrote:
           | Since in AM5 all CPUs have a basic iGPU, for a home server
           | all the normal CPUs already work fine. The advantage of the
           | APUs is they're on a monolithic die so should have quite a
           | bit lower idle power usage which is important if you have a
           | NAS or other homelab server running 24/7.
        
       | kokonoko wrote:
       | I hope (but I doubt) that this will be more than a marketing
       | stand to include something "AI" in their product line. Every
       | vendor has their own hardware that is badly supported by tools,
       | and even if it is supported, only a fraction of the available
       | software uses it. In the meantime it takes precious die area and
       | resources.
        
       | Aissen wrote:
       | A quick search into it shows that this Ryzen AI NPU's support
       | isn't integrated into upstream inference frameworks yet -- so
       | right now it's just useless silicon surface you pay for :-/
        
         | Rinzler89 wrote:
         | Some AMD laptops haven't even yet enabled the NPU in firmware
         | even on the 7000 series wich are about a year old. Meaning it's
         | still useless.
         | 
         | I was kinda bummed out they released the 8000 series after I
         | just bought a laptop with 7000 series, but I think I actually
         | dodged a bullet here since it doesn't look like much of an
         | upgrade and the AI silicone screams of very early first gen
         | product to me, as if they rushed it out the door because
         | everyone else was doing "AI" and they needed to also cash in on
         | the hype, kinda like the first gen RTX cards.
         | 
         | I think by the time I'll actually upgrade, the AI/NPU tech
         | would have matured considerably and actually be useful.
        
           | robocat wrote:
           | Does anyone have any mental heuristics for judging how
           | "useless" a feature is?
           | 
           | Over decades I have a growing antipathy towards products with
           | too many features. Especially new versions/models where the
           | vaunted features of the previous version/model seem to never
           | have been used by anyone.
        
             | Rinzler89 wrote:
             | _> Does anyone have any mental heuristics for judging how
             | "useless" a feature is? _
             | 
             | My favorite example is the story I got to live through of
             | the first generations of consumer 64 bit CPUs.
             | 
             | When the first AMD Athlon 64 came out, everyone I knew was
             | buying them because they though they were getting something
             | totally future proof by jumping early on the 64 bit
             | bandwagon, in 2003, when nobody yet had 4GB+ of RAM and
             | neither Windows nor any software would see 64bit releases
             | till several years later when Vista came out which everyone
             | avoided and staid on Windows XP 32bit waiting for Windows
             | 7.
             | 
             | And by the time RAM sizes over 4GB and 64 bit software
             | became even remotely mainstream, we already had dual- and
             | quad-core CPUs miles ahead of those original 64 bit CPUs
             | which were now obsolete (tech progress back then was wild).
             | 
             | So just like how 64bit silicone was a useless feature on
             | consumer CPUs, and like the first GPUs with raytracing, I
             | feel like now we're in the same boat with AI silicone in
             | PCs, no much SW support for them and when it does come,
             | these early chips will be obsolete. It's the price of being
             | an early adopter.
        
           | sva_ wrote:
           | > Some AMD laptops haven't even yet enabled the NPU in
           | firmware
           | 
           | This is entirely the fault of the OEMs though, not AMD. It is
           | activated on mine for example. But pretty much unusable under
           | Linux at the moment (unless you're willing to run a custom
           | kernel for it[0].)
           | 
           | 0. https://github.com/amd/xdna-driver
        
             | Rinzler89 wrote:
             | _> This is entirely the fault of the OEMs though, not AMD._
             | 
             | Not true. AMD can demand how OEMs integrate and use their
             | chips in their products as part of the sales agreement,
             | same how Nvidia does.
             | 
             | AMD could have said to every system integrator buying 7000
             | series chips and up, that the NPU must be active in the
             | final product.
             | 
             | So if the end products suck, AMD bares most of the blame
             | for not ensuring a minimum level of QA with its integrators
             | who release half-assed stuff since it all reflects poorly
             | on them in the end. It's one of the reason why Nvidia keeps
             | such a tight grip over its integrators on how their chips
             | are to used.
        
         | dhruvdh wrote:
         | There is a VitisAI execution provider for ONNX, and you can use
         | ONNX backends for inference frameworks that support it. More
         | info here - https://ryzenai.docs.amd.com/en/latest/
         | 
         | But regardless, 16 TOPs is no good for LLMs. Though there is a
         | Ryzen AI demo that shows Llama 7B running on these at 8
         | tokens/sec. A sub-par experience for a sub-par LLM.
        
           | Aissen wrote:
           | Thanks, I was looking for information on this, it seems to be
           | lower speed than pure-CPU inference on M2, and probably much
           | worse than a ROCm GPU-based solution?
        
             | p_l wrote:
             | Because the NPU isn't for high-end inferencing. It's a
             | relatively small coprocessor that is supposed to do bunch
             | of tasks with high TOPS/watt without engaging the way more
             | power hungry GPU.
             | 
             | At release time, the windows driver for example included
             | few video processing offloads used by Windows Frameworks
             | used for example by MS Teams for background removal - so
             | that such tasks use less battery on laptops and free up
             | CPU/GPU for other tasks on desktop.
             | 
             | For higher end processing you can use the same AIE-ML
             | coprocessors various chips available previously from Xilinx
             | and now under AMD brand.
        
               | fpgamlirfanboy wrote:
               | > the same AIE-ML coprocessors
               | 
               | they're not the same - versal acaps (whatever you want to
               | call them) have AIE1 arch while phoenix has AIE2 arch.
               | there are significant differences between the two arches
               | (local memory, bfloat16, etc.)
        
               | p_l wrote:
               | Phoenix has AIE-ML (what you call AIE2), Versal has
               | choice of AIE (AIE1) and AIE-ML (AIE2) depending on chip
               | you buy.
               | 
               | Essentially, AMD is making two tile designs optimized for
               | slightly different computations and claims that they are
               | going to offer both in Versal, but NPUs use exclusively
               | the ML-optimized ones.
        
           | markdog12 wrote:
           | Wow, that's simply embarrassing.
        
       | myself248 wrote:
       | What does "commercial market" mean here? The article says these
       | features were already available in the "consumer" versions of
       | these chips -- were those given away for free or something? What
       | about them was not "commercial"?
       | 
       | I'm sure there's some market segmentation thing at work here, but
       | this just sounds like a Hallmark-holiday excuse to rehash an old
       | press-release and pretend it's another revolution all over again.
        
         | dhruvdh wrote:
         | You can't buy these pro variants from Microcenter for example,
         | but you can buy them from pre-built OEM desktops. Mostly meant
         | for enterprise customers who buy in bulk, I think.
        
         | transpute wrote:
         | _> the article says these features were already available_
         | 
         | Actually, the article says they were _not_ available:
         | the Pro series is based on AMD's existing consumer-oriented
         | processor models but comes with additional features
         | 
         | Pro (enterprise) CPUs include remote management, memory
         | encryption and other security features:
         | https://www.amd.com/en/ryzen-pro
        
       | baarsh wrote:
       | How come new series comes with options only up to 8 cores, while
       | 5900X and 5950X were already 12 and 16 cores few years ago?
        
         | Arrath wrote:
         | Better sustained performance with the thermal envelope afforded
         | to only 8 cores vs 12+?
        
         | wmf wrote:
         | Desktop Ryzen = up to 16 cores
         | 
         | Laptop Ryzen = up to 8 cores (the 8000G are laptop CPUs in a
         | desktop socket)
        
           | protastus wrote:
           | Dragon Range (e.g., 7945HX) is a laptop/mobile workstation
           | Ryzen with up to 16 cores, but power inefficient compared to
           | the 8000-series due to the chiplet design. Already in market,
           | mostly in gaming laptops.
        
       | JonChesterfield wrote:
       | Zen4 is _very_ pretty, see
       | https://news.ycombinator.com/item?id=32983406.
       | 
       | I like their APUs a lot. Using a 4800U in the cable tray under
       | the desk to drive screens on which I'm writing this. One in a
       | laptop for whenever I'm away from the desk.
       | 
       | If you're sufficiently determined the compute units on these
       | things are totally usable for running arbitrary code. As in a
       | program that spawns a bunch of threads to work stuff out could
       | have some of those "threads" running on the GPU cores.
       | 
       | I wouldn't say the software stack is totally there for out of the
       | box convenience. As in you'll be writing in freestanding ~C and
       | maybe a bit of assembly. I got partway through implementing that
       | and got sidetracked. The GPU libc in LLVM is roughly the
       | production version of some of that hacking. Between these
       | machines coming out and the MI300A landing I really should put
       | something up on github which looks like a pthread_create that
       | executes on the GPU instead.
        
       ___________________________________________________________________
       (page generated 2024-04-16 23:02 UTC)