https://www.tomshardware.com/news/nvidia-unveils-dgx-gh200-supercomputer-and-mgx-systems-grace-hopper-superchips-in-production Skip to main content (*) ( ) Open menu Close menu Tom's Hardware [ ] Search Search Tom's Hardware [ ] RSS US Edition flag of US flag of UK UK flag of US US flag of Australia Australia flag of Canada Canada * * Reviews * Best Picks * Raspberry Pi * CPUs * GPUs * Coupons * Newsletter * More + Laptops + SSDs + Motherboards + Cooling + Desktops + PC Builds + Monitors + RAM + PC Cases + Keyboards + Headsets + Mice + Power Supplies + 3D Printers + Gaming Chairs + Webcams + Microphones + VR Headsets + About Tom's Hardware Forums Trending * Computex 2023 * Memorial Day Deals * Try Our AI Chatbot * Where to Buy RTX 4060 Ti When you purchase through links on our site, we may earn an affiliate commission. Here's how it works. 1. Home 2. News Nvidia Unveils DGX GH200 Supercomputer, Grace Hopper Superchips in Production By Paul Alcorn published 29 May 2023 Nvidia's new supercomputing silicon fuels the fires of AI. * * * * * * * Comments (3) Nvidia Grace Hopper (Image credit: Tom's Hardware) Nvidia CEO Jensen Huang announced here at Computex 2023 in Taipei, Taiwan that the company's Grace Hopper superchips are now in full production, and the Grace platform has now earned six supercomputer wins. These chips are a fundamental building block of one of Huang's other big Computex 2023 announcements: The company's new DGX GH200 AI supercomputing platform, built for massive generative AI workloads, is now available with 256 Grace Hopper Superchips paired together to form a supercomputing powerhouse with 144TB of shared memory for the most demanding generative AI training tasks. Nvidia already has customers like Google, Meta, and Microsoft ready to receive the leading-edge systems. Nvidia also announced its new MGX reference architectures that will help OEMs build new AI supercomputers faster with up to 100+ systems available. Finally, the company also announced its new Spectrum-X Ethernet networking platform that is designed and optimized specifically for AI server and supercomputing clusters. Let's dive in. Image 1 of 8 Nvidia (Image credit: Tom's Hardware) Nvidia (Image credit: Tom's Hardware) Nvidia (Image credit: Tom's Hardware) Nvidia (Image credit: Tom's Hardware) Nvidia (Image credit: Tom's Hardware) Nvidia (Image credit: Tom's Hardware) Nvidia (Image credit: Tom's Hardware) Nvidia (Image credit: Tom's Hardware) Nvidia Grace Hopper Superchips Now in Production Nvidia (Image credit: Nvidia) We've covered the Grace and Grace Hopper Superchips in depth in the past. These chips are central to Nidia's new systems that it announced today. The Grace chip is Nvidia's own Arm CPU-only processor, and the Grace Hopper Superchip combines the Grace 72-core CPU, a Hopper GPU, 96GB of HBM3, and 512 GB of LPDDR5X on the same package, all weighing in at 200 billion transistors. This combination provides astounding data bandwidth between the CPU and GPU, with up to 1 TB/s of throughput between the CPU and GPU offering a tremendous advantage for certain memory-bound workloads. With the Grace Hopper Superchips now in full production, we can expect systems to come from a bevy of Nidia's systems partners, like Asus, Gigabyte, ASRock Rack, and Pegatron. More importantly, Nvidia is rolling out its own systems based on the new chips and is issuing reference design architectures for OxMs and hyperscalers, which we'll cover below. Nvidia DGX GH200 Supercomputer Image 1 of 2 Supercomputer (Image credit: Nvidia ) Supercomputer (Image credit: Nvidia ) Nvidia's DGX systems are its go-to system and reference architecture for the most demanding AI and HPC workloads, but the current DGX A100 systems are limited to eight A100 GPUs working in tandem as one cohesive unit. Given the explosion of generative AI, Nvidia's customers are eager for much larger systems with much more performance, and the DGX H200 is designed to offer the ultimate in throughput for massive scalability in the largest workloads, like generative AI training, large language models, recommender systems and data analytics, by sidestepping the limitations of standard cluster connectivity options, like InfiniBand and Ethernet, with Nvidia's custom NVLink Switch silicon. Details are still slight on the finer aspects of the new DGX GH200 AI supercomputer, but we do know that Nvidia uses a new NVLink Switch System with 36 NVLink switches to tie together 256 GH200 Grace Hopper chips and 144 TB of shared memory into one cohesive unit that looks and acts like one massive GPU. The new NVLink Switch System is based on its NVLink Switch silicon that is now in its third generation. The DGX GH200 comes with 256 total Grace Hopper CPU+GPUs, easily outstripping Nvidia's previous largest NVLink-connected DGX arrangement with eight GPUs, and the 144TB of shared memory is 500X more than the DGX A100 systems that offer a 'mere' 320GB of shared memory between eight A100 GPUs. Additionally, expanding the DGX A100 system to clusters with more than eight GPUs requires employing InfiniBand as the interconnect between systems, which incurs performance penalties. In contrast, the DGX GH200 marks the first time Nvidia has built an entire supercomputer cluster around the NVLink Switch topology, which Nvidia says provides up to 10X the GPU-to-GPU and 7X the CPU-to-GPU bandwidth of its previous-gen system. It's also designed to provide 5X the interconnect power efficiency (likely measured as PJ/bit) than competing interconnects, and up to 128 TB/s of bisectional bandwidth. The system has 150 miles of optical fiber and weighs 40,000 lbs, but presents itself as one single GPU. Nvidia says the 256 Grace Hopper Superchips propel the DGX GH200 to one exaflop of 'AI performance,' meaning that value is measured with smaller data types that are more relevant to AI workloads than the FP64 measurements used in HPC and supercomputing. This performance comes courtesy of 900 GB/s of GPU-to-GPU bandwidth, which is quite impressive scalability given that Grace Hopper tops out at 1 TB/s of throughput with the Grace CPU when connected directly together on the same board with the NVLink-C2C chip interconnect. Nvidia provided projected benchmarks of the DGX GH200 with the NVLink Switch System going head-to-head with a DGX H100 cluster tied together with InfiniBand. Nvidia used varying numbers of GPUs for the above workload calculations, ranging from 32 to 256, but each system employed the same number of GPUs for each test. As you can see, the explosive gains in interconnect performance are expected to unlock anywhere from 2.2X to 6.3X more performance. Nvidia will provide the DGX GH200 reference blueprints to its leading customers, Google, Meta, and Microsoft, before the end of 2023, and will also provide the system as a reference architecture design for cloud service providers and hyperscalers. Nvidia is eating its own dogfood, too; the company will deploy a new Nvidia Helios supercomputer comprised of four DGX GH200 systems that it will use for its own research and development work. The four systems, which total 1,024 Grace Hopper Superchips, will be tied together with Nvidia's Quantum-2 InfiniBand 400 Gb/s networking. Nvidia MGX Systems Reference Architectures Image 1 of 1 MGX (Image credit: Nvidia) While DGX steps in for the highest-end systems, Nvidia's HGX systems step in for hyperscalers. However, the new MGX systems step in as the middle point between these two systems, and DGX and HGX will continue to co-exist with the new MGX systems. Nvidia's OxM partners face new challenges with AI-centric server designs, thus slowing design and deployment. Nvidia's new MGX reference architectures are designed to speed that process with 100+ reference designs. The MGX systems comprise modular designs that span the gamut of Nvidia's portfolio of CPUs and GPUs, DPUs, and networking systems, but also include designs based on the common x86 and Arm-based processors found in today's servers. Nvidia also provides options for both air- and liquid-cooled designs, thus providing OxMs with different design points for a wide range of applications. Image 1 of 7 Nvidia (Image credit: Tom's Hardware) Nvidia (Image credit: Tom's Hardware) Nvidia (Image credit: Tom's Hardware) Nvidia (Image credit: Tom's Hardware) Nvidia (Image credit: Tom's Hardware) Nvidia (Image credit: Tom's Hardware) Nvidia (Image credit: Tom's Hardware) Naturally, Nvidia points out that the lead systems from QCT and Supermicro will be powered by its Grace and Grace Hopper Superchips, but we expect that x86 flavors will probably have a wider array of available systems over time. Asus, Gigabyte, ASRock Rack and Pegatron will all use MGX reference architectures for systems that will come to market later this year into early next year. The MGX reference designs could be the sleeper announcement of Nvidia's Computex press blast - these will be the systems that mainstream data centers and enterprises will eventually deploy to infuse AI-centric architectures into their deployments, and will ship in far greater numbers than the somewhat exotic and more costly DGX systems - these are the volume movers. Nvidia is still finalizing the spec, which will be public, and will release a whitepaper soon. Nvidia Spectrum-X Networking Platform Image 1 of 4 Switch (Image credit: Nvidia) Switch (Image credit: Nvidia) Switch (Image credit: Nvidia) Switch (Image credit: Nvidia) Nvidia's purchase of Mellanox has turned out to be a pivotal move for the company, as it can now optimize and tune networking componentry and software for its AI-centric needs. The new Spectrum-X networking platform is perhaps the perfect example of those capabilities, as Nvidia touts it as the 'world's first high-performance Ethernet for AI' networking platform. Image 1 of 8 Nvidia (Image credit: Tom's Hardware) Nvidia (Image credit: Tom's Hardware) Nvidia (Image credit: Tom's Hardware) Nvidia (Image credit: Tom's Hardware) Nvidia (Image credit: Tom's Hardware) Nvidia (Image credit: Tom's Hardware) Nvidia (Image credit: Tom's Hardware) Nvidia (Image credit: Tom's Hardware) One of the key points here is that Nvidia is pivoting to Ethernet as an interconnect option for high-performance AI platforms, as opposed to the InfiniBand connections often found in high-performance systems. The Spectrum-X design employs Nvidia's 51 Tb/s Spectrum-4 400 GbE Ethernet switches and the Nvidia Bluefield-3 DPUs paired with software and SDKs that allow developers to tune systems for the unique needs of AI workloads. In contrast to other Ethernet based systems, Nvidia says Spectrum-X is lossless, thus providing superior QoS and latency. It also has new adaptive routing tech, which is particularly helpful in multi-tenancy environments. The Spectrum-X networking platform is a foundational aspect of Nvidia's portfolio, as it brings high-performance AI cluster capabilities to Ethernet-based networking, offering new options for wider deployments of AI into hyperscale infrastructure. The Spectrum-X platform is also fully interoperable with existing Ethernet-based stacks and offers impressive scalability with up to 256 200 Gb/s ports on a single switch, or 16,000 ports in a two-tier leaf-spine topology. The Nvidia Spectrum-X platform and its associated components, including 400G LinkX optics, are available now. Nvidia Grace and Grace Hopper Superchip Supercomputing Wins Nvidia's first Arm CPUs (Grace) have already been in production and made an impact with three recent supercomputer wins, including the newly announced Taiwania 4 which will be built by computing vendor ASUS for the Taiwan National Center for High-Performance Computing. This system will feature 44 Grace CPU nodes, and Nvidia claims it will rank among the most energy-efficient supercomputers in Asia when deployed. The supercomputer will be used to model climate change issues. Nvidia also shared details of its new Taipei 1 supercomputer that will be based in Taiwan. This system will have 64 DGX H100 AI supercomputers and 64 Nvidia OVX systems tied together with the company's networking kit. This system will be used to further unspecified local R&D workloads when it is finished later this year. Stay on the Cutting Edge Join the experts who read Tom's Hardware for the inside track on enthusiast PC tech news -- and have for over 25 years. We'll send breaking news and in-depth reviews of CPUs, GPUs, AI, maker hardware and more straight to your inbox. [ ][ ]Contact me with news and offers from other Future brands[ ]Receive email from us on behalf of our trusted partners or sponsors[Sign me up] By submitting your information you agree to the Terms & Conditions and Privacy Policy and are aged 16 or over. Paul Alcorn Paul Alcorn Social Links Navigation Deputy Managing Editor Paul Alcorn is the Deputy Managing Editor for Tom's Hardware US. He writes news and reviews on CPUs, storage and enterprise hardware. More about cpus Meteor Lake Intel Demos Meteor Lake's AI Acceleration for PCs, Details VPU Unit TSMC 2023 Technology Symposium Pumped Up Procs: TSMC Planning Chips 3x Bigger Than Today Latest Hyte THICC Q60 AiO cooler Hyte THICC Q60 AiO Water Block Includes 5-Inch IPS Display See more latest > Topics CPUs See all comments (3) [ ] 3 Comments Comment from the forums * bit_user Thanks for the writeup! up to 1 TB/s of throughput between the CPU and GPU offering a tremendous advantage for certain memory-bound workloads. The way they have a dual-CPU configuration option can lead to some confusion about whether the stats are referring to a single CPU or dual-CPU config. As you can see from their developer blog post, the 1 TB/s is the memory bandwidth across 2 Grace CPUs. The Chip-to-Chip link is only 900 GB/s. When a GPU is paired with a CPU, the GPU would be bottlenecked by the CPU's 500 GB/s memory interface. Then again, I'd bet Nvidia is arriving at 900 GB/s by summing throughput in each direction, in which case it could only read or write CPU memory at a peak of @ 450 GB/s. Source: https://developer.nvidia.com/blog/ nvidia-grace-cpu-superchip-architecture-in-depth/ The system has 150 miles of optical fiber and weighs 40,000 lbs, but presents itself as one single GPU. Nvidia says the 256 Grace Hopper Superchips ... Really? That's 156 pounds per "superchip". That sounds like pretty bad scaling, actually. @PaulAlcorn , in the 7-slide album under "Nvidia MGX Systems Reference Architectures", the last 2 slides appear to be duplicates of slides 2 & 3. Not that I really care that much about the details of MGX, but I did notice. ; ) The MGX reference designs could be the sleeper announcement of Nvidia's Computex press blast Yeah, but let's call it out for what it is: an end-run around OCP (Open Compute Project). The industry has a nice, open multi-vendor ecosystem that already supports modular accelerators. Nvidia is pivoting to Ethernet as the interconnect for high-performance AI platforms "Never bet against Ethernet." @PaulAlcorn, in the album under "Nvidia Spectrum-X Networking Platform", the second image seems to be missing? Nvidia's first Arm CPUs (Grace) ... Server CPUs, that is. They've been making ARM-based SoCs for embedded applications since about 1.5 decades ago. Reply * zecoeco Nvidia throwing the new buzzword "AI" in every new product they announce. They just won't stop overhyping AI.. Reply * ngilbert zecoeco said: Nvidia throwing the new buzzword "AI" in every new product they announce. They just won't stop overhyping AI.. Given that it is the buzzword of the month across the industry, and they have a massive lead in that type of workload, they would be foolish not to push it as far as they can. Reply * View All 3 Comments Most Popular [missing-im]Lord of the Rings: Gollum Benchmarked. You'll Need a Precious GPU. By Jarred WaltonMay 29, 2023 [missing-im]Intel Demos Meteor Lake's AI Acceleration for PCs, Details VPU Unit By Paul AlcornMay 29, 2023 [missing-im]Nvidia Unveils DGX GH200 Supercomputer, Grace Hopper Superchips in Production By Paul AlcornMay 29, 2023 [missing-im]Nvidia ACE Brings AI to Game Characters, Allows Lifelike Conversations By Avram PiltchMay 29, 2023 [missing-im]How to Watch Nvidia's Computex 2023 Keynote By Avram PiltchMay 29, 2023 [missing-im]AMD Ryzen 5 7600X CPU Down to Just $209 at Newegg By Ash HillMay 28, 2023 [missing-im]Acer Testing Distinctive GeForce RTX 4090 With Integrated Liquid Cooler By Mark TysonMay 28, 2023 [missing-im]MSI Mech Radeon RX 6750 XT 12GB GPU Only $319 at Newegg By Ash HillMay 28, 2023 [missing-im]2TB SSD For PS5 Drops to $99 at Newegg By Ash HillMay 28, 2023 [missing-im]Sony Project Q Handheld Rumored to Offer Just 3 or 4 Hours Battery Life By Mark TysonMay 28, 2023 [missing-im]Raspberry Pi Adapter Sends Keyboard Input From iPad via HID to Devices By Ash HillMay 28, 2023 MOST POPULARMOST SHARED 1. GeForce RTX 4070 Megalodon 1 Asus Demos RTX 4070 GPU With Zero Power Connectors 2. 2 34-Inch LG QHD Curved Gaming Monitor Drops to $299 3. 3 Cooler Master Unveils Full Line of Street Fighter Components, Peripherals 4. 4 Nvidia Introduces ULMB 2, Boosting Motion Clarity To 1400 Hz 5. 5 Raspberry Pi Malware Infects Using Default Username and Password 1. Hyte THICC Q60 AiO cooler 1 Hyte THICC Q60 AiO Water Block Includes 5-Inch IPS Display 2. 2 Cooler Master's Tiny NCore 100 Max Case Finds Room for RTX 4090 3. 3 Asus Demos RTX 4070 GPU With Zero Power Connectors 4. 4 34-Inch LG QHD Curved Gaming Monitor Drops to $299 5. 5 Cooler Master Unveils Full Line of Street Fighter Components, Peripherals Tom's Hardware is part of Future US Inc, an international media group and leading digital publisher. Visit our corporate site. * Terms and conditions * Contact Future's experts * Privacy policy * Cookies policy * Accessibility Statement * Advertise * About us * Coupons * Careers (c) Future US, Inc. Full 7th Floor, 130 West 42nd Street, New York, NY 10036.