https://www.tomshardware.com/tech-industry/artificial-intelligence/exacluster-with-144-nvidia-h200-ai-gpus-detailed-by-its-designer-hydra-host-enters-the-scene Skip to main content (*) ( ) Open menu Close menu Tom's Hardware [ ] Search Search Tom's Hardware [ ] RSS US Edition flag of US flag of UK UK flag of US US flag of Australia Australia flag of Canada Canada * * Best Picks * Raspberry Pi * CPUs * GPUs * 3D Printers * News * Coupons * More + Newsletter + Reviews + PC Components + Motherboards + SSDs + PC Building + Monitors + Laptops + Gaming + Cooling + RAM + Power Supplies + Cases + 3D Printers + Desktops + Overclocking + Peripherals + About Us Forums Trending * RTX 5080 * Arrow Lake's Fix * Arc B570 * Ryzen 9 9950X3D * Snapdragon X * Blackwell 1. Tech Industry 2. Artificial Intelligence Exacluster with 144 Nvidia H200 AI GPUs detailed by its designer: Hydra Host enters the scene News By Anton Shilov published 31 January 2025 Best performance at best price enabled by Hydra Host. * * * * * * * Comments (2) When you purchase through links on our site, we may earn an affiliate commission. Here's how it works. ExaAI's H200 cluster (Image credit: Will Bryk/X) Earlier this month, we reported on ExaAILabs's Exacluster, a cluster of 18 machines running 144 Nvidia H200 GPUs, which happens to be one of the first clusters based on these processors. Since then, Hydra Host, the company that facilitated the construction of the cluster, has given us additional details about the system. The cluster uses Lenovo systems with multiple customizations from Hydra Host, which played a significant role. The machine can also be rented -- when not in use by the owner -- through Hydra's Brokkr platform. A Lot of Compute Power The cluster's backbone consists of 18 Lenovo nodes equipped with 144 Nvidia H200 GPUs and 20TB of HBM3E memory -- or eight per system -- enabling compute performance of 570 FP8 PetaTOPS for AI. 16 nodes are configured and fine-tuned by HydraHost for training, which requires massive computation and memory performance, while the remaining two serve as inference nodes. In addition, Hydra Host installed its Brokkr platform for GPU provisioning, management, and remote renting (more on this later). Hydra Host collaborated with Computacenter to design a high-performance networking architecture tailored to the cluster's needs. The setup uses 3.2Tbps InfiniBand for east-west traffic and 400Gbps Ethernet for north-south communication, including dual 200Gbps connections per server and 400Gbps Dell Ethernet switches. Computacenter's networking engineers ensured all components aligned with Nvidia's reference architecture for seamless compatibility. "We supplied the 18 Lenovo nodes with H200 GPUs (16 interconnected and two inference nodes), designed the networking architecture in collaboration with Computacenter, and facilitated colocation through Patmos," explained Andrea Holt, a spokesperson for Hydra Host. The cluster itself is quite powerful, even in terms of general-purpose computing. The servers feature 192 96-core processors (for a total of 3,456 cores) paired with 36TB of DDR5 memory and 270TB of NVMe solid-state storage. There are spare bays so that storage space can be expanded easily. The supercomputer uses a network custom-built by HydraHost. The company also brought in Patmos to handle colocation, providing enough power (around 100kW) and cooling for the power-hungry and hot machines. Best Performance at Best Price The Exacluster costs $5 million, averaging $277,777 per machine, comparable to a single 8-way H200 baseboard rather than a full server. Here is where it gets interesting. Who facilitated that price? Stay On the Cutting Edge: Get the Tom's Hardware Newsletter Get Tom's Hardware's best news and in-depth reviews, straight to your inbox. [ ][ ]Contact me with news and offers from other Future brands[ ]Receive email from us on behalf of our trusted partners or sponsors[Sign me up] By submitting your information you agree to the Terms & Conditions and Privacy Policy and are aged 16 or over. On the one hand, Hydra Host is a close Nvidia partner and only offers Nvidia GPUs as a service. In addition, its Brokkr software is optimized primarily for CUDA. On the other hand, ExaAI is a company backed by Nvidia, so it can potentially get preferential pricing. "We are best in market at getting our customers the right GPU for their needs and at the best price," said Ryan Horjus, Lead Sales Engineer at Hydra. "This cluster was supported by Nvidia from an architecture design and their Inception program. Hydra handled it for Exa, as we do for other companies." Hydra also specializes in building custom solutions for startups and even monetizes their machines when not in use. "Hydra has helped startups get into their own clusters for better pricing through bulk purchasing," Horjus added. "They can achieve ideal pricing through our network. They are also able to monetize the servers when not in use via the Brokkr management platform." Speaking of Brokkr, it is a GPU management and provisioning software and a monetization platform for GPUs. It provides datacenters and startups with a turnkey software solution for getting their hardware into customers' hands and getting them paid for, explained Ariel Deschapell, chief technology officer and co-founder of Hydra. "One of its key features is automated bare metal provisioning and lifecycle management," described Deschapell. "That means the platform does all the work of configuring and managing the base server OS and firmware, setting up drivers and other supporting software, and running tests on the GPUs and other components. That speeds up and standardizes the delivery process significantly, reducing idle time on servers and GPUs. It also makes it easy to resell unused servers later to other users on the Brokkr platform looking for bare metal GPUs, if capacity needs change." See all comments (2) Anton Shilov Anton Shilov Social Links Navigation Contributing Writer Anton Shilov is a contributing writer at Tom's Hardware. Over the past couple of decades, he has covered everything from CPUs and GPUs to supercomputers and from modern process technologies and latest fab tools to high-tech industry trends. More about artificial intelligence Nvidia CEO Jensen Huang holding Project Digits Nvidia CEO Jensen Huang to meet President Trump at White House today Nvidia Hopper HGX H200 U.S. investigates whether DeepSeek smuggled Nvidia AI GPUs via Singapore Latest RTX 5090 supply at Newegg Most RTX 50-series GPUs sold out in five minutes at Newegg -- entire inventory evaporated in just 20 minutes See more latest > [ ] 2 Comments Comment from the forums * bit_user The article said: The setup uses 3.2Tbps InfiniBand for east-west traffic and 400Gbps Ethernet for north-south communication In this context, what do "east-west" and "north-south" refer to? I'd guess east-west means communication among peer nodes and north-south is referring to communication with clients and storage. Reply * bit_user The article said: On the one hand, Hydra Host is a close Nvidia partner and only offers Nvidia GPUs as a service. In addition, its Brokkr software is optimized primarily for CUDA. On the other hand, ExaAI is a company backed by Nvidia, so it can potentially get preferential pricing. LOL, it's the same hand! The article said: Hydra also specializes in building custom solutions for startups and even monetizes their machines when not in use. This is a nice idea, in theory. However, the concern I'd have is that AI training tends to be so data-intensive that it would take a long time to upload all of your training data to their servers and then you get to actually use the cluster for how long??? Plus, once you get bumped, because they want to use it, what do you do? I guess you have to transfer your partially-trained model + training data to some other cluster and continue there? Sounds inefficient, to me. I guess if the cost savings are substantial vs. one of the big cloud operators, then it might be worth the downsides. Reply * View All 2 Comments Most Popular [missing-im] Asus publishes GeForce RTX 5090 prices: $3,099 for range-topping model [missing-im] Intel delays key Xeon data center processor amid massive losses -- Clearwater Forest pushed back to 1H 2026 [missing-im] Intel cancels Falcon Shores GPU for AI workloads -- Jaguar Shores to be successor [missing-im] Seagate responds to fraudulent hard drives scandal, says resellers should only buy from certified partners [missing-im] RTX 50-series paper launch saw RTX 5090, RTX 5080 fly off the shelves -- Micro Center, Best Buy, and Newegg are all out of stock [missing-im] Facebook admits that the Linux topic crackdown was 'in error' and has been fixed [missing-im] Reviewer reports RTX 5080 FE instability -- PCIe 5.0 signal integrity likely the culprit [missing-im] Microsoft Snapdragon X Copilot+ PCs get local DeepSeek-R1 support -- Intel, AMD in the works [missing-im] Chinese chipmaker ships record-breakers: YMTC quietly begins shipping 5th Gen 3D TLC NAND [missing-im] Asus comments on Q-Release Slim damaging GPUs, issues official removal guidelines [missing-im] Aurora supercomputer is now fully operational, available to researchers Tomshardware is part of Future US Inc, an international media group and leading digital publisher. Visit our corporate site. * Terms and conditions * Contact Future's experts * Privacy policy * Cookies policy * Accessibility Statement * Advertise with us * About us * Coupons * Careers (c) Future US, Inc. Full 7th Floor, 130 West 42nd Street, New York, NY 10036. []