https://www.tomshardware.com/tech-industry/artificial-intelligence/sohu-ai-chip-claimed-to-run-models-20x-faster-and-cheaper-than-nvidia-h100-gpus Skip to main content (*) ( ) Open menu Close menu Tom's Hardware [ ] Search Search Tom's Hardware [ ] RSS US Edition flag of US flag of UK UK flag of US US flag of Australia Australia flag of Canada Canada * * Reviews * Best Picks * Raspberry Pi * CPUs * GPUs * News * Coupons * More + Newsletter + PC Components + SSDs + Motherboards + PC Building + Monitors + Laptops + Desktops + Cooling + Cases + RAM + Power Supplies + 3D Printers + Peripherals + Overclocking + About Us Forums Trending * What is an AI PC * Copilot+ PCs * Snapdragon X Elite * Lunar Lake * Ryzen 9000 * Blackwell When you purchase through links on our site, we may earn an affiliate commission. Here's how it works. 1. Tech Industry 2. Artificial Intelligence Sohu AI chip claimed to run models 20x faster and cheaper than Nvidia H100 GPUs News By Jowi Morales published 26 June 2024 Startup Etched has created this LLM-tuned transformer ASIC. * * * * * * * Comments (4) Sohu AI Chip (Image credit: Etched) Etched, a startup that builds transformer-focused chips, just announced Sohu, an application-specific integrated circuit (ASIC) that claims to beat Nvidia's H100 in terms of AI LLM inference. A single 8xSohu server is said to equal the performance of 160 H100 GPUs, meaning data processing centers can save both on initial and operational costs if the Sohu meets expectations. Nvidia H100 vs H200 vs Sohu (Image credit: Etched) According to the company, current AI accelerators, whether CPUs or GPUs, are designed to work with different AI architectures. These differing frameworks and designs mean hardware must be able to support various models, like convolution neural networks, long short-term memory networks, state space models, and so on. Because these models are tuned to different architectures, most current AI chips allocate a large portion of their computing power to programmability. Most large language models (LLMs) use matrix multiplication for the majority of their compute tasks and Etched estimated that Nvidia's H100 GPUs only use 3.3% percent of their transistors for this key task. This means that the remaining 96.7% silicon is used for other tasks, which are still essential for general-purpose AI chips. However, the transformer AI architecture has become very popular as of late. For example, ChatGPT, arguably the most popular LLM today, is based on a transformer model. In fact, it's in the name -- Chat generative pre-trained transformer (GPT). Other competing models like Sora, Gemini, Stable Diffusion, and DALL-E are all also based on transformer models. Image 1 of 2 Transformer AI models (Image credit: Etched) Transformer AI models (Image credit: Etched) Etched made a huge bet on transformers a couple of years ago when it started the Sohu project. This chip bakes in the transformer architecture into the hardware, thus allowing it to allocate more transistors to AI compute. We can liken this with processors and graphics cards -- let's say current AI chips are CPUs, which can do many different things, and then the transformer model is like the graphics demands of a game title. Sure, the CPU can still process these graphics demands, but it won't do it as fast or as efficiently as a GPU. A GPU that's specialized in processing visuals will make graphics rendering faster and more efficient, that's because its hardware is specifically designed for that. This is what Etched did with Sohu. Instead of making a chip that can accommodate every single AI architecture, it built one that only works with transformer models. When it started the project in 2022, ChatGPT didn't even exist. But then it exploded in popularity in 2023, and the company's gamble now looks like it is about to pay off -- big time. Nvidia is currently one of the most valuable companies in the world, posting record revenues ever since the demand for AI GPUs surged. It even shipped 3.76M data center GPUs in 2023, and this is trending to grow more this year. But Sohu's launch could threaten Nvidia's leadership in the AI space, especially if companies that exclusively use transformer models move to Sohu. After all, efficiency is the key to winning the AI race, and anyone who can run these models on the fastest, most affordable hardware will take the lead. Stay On the Cutting Edge: Get the Tom's Hardware Newsletter Get Tom's Hardware's best news and in-depth reviews, straight to your inbox. [ ][ ]Contact me with news and offers from other Future brands[ ]Receive email from us on behalf of our trusted partners or sponsors[Sign me up] By submitting your information you agree to the Terms & Conditions and Privacy Policy and are aged 16 or over. Ever since AI data centers started popping up left and right, many experts have raised their concerns over the power consumption crisis this power-hungry infrastructure will lead us to. Meta founder Mark Zuckerberg says electricity supply will constrain AI growth, and even the U.S. government has stepped in to discuss AI power demands. All the GPUs sold last year consume more power than 1.3 million homes, but if Etched's approach to AI computing with Sohu takes off, we can perhaps reduce AI power demands to more manageable levels, allowing the electricity grid to catch up as our computing needs grow more sustainably. Jowi Morales Social Links Navigation Freelance News Writer More about artificial intelligence LED lightbulbs, which usually consume about 10 Watts of power a piece. AI researchers run AI chatbots at a lightbulb-esque 13 watts with no performance loss -- stripping matrix multiplication from LLMs yields massive gains Nvidia Nvidia to sell its advanced AI processors to Middle East countries amid tough US export rules Latest KT Corporation logo on building South Korean telecom company attacks torrent users with malware -- over 600,000 customers report missing files, strange folders, and disabled PCs See more latest > See all comments (4) [ ] 4 Comments Comment from the forums * Diogene7 In theory, Intel seems to have shown that their Intel MESO (Magnetoelectric spin-orbit) logic, a post-CMOS technology related to spintronics, is much more amenable to both low-power logic and low-power neuromorphic computing. For example, emerging Non-Volatile-Memory (NVM) / Persistent Memory MRAM is already based on spintronics phenomena Therefore much more R&D ressources should be allocated to develop new manufactuting tools to improve and lower MRAM manufacturing cost, and then improve those tools to evolve to MESO manufacturing : this would be much, much more groundbreaking !!! https://www.techspot.com/news/ 77688-intel-envisions-meso-logic-devices-superseding-cmos-tech.html https://www.imec-int.com/en/press/ imecs-extremely-scaled-sot-mram-devices-show-record-low-switching-energy-and-virtually Reply * JTWrenn This was an interesting bit of tech news but....from what I can tell they have yet to actually build a working chip. They just have an idea and not even a fully fleshed out design which they will outsource. Seems like it could be amazing but can anyone find anything saying they have built and tested one? Right now this feels like vaporware with a huge upside if it works but s ton of hurdles to get to there. Am I wrong on that? Or does this thing only exist on paper right now? Reply * Diogene7 JTWrenn said: This was an interesting bit of tech news but....from what I can tell they have yet to actually build a working chip. They just have an idea and not even a fully fleshed out design which they will outsource. Seems like it could be amazing but can anyone find anything saying they have built and tested one? Right now this feels like vaporware with a huge upside if it works but s ton of hurdles to get to there. Am I wrong on that? Or does this thing only exist on paper right now? To my knowledge, the European semiconductor research center IMEC was able to build a prototype but not with all needed key metrics. So indeed there isn't yet any hard proof yet. That the reason why I would first allocate significant AI ressources to help find appropriate materials that would help build some prototypes. From there, more ressources will be needed to build manufacturing tools but some of them may be able to piggy back on MRAM manufacturing tools. That's the reason why it is also important to scale up MRAM manufacturing to accelerate the transition to spintronics technology. My personal belief is that we are currently at the start of the transition from silicon transistors to spintronics, and we are like in the 1950's / 1960's when solid state silicon transistors were emerging to replace vacuum tubes... Reply * CmdrShepard This is so stupid, the time of fixed function hardware was in the 90s. At that time the software technology developed much slower so it made sense to use fixed function hardware to accelerate stuff that wasn't about to change with the next big fad. Today we have transformers, but tomorrow? And what are datacenters going to do with all those useless ASICs then? More e-waste. At least with H100 you can program and run whatever you want. Reply * View All 4 Comments Most Popular [missing-im]Raspberry Pi 5 patch boosts performance up to 18% via NUMA emulation -- Geekbench tests reveal gains in both single and multi-threaded performance [missing-im]Olimex Neo6502pc delivers Apple II emulation with modern peripherals -- W65C02 and RP2040 co-processor in one $32 SBC [missing-im]Analysts say average laptop RAM quota will reach 11.8GB in 2024 -- up 12% year-on-year [missing-im]Sohu AI chip claimed to run models 20x faster and cheaper than Nvidia H100 GPUs [missing-im]AMD MI300X performance compared with Nvidia H100 -- low-level benchmarks testing cache, latency, inference, and more show strong results for a single GPU [missing-im]Taiwanese research firm says TSMC will maintain foundry lead until at least 2032 -- US expected to account for 28% of leading edge manufacturing [missing-im]Nvidia RTX 4070 Ti Super using AD102 GPU appears -- a fresh variant surfaces with a harvested RTX 4090 die [missing-im]Raspberry Pi Connect expanded to provide SSH access, support for older models [missing-im]AMD MI300X posts fastest ever Geekbench 6 OpenCL score -- 19% faster than RTX 4090, and only eight times as expensive [missing-im]Qualcomm's flagship Snapdragon X Elite in PassMark: Fails to Beat Apple's M3 [missing-im]AMD talks 1.2 million GPU AI supercomputer to compete with Nvidia -- 30X more GPUs than world's fastest supercomputer Tom's Hardware is part of Future US Inc, an international media group and leading digital publisher. Visit our corporate site. * Terms and conditions * Contact Future's experts * Privacy policy * Cookies policy * Accessibility Statement * Advertise with us * About us * Coupons * Careers (c) Future US, Inc. Full 7th Floor, 130 West 42nd Street, New York, NY 10036. []