[HN Gopher] UK to invest PS900M in supercomputer in bid to build...
___________________________________________________________________
UK to invest PS900M in supercomputer in bid to build own 'BritGPT'
Author : whyte
Score : 9 points
Date : 2023-03-15 20:46 UTC (2 hours ago)
(HTM) web link (www.theguardian.com)
(TXT) w3m dump (www.theguardian.com)
| Havoc wrote:
| I struggle to see the strategic benefit? Are they going to do
| restrictive license by geography or something?
|
| Feels very "me too" without much of a plan frankly
| Bellend wrote:
| The country runs on old people voting to "level up" and "build
| back better". 2-3 word slogans so hearing "Quantum" and "AI"
| during the budget today is just bravo.
| fatjokes wrote:
| They must really hate that ChatGPT spells color without a u.
| klipt wrote:
| I assume ChatGPT can spell the UK way if carefully prompted.
| Its training corpus surely includes both spellings.
| kwere wrote:
| lets hope their lead engineers won't get the turing treatment,
| expecially when cadres will be dissatisfied about their ego
| petproject managed by their buddies.
|
| UK had its chance to be a tech powerhouse, and wasted it. Nobody
| ever cared about why.
| ProjectArcturis wrote:
| Is "supercomputer" still even a thing? It's not as if they're
| building something with a single, super-powerful core. They're
| combining lots and lots of CPUs and GPUs. So by that measure,
| shouldn't AWS and Azure be the world's biggest supercomputers by
| far?
| dragontamer wrote:
| A "supercomputer" will have far stronger communications between
| the CPUs / GPUs than the typical AWS network.
|
| For example, a typical DGX supercomputer system from NVidia is
| pushing 7.2TBps (https://www.nvidia.com/en-us/data-
| center/dgx-h100/) GPU-to-GPU communications.
|
| In contrast, a typical DDR4 RAM on your typical desktop is 0.05
| TBps, so yeah, 7.2TBps external bandwidth between GPUs is quite
| a lot.
|
| -----------
|
| For Frontier, the Slingshot NICs are 100GBps each. So each node
| can communicate with more bandwidth to each other than your
| typical Desktop computer has RAM-bandwidth.
|
| The diagrams imply that there's 4x Slingshot NICs per node on
| Frontier, suggesting 400GBps bandwidth to the interconnect.
| (https://www.olcf.ornl.gov/wp-
| content/uploads/2020/02/frontie...)
| quesomaster9000 wrote:
| I wouldn't even consider the house-priced DGX to be a
| supercomputer in its own right, you need multiple racks of
| them otherwise the I/O bandwidth would be under utilized.
|
| It's not just the intra-GPU bandwidth that's impressive, each
| H100 card is essentially paired with a 400gbit full duplex
| NIC, which when configured in a hyper-torus (or similar) with
| the appropriate cross-node scheduling software the
| infrastructure enables tensor processing on potentially huuge
| matrices (trillions of parameters).
|
| These are exceptional beasts, but as with most supercomputing
| it needs a country full of universities to find and engineer
| problems which can fully exploit it.
| dekhn wrote:
| I think there's a number of math errors in your comparisons.
| For example, you're comparing RAM bandwidth to network
| performance. And those network numbers are summed over a
| large number of NICs.
|
| You mixed up GBps (gigaBYTES per second) with Gbps (gigaBITs
| per second)- divide your numbers by ten to get a good idea of
| gigabytes for comparing to RAM.
|
| RAM is about 50GBytes/sec, fast individual nics are 400Gbps
| (or about 40GB/sec). Unless you have special caches or very
| large ram, you will run out of bits to send very quickly on a
| network like that.
|
| Typically supercomputers don't have RAM or busses that are
| more than 2X faster than conventional machines because it's
| not economic.
| ccday wrote:
| You can rent a supercomputer from AWS:
| https://aws.amazon.com/ec2/instance-types/p4/
| dekhn wrote:
| No, most of the cloud and internal infra at those orgs isn't
| supercomputers. Each of the major cloud vendors has
| supercomputer-like products. Supercomputers today are defined
| by having networks that allow the processing elements to reach
| a significant level of utilization even though the problem is
| partitioned over millions of processing elements.
|
| The cluster that MSFT build for OpenAI is definitely a
| supercomputer.
___________________________________________________________________
(page generated 2023-03-15 23:03 UTC)