[HN Gopher] So you want to rent an NVIDIA H100 cluster? 2024 Con...
       ___________________________________________________________________
        
       So you want to rent an NVIDIA H100 cluster? 2024 Consumer Guide
        
       Author : ea016
       Score  : 213 points
       Date   : 2024-07-09 11:02 UTC (3 days ago)
        
 (HTM) web link (www.photoroom.com)
 (TXT) w3m dump (www.photoroom.com)
        
       | 8organicbits wrote:
       | Pretty sparse on pricing data, I guess everyone asked them to
       | keep it private.
        
         | Tepix wrote:
         | Genesiscloud.com mentions ,,starting at @2.00/hr" for HGX100
         | H100.
        
           | spott wrote:
           | Only 200Gbps network per node.
           | 
           | The infiniband stuff is 400Gbps per GPU (3.2Tbps per node).
        
             | GC_tris wrote:
             | Each GPU is connected with 400Gbps. The rest ist just the
             | normal dataplane which is independent from the GPUs.
             | 
             | Source: Was personally involved in design of that
             | deployment.
        
               | spott wrote:
               | Good to know. I went to recheck, because I swore I didn't
               | see anything about that when I looked earlier, but they
               | now say 3.2Tbps infiniband... not sure if they changed it
               | or I was just blind.
               | 
               | Thanks!
        
         | latchkey wrote:
         | Compute pricing isn't really private. When you're talking about
         | high end compute, pricing is very much case by case. What is
         | the point of posting it if changes due to everyone having
         | different needs?
         | 
         | We have base pricing on our website, but I guarantee that if
         | someone comes to me asking for a year reservation, I'm not
         | going to give the quoted price. What I have there is just a
         | good starting point to get the discussion going.
         | 
         | I also had a great dialog with GC on LI over their version of
         | this posting, it seems they really value this customer and it
         | is long term relationship. My assumption is the actual pricing
         | reflects that.
         | 
         | One other thing on the special needs, GC mentioned they had 3
         | extra spare chassis in play as uptime was critical. That is not
         | an insignificant amount of investment to have just laying
         | around.
        
         | spott wrote:
         | https://gpulist.ai
         | 
         | No idea about how accurate that is, but if you want cluster
         | pricing...
        
           | latchkey wrote:
           | I got them to add the Verified badge, but the rest is pretty
           | much about as accurate as craigslist.
        
         | Der_Einzige wrote:
         | I hate how all high end markets wage wars on price discovery.
         | Most sell their products through middle-men instead of directly
         | for a reason.
         | 
         | High end furniture? Suddenly prices go away and you have to
         | "get quotes".
         | 
         | High end GPUs? Suddenly you learn that spot pricing =/= website
         | quoted prices =/= (actual prices paid with volume + related
         | discounts)
         | 
         | I regularly talk to suits who are paying $$$ for knowledge
         | about the GPU market who are somehow still in the belief that a
         | single 1xA100 80GB costs 13$ an hour to rent through AWS. When
         | I tried to correct them, they almost seemed not to believe me.
         | 
         | Things that take us tech bros minutes (checking the up-to-date
         | price data by going to the screen in your cloud console to spin
         | one up) or hours (emailing your cloud rep for pricing data with
         | discounts) take suits years to poorly approximate knowledge of.
         | 
         | If I go into a market, and price discovery isn't easy on a
         | product, I know I'm dancing with a good chance of being
         | scammed, and by the most greedy, comic-book evil kind of rich
         | people. Suits aren't immune to this, and I'm certainly not
         | either.
        
           | lotsofpulp wrote:
           | Because it is always in a sellers best interest to "price
           | discriminate". A seller benefits most by selling at the
           | highest price that each buyer is willing to buy at, and each
           | buyer has a different highest price they can or are willing
           | to pay, so price transparency works against being able to
           | price discriminate.
           | 
           | https://en.wikipedia.org/wiki/Price_discrimination
           | 
           | Buyers obviously benefit from price transparency. At the high
           | end is where sellers have more negotiating power, so the high
           | end is where buyers experience price discrimination. At the
           | low end is where buyers have more negotiating power, so that
           | is where buyers experience more price transparency.
        
             | Der_Einzige wrote:
             | If I put these sellers into a "fake" bidding war with a
             | "fake" invoice from someone purporting that they will sell
             | me something for cheaper than they would, the state may in-
             | fact call that fraud if anyone sued. Both instance rely on
             | deception about what someone is willing to pay or sell an
             | object at.
             | 
             | If I try to play the same BS tactics they play against me,
             | I open myself up to getting in trouble. Heads you win,
             | tails I lose.
             | 
             | Price discrimination as an idea should be rooted out. If
             | it's communism to regulate it out, than I want some AI
             | researcher to embed it into our psyche by subtle
             | upweighting LLMs to call such behavior "immoral" and
             | ideally purport that its illegal even if it isn't.
        
               | lotsofpulp wrote:
               | >If I try to play the same BS tactics they play against
               | me, I open myself up to getting in trouble. Heads you
               | win, tails I lose.
               | 
               | What? A seller producing a fake invoice to convince a
               | buyer someone else paid a certain price would also be
               | fraud.
               | 
               | You are free to use the exact same tactics as the seller.
               | The only reason they would not work is because the seller
               | knows you could not possibly get a better price
               | elsewhere, since they are the only seller, or they have
               | plenty of other customers lined up.
               | 
               | Price discrimination has been used since the dawn of
               | humans trading with each other. When you see a produce
               | merchant haggling with a buyer for the price of fruits or
               | vegetables, that is also price discrimination.
        
       | latchkey wrote:
       | Great post. The ethernet section is especially interesting to me.
       | 
       | I'm building a cluster of 16x Dell XE9680's (128 AMD MI300x GPUs)
       | [0], with 8x 2p200G broadcom cards (running at 400G), all
       | connected to a single Dell PowerSwitch Z9864F-ON, which should
       | prevent any slowness. It will be connected over rocev2 [1].
       | 
       | We're going with ethernet because we believe in open standards,
       | and few talk about the fact that the lead time on IB was last
       | quoted to me at 50+ weeks. As kind of mentioned in the article,
       | if you can't even deploy a cluster the speed of the network means
       | less and less.
       | 
       | I can't wait to do some benchmarking on the system to see if we
       | run into similar issues or not. Thankfully, we have a great Dell
       | partnership, with full support, so I believe that we are well
       | covered in terms of any potential issues.
       | 
       | Our datacenter is 100% green and low PUE and we are very proud of
       | that as well. Hope to announce which one soon.
       | [0] https://hotaisle.xyz/compute/            [1]
       | https://hotaisle.xyz/networking/
        
         | omneity wrote:
         | Did you procure the servers directly from Dell or through a
         | distribution partner?
        
           | latchkey wrote:
           | Kind of both. We started talking to Dell first and then they
           | introduced us to Advizex. We are now effectively partnered
           | with both companies, which is fantastic as the Advizex team
           | have ex-Dell people working directly with us. We are lucky to
           | have two whole teams of super talented people helping us out
           | on this journey.
        
         | logicchains wrote:
         | Meta had success building an ethernet cluster on Arista 7800
         | with Wedge400 and Minipack2 OCP rack switches.:
         | https://www.datacenterdynamics.com/en/news/meta-reveals-deta...
        
           | latchkey wrote:
           | I actually have a call with Arista next week to learn more
           | about their solutions. Especially those 7800's. They look
           | awesome.
           | 
           | One "problem" we have right now is that our cluster cannot
           | support more than 128 GPUs. If we wanted to scale with Dell,
           | we'd have to buy 6x more Z9864F to add one more cluster,
           | which is crazy expensive and complicated.
           | 
           | I want to see if Arista has something that can help us. That
           | said, I also have to find a customer that wants more than 128
           | MI300x and that hasn't happened... yet.
        
         | csmpltn wrote:
         | > "Our datacenter is 100% green"
         | 
         | Cool, where can I read more about this? How do you power your
         | DC?
        
           | walrus01 wrote:
           | Plenty of datacenters that are somewhere generally in the
           | pacific NW can claim to be "Green" because their power supply
           | is entirely hydroelectric.
           | 
           | https://www.nwd.usace.army.mil/CRWM/CR-Dams/
           | 
           | Many of those areas also happen to have the lowest $ per kWh
           | electricity in North American, the only lower rate is
           | available near a few hydroelectric dams in Quebec.
        
             | latchkey wrote:
             | Previously, I had multiple data centers in Quincy, WA.
             | Those were hydro green. It is an area that hosts a whole
             | multitude of big hyperscaler companies.
        
             | mulmen wrote:
             | As anyone who has driven through the Columbia River Basin
             | can tell you wind power is also abundant in Washington. The
             | grid is very clean here but it's certainly not purely
             | hydro.
        
             | dlkf wrote:
             | Why is green in scare-quotes?
        
               | uncertainrhymes wrote:
               | People make the argument that is a giant datacenter is
               | consuming 50% of some local hydro installation, everyone
               | else is town is buying something else that is less green.
               | 
               | It opens up questions about grids and market efficiency,
               | so your mileage may vary.
        
               | sangnoir wrote:
               | > People make the argument that is a giant datacenter is
               | consuming 50% of some local hydro installation, everyone
               | else is town is buying something else that is less green.
               | 
               | I don't think that's a cogent argument. It's akin to
               | saying a vegan commune in a small is buying is buying up
               | 50% of the vegan food, "forcing" others to buy meat-
               | products, and framing this to cast doubts on whether they
               | are _truly_ vegan. Consumers aren 't in a position to
               | solve supply problems.
        
               | netrus wrote:
               | Hydro power is a great thing, it was the first renewable
               | energy that was available in meaningful quantities.
               | However, great sites for hydro power are definitely
               | limited. We will not suddenly find a great spot for a new
               | huge dam. Imagine the only source of vegan B12 to be some
               | obscure plant that can only be grown on a tiny island. In
               | this scenario, the possible extend of vegan consumption
               | is fixed.
        
               | dlkf wrote:
               | In the regions where it works (PNW, Quebec, etc) we could
               | easily build more. The hurdles are regulatory. The
               | regulation isn't baseless - a dam will affect the local
               | ecosystem adversely. But that's a tradeoff we choose
               | rather than a fundamental limitation.
        
               | aziaziazi wrote:
               | The other trade off is correlated with the energy stored
               | : potential for catastrophic disaster in case of failure.
               | as a society, living below is not risk free in the long
               | run.
        
               | sangnoir wrote:
               | You've just painted a clearer picture than I did: the
               | crux is that it is a supply-side problem.
        
               | saagarjha wrote:
               | There's also discussion of environmental damage from
               | damming rivers
        
           | latchkey wrote:
           | We haven't announced the dc yet, but will soon. Very well
           | known. I'm actually pretty excited about it.
        
             | tamiral wrote:
             | waiting to see it posted on HN!
        
           | epistasis wrote:
           | I have no idea how they are powering it, but with the speed
           | with which solar and battery prices are falling, and the
           | slowness of getting a new big grid interconnection, I would
           | not be surprised to see new data centers that are primarily
           | powered by their own solar+batteries. Perhaps with a small,
           | fast and cheap grid connection for small bits of backup.
           | 
           | If not this year, definitely in the 2030s.
           | 
           | Edit: for a much smaller scale version of this, here's a
           | titanium plant doing this, instead of a data center. The nice
           | thing about renewables is that they easily scale; if you can
           | do it for 50MW you can do it for 500MW or 5GW with a linear
           | increase in the resources.
           | https://www.canarymedia.com/articles/clean-industry/in-a-
           | fir...
        
             | spywaregorilla wrote:
             | I don't think you could cover a data center with enough
             | solar panels to power it
        
               | FrojoS wrote:
               | They can be to the side of the data center, too.
        
             | rlupi wrote:
             | For hundreds of MW?
             | 
             | https://www.visualcapitalist.com/cp/top-data-center-
             | markets/
             | 
             | You can do that but they are the size of a mountain
             | (literally: https://www.swissinfo.ch/eng/sci-tech/inside-
             | switzerland-s-g... this is 900MW)
        
               | epistasis wrote:
               | Why not? A typical value for land is 10 acres/MW, so
               | 15-30 sq mi for 1-2 GW, which will average out to
               | hundreds of megawatts over the course of 24 hours, even
               | in shady days.
               | 
               | Including land costs, solar is the cheapest source of
               | energy, at <$1/W, which is a tiny fraction of the cost of
               | the rest of the data center equipment, and has a 30-50
               | year lifetime. For less than $5B you could have hundreds
               | of megawatts of continuous solar power backed by 24+
               | hours (5+GWh) of batteries. Hydro storage really can't
               | compete with batteries for this sort of application, at
               | least for new storage capacity. Existing hydro certainly
               | is great, just building new stuff is hard.
               | 
               | And GW-scale solar installations are fairly commonplace.
               | Far easier to procure than a matching number of H100s.
        
               | rlupi wrote:
               | Maybe it's possible, I haven't seen it done yet. I guess
               | there were better alternatives.
               | 
               | Well, I can't share numbers about datacenter MW sizes...
               | the fact that I misread some of those numbers as per
               | datacenter MW is telling :0
               | 
               | In any case, Meta (not my employer) has 24k GPU clusters.
               | In the most dense (and less power hungry) setup, Nvidia
               | superpods have 4x DGX per rack, 8 GPU per DGX
               | (hyperscaler use HGX to build their stuff, but it's the
               | same), and each rack uses ~40kW. That's 750 racks and
               | 30MW of just ML, you need to add some 10-20% for
               | supporting compute&storage, and other DC infrastructure
               | (cooling, etc.).
               | 
               | 24k GPU is likely one building, or even just one floor.
               | Meta will likely have multiple clusters like that in the
               | same POP.
               | 
               | That's in the ballpark of 100+MW per datacenter, as the
               | starting point.
        
               | epistasis wrote:
               | Oh I don't dispute your hundreds of MW at all. Other
               | freely available information definitely supports that for
               | the hypescalers. Browsing recent construction projects, I
               | see 250,000 sq ft projects from Apple, and 700,000+ from
               | the hyperscalers, and typically consumption is
               | 150-300W/sqft, with the hyper dense DGX systems at that
               | or above.
               | 
               | There are _lots_ of large scale renewable power projects
               | out there waiting to get onto the grid, stuck in long
               | interconnection queues, more than 1TW projects last I
               | heard. There are also lots of data centers wanting to get
               | big power connections, enough that utilities are able to
               | scare their regulatory bodies to make bad short-term
               | decisions to try to support the new load centers.
               | 
               | Connecting the builders of these large projects directly
               | to the new demand, and going outside the slow, corrupt,
               | and inept utilities would solve a lot of problems. And
               | you could still _eventually_ get that big interconnection
               | to the grid installed, and in the interim 3-5 years,
               | power the data center mostly off-grid. Because that
               | massive battery plus solar resource would eventually be a
               | massive grid asset that could benefit everyone too, if
               | the utilities weren 't so slow.
        
               | rapsey wrote:
               | I don't know what the power consumption of an individual
               | data center is, but your link talks about the consumption
               | of all datacenters in individual states/countries.
        
               | epistasis wrote:
               | A couple hundred megawatts is what I would expect from a
               | data center, and had planned out before commenting, so
               | that's on the mark. An H100 is very roughly a kW after
               | adding in all the rest of the overhead, so a 200MW data
               | center would have 200,000 H100s. That's massive cluster,
               | but not inconceivable.
        
         | dpflan wrote:
         | What are you using these for? Providing compute for customers
         | that want to train/infer? What is the level of interest and
         | level of success customers are seeing using these services Hot
         | Aisle offers?
        
           | latchkey wrote:
           | We are a bare metal compute offering. They can be used for
           | whatever people want to use them for (within legal limits, of
           | course). Interest is much higher now that we've started to
           | work with teams publishing benchmarks which show that H100's
           | have a nice competitor [0].
           | 
           | I'll admit, it is still early days. We just finished up
           | another free compute [1] two week stint with a benchmarking
           | team. One thing we discovered is that saving checkpoints is
           | slow AF. I'm guessing an issue with ROCm. Hopefully get that
           | resolved soon. Now we are in the process of onboarding the
           | next team.                 [0]
           | https://hotaisle.xyz/benchmarks-and-analysis/            [1]
           | https://hotaisle.xyz/free-compute-offer/
        
             | rvnx wrote:
             | > We would love to offer hourly on-demand rates for
             | individual GPUs, but we can't do so at this time due to a
             | limitation in the ROCm/AMD drivers. This limitation
             | prevents PCIe pass-through to a virtual machine, making
             | multi-tenancy impossible. AMD is aware of this issue and
             | has committed to resolving it.
             | 
             | One idea to help you: Are you sure you need a virtual
             | machine ? Couldn't you boot the machines under PXE to solve
             | the imaging problem ?
             | 
             | Essentially you have TFTP server that gives a Linux image
             | and boot on it directly
        
               | latchkey wrote:
               | 1 chassis, 8 gpus.
               | 
               | We want to be able to break that chassis up into
               | individual GPUs and allocate 1 GPU to 1 "machine". I
               | previously PXE booted 20,000 individual playstation 5
               | diskless blades and I'm not sure how PXE would solve
               | this.
               | 
               | The only alternative right now is to do what runpod (and
               | AMD's aac) are doing and do docker containers. But that
               | has the limitation of docker in docker, so people end up
               | having to repackage everything. You also can't easily run
               | different ROCm versions since that comes from the host,
               | and if you have 8 people on a single chassis... it
               | becomes a nightmare to manage it.
               | 
               | We're just patiently waiting for AMD to fix the problem.
        
               | rvnx wrote:
               | Got it it's clear, I thought you had 1 GPU in one chassis
               | in some cases.
        
         | rlupi wrote:
         | https://docs.nvidia.com/dgx-superpod/reference-architecture-...
         | 
         | NVIDIA large GPU supercomputers have separate compute-
         | networking (between GPUs) and storage-networking (storage to
         | GPUs, or storage to SSD, SSD to GPUs with CPU assistance). This
         | helps avoid networking issues, even more if not using
         | Infiniband.
         | 
         | From what I read here and on your website, you don't go that
         | route. I haven't found the equivalent system level reference
         | architecture for MI300x from AMD. I wonder if you have a link
         | to a public document where AMD provides guidance about this
         | choice?
        
           | derefr wrote:
           | > even more if not using Infiniband
           | 
           | It's interesting that the above "HPC reference architecture"
           | shows a GPU-to-GPU Infiniband fabric, despite Nvidia also
           | nominally pushing NVLink Switch (https://www.nvidia.com/en-
           | us/data-center/nvlink/) for the HPC use-case.
        
             | bee_rider wrote:
             | How does NVLink work? Because I already know MPI and I'm
             | not going to learn anything else, lol.
             | 
             | Edit: after googling it looks like OpenMPI has some NVLink
             | support, so maybe it is OK.
        
               | zxexz wrote:
               | I use OpenMPI with no issues over multiple H100 nodes and
               | A100 nodes, with multiple infiniband 200G and ethernet
               | 100G/200G networks, and RDMA (though using mellanox
               | instead of broadcom cards, but afaik broadcom supports
               | this just the same). Side note, make sure you compile
               | nvidia_peermem correctly if you want GDRMA to work :)
        
           | latchkey wrote:
           | We have a separate OOB/east-west network which is 100G and
           | would be used for external storage. We're spending an absurd
           | amount of money on just cables.
           | 
           | It is documented on the website [0], but I do see that I did
           | not document the actual cards for that, will add when I wake
           | up tomorrow. The card is:
           | 
           | Broadcom 57504 Quad Port 10/25GbE,SFP28, OCP NIC 3.0
           | 
           | As far as I know, AMD doesn't really have the docs, it is
           | Dell. Their team actively helped us design this whole
           | cluster.
           | 
           | We haven't decided on which type storage we want to get yet.
           | It'll really depend on customer demand and since we haven't
           | deployed quite yet, we are punting that can down the road a
           | bit. Our boxes do all have 122TB in them and we have some
           | additional servers not listed as well with 122TB... so for
           | now I think we can cobble something useful together.
           | 
           | [0] https://hotaisle.xyz/networking/
        
         | RobRivera wrote:
         | Nice! Whats benchmark standard these days? Still Superbench or
         | yall have something inhouse?
         | 
         | Re: lead time quote :O I guess I got spoiled working for one of
         | the major cloud vendors. The thought of poor b2b vendor support
         | never entered my risk matrix.
         | 
         | If you own your own cluster, the network bottleneck becomes
         | less a dollar cost I suppose, since you arent being charged a
         | premium to rent someone elses compute
        
         | pheatherlite wrote:
         | Being out of the loop for awhile. Has amd made anything similar
         | to cuda? Are cots frameworks such as pytorch and tensorflow on
         | par when running on amd hardware? What makes investing in amd
         | cluster/chips worthwhile?
        
         | JackYoustra wrote:
         | Why xeon instead of epyc?
        
         | startupsfail wrote:
         | It seems the reliability, speed and scalability drops with the
         | Ethernet are somewhat manageable.
         | 
         | To quote the article - From our tests, we found that Infiniband
         | was systematically outperforming Ethernet interconnects in
         | terms of speed. When using 16 nodes / 128 GPUs, the difference
         | varied from 3% to 10% in terms of distributed training
         | throughput[1]. The gap was widening as we were adding more
         | nodes: Infiniband was scaling almost linearly, while other
         | interconnects scaled less efficiently.
         | 
         | And then they do mention that the research team needs to debug
         | unexplained failures on Ethernet that they've not seen on
         | Infiniband. This actually can be the expensive part.
         | Particularly if the failures are silent and cause numerical
         | errors only.
        
         | teaearlgraycold wrote:
         | We just set up a small cluster of our own. We're not using
         | infiniband but it didn't seem like it would be a 50 week lead
         | time to setting it up. Where did you get that number?
        
       | huqedato wrote:
       | The only essential aspect this article doesn't answer: How much
       | does it cost? All the rest is metadata. I would have preferred a
       | clear table with vendors, prices and features. And less bla-bla.
        
         | ea016 wrote:
         | I couldn't share any pricing data since the discussions with
         | providers are private. Instead, I added a graph of prices from
         | gpulist.ai. For an Infiniband cluster, median is $2.3 per H100
         | hour, average is $2.47.
        
           | ilaksh wrote:
           | $2.47 * 256 * 24 * 30 = $455k ?
        
             | dijit wrote:
             | Based solely on my own calculations that I made for the
             | board of my company; this is within the parameters I would
             | expect, yeah.
        
       | silverlake wrote:
       | Good info! I use an HPC with SLURM. 40k GPUs shared by hundreds
       | of users. It works well enough. I don't know how the market for
       | cloud-based clusters works. Why didn't OP use AWS or Google for
       | on-demand training? Is it just down to cost?
        
         | macksd wrote:
         | If you do, in fact, need H100s, they can be very hard to get.
         | Even the smaller flavors of A100 you sometimes request, wait
         | days for, and then 1 node might show up during a weekend. And
         | for the reasons described in the article and the fact that
         | large training jobs can be network-limited, nicer networks can
         | be a big deal.
        
       | eigenvalue wrote:
       | Lots of good and detailed information here, thanks. I'm curious
       | why Ethernet interconnect is so unreliable in practice compared
       | to the Infiniband. I would think that at this point, after a
       | decade or more of current Ethernet standards, all the kinks would
       | be worked out and the worst that would happen would be occasional
       | latency spikes and a few lost packets that could be retransmitted
       | quickly. Shouldn't the training frameworks be more robust to that
       | sort of thing?
        
         | rlupi wrote:
         | Infiniband and ethernet are very different at the lowest
         | levels. Ethernet interconnects use RoCE (RDMA over converged
         | ethernet), which actually encapsulated infiniband transport in
         | ethernet, but you still pay for higher routing latency, and you
         | need separate compute-network and storage-network to avoid
         | queueing (lossless ethernet).
         | 
         | https://community.fs.com/article/infiniband-vs-ethernet-whic...
         | 
         | https://en.wikipedia.org/wiki/RDMA_over_Converged_Ethernet
         | 
         | Also... don't underestimate the PCI bus bottlenecks when you
         | put 8x 400GB networking + 8x GPUs. There are ways now to have a
         | tree of PCI switches and avoid overloading the main one, each
         | GPU gets its own networking card and PCI switch.
        
           | latchkey wrote:
           | This is a great comment.
           | 
           | Our cluster is 128 GPUs into a single Dell switch... should
           | help with the queuing. We also have a separate e-w 100G
           | network.
           | 
           | This is why we went with Dell XE9680 chassis... people forget
           | that PCI switches are quite important with this level of
           | compute. Dell has done a good job here.
        
           | eigenvalue wrote:
           | Interesting, thanks. From the wikipedia link, this seems like
           | the probable culprit for why things break:
           | 
           | "Although in general the delivery order of UDP packets is not
           | guaranteed, the RoCEv2 specification requires that packets
           | with the same UDP source port and the same destination
           | address must not be reordered."
        
             | wmf wrote:
             | In practice it's easy to design an Ethernet network that
             | doesn't reorder packets.
        
       | ec109685 wrote:
       | How do the large clouds compare from an availability and cost
       | perspective compared to finding a smaller provider and renting a
       | dedicated cluster?
        
         | choppaface wrote:
         | The large clouds will often be able to give the biggest players
         | a big discount. Perhaps not on raw GPU prices (maybe extended
         | trial), but if you already have e.g. 1PB in object storage then
         | they might give you 20-30% discount. Moreover if your contract
         | is jumbo, they'll give you not only a Slack channel but send
         | Forward Deployed Engineers to your office and/or help you build
         | part of the model training software.
         | 
         | But for a deployment the size of the OP (Photoroom) I doubt any
         | of the big clouds would offer a discount. Especially if they
         | were not already negotiating with multiple clouds.
         | 
         | Probably the best argument for going with a large cloud
         | provider on a smaller budget is that you already use some of
         | their other services significantly and your MLE-to-devops
         | headcount makes something like Photoroom's test infeasible.
        
       | barbazoo wrote:
       | > Electricity sources and CO2 emissions
       | 
       | I love that they included this in their consideration and pointed
       | out the impact running these GPUs has on the environment.
        
         | storyinmemo wrote:
         | Shameless employer promotion here while I work on H100 clusters
         | today: https://www.datacenterdynamics.com/en/news/crusoe-puts-
         | cpus-..., https://crusoe.ai/cloud/
         | 
         | Iceland seems to have an excess of energy to population and
         | it's very green.
        
           | barbazoo wrote:
           | I applaud the vision, I wish you folks hired remote software
           | devs
        
             | irq wrote:
             | Crusoe is unabashedly anti-remote work, which is curious
             | considering their company's environmental and energy
             | locality focus. I interviewed with them, they are very much
             | an old school "you must all work physically in San
             | Francisco" company. I work for one of their competitors
             | now, one that embraces remote work.
        
         | 123yawaworht456 wrote:
         | it has fuck all any impact if the electricity is sourced from
         | nuclear/hydro/solar/wind/geothermal.
        
           | barbazoo wrote:
           | Well, exactly, that's their point. If possible, choose a data
           | center location with access to renewable energy.
        
             | joe_the_user wrote:
             | The idea that consumer choice can change total carbon
             | output is so absurd that it's better to boycott anything
             | claiming some special source.
             | 
             | Yes, the present CO2 output is a planet wide-catastrophe.
             | No, your little "contribution" can't change that - if you
             | don't buy cheap coal energy, someone will, etc. Strong
             | state regulation forcing a decrease in total CO2 output
             | with no exceptions is the only force that could save us in
             | the near term (and I'm not holding my breath here but I had
             | to say that). All your "market" and "choice" solutions are
             | burning up like the California vegetation.
        
               | barbazoo wrote:
               | Wouldn't that be great, we wouldn't have to feel bad
               | about destroying the planet anymore :) I don't think it's
               | true though. Do you have any evidence that consumer
               | choice does not have any impact on carbon emissions? In
               | other areas it definitely does [1]
               | 
               | > if you don't buy cheap coal energy, someone will
               | 
               | No, not if everyone starts asking for clean energy
               | instead. There isn't an infinite number of people, at
               | some point demand for clean energy will make it
               | unattractive to offer dirty energy.
               | 
               | [1] https://news.climate.columbia.edu/2020/12/16/buying-
               | stuff-dr...
        
               | abdullahkhalids wrote:
               | Generally, globally sub-optimal many-agent game-theoretic
               | equilibria can't be escaped by individual agents shooting
               | themselves in the foot. There is a reason why it is an
               | equilibrium in the first place. If there a traffic jam
               | every day that adds 30 min to everyone's drive, you can't
               | solve it by asking some people to voluntarily take the
               | alternate path that adds 1 hour to their drive.
               | 
               | The correct way to move from the sub-optimal equilibrium
               | to a better one is via global action i.e. changing the
               | rules of the game. In the real world that means state
               | level actions as GP talks about.
        
       | Jun8 wrote:
       | Say you want to burn about $500 as a curiosity project for 8
       | nodes for a day. Any suggestions for what job to run?
        
         | teaearlgraycold wrote:
         | If you're just burning money you might as well mine crypto.
        
       ___________________________________________________________________
       (page generated 2024-07-12 23:00 UTC)