[HN Gopher] OVH CEO Octave Klaba speaking about the incident [vi...
___________________________________________________________________
OVH CEO Octave Klaba speaking about the incident [video]
Author : mxschmitt
Score : 100 points
Date : 2021-03-11 19:28 UTC (3 hours ago)
(HTM) web link (www.ovh.com)
(TXT) w3m dump (www.ovh.com)
| jacquesm wrote:
| I love the lack of fluff, lack of political bs and the honesty on
| display here.
| tyingq wrote:
| Two interesting pictures from the earlier story...
|
| How close the 3 data centers are in SBG:
| https://cdn.baxtel.com/data-center/ovh-strasbourg-campus/pho...
|
| How hot that fire was. I'm pretty sure the orange spots are holes
| melted in the walls that are made from metal shipping containers:
| https://pbs.twimg.com/media/EwGqV17XMAMF_wa?format=jpg&name=...
| iptrans wrote:
| Can somebody add [video] to the title?
| 5h wrote:
| [video - poor audio] maybe
|
| edit: That comment was snide, my heart goes out to the OVH
| team, the message within the video was good, forthright &
| honest. I hope it will be well received by their customers -
| just a shame it's a bit difficult to listen to!
| FDSGSG wrote:
| Isn't that already implied by the title?
| dang wrote:
| Sure. Done.
| gautamcgoel wrote:
| What incident?
| aidos wrote:
| They had a fire.
| kevinmgranger wrote:
| https://news.ycombinator.com/item?id=26407323
| sergiotapia wrote:
| Is it true that OVH has literally all data in that single
| datacenter?
| FDSGSG wrote:
| No. OVH has 27(-1) datacenters.
| tyingq wrote:
| I'm guessing the other two SBG datacenters took some kind of
| hit.
|
| The buildings are very close together:
| https://cdn.baxtel.com/data-center/ovh-strasbourg-
| campus/pho...
|
| And the fire looked really hot...like melting steel
| containers hot: https://pbs.twimg.com/media/EwGqV17XMAMF_wa?f
| ormat=jpg&name=...
| Nomikos wrote:
| No, they have 26 left https://www.ovh.com/world/us/about-
| us/datacenters.xml
| sergiotapia wrote:
| What a stupid rumor thanks for clarifying
| batmansmk wrote:
| Nope. 27 in the world, 10 in France, 1 caught fire, 2 others
| are stopped for inspection.
|
| https://www.ovh.com/world/us/about-us/datacenters.xml
| anyfoo wrote:
| Genuine question: How did that question come to be?
|
| Knowing nothing about OVH, I just typed "ovh datacenters" into
| Google and the first hit was this:
| https://www.ovh.com/world/us/about-us/datacenters.xml with the
| first sentence being "27 data centers around the world,
| including 2 of the largest ones".
| numpad0 wrote:
| It is usually true that data literally on site is on site
| tormeh wrote:
| The raw video link is https://www.ovh.com/fr/images/sbg/Octave-
| Klaba-speaking-en-v...
| dang wrote:
| That seems like a more precise URL so I've changed to that from
| https://www.ovh.com/fr/images/sbg/index-en.html for now.
| Thanks!
| lovedswain wrote:
| tl;dr UPS maintenance was performed by a vendor the day before
| the fire. Fire department used a thermal camera to isolate source
| of fire, it seemed to originate with 2 UPSes, one of which was
| the recently maintained UPS
| mgbmtl wrote:
| Other random bits: SBG-2 was an older generation datacenter,
| had ventilation issues? They have 4 other datacenters who have
| a similar design. Others, including SBG-3, have newer designs.
|
| They're building 2500 servers per week.
|
| For the offline buildings that are not destroyed, they have to
| rebuild the electrical distribution and network. It was not
| clear if they are also moving servers physically.
| verytrivial wrote:
| I read his description and literal hand waving as saying that
| for efficiency reasons, in 2011 they were sort of build like
| how you'd build a camp fire: convective airflow drawn through
| and out the top.
| fanf2 wrote:
| Right, this is the OVH "tower" data center design. Here's a
| google maps link to the OVH Strasbourg site
| https://goo.gl/maps/rQjir8byKNwMDZrm9 with (from south to
| north) SBG1 (a collection of shipping containers), SBG2 (a
| figure 8 shape), and SBG3 (a box).
|
| The idea of the tower design is to take ambient air in
| through the outside walls, through I think just one row of
| servers, then into a central void where convection pulls
| the air out.
| jeffbee wrote:
| This is taking to extremes my maxim that UPSes cause more
| outages than they prevent.
| justicezyx wrote:
| This would be one of inherient difference between smaller vs.
| giga players in cloud hosting.
|
| AWS/Google/Azure, if this happens, there should only be limited
| outage to a small fraction of customers. As a matter of fact,
| Google had such an incident before, and literally no customers
| (internal and external) noticed.
| stefan_ wrote:
| This is a difference in what you are buying. When you are
| buying a dedicated server, there isn't exactly a good way to
| hide that the thing has just gone up in smokes.
|
| When you buy a storage API, sure, failure rates go up, latency
| increases 100x, but after a few hours its probably back to
| normal.
|
| Of course, with the increased abstraction, you get _more_
| problems. "Availability zones" are useless when most cloud
| outages are because of configuration or systemic issues that
| tend to bring the whole thing down, no matter which AZ you are.
| But apparently it's now considered "good enough" to just go "oh
| we are down because AWS is down".
| ev1 wrote:
| This is an apples to oranges comparison. OVH largely sells bare
| metal; their public cloud wasn't really impacted.
|
| If you are using AWS, Google, or Azure, ran a single (or
| multiple machines) inside a single AZ with no backups and opted
| out of snapshots, you would face the exact same situation.
|
| I can definitely say I see people complaining about how
| everything they have is down on AWS when us-east-1 goes down
| periodically, while large players that deploy sanely like
| Netflix fail over to another region seamlessly.
|
| This [only owning a single machine at all] is what most of
| their customers whinging the most were doing. People that have
| actual sane production workloads on AWS or GCP are not going to
| be running 100% of their workload on a single EC2 instance with
| no backups.
|
| People that are running on OVH are running often things like
| gameservers etc that monopolise 100% of a physical machine and
| don't support horizontal scaling. You quite literally cannot
| force a srcds/hlds server to "load balance" dynamically and
| fail over on heartbeat.
|
| Often they are kids or students too, and the $30/m for a
| machine with 32-64GB ram is all they can afford (though this
| doesn't absolve them of paying $1-2/m more for offsite backups
| elsewhere)
|
| You can provision more physical machines with the OVH API and
| have them be up in a different city in a minute or two. You get
| linespeed bandwidth between OVH DCs. It's up to you to use it.
| jonas21 wrote:
| On the other hand, just about every month, there's a story on
| HN saying why are you wasting your money on AWS when OVH is
| so much cheaper (for example
| https://news.ycombinator.com/item?id=24966028).
|
| And well, I guess this is one of the reasons.
| kazen44 wrote:
| how so? using OVH and their bare metal servers doesn't
| absolve you from doing your own due dilligence.
|
| As said earlier, their cloud service is unaffected.
| ev1 wrote:
| If you choose to run 100% of your workload on a single EC2
| VM in us-east-chaos-monkey and put nothing in S3, only
| local mounted block storage that also disappears when you
| reboot your on-demand EC2, that is on you.
| kuschku wrote:
| 2 OVH dedicated servers in different countries are still
| cheaper than one AWS instance.
|
| e.g. I've got servers at OVH SBG-2, Hetzner's
| Falckenstein, and Online.net's AMS datacenters -- the
| total of which is still almost a magnitude less than the
| same cost on AWS or GCP (granted, that's including
| traffic)
| ev1 wrote:
| Pretty much. Use 5% of the money you saved moving your
| workload from AWS to OVH to support failover, DR, and
| backups. You can probably buy like five or six machines
| of equivalent spec or more.
| dvfjsdhgfv wrote:
| No, it doesn't work like this. I have several bare-metal
| (with Heztner, I use OVH for DNS), it's been over 10 years
| already. I know that if I only rent one machine in one
| location, I'm asking for trouble. Based on my experience, I
| would say that every 2-6 years something dies in a server.
| A disk, a controller, a fan, you name it. It's rare to have
| servers running for longer than 7 years without any issues,
| and they're outdated by that time anyway so they need to be
| migrated to a new machine.
|
| So, as a bare minimum, you rent at least two different
| machines at two different locations for each project and
| make offsite backups. It's still way less expensive than
| AWS.
|
| If I don't need a powerful server and just need to spin
| some instances for testing or small projects, I use Hetzner
| Cloud, it's ridiculously cheap.
| justicezyx wrote:
| Oh good to know. I don't use OvH, and my limited
| understanding were from their products page which lists VM
| style offerings. I had assumed VMs were the major use cases
| on OvH.
| ev1 wrote:
| The pricing on their dedicated servers are cheaper than
| most 1-2GB VMs on cloud: https://www.ovhcloud.com/en/bare-
| metal/ - this is their flagship and most expensive brand,
| the cheaper ones are even less
|
| The funniest tweets demanding their data and saying they'll
| lose everything are the people running:
|
| https://twitter.com/Sensity_RP/status/1369496048998223873 -
| GTA5 multiplayer gameserver that begs for donations.
| Running on one of the cheaper sub-brands of OVH probably
| (soyoustart, for GAME ddos-protection). $30-40/machine.
|
| https://twitter.com/pdfshift/status/1369550522479480833 - I
| can't tell if this is a troll
|
| https://twitter.com/KatsanosAlex/status/1369501497348812801
| - I can't tell if this is a troll
| jeffbee wrote:
| I can't even find any press articles about the Google incident.
| Nomikos wrote:
| Are you googling? ;-)
| Saris wrote:
| It also depends if you're renting a dedicated server, vs
| cloud/VPS. AWS/Google/Azure deal with virtualized systems that
| can be moved around to another server easily.
|
| OVH has a lot of dedicated servers as well though, so if you're
| using one of those then it can't be moved very easily to avoid
| downtime.
| nickdothutton wrote:
| Lucky it wasn't a UPS explosion.
| https://h2tools.org/lessons/battery-room-explosion
| riffic wrote:
| There are a couple good bell system practice documents covering
| lead-acid battery hazards:
|
| * http://etler.com/docs/bsp-archive/157/157-601-701_I19.pdf
|
| * http://etler.com/docs/bsp-archive/157/157-601-101_I7.pdf
| deftnerd wrote:
| Are there industry options or methods of wiring to allow for a
| UPS room separate from the actual rooms the racks are stored in?
|
| It's almost tradition to have a rack with UPS's in the bottom and
| then the rest of the space filled with servers or drive arrays.
|
| We wouldn't ever think of putting a tiny backup generator in the
| bottom of every rack, so why do we put a battery storage system
| there? Also, with the advances in battery chemistry technology
| that improve reliability and density, it's only a matter of time
| until Lithium chemistry batteries are available and that also
| increases the risk of fire.
|
| Is there any reason not to move backup power to another room, or
| even to a separate structure like how they put backup generators
| on a pad outside of the building?
| kazen44 wrote:
| in most datacenters, the UPS's are stored in the same collum
| like structure as their generators. A Datacenter is divided
| into vertical columns, with eacht collum being powered by
| redundant power feeds, generator(s) and UPS's.
| hinkley wrote:
| The farther your UPS is away from the server the fewer causes
| of power loss you can prevent.
|
| I've seen people use UPSes to allow them to rearrange wiring.
| I've seen them _fail_ by relying on the UPSes as well, of
| course.
|
| If you wire the entire room with 2-3 separate electrical
| systems all powered off of separate remote UPSes, you can do
| whatever you want, but it's harder to change your mind or build
| out incrementally if you do.
| bayindirh wrote:
| Our current DC has generation and UPSes in different rooms, in
| isolated places from each other. Both are pretty far from the
| actual DC itself.
| throw0101a wrote:
| > _Are there industry options or methods of wiring to allow for
| a UPS room separate from the actual rooms the racks are stored
| in?_
|
| Yes: longer cables.
|
| See Figure 1 in the Schneider-APC white paper, where they have
| "Electrical Space", "Mechanical Space" (HVAC), and IT Space:
|
| * https://download.schneider-
| electric.com/files?p_File_Name=VA...
|
| Power is generated hundreds of kilometres from where it is
| used, so having your UPS room a few dozen metres from your
| actual DC room isn't a big deal. I-squared-R losses aren't
| going to be that huge.
|
| Europe uses 400Y/230 for nominal low-voltage distribution (see
| Table 1 in above), so stringing some 400V extra copper to the
| PDUs, which then have 230V at the plugs, isn't a big deal.
| walshemj wrote:
| Depends for a DC / Telco your normally feeding 48v DC from the
| UPS which is normally separate.
|
| The lack of fire suppression is also very worrying.
| tyingq wrote:
| There are plenty of datacenters with a separate battery room,
| sure.
| [deleted]
| mikepurvis wrote:
| For the extreme opposite of that, Google famously trolled
| everyone 10 years ago by announcing that every one of their
| servers had its own in-chassis 12V battery:
|
| https://www.cnet.com/news/google-uncloaks-once-secret-server...
| jeffbee wrote:
| Yes but Google also later moved to a 48VDC power architecture
| with the batteries at the bottom of the rack.
|
| See http://apec.dev.itswebs.com/Portals/0/APEC%202017%20Files
| /Pl... page 6.
| tyingq wrote:
| That presentation is pretty interesting, including Google
| inventing their own "Switched Tank" DC-DC converters
| because the existing ones weren't efficient or reliable
| enough.
| jeffbee wrote:
| There's a whole separate paper on that:
|
| https://storage.googleapis.com/pub-tools-public-
| publication-...
| LinuxBender wrote:
| And Facebook had small UPS/ATS units at the end of each row.
| Not sure if they still do that today it was like that when I
| walked through their datacenter. They did that for the
| purpose of power efficiency. They lost far less power by
| having many smaller units.
| tpmx wrote:
| I think the key issue here is that there wasn't a functioning
| fire suppresssion system in place.
|
| Second question: is such a system required for this kind of
| operation? Maybe?
| bayindirh wrote:
| If you're running an IT operation this big, you should have
| both fire detection units and oxygen replacing suppressants.
|
| That should be mandatory. Otherwise, it'd be very hard to
| contain. Especially when all your servers have RAID controllers
| with Li-Ion batteries or supercapacitors or other extremely
| trigger-happy components.
|
| Oh, and cooling systems. You're just kindling the fire with it
| at the beginning.
| makkesk8 wrote:
| If it turns out it's the upses that caught fire, one can't help
| but wonder if it would be a better idea to house the upses/backup
| power solutions in an adjacent smaller building outfitted with
| sprinklers perhaps?
| yholio wrote:
| The problem with water in the same place with high power
| equipment is that it instantly turns the room into a death trap
| for any personnel, now everything is potentially live.
|
| Also, in the first part of a lithium battery fire, dropping
| water on them is quite explosive. It will eventually quench the
| fire but on the short run it will make it worse, filling the
| room with explosive hydrogen and poisonous lithium hydroxide.
| So when your water sprinklers engage over your UPS, you better
| be sure there's nobody around:
| https://www.youtube.com/watch?v=cTJh_bzI0QQ
| jacquesm wrote:
| You would be using Halon injection in a normal DC, after
| sufficient warning to employees to gtfo if they hadn't done
| that already.
| fanf2 wrote:
| Not halon any more, but instead an ozone-friendly fire
| suppression gas such as argonite
| jacquesm wrote:
| Thank you! Totally missed that Halon was 'out'.
| throw0101a wrote:
| > _The problem with water in the same place with high power
| equipment_ [...]
|
| Most high-end UPSes have a relay where you can run an active-
| high or active-low emergency power off (EPO) signal. The EPO
| can either be a button that is pressed manually by the staff,
| automatically via fire suppression system, or both/either.
|
| Schneider-APC white paper (PDF)
|
| * https://download.schneider-
| electric.com/files?p_File_Name=AS...
|
| The EPO can also cut-off the HVAC so oxygen is no longer fed
| into the area, and smoke isn't (re-)circulated.
|
| In the US, this is probably covered in NFPA 75, "Standard for
| the Fire Protection of Information Technology Equipment":
|
| * https://www.nfpa.org/codes-and-standards/all-codes-and-
| stand...
| walshemj wrote:
| You would not be using water here
| notJim wrote:
| Not to be rude, but it's really wild to me that even after all
| this time during the pandemic, the CEO doesn't have a headset he
| can use so that the audio is intelligible. It's gotta be one of
| the highest ROI investments you can possibly make at this point.
| tpmx wrote:
| It's quite consistent with how the entire operation is run, at
| least when witnessed from the outside as a customer.
| bigyikes wrote:
| This also frustrates me so much with all the newly-live-
| streamed events this year. So many companies are spending so
| much money putting together virtual conferences, but can't be
| bothered to ship their speakers a decent mic or webcam. Heck,
| Apple's $20 headphones would make a huge difference. Instead we
| get audio that sounds like it was recorded in my shower.
| gruez wrote:
| The audio was mostly fine, it was just randomly crapping out
| for a few seconds at a time.
| sneak wrote:
| Headset with mic, a neutral backdrop, a key light, a good
| camera, and a skincare routine have all paid for themselves 10x
| since this thing started, in my case.
| abluecloud wrote:
| a functioning fire suppression system would have probably been
| better but second to that i guess a mic would be a good
| investment
| Thaxll wrote:
| So putting water on servers don't destroy them?
| [deleted]
| LinuxBender wrote:
| This is typically a FM200 system in the datacenter [1]
| which is a pressurized gas. There are also often water
| "dry-pipes" in place that can pressurize if the FM200
| fails. Fire suppression codes vary by country.
|
| [1] -
| https://en.wikipedia.org/wiki/Automatic_fire_suppression
| __turbobrew__ wrote:
| There are gas based fire suppression system such as FM-200
| and Novec 1230.
|
| If you aren't using gas based fire suppression systems it
| is basically an amateur operation. The small data centre at
| my work (10-ish racks) has a FM-200 fire suppression
| system, it really isn't that expensive to set up.
| subssn21 wrote:
| Most Fir suppression systems used in Electronics are Halon,
| CO2, or some other O2 replacing substance
| tpmx wrote:
| They had a DIY water sprinkler system in a wood-based
| structure:
|
| https://lafibre.info/ovh-datacenter/ovh-et-la-protection-
| inc... (posted in previous threads)
| dtx1 wrote:
| Good to get out a response out quick, but this is too quick, the
| audio is garbage.
___________________________________________________________________
(page generated 2021-03-11 23:02 UTC)