[HN Gopher] Google Cloud: Major packet loss between multiple reg...
       ___________________________________________________________________
        
       Google Cloud: Major packet loss between multiple regions - Europe
       and Asia
        
       Author : yla92
       Score  : 176 points
       Date   : 2022-06-07 14:28 UTC (8 hours ago)
        
 (HTM) web link (status.cloud.google.com)
 (TXT) w3m dump (status.cloud.google.com)
        
       | nanankcornering wrote:
       | affecting other T1 providers too, NTT, Arelion (ex-Telia) also
       | cloud providers OVH who have dark fibers between EU <-> Asia
        
         | nanankcornering wrote:
         | outage caused by AAE-1 submarine cable cut in Egypt.. well..
         | let's see how long it takes
         | 
         | https://twitter.com/CloudflareRadar/status/15342010304303636...
        
           | ArtWomb wrote:
           | actually logged in to make a joke about how the everyday vpc
           | hiccups we live with are becoming indistinguishable from
           | reported downage but of course hn has more current info than
           | the gcp team i was chatting with and if this trend of cutting
           | undersea cables continues i'm getting a shortwave ;)
        
             | RF_Savage wrote:
             | It's a fun hobby even without cut fibers. :)
        
               | seabea wrote:
        
           | Victerius wrote:
           | The Internet really is a series of tubes after all.
        
             | mgdlbp wrote:
             | But it's now a bit stuck.
        
               | randmeerkat wrote:
               | The tubes are full.
        
           | iso1631 wrote:
           | And SEA-ME-WE-5 too apparently.
           | 
           | I had a Wiocc outage on London-Kenya last night, didn't think
           | much about it, maybe they rerouted from break last night, and
           | that's why I had such big problems today
           | 
           | If they both went together than that feels like an anchor, if
           | one went last night and another today that's a nasty
           | coincidence. I'm not sure which is worse.
        
         | dontbenebby wrote:
         | >affecting other T1 providers too
         | 
         | Then maybe the headline should be more accurate rather than
         | give free marketing to what some might term a monopoly.
        
           | [deleted]
        
           | sophacles wrote:
           | This was reported before it was clear that cables were cut.
           | As with many things symptoms appeared before the cause was
           | understood and that knowledge filter through the NOCs of the
           | world, let alone out to users.
           | 
           | Also, how is the headline not accurate? Did Google Cloud not
           | in fact have major packet loss between multiple regions? How
           | does "this service is down" count as marketing, isn't it
           | giving people bad impressions?
        
             | dontbenebby wrote:
             | >This was reported before it was clear that cables were
             | cut.
             | 
             | That's rude. Who did that?
        
               | mynameisvlad wrote:
               | What is this snarky attitude? Be civil.
               | 
               | As the parent comment mentions, the service providers are
               | obviously going to see an issue the second it happens,
               | but it might take minutes/hours for information to
               | disseminate on the root cause.
               | 
               | Just like... Practically any incident. If you immediately
               | know the root cause, it shouldn't have been an incident
               | in the first place.
        
               | dontbenebby wrote:
               | >What is this snarky attitude? Be civil.
               | 
               | How am I beint uncivil? If I reply in short sentences I'm
               | uncivil, longer ones have my words twisted and picked
               | apart.
               | 
               | "That's rude. Who cut the cable?" is a legitimate
               | question and perfectly civil.
               | 
               | If you are perceiving something as "snarky", that is on
               | you. You cannot see into my mind.
               | 
               | (If I didn't care, I'd just smirk to myself and not even
               | ask who cut the cable.)
               | 
               | Now, do you think the above is uncivil? If so why?
               | 
               | And, as I asked previously, do you know who cut the
               | cable?
        
               | mynameisvlad wrote:
               | You actually asked "That's rude. Who did that?" after
               | commenting "Then maybe the headline should be more
               | accurate rather than give free marketing to what some
               | might term a monopoly."
               | 
               | The other interpretation of that statement, given your
               | antagonistic first comment, is asking "Who issued the
               | incident", as evidenced by the second half of my reply.
        
               | [deleted]
        
               | joatmon-snoo wrote:
               | I think there's a /s on the comment you're responding to
               | :)
        
               | wtch98 wrote:
               | With 2 cables at the same time it could well be a ship
               | dragging an anchor
        
         | Tinkeringz wrote:
         | Maybe a shared undersea cable or landing point issue since it's
         | multiple
        
         | RedShift1 wrote:
         | Well they are clearly not dark fibers as the cut is causing
         | issues?
        
           | jthrowsitaway wrote:
           | Indeed. Not sure why that term is misused so much. There's
           | nothing dark about fiber that's being used.
           | 
           | > A dark fibre or unlit fibre is an unused optical fibre,
           | available for use in fibre-optic communication
        
             | escapecharacter wrote:
             | oh, so dark matter is like unused matter. Got it.
        
             | zw123456 wrote:
             | My favorite smart ass response when someone asks me what
             | dark fiber is: Evil Fiber
             | 
             | But seriously, it is just a term of art, when one carrier
             | sub-leases fiber without any terminating equipment and they
             | provide their own equipment to light the fiber. It is
             | "dark" from the standpoint of the provider, they are not
             | "lighting" it. The customer is. Typically the fiber is
             | leased using something called an IRU indefeasible right to
             | use. Basically "condo-ised" fiber.
        
               | exikyut wrote:
               | Ooooooh. I thought "dark" was some weird domain-specific
               | nomenclature to describe "not available on the open
               | market", with the open market being like a transport
               | network of (lit up) lines, and dark = not lit up = not on
               | the open market.
               | 
               | Now I actually understand the definition I'd say "unlit
               | fiber" feels like a better-resolved way of putting it,
               | although I could see the uninitiated imagining that
               | describing a service that needed an additional
               | subscription on top or something, which isn't quite the
               | right nuance.
        
               | rstat1 wrote:
               | At least "dark fiber" isn't a complete lie, like some
               | other tech terms are (Looking at you "serverless")
        
               | Jenk wrote:
               | Serverless is from the notion that you are not concerned
               | with the server. I.e., that your product/ infrastructure
               | inventory is... Serverless.
        
               | Frost1x wrote:
               | Don't forget the "dark" web, wharever that means anymore.
               | These terms are great for marketing groups to paint an
               | image but often very quickly lose any technical
               | relationships they once had.
        
             | toast0 wrote:
             | More or less, someone with a cable leases you dark fiber,
             | and then you light it. What are you using? lit fiber
             | (obviously), what did you lease? dark fiber.
             | 
             | When it's severed, it's not exactly lit anymore either ;)
        
               | dpe82 wrote:
               | By that logic, all fiber is "dark fiber". :)
               | 
               | As I recall, "dark fiber" came into use after the dotcom
               | bust left _lots_ of overbuilt and unused network
               | infrastructure around that companies could buy up for
               | years after for pennies on the dollar. Buying  "dark
               | fiber" in that context had meaning - it meant you were
               | buying already built-out and unused fiber, compared to
               | running your own fiber lines at full cost as had
               | previously been more common.
        
               | toast0 wrote:
               | Well, I was joking a bit, but I would consider it dark
               | fiber when you lease the fiber and run the equipment on
               | both ends, and lit fiber when the owner of the fiber is
               | running the equipment on both ends and selling you IP
               | transit.
        
               | sophacles wrote:
               | > When it's severed, it's not exactly lit anymore either
               | ;)
               | 
               | It literally is lit though, the lasers at both ends are
               | still trying to send bits, or at least sending pulses to
               | do fault location.
        
         | njibhu wrote:
         | Source for OVH: https://network.status-
         | ovhcloud.com/incidents/pphdyqq9cgyl Couldn't find for Arelion
        
           | gundmc wrote:
           | This link has been updated to say the root cause is a fiber
           | cut.
        
       | nixcraft wrote:
       | I am pretty sure Google Ads, search core business, Gmail
       | (workspace), YouTube, and their main web properties are not
       | affected. Does that mean they use a different network/IDC for
       | Google properties? Maybe all of their web properties are designed
       | with multiple HA zones. I don't know.
        
         | dpe82 wrote:
         | I used to work at YouTube so can speak to that; I assume the
         | other high-value properties are similar.
         | 
         | YouTube runs many geographically distributed datacenters and
         | can each more-or-less independently serve user requests;
         | there's also sufficient redundancy that a full DC can go down
         | and the rest are able to carry the redirected load. When a link
         | goes down like today it's often not user visible as most
         | requests are already served at least from within the same
         | continent. If a whole data center goes down, users might
         | experience higher latency from longer round-trips and colder
         | caches because their traffic is being routed to a DC further
         | away that now has more diverse traffic, but they also won't
         | experience an outage.
         | 
         | This stuff was fun to work on. :)
        
           | fdgsdfogijq wrote:
           | Thanks for the comment! Always nice to hear snippets about
           | the resilience of these types of systems
        
         | dontbenebby wrote:
         | Remember when Google was a search engine that tried to find an
         | ethical way to do ads to support the search engine?
         | 
         | Those were fun times on the web, when the alternative was
         | curated lists on Yahoo or... whatever the hell Ask Jeeves was
         | aside from a collection of Easter Eggs directed at rude middle
         | schoolers.
        
         | javajosh wrote:
         | Kind of interesting to speculate what kind of traffic needs to
         | transit and what doesn't. It would be a major efficiency
         | failure if the services you mentioned needed to transit the
         | atlantic for any reason. "Maybe all of their web properties are
         | designed with multiple HA zones." I don't think it's physically
         | possible that Google could do what it does without a
         | comprehensive edge network.
        
           | the-rc wrote:
           | I did capacity planning for a large user facing service more
           | than a decade ago. Not only we had global N+2 redundancy (we
           | could still serve peak traffic if we lost the two largest
           | clusters worldwide), but I also made sure we had at least N+1
           | per "continent". That's how I managed to get resources for us
           | in the infamous, reserved and now long gone Budapest
           | cluster(s), whose only other tenant was Gmail.
        
           | Chabsff wrote:
           | They sure do have one:
           | https://peering.google.com/#/infrastructure
        
             | nixcraft wrote:
             | What fantastic info that answers my questions. Thanks for
             | sharing it.
        
           | jeffbee wrote:
           | I would expect mail delivery flows cross the entire network
           | on average because the mail arrives from places that are in
           | no way related to where the user account is homed.
        
         | tex0 wrote:
         | What's an HA zone?
        
           | alex_young wrote:
           | High Availability
        
           | aaaaaaaaata wrote:
           | What's an HAProxy? -_~
        
           | neoncontrails wrote:
           | High Availability. A service with (hopefully) a lot of 9s in
           | its percent uptime.
        
       | NeckBeardPrince wrote:
       | Another day, another huge GCP outage that impacts every region.
        
         | johndfsgdgdfg wrote:
         | Yes, exactly my thinking. Google invades our privacy and keeps
         | users hostages for more money. We need to boycott Google at any
         | cost.
        
         | davidkuennen wrote:
         | I don't remember having any outage so far on GCP and I'm
         | hosting there for years now. Even now no disruption for my
         | service (europe-west-1)
        
           | wussboy wrote:
           | I can't recall a GCP outage in my 2 years using it. I can
           | recall 3 separate AWS outages that affected one of our SaaS
           | platforms.
        
             | davidkuennen wrote:
             | You can say about Google what you want, but GCP has been an
             | awesome product so far and the status page game is top-
             | notch compared to other players in the field.
        
         | mhoad wrote:
         | Are you thinking of AWS maybe?
        
       | gfv wrote:
       | I think the issue started a bit earlier. If you'd tracerouted
       | from Asia to Europe at roughly 13UTC today, you could see that
       | multiple ISPs started carrying traffic the wrong way around the
       | world, eastward through North America, instead of the usual
       | direct undersea cable.
       | 
       | The AMS-IX[0] traffic graph shows a non-insignificant drop too.
       | 
       | [0]: https://www.ams-ix.net/ams
        
       | iso1631 wrote:
       | My links from London to Delhi (Tata) and Nairobi (WIOCC) had
       | issues starting 12:24:21 GMT, losing leased lines and internet
       | over the course of a couple of seconds then. Delhi internet
       | resumed quickly (less than 2 seconds) with an extra 9ms rtt.
       | Private wire resumed about 15 minutes later.
       | 
       | Nairobi however remained unreachable for hours -- outgoing
       | internet traffic to some places worked -- if I tracerouted from
       | the office it would die in LINX, but from home it would run from
       | London, New York, MSP, Seatle, Tokyo, Singapore, Djubouti and
       | down to Nairobi, mostly on Hurricane Electric
       | 
       | WIOCC tell me a fibre cut between Marseille and Djibouti. They've
       | got my circuits back up at least, although still routing on WIOCC
       | as far as London.
        
       | 908B64B197 wrote:
       | What are the odds this brings GCP closer to
       | https://killedbygoogle.com/ ?
       | 
       | Nobody pays retail for GCP anyways. Always wondered how much they
       | were losing on it.
        
         | toast0 wrote:
         | A packet loss incident due to cut fiber that lasted just over
         | three hours according to the status page is not really that big
         | of a deal. Fiber cuts happen, and are the only reliable method
         | to determine if your redundant fibers are in the same bundle.
         | As a non-customer, their status updates seem ok, one hour to
         | open the incident isn't great, and transparency is near zero,
         | but at least they provided regular updates once it was open.
        
         | asjkaehauisa wrote:
         | They won't. GCP is a investment. Building DC/backbone
         | infrastructure is a lot harder than x application. Still they
         | need computing power to own products. Maybe in the paper
         | they're losing money but in reality they're securing position
         | in the market
        
         | oneepic wrote:
         | IMO those odds are literally zero. Do you think AWS or Azure
         | would get canned because of a single outage? There are lots of
         | customers on GCP, and Google also appears to have a pretty big
         | buy-in already -- just look at all the regions they have
         | available in GCP.
        
         | nknealk wrote:
         | They actually break it out now in their financial reporting.
         | It's a loss of 931 million for the most recent quarter on
         | revenues of 5.8 billion. Likely most of the loss stems from R&D
         | to develop the product offering.
        
       | davidkuennen wrote:
       | Seems like it's been resolved.
        
       | m1keil wrote:
       | I had issues between Sydney and Eu earlier in AWS.
        
         | dontbenebby wrote:
        
       | olliej wrote:
       | Maybe should be retitled? It seems to be a problem with the AAE1
       | and SMW5 submarine cables[1] and effects multiple providers, not
       | just GCP
       | 
       | [1]
       | https://twitter.com/netblocks/status/1534211064673587204?s=2...
        
       | [deleted]
        
       | Linda703 wrote:
        
       ___________________________________________________________________
       (page generated 2022-06-07 23:02 UTC)