[HN Gopher] Google Cloud Networking reporting issues
___________________________________________________________________
Google Cloud Networking reporting issues
Author : xur17
Score : 87 points
Date : 2021-03-17 17:25 UTC (5 hours ago)
(HTM) web link (status.cloud.google.com)
(TXT) w3m dump (status.cloud.google.com)
| freedomben wrote:
| Does Cloudflare use Google Cloud? I know they have a ton of their
| own hardware, but maybe they outsource some things?
|
| I had my first ever Cloudflare outage earlier today. It's just a
| free site that has a single page rule that redirects to a
| different domain, but it started refusing connections for a
| little while today.
| pubg wrote:
| No.
| alexdumitru wrote:
| All Google services have been down for me for the past 30
| minutes, at least.
| toomuchtodo wrote:
| Miami: Google services (Mail, Drive, Calendar, Maps, DNS)
| available to me.
| tpmx wrote:
| Sweden: Google's public web-based services all seem to be
| available, GCP services are not. Except the console.
|
| Edit: things started working now.
| gcbirzan wrote:
| Try not using 8.8.8.8, it's down from some places.
| 1023bytes wrote:
| My home internet connection is down, maybe some larger scale
| backbone problem?
| NicoJuicy wrote:
| Dns on 8.8.8.8 perhaps?
| 1023bytes wrote:
| Oh that's it, thanks
| opsunit wrote:
| If you are unsure of the tenancy of any given site
| https://runson.cloud may be of some help.
| booleanbetrayal wrote:
| Seeing API calls to Sentry fail - https://status.sentry.io/
| [deleted]
| Thaxll wrote:
| That's why WoW is down probably.
| mywacaday wrote:
| Also LinkedIn android app
| packetslave wrote:
| WoW is self-hosted, or at least they were a few years ago.
| Racks and racks of HP blade servers.
| ihsw wrote:
| Blizzard leans heavily on GCP now.
|
| https://cloud.google.com/press-releases/2020/0124/abk-and-
| go...
| packetslave wrote:
| Nothing in that press release mentions anything about WoW
| vntok wrote:
| OVH last Wednesday, Oracle the very next day, Azure yesterday,
| GCP today... will a crime syndicate employing arsonists as well
| as hackers claim all these in a month as a demonstration of their
| capabilities?
| throwaway53453 wrote:
| These comments crop up regularly on outage posts. There's
| probably some psychological bias that explains it but I can't
| remember... I think the last time Google went down there were
| comments blaming it on Summer Interns :)
| dmlittle wrote:
| No, it was SolarWinds that blamed the intern for their
| security issues.
|
| In all seriousness... if a single employee can accidentally
| single-handedly screw things up you've got a process problem.
| There should be checks and balances to make sure that what
| you're trying to do make sense. These might be in the form of
| a code-review, automated validation, a staging environments,
| etc. Yes, these might be things that don't exist at a scrappy
| startup and that is fine as they'll get there eventually but
| for any large company not having good processes is reckless.
| throwaway53453 wrote:
| These issues happen despite all those things. You can bet
| that Google has that and more in place. Just like in
| aerospace engineering, things go wrong when dozens of
| failures happen across the critical path, not any
| individual piece.
| dmlittle wrote:
| Yes, and issues will always occur. What I was trying to
| point out that it's never an individual's fault alone.
| Sure someone might have tripped the final domino but they
| might not have had anything to do with everything else
| that happened. That's why blameless postmortems are a
| thing. Blameless (to me) doesn't mean that you don't name
| any person but rather that there's no blame assigned to
| anyone. I want to know what Person A did to cause the
| incident but there shouldn't be any retribution against
| Person A.
| rzzzt wrote:
| Frequency illusion aka. Baader-Meinhof phenomenon?
| https://en.wikipedia.org/wiki/Frequency_illusion
| dylan604 wrote:
| Anybody know the current where abouts of Mr. Robot?
| [deleted]
| dboreham wrote:
| s/reporting/is reporting/ ?
| CSDude wrote:
| Why the most GCP outages are global? Even though it says issue
| started in US/Pacific, I'm not sure if it affects other regions
| and the status page of them does not make it clear.
| notyourday wrote:
| Because Google engineers think that they are the smartest
| people in the world which means they don't actually engineer
| for minimization of a blast radius.
|
| You should just see how Google builds the prefix lists that it
| allows customers to advertise to Google via PNIs. You would
| think they are building it off some registry lists because why
| else would Google insists that the routes are registered? Oh
| no, that's for small people. At Google they just bring up a
| session to you, let you advertise and take whatever you
| advertised to them during the setup as the allow list because
| everything is done via custom automated software!
| merb wrote:
| they aren't, everything from germany works.
| advisedwang wrote:
| "US/Pacific" is detailing the timezone they are using for
| start/end times. It's not saying the issue is only in that
| location.
| smithcoin wrote:
| Brutal. We are seeing requests that originate from heroku time
| out because response times are well over 30 seconds.
| endlessvoid94 wrote:
| Same. It took over an hour for this to show up as an incident
| somewhere. I was going crazy.
| gundmc wrote:
| Seems to be resolved according to the same source.
| mavelikara wrote:
| Details here: https://status.cloud.google.com/incident/cloud-
| networking/21...
| [deleted]
| alexdumitru wrote:
| Everything Google-related is down for me, so I can't open that
| page.
| paxys wrote:
| Hosting your status page on the same infrastructure it is
| reporting on is the most idiotic thing a service can do.
| macintux wrote:
| It's tough, though. The PR embarrassment of hosting your
| Google Cloud status page on AWS, say, would be substantial.
| dilyevsky wrote:
| I'm sure google can afford some colo space (probably
| would be cheaper too)
| macintux wrote:
| Sure, but what's more likely to be reliable: GCP/AWS or a
| colo?
|
| It's also embarrassing (and can cause stress for your
| customers) when your status page is down.
| closeparen wrote:
| A separate failure domain / uncorrelated failure can be
| more important than the absolute rate of failure.
| dilyevsky wrote:
| Imo for simple deployment colo with a major provider
| (equinix, coresite, etc) and redundant transit beats any
| cloud on reliability hands down
| jeffbee wrote:
| How do you know that there isn't a disaster plan under
| which Google routes requests for status.cloud.google.com
| to some other, non-Google host?
| dilyevsky wrote:
| I honestly have no idea, hope they do. My point is aws is
| not the only alternative (and not even the best one) for
| setting this up.
| dylan604 wrote:
| Makes me think of a calculus word problem: As the things
| Google cannot afford approaches 0...
| [deleted]
| tpmx wrote:
| _Seems_ like the status page was on separate infra (I could
| access it while the GCP services were down), but Google
| Public DNS (8.8.8.8) was also down.
|
| Perhaps alexdumitru was using 8.8.8.8?
| pentium10 wrote:
| The issue started occurring intermittently at 08:26
| US/Pacific. The issue is impacting Google's Backbone network
| and may impact various services when accessing them from a
| different region or from the internet. Impacted services
| include Cloud Services (Workspace, Firebase, GCP) as well as
| other Google properties.
|
| Connectivity within a zone should not be impacted.
|
| Our engineering team has implemented a mitigation and is now
| monitoring the effectiveness of the change.
| naranha wrote:
| > Workaround: None at this time.
|
| I wish I could say that to my customers...
___________________________________________________________________
(page generated 2021-03-17 23:02 UTC)