[HN Gopher] Cloudflare was down
       ___________________________________________________________________
        
       Cloudflare was down
        
       Author : datadrivenangel
       Score  : 254 points
       Date   : 2025-06-12 18:24 UTC (4 hours ago)
        
 (HTM) web link (www.cloudflarestatus.com)
 (TXT) w3m dump (www.cloudflarestatus.com)
        
       | aranchelk wrote:
       | Seems to be affecting functionality of their "Verify you are
       | human" dialogs as well as Workers.
        
         | clairegraham wrote:
         | Yep, KV is broken too. Any worker that depends on KV is
         | throwing exceptions. I was able to get into the dash, but it's
         | very slow. Error rates started to go up significantly around
         | 18:00 UTC.
         | 
         | Edit: The CF status page has acknowledged it's a broad outage
         | across many services:
         | https://www.cloudflarestatus.com/incidents/25r9t0vz99rp
        
           | aranchelk wrote:
           | After many tries I also got into the dashboard, but it's not
           | that usable, constant error pop-ups.
        
         | bgwalter wrote:
         | It does. Another question is why do we get these dialogues
         | always from Cloudflare and never from Akamai in the first
         | place?
        
           | bgwalter wrote:
           | Downvoting this comment and flagging the submission does not
           | address the serious issue. These verification dialogues make
           | the Internet unusable.
        
             | perching_aix wrote:
             | Nor does venting about it in unrelated threads, or
             | asserting your opinion as fact.
        
               | scubbo wrote:
               | It's not much of a reach to go from "discussion about
               | impact on human-verification dialogs" to 'discussion
               | about human-verification dialog policy". This isn't an
               | incident-management channel, it's a discussion forum -
               | tangents are fine!
        
               | bgwalter wrote:
               | I complained in the apnews.com thread, because the
               | apnews.com verification, which is annoying by itself, did
               | not work at all this time. That is hardly unrelated.
        
       | iimblack wrote:
       | They updated the incident noting that it's not just
       | authentication affected.
        
       | pier25 wrote:
       | Workers KV has been down for like +30mins. This is impacting us
       | seriously.
       | 
       | Their API is down too.
       | 
       | Amazing that something can impact their whole infrastructure like
       | this given how much redundance they have.
        
         | stri8ted wrote:
         | The same is true for Google.
        
         | nijave wrote:
         | >can impact their whole infrastructure
         | 
         | CDN and WAF seem to be working fine. I think CF rushed a lot of
         | newer services out without the reliability some of their
         | older/core services enjoy
        
         | kenhwang wrote:
         | From their incident page
         | (https://www.cloudflarestatus.com/incidents/25r9t0vz99rp):
         | 
         | > Cloudflare's critical Workers KV service went offline due to
         | an outage of a 3rd party service that is a key dependency.
         | 
         | I bet that 3rd party service is GCP.
         | 
         | I would be pretty pissed if I were a CF customer that used
         | Workers KV for redundancy because it was heavily marketed as
         | running on CF data centers.
        
       | jerrygoyal wrote:
       | GCP is also down https://news.ycombinator.com/item?id=44260810
        
         | ipsum2 wrote:
         | Odd coincidence. Wonder if Cloudflare uses GCP?
        
           | ikiris wrote:
           | It's likely their auth infra based on what the Google outage
           | is
        
             | devmor wrote:
             | What do you mean by this? The Google outage is a widespread
             | outage of most GCP services.
        
               | pageandrew wrote:
               | Google is claiming the root cause is with some of their
               | central IAM services, which would have a cascading effect
               | to the rest of their services.
        
               | devmor wrote:
               | Where did you see this information? Was it on a social
               | media channel? I do see the IAM services in the list of
               | affected services in the incident report.
        
               | tom1337 wrote:
               | check https://status.cloud.google.com/incidents/ow5i3PPK9
               | 6RduMcb1S...
               | 
               | > Multiple GCP products are experiencing impact due to
               | Identity and Access Management Service Issue
        
               | ikiris wrote:
               | Scroll up. Its literally in this HN comment section
               | highly upvoted.
        
               | ikiris wrote:
               | The comment was self explanatory, and no, it wasn't a
               | widespread GCP outage. Most everything was up except for
               | GCS and firebase, and later on identity stuff started
               | causing cascading issues but not when this was posted.
        
               | zerd wrote:
               | > it wasn't a widespread GCP outage.
               | 
               | If this wasn't widespread, what is?
               | 
               | Incident affecting API Gateway, Agent Assist, AlloyDB for
               | PostgreSQL, Apigee, Apigee Edge Private Cloud, Apigee
               | Edge Public Cloud, Apigee Hybrid, Cloud Data Fusion,
               | Cloud Firestore, Cloud Logging, Cloud Memorystore, Cloud
               | Monitoring, Cloud Run, Cloud Security Command Center,
               | Cloud Shell, Cloud Spanner, Cloud Workstations, Contact
               | Center AI Platform, Contact Center Insights, Data
               | Catalog, Database Migration Service, Dataform, Dataplex,
               | Dataproc Metastore, Datastream, Dialogflow CX, Dialogflow
               | ES, Google App Engine, Google BigQuery, Google Cloud
               | Bigtable, Google Cloud Composer, Google Cloud Console,
               | Google Cloud DNS, Google Cloud Dataflow, Google Cloud
               | Dataproc, Google Cloud Pub/Sub, Google Cloud SQL, Google
               | Cloud Storage, Google Compute Engine, Identity Platform,
               | Identity and Access Management, Looker Studio, Managed
               | Service for Apache Kafka, Memorystore for Memcached,
               | Memorystore for Redis, Memorystore for Redis Cluster,
               | Persistent Disk, Personalized Service Health, Pub/Sub
               | Lite, Speech-to-Text, Text-to-Speech, Vertex AI Search
        
             | artursapek wrote:
             | Their KV store was definitely down.
        
         | tete wrote:
         | When being down scales. :D
        
       | ineedaj0b wrote:
       | solar flare?
        
         | CoopaTroopa wrote:
         | No, Cloudflare.
        
       | pier25 wrote:
       | They've changed the title to "Broad Cloudflare service outages"
        
       | paxys wrote:
       | Let me guess, someone pushed out a bad BGP config?
        
         | CSMastermind wrote:
         | For an outage this large and widespread that would have to be
         | the main culprit.
        
       | neo_doom wrote:
       | Yeah this is going to be a problem. I haven't seen an issue this
       | widespread across so many services in a while.
        
         | tete wrote:
         | Seems to be semi regular now that everyone puts all their eggs
         | in only a few baskets.
        
       | joduplessis wrote:
       | Hopefully they also publish the prompt that did this.
        
         | tough wrote:
         | i was thinking about this too
        
         | daxfohl wrote:
         | They should make the AI lead the postmortem.
        
         | vsgherzi wrote:
         | They're just moving fast and breaking things 100x faster. Who
         | cares what code does just vibe it all away /s
        
       | b0a04gl wrote:
       | distributed systems break, that's the whole point what actually
       | matters is how fast they localize damage and how invisible that
       | feels to the end user if kv failing takes down auth, ui, and
       | workers, then failure isolation's missing recovery is fine, but
       | if your fix needs global coordination to unbreak local flows,
       | that's a design smell not saying perfect uptime, but the post-
       | outage ux should feel smoother, not shakier right now it feels
       | like the system survived but the interface didn't
        
       | vimwizard wrote:
       | proxy seems available in general, must just be local to workers
       | because only one of my sites going thru ZT tunnel with identity
       | access rules is affected
        
       | ourmandave wrote:
       | Is it coincidence that there's a Scheduled Maintenance in Tokyo
       | for 18:00 UTC in progress, and the problems started at 18:19 UTC?
        
         | perching_aix wrote:
         | Guess we'll find out from the postmortem. Always the silver
         | lining with these, get to learn from and enjoy a good writeup.
        
           | solarmist wrote:
           | Do these get posted publicly?
        
             | perching_aix wrote:
             | > Do these get posted publicly?
             | 
             | Yes.
        
         | bhaney wrote:
         | Probably
        
         | jonfw wrote:
         | There is always scheduled maintenance on that page, so that's
         | not much of a signal in my experience
        
         | alexcroox wrote:
         | Unrelated, they have a few services that rely on GCP which is
         | down. Still, I imagine the people working on the maintenance
         | for Tokyo turned white during that job worried it was caused by
         | them...
        
       | koliber wrote:
       | https://downdetector.com/ is showing outages at many major
       | companies including Google, CloudFlare, AWS and more.
       | 
       | Word on the street is that there are large BGP routing issues
       | behind all of this.
        
         | cogman10 wrote:
         | Would make sense. I think the last time I saw this sort of
         | thing it was BGP causing a bunch of traffic to route through
         | Iran or china IIRC.
        
           | koliber wrote:
           | I vaguely recall that incident. But it did not feel like it
           | affected this many services.
           | 
           | At the same time I have not noticed anything being down
           | firsthand. I am in Europe.
        
             | cogman10 wrote:
             | Here's the case [1]. Looks like they targeted a single /24
             | so that's likely why it wasn't a bigger issue.
             | 
             | [1] https://bishopfox.com/blog/bgp-hijacking-technical-
             | post-mort...
        
           | NooneAtAll3 wrote:
           | so this is related to Israel's escalation that everyone is
           | expecting?
        
             | CoopaTroopa wrote:
             | The Pentagon Pizza Report has been having a lot of activity
             | the past 24 hours. Maybe just a coincidence
        
           | nijave wrote:
           | There was also an older instance with China
           | https://www.cyberdefensemagazine.com/experts-detailed-how-
           | ch...
        
         | ramesh31 wrote:
         | Anthropic down/degraded as well. Time to go for a walk.
        
         | Animats wrote:
         | Internet Health Report is reporting "No data to show".
         | 
         | [1] https://www.ihr.live/
        
       | pier25 wrote:
       | "All locations except us-central1 have fully recovered. us-
       | central1 is mostly recovered. We do not have an ETA for full
       | recovery in us-central1."
        
       | tete wrote:
       | Big blog post about how they saved the internet upcoming. ;)
       | 
       | Currently down, but reference: https://blog.cloudflare.com/the-
       | ddos-that-almost-broke-the-i...
        
       | claudex wrote:
       | > Cloudflare's critical Workers KV service went offline due to an
       | outage of a 3rd party service that is a key dependency.
       | 
       | So they depend on GCP for (some of) their services
        
         | reimertz wrote:
         | wrote a similar comment - good to know for the future.
        
         | its-kostya wrote:
         | If that is true, and there is no other BGP shenanigans, then I
         | suspect this dependency will not be around for long
        
           | beastman82 wrote:
           | My WAG is it comprises 95% of the company infrastructure
        
           | pizzafeelsright wrote:
           | ceo just said not for long
        
         | asteroidburger wrote:
         | Sub-processor pages are an easy way to verify that sort of
         | thing.
         | 
         | https://www.cloudflare.com/gdpr/subprocessors/cloudflare-ser...
        
       | pier25 wrote:
       | Our Workers apps are up again
       | 
       | edit:
       | 
       | It works in the US but EU customers are still reporting our
       | services as down.
       | 
       | edit:
       | 
       | EU customers are reporting ok
        
       | poorman wrote:
       | Can't wait to read this post-mortem. Seems odd that a Google
       | Cloud outage would bring down Cloudflare services.
        
       ___________________________________________________________________
       (page generated 2025-06-12 23:01 UTC)