[HN Gopher] Bluesky April 2026 Outage Post-Mortem
       ___________________________________________________________________
        
       Bluesky April 2026 Outage Post-Mortem
        
       Author : jcalabro
       Score  : 123 points
       Date   : 2026-04-10 15:51 UTC (7 hours ago)
        
 (HTM) web link (pckt.blog)
 (TXT) w3m dump (pckt.blog)
        
       | threecheese wrote:
       | > What I had missed is that we deployed a new internal service
       | last week that sent less than three GetPostRecord requests per
       | second, but it did sometimes send batches of 15-20 thousand URIs
       | at a time. Typically, we'd probably be doing between 1-50 post
       | lookups per request.
       | 
       | That'll do it.
        
         | bombcar wrote:
         | Zero, one, many, many thousands.
        
         | 98codes wrote:
         | Ahh, the three relevant numbers in development: 0, 1, and
         | infinity.
        
         | htx80nerd wrote:
         | less than ideal if I had to be frank.
        
         | jandrese wrote:
         | The incredible part about this is because their backend is all
         | TCP/IP they were literally exhausting the ports by leaving all
         | 65k of them in TIME_WAIT, and the workaround was to start
         | randomizing the localhost address to give them another trillion
         | ports or so.
        
           | Night_Thastus wrote:
           | I mean, it's one GetPostRecord, Michael. What could it cost?
           | 1 trillion ports?
        
           | kyledrake wrote:
           | This is a pretty interesting solution. I could see how this
           | could useful for certain kinds of problems (as part of a ddos
           | attack mitigation for example).
        
         | LoganDark wrote:
         | And then they fix the issue by using multiple localhost IPs
         | rather than, perhaps, _not sending 15-20 thousand URIs at a
         | time_
        
           | odo1242 wrote:
           | They mentioned it was a temporary fix that they removed after
           | finding and fixing the true root cause, though.
        
       | goekjclo wrote:
       | > The timing of these log spikes lined up with drops in user-
       | facing traffic, which makes sense. Our data plane heavily uses
       | memcached to keep load off our main Scylla database, and if we're
       | exhausting ports, that's a huge problem.
       | 
       | I expect this is common.
        
       | jmclnx wrote:
       | Lite Blue on a dark Blue background. That is a new one, I have
       | seen grey text on lite grey, but blue on blue ?
       | 
       | The article does work in lynx, at least I can read it.
        
       | rvz wrote:
       | Thank you for the post mortem on this outage.
        
       | jonstaab wrote:
       | nostr never goes down
        
         | pfraze wrote:
         | All support to other decentralizers but nothing never goes
         | down.
        
           | jonstaab wrote:
           | 1000x redundancy makes it vanishingly unlikely. Although I
           | know we're due for a pole shift so all bets are off I
           | suppose.
        
             | numpad0 wrote:
             | Wasn't aware there are ~2k relays now. Have inter-relay
             | sharing situation improved?
             | 
             | When I tried it long time ago, the idea was just a
             | transposed Mastodon model that the client would just multi-
             | post to dozen different servers(relays) automatically to be
             | hopeful that the post would be available in at least one
             | shared relays between the user and their followers. That
             | didn't seem to scale well.
        
               | jonstaab wrote:
               | Getting clients to do the right thing is like herding
               | cats, but there has been some progress. Early 2023 Mike
               | Dilger came up with the "gossip model" (renamed "outbox
               | model" for obvious reasons). Here's my write-up:
               | https://habla.news/hodlbod/8YjqXm4SKY-TauwjOfLXS
               | 
               | The basic idea is that for microblogging use cases users
               | advertise which relays their content is stored on, which
               | clients follow (this implies that there are less-
               | decentralized indexes that hold these pointers, but it
               | does help distribute content to aligned relays instead of
               | blast content everywhere).
               | 
               | Also, relays aside, one key difference vs ActivityPub is
               | that no third party owns your identity, which means you
               | can move from one relay to another freely, which is not
               | true on Mastodon.
        
           | nout wrote:
           | The comparison here is to something like TCP/IP. TCP/IP never
           | goes down. TCP/IP is a protocol, the servers may go down and
           | cause disruption, but the protocol doesn't really have the
           | ability to "go down". Nostr is also a protocol. The
           | communication on top of Nostr is pretty resilient compared to
           | other solutions though, so that's the main highlight here.
           | 
           | If tens of servers go down, then some people may start
           | noticing a bit of inconvenience. If hundreds of servers go
           | down, then some people may need to coordinate out of bound on
           | what relays to use, but it still generally speaking works ok.
        
         | jandrese wrote:
         | If nostr went down would people even notice?
        
           | jonstaab wrote:
           | probably not
        
           | nout wrote:
           | If any major nostr relay goes down, no one notices. That has
           | happened many times, the network is very resilient to that.
        
         | emidoots wrote:
         | There's stark contrast for an average human visiting the
         | landing page of bsky.app vs nostr.org
        
       | electrondood wrote:
       | Great write up... curious about the RCA. Thanks!
        
       | tapoxi wrote:
       | I don't really understand this architecture, but I thought
       | Bluesky was distributed like Mastodon? How can it have an outage?
        
         | Retr0id wrote:
         | Mastodon infra can have outages, too.
        
           | tapoxi wrote:
           | It's just confined to one instance if it goes down, not all
           | of Mastodon.
        
         | pfraze wrote:
         | This writeup is useful for backend engineers:
         | https://atproto.com/articles/atproto-for-distsys-engineers
         | 
         | The simple answer is that atproto works like the web & search
         | engines, where the apps aggregate from the distributed
         | accounts. So the proper analogy here would be like yahoo going
         | down in 1999.
        
           | isodev wrote:
           | Google and MSN Search were already available at this time.
           | Also websites used to publish webrings and there was IRC and
           | forums to ask people about things.
        
           | tapoxi wrote:
           | This is a fantastic write-up, thanks for sharing!
        
           | fiatjaf wrote:
           | Sorry, but this analogy is very misleading, no one browses
           | websites through Google's servers.
           | 
           | For example, right now in my URL bar I read
           | "news.ycombinator.com", not
           | "google.com/profile/news.ycombinator.com".
           | 
           | If Google goes down now I can keep browsing this website and
           | all the other websites I have in all my other tabs as if
           | nothing had happened.
        
             | danabramov wrote:
             | Does Google Reader help you make sense of it? It's more
             | like each app is like its own Google Reader. And indeed you
             | were able to access the same posts via other apps at that
             | time of outage.
        
             | evbogue wrote:
             | Do you have ideas about how Bluesky could decentralize?
        
               | wmf wrote:
               | Not the original poster but I do have some ideas.
               | Official Bluesky clients could randomly/round-robin
               | access 3-4 different appview servers run by different
               | organizations instead of one centralized server. Likewise
               | there could be 3-4 relays instead of one. Upgrades could
               | roll across the servers so they don't all get hit by bugs
               | immediately.
        
               | evbogue wrote:
               | If multiple personal data servers (pdses) share the same
               | set of posts how would we guarantee that they are tamper
               | resistant to third parties?
        
               | wmf wrote:
               | PDSes should be sharded not replicated. Your posts live
               | on your PDS which lives in one place (although it can
               | move).
        
               | evbogue wrote:
               | What's stopping us from doing both?
        
               | wmf wrote:
               | Cost and complexity tradeoffs. IMO the relay/appview is
               | the current bottleneck.
        
               | evbogue wrote:
               | This is why I'm hoping fiatjaf has a recommendation here.
               | I have a feeling he might have a proposal that solves
               | this. But doesn't solve all of it, just some of it.
        
             | pfraze wrote:
             | Technically you can still view the posts directly from the
             | PDS. It's just uninteresting compared to web pages
        
         | isodev wrote:
         | It's more of a concept of a plan for being distributed. I even
         | went through the trouble of hosting my own PDC and still, I was
         | unable to use the service during the outage
        
         | LoganDark wrote:
         | A web interface and home server can have an outage. Bluesky is
         | just a web interface and home server.
        
         | direwolf20 wrote:
         | It's not really distributed. It's a centralised service that
         | pulls some parts of 0.01% of user profiles from their own
         | servers.
        
       | gsibble wrote:
       | Did all 3 users notice?
        
         | ffsm8 wrote:
         | Naw, only one did. Turns out the other two were his socket
         | accounts he used to upvote and comment on his own content.
         | 
         | Okay, nuff trolling for today
        
       | mwkaufma wrote:
       | Tell us more about this buggy "new internal service" that's
       | scraping batch data :P
        
       | drewg123 wrote:
       | Golang's use of a potentially unbounded number of threads is just
       | insane. I used to be fairly bullish on golang, but this, combined
       | with the fact that its garbage collected, makes me feel its just
       | unsuitable for production use.
        
         | tombert wrote:
         | Why does garbage collection make it unsuitable for production
         | use? A _lot_ of production software is written in garbage
         | collected languages like Java. Pretty much the entire backend
         | for iTunes /Apple Music is written in Java, and it's not doing
         | any kind of fancy bump allocator tricks to avoid garbage. In my
         | mind, kind of hard to argue that Apple Music is not "production
         | use".
         | 
         | There are certainly plenty of projects where garbage collection
         | is too slow, but I don't know that they're the majority, and
         | more people would likely prefer memory safety by default.
        
           | slopinthebag wrote:
           | Everything is understood by comparison. Unsuitable for
           | production use, _compared to what_ is the more apt question.
        
           | madeofpalk wrote:
           | Based on my experience of Apple Music being pretty bad at
           | streaming music, i would say that it's not ready for
           | 'production use'.
        
         | floating-io wrote:
         | You can have this problem with any kind of thread -- including
         | OS threads -- if you do an unbounded spawn loop. Go is hardly
         | unique in this.
         | 
         | Goroutines are actually better AFAIK because they distribute
         | work on a thread pool that can be much smaller than the number
         | of active goroutines.
         | 
         | If my quick skim created a correct understanding, then the
         | problem here looks more like architecture. Put simply: does the
         | memcached client really require a new TCP connection for every
         | lookup? I would think you would pool those connections just
         | like you would a typical database and keep them around for
         | approximately forever. Then they wouldn't have spammed memcache
         | with so many connections in the first place...
         | 
         | (edit: ah, it looks like they do use a pool, but perhaps the
         | pool does not have a bounded upper size, which is its own kind
         | of fail.)
        
           | slopinthebag wrote:
           | Rust's async doesn't have this issue. Or at least, it's the
           | same issue as malloc in an unbounded loop, but that's a more
           | general issue not related to async or threading.
           | 
           | 15-20 thousand futures would be trivial. 15-20 thousand
           | goroutines, definitely not.
        
             | floating-io wrote:
             | I don't know enough about rust to confirm or deny that --
             | but unless rust somehow puts a limit on in-flight async
             | operations, I don't see how it would help.
             | 
             | The problem is not resource usage in go. The problem is
             | that they created umpteen thousand TCP connections, which
             | is going to kill things regardless of the language.
        
       | pembrook wrote:
       | Distributed social media goes down? hrmmm.
       | 
       | Email and the internet don't have "downtime." Certain key infra
       | providers do of course. ISPs can go down. DNS providers can go
       | down. But the internet and email itself can't go down absent a
       | global electricity outage.
       | 
       | You haven't built a decentralized network until you reach that
       | standard imo. Otherwise its just "distributed protocol" cosplay.
       | Nice costume. Kind of like how everybody has been amnesia'd into
       | thinking Obsidian is open source when it really isn't.
        
         | iAMkenough wrote:
         | Bluesky is a provider. Blacksky didn't go down.
        
           | pembrook wrote:
           | Is there anything running on Blacksky other than Bluesky with
           | more than say, 100 active users?
           | 
           | AOL never even got to that level of dominance in the internet
           | 1.0 era.
           | 
           | The point is it's not a distributed network if one node is
           | 99.9% of all traffic.
        
       | opem wrote:
       | At least they aren't hiding and transparent about it unlike the
       | big tech corps with so called SLAs
        
         | tmpz22 wrote:
         | There are no outages in Azure sing se.
        
       | streetfighter64 wrote:
       | > They represent real user-facing downtime
       | 
       | Off-topic, but "real" feels like the new "delve". Is there such a
       | thing as "fake" or "virtual" downtime, or why do people feel the
       | need to specify that all manner of things are "real" nowadays?
        
       | mwagstaff wrote:
       | With my SRE hat on, dare I ask... could/should this have been
       | picked up in testing?
       | 
       | And then normally there's a nice discussion about how production
       | is very different to the test environment.
        
       ___________________________________________________________________
       (page generated 2026-04-10 23:00 UTC)