[HN Gopher] Heroku was Down
       ___________________________________________________________________
        
       Heroku was Down
        
       Incident: https://status.heroku.com/incidents/2402  Update: our
       apps appear back up after 23 minutes total downtime. Others are
       reporting applications still down.  Update: it appears most or all
       services have been restored.
        
       Author : ericpauley
       Score  : 245 points
       Date   : 2022-02-24 16:33 UTC (6 hours ago)
        
       | leetrout wrote:
       | Heroku just sent an email to us confirming issues. Links to a
       | status page but that isn't loading.
       | 
       | https://status.heroku.com/incidents/2402
        
       | neilobremski wrote:
       | https://status.heroku.com/incidents/2402
       | 
       | (It took a few minutes to get this link to work for me)
        
       | jamil7 wrote:
       | And we're back? I think.
        
       | iamricks wrote:
       | So... what are the chances this is related to a cyber attack from
       | Russia? Also our apps are down
        
         | adamrezich wrote:
         | what would lead you to this conclusion?
        
           | 2OEH8eoCRo0 wrote:
           | https://www.cisa.gov/shields-up
        
             | adamrezich wrote:
             | what does this have to do with Heroku? why would it be
             | targeted? especially given that they're running on AWS?
        
               | 2OEH8eoCRo0 wrote:
               | > Every organization in the United States is at risk from
               | cyber threats
               | 
               | Heroku and AWS are organizations no?
               | 
               | > While there are not currently any specific credible
               | threats to the U.S. homeland, we are mindful of the
               | potential for the Russian government to consider
               | escalating its destabilizing actions in ways that may
               | impact others outside of Ukraine.
        
               | adamrezich wrote:
               | this morning at work we had a database connection issue
               | for a few hours. I work for a school district, which is
               | an organization. therefore, it was probably a Russian
               | cyberattack.
               | 
               | like what
        
               | 2OEH8eoCRo0 wrote:
               | > what would lead you to this conclusion?
               | 
               | I'm just posting what would lead somebody to that
               | conclusion. Nobody is definitively saying anything right
               | now.
        
               | adamrezich wrote:
               | yes, and I'm just pointing out how ridiculous of a
               | conclusion this is.
        
         | driverdan wrote:
         | Practically zero. Why would Russia care about Heroku?
        
           | Trasmatta wrote:
           | Heroku runs on AWS, so it would be an AWS attack of some
           | sort. But AWS in general seems to be up from what I can tell,
           | so I'm feeling like it's not a cyberattack, and just a Heroku
           | specific problem.
        
           | briandear wrote:
           | https://www.cnn.com/2022/02/24/tech/russia-ukraine-us-
           | sancti...
        
           | adamnemecek wrote:
           | Maybe not heroku per se but something hosted on heroku?
        
           | dc-programmer wrote:
           | Collateral damage (but probably DNS lol)
        
         | ezekg wrote:
         | I wasn't gonna say it...
        
           | 2OEH8eoCRo0 wrote:
        
         | brimble wrote:
         | Now is a good time for everyone to double-check their backups
         | and go over their disaster recovery plans. I know I am.
         | 
         | I mean, yesterday was a good day. But today's the best you've
         | got if you didn't do it then.
        
       | jrochkind1 wrote:
       | My heroku app was back up again as of about 100 seconds before
       | this comment timestamp.
       | 
       | My app was down about 35 minutes total, according to my
       | monitoring.
        
       | seannui wrote:
       | No update to their own status or their Twitter (@herokustatus)
       | after 10+ minutes. Pro.
        
       | canadianwriter wrote:
       | Gotta love asking in slack "is uh... literally everything down
       | right now?"
        
       | craigkerstiens wrote:
       | It appears to only be affecting apps that have multiple dynos.
       | 
       | While Heroku is busy fixing the issue if you scale your app to a
       | single dyno (assuming it can handle the traffic) it should
       | restore availability.
        
         | jrochkind1 wrote:
         | My single-dyno app was affected.
         | 
         | However, while I could not connect to my app, nobody I know
         | could connect to my app, and my monitoring service could not
         | connect to my app for a ping... my app logs showed that _some_
         | traffic continued to connect throughout the outage. So it was
         | not entirely universal. And was clearly a routing problem of
         | some kind.
        
         | bisRepetita wrote:
         | I had single dynos affected.
        
         | ericpauley wrote:
         | We also had single-dyno apps affected.
        
           | craigkerstiens wrote:
           | It may have been during an earlier part of the instance, but
           | we just scaled multiple affected apps to a single dyno and
           | all recovered. ps restart with multiple dynos and no effect.
        
           | jeremyjh wrote:
           | I did too. I'm curious if you also have pre-boot enabled? I
           | think that may set the routing configuration the same as if
           | multiple dynos are up.
        
         | reubano wrote:
         | My 6 apps are all single dyno. Not sure if I have pre-boot
         | enabled or not though....
        
       | ageitgey wrote:
       | Same here
        
       | ezekg wrote:
       | I've been with Heroku for 7 years and this is the first time in
       | recent memory that I've seen Heroku completely go down. Nothing
       | at all works.
        
         | digerata wrote:
         | Ahh... were you sleeping in 2021? They went down in September,
         | November, and December. Probably some other times earlier in
         | the year as well. I stopped keeping track at how bad their
         | service is because of how bad AWS's service is.
        
           | ezekg wrote:
           | I'm well aware of their other outages. But I don't remember
           | another time when they were _completely_ down like this, even
           | down to their status page.
        
             | digerata wrote:
             | Well, they aren't completely down now either. Heroku Shield
             | is up. Dashboard is up. EU is reported as up.
             | 
             | As an aside, I spoke with our AE in early January on where
             | they are going in dealing with AWS unreliability. One would
             | think they would have a good answer. They don't.
        
           | jrochkind1 wrote:
           | My app didn't go down during any of those times.
           | 
           | I recall at least one of them I could not deploy new versions
           | of my app though, which is pretty bad. I don't believe I had
           | any downtime at all due to heroku platform outages in 2021. i
           | believe you if you say you did though!
           | 
           | This particular outage definitely seems to be of a rare level
           | of severity.
        
             | heyoni wrote:
             | Their authentication system completely broke I think in
             | January...for hours.
        
               | jrochkind1 wrote:
               | Oh right, I remember that now.
               | 
               | Didn't bring my deployed app down though! No user-facing
               | outage for me.
               | 
               | This time my app was down for about 35 minutes.
        
               | heyoni wrote:
               | I think apps still worked but I'm not sure. I know I
               | couldn't create a privatelink to our AWS environment
               | because the CLI kept failing, so we were locked out of a
               | dev database. Not too bad.
        
         | driverdan wrote:
         | You haven't been paying attention. We're very heavy Heroku
         | users. Their dashboard/API has gone down a few times in the
         | past year. Overall they've been reasonably good.
         | 
         | It's much better than it was 5+ years ago. Back then they had
         | almost weekly downtime.
        
           | ezekg wrote:
           | I am paying attention. My business runs on Heroku. I never
           | said they don't go down... they go down quite a lot (much to
           | my dismay), but never _completely_ down like this one. Even
           | their status page is down, which is a first for me. Must be
           | bad.
        
           | bin_bash wrote:
           | they had their big layoff in 2020 or 2021 though so I
           | wouldn't expect it to improve
        
       | moeadham wrote:
       | US down for us, EU is up
        
       | jlangenauer wrote:
       | All our apps are down as well. And it seems (and I really have to
       | almost laugh here) that https://status.heroku.com no longer
       | loads.
        
         | dspillett wrote:
         | As always, a note to infrastructure providers: HOST YOUR LIVE
         | STATUS DETAILS ON OTHER INFRASTRUCTURE.
         | 
         | Of course, they won't. If they host is on someone else's then
         | that might look bad (tacitly saying that a competitor is
         | reliable and might be up when they are down) and if they hive
         | off an extra copy of some of their infrastructure there will
         | still be single points of failure either accidentally, by human
         | error (someone somehow messing up both segments at once), or by
         | design (possibly through management trying to save pennies when
         | they noticed this extra bit of infrastructure on the balance
         | sheet).
        
           | wuputah wrote:
           | I agree with your lead statement and argued as such, but was
           | overruled. Last I knew and understood, Heroku Status is
           | static pages pushed out to Fastly, with the internal admin
           | site (that does that work) running in a Heroku Private Space.
           | If you look at the DNS, it still appears to be served by
           | Fastly, and Heroku Private Spaces are generally pretty
           | isolated infra, so I would be curious what the failure mode
           | was here. But ultimately this is the fire you play with when
           | you self-host your status site...
        
           | bin_bash wrote:
           | I am pretty sure Heroku has a Linode box or something for
           | their status page. I may be wrong though. They may also have
           | moved it but that would seem like an incredibly dumb idea.
           | 
           | I worked there and I vaguely remember something like this but
           | it's been a long time.
           | 
           | If I had to guess this is probably a DNS issue (it's always
           | DNS).
        
       | amotinga wrote:
       | yes, site is down
        
       | reubano wrote:
       | Yep, all 6 of my Heroku apps are down.
        
         | reubano wrote:
         | Interestingly enough, my 2 Squarespace sites weren't affected.
         | Anyone know who they host with?
        
       | andygcook wrote:
       | Our app is currently down as well and we host on Heroku. We
       | received a down alert from our monitoring service at 11:32AM EST.
        
       | smoldesu wrote:
       | Rust's crate repository (crates.io) seems to be down too. I
       | wonder if they're connected...
        
       | tempnow987 wrote:
       | Status Page was running a bit slowly for me connecting to
       | developers.salesforce.com.
       | 
       | Link here to latest snapshot:
       | 
       | https://postimg.cc/McZHpCWg
       | 
       | No issues noted there.
        
       | jrochkind1 wrote:
       | At first the heroku status page was still all green for me...
       | although took 30 seconds to load. I guess when the status page
       | takes 30 seconds to load, that's an indicator!
       | 
       | I did figure, okay, a 30-second-to-load status page probably
       | means my app outage is a heroku platform problem.
       | 
       | (Also an indicator the status page is sharing too much platform
       | with the platform it's supposed to be reporting on? Also in this
       | case an indication that the platform problems are pretty deep?)
       | 
       | Interestingly, my app logs (via papertrail, which is still up)
       | show that _some_ traffic is getting through continually through
       | this current outage, although I can 't (and my monitoring app
       | can't either, which pinged me).
        
       | perryraskin wrote:
       | My Heroku app is completely down as well
        
       | [deleted]
        
       | reubano wrote:
       | AWS is down too, this is likely the cause since Heroku runs on
       | it. https://downdetector.com/status/aws-amazon-web-services/
        
         | jrochkind1 wrote:
         | AWS status page is all green.... Just kidding! I know the
         | information content of the AWS status page is literally zero,
         | it's always green!
        
           | somenewaccount1 wrote:
           | Nah, it's just on a 12 hour delay between updates.
        
         | nickphx wrote:
         | No, AWS is not down.
        
         | jermaustin1 wrote:
         | AWS status page is up, but going VERY slow, and says no issues.
        
           | reubano wrote:
           | I still don't get why they use images                   <td
           | class="bb pad04 top center" style="width: 32px">
           | <img src="/images/status0.gif">         </td>
           | 
           | Instead of
           | unicode:https://www.htmlsymbols.xyz/unicode/U%2b2705
        
             | cure wrote:
             | It's because they can host the images in S3 buckets, so
             | that their status page goes down when S3 is down.
             | 
             | /s
             | 
             | But, this really happened in 2017.
        
             | jermaustin1 wrote:
             | if the image tag is drawn dynamically then it is actually
             | probably less bandwidth than the unicode character since
             | the image can be cached at multiple locations including the
             | browser.
        
               | cure wrote:
               | .... no way.
               | 
               | Cache freshness checks involve a lot of headers, which
               | take up way more bandwidth than one unicode character.
        
               | jermaustin1 wrote:
               | if you care about cache freshness - you could set the
               | cache header to 10 years and it never be an issue again.
        
               | gjs278 wrote:
        
               | jrochkind1 wrote:
               | Less bandwidth than the 1-4 bytes of a codepoint in
               | UTF-8? How do you figure?
        
               | jermaustin1 wrote:
               | Because if it is cached, then it is 0 bytes transferred.
               | First request could be probably a few hundred bytes, but
               | never needed again. And once it is at a CDN, there is
               | never another request to the server.
        
           | XzAeRosho wrote:
           | As usual whenever there's an AWS outage
        
         | mattwad wrote:
         | AWS must be only partially down, all of it is running fine for
         | me in us-east-1. Elastic beanstalk ftw! :)
        
           | throwthere wrote:
           | If us-east-1 is up I have a heard time believing ANYTHING aws
           | is down
        
         | [deleted]
        
         | jmartens wrote:
         | What is the proof that AWS is down? Functional monitoring of
         | AWS by metrist.io (I'm a co-founder) shows no AWS problems.
         | Downdetector is not a reliable source.
        
           | reubano wrote:
           | What makes Downdetector unreliable? It's showing a huge spike
           | right now.
        
             | djbusby wrote:
             | Their (metrist) claim is that DD is human reported and
             | therefore unreliable.
             | 
             | Metrist monitors via bots
        
               | jmartens wrote:
               | It can be useful, but you have to take it with a grain of
               | salt. A perfect example is the recent Facebook (Meta)
               | outage. When that happened, Downdetector showed that ATT,
               | Verizon, and T-Mobile all had issues. They didn't, it was
               | just Facebook and users mentioned or otherwise claimed
               | that it was their mobile carrier.
        
             | zorpner wrote:
             | Downdetector relies on user reports, so e.g. if a user's
             | ISP is down and they can't get to Facebook, they might
             | report Facebook being down (or vice versa). DD spikes are
             | typically indicative of _something_, but it's not always
             | the actual down service.
        
               | reubano wrote:
               | Gotcha. Although for a spike this large (over 1000) for a
               | tech service (AWS vs Facebook) I'd give some credence to
               | it. It _could_ be that everyone who reported AWS as down
               | is running on Heroku. Definitely possible. For comparison
               | Azure [0] and Google Cloud [1] have spikes under 30.
               | 
               | [0] https://downdetector.com/status/windows-azure/ [1]
               | https://downdetector.com/status/google-cloud/
        
             | jaywalk wrote:
             | It's solely based off social media and user reports. It's
             | the "smoke" in the saying "where there's smoke, there's
             | fire" with the caveat that in some cases there's actually
             | no fire even if there's a decent amount of smoke.
        
           | djbusby wrote:
           | Neat! Wish you had a clickable demo rather than just
           | screenshots.
        
             | jmartens wrote:
             | Thanks for the feedback, we can do that soon.
        
           | nuggien wrote:
           | maybe your service isn't as reliable as you think it is ;)
        
       | mabsboza wrote:
       | reporting from Nicaragua, our apps are down!
        
       | rlopezc wrote:
       | Down for us too
        
       | yekurtal wrote:
       | Mine is still up.
        
       | mhale wrote:
       | Heroku apps are down, but API access is available.
       | 
       | Might be a good time to run: heroku pg:backups:download
        
       | jpmoyn wrote:
       | Down along with the status page right now
        
       | VWWHFSfQ wrote:
       | It is hard-down for me
        
       | [deleted]
        
       | emilsedgh wrote:
       | Yes it appears to be down for us as well.
        
       | mmarcant wrote:
       | 2/3 of our apps have been timing out of synthetic checks since
       | 8:30a PT
        
       | leros wrote:
       | I was able to get my apps up faster by restarting the dynos
        
       | pwned1 wrote:
       | Back up just now.
        
       | pedroborges wrote:
       | Our app is down too :(
        
       | henryaj wrote:
       | EU outage towards the end of last year was similarly bad, but
       | lasted much longer. Asked Heroku for at least a refund for the
       | dyno time when our apps were unavailable and was flatly told no.
        
       | reubano wrote:
       | Back up now!
        
       | ericpauley wrote:
       | Our app is currently 100% down at the router layer, and dashboard
       | requests are mostly failing.
        
         | VWWHFSfQ wrote:
         | Routing and logs are down for me as well
        
       | noodle wrote:
       | Things just came back up for me
        
       | nordec wrote:
       | Same here
        
       | dec0dedab0de wrote:
       | I've only ever used heroku for free tier level personal projects,
       | but as I understand it they use AWS to do the actual hosting. I
       | can understand an outage affecting their deployment process, but
       | what could cause running servers to go down?
       | 
       | As I typed that out I remembered that they handle DNS, load
       | balancing, and databases, so I guess any one of them.
        
         | alpb wrote:
         | All it takes is a load balancer misconfiguration on a core
         | service to take a large scale service down.
        
           | jrochkind1 wrote:
           | it indeed behaved like a routing issue of some kind (my app
           | was still UP, and was still logging, just no traffic could
           | get to it), and a heroku incident status line said "Engineers
           | are recovering affected routing components," so, yup.
        
       | escot wrote:
       | Mines back up, at least 20min down time
        
       | vendiddy wrote:
       | Down for us as well.
        
       | btown wrote:
       | Same here, hard down for us
        
       | bisRepetita wrote:
       | Some of my apps are back up. Others are not.
        
       | digerata wrote:
       | Ours are up. US East. Heroku Shield.
        
       | forgingahead wrote:
       | Our apps are all down as well, Heroku Status not loading either
       | (or just incredibly sluggish).
        
       | [deleted]
        
       | bovage wrote:
       | my apps are back up now.
        
       | kennysmoothx wrote:
       | Our Heroku applications are all down, it seems to be that it is a
       | Heroku Router/DNS issue.
       | 
       | We can access our logs and the application is running, just no
       | incoming requests.
        
       | jpmw wrote:
       | Let me bet: it's DNS.
       | 
       | It's always DNS.
        
         | ne0n wrote:
         | It was, but I don't know why. I'm curious to hear if Heroku
         | releases any information about how this happened. Heroku's DNS
         | was returning a single 100.64.x.x address which is in a
         | reserved range.
        
           | jpmw wrote:
           | They typically publish post-mortems, I think? Not 100% sure.
           | They definitely do in-house.
           | 
           | We'll have to wait and see I guess.
        
       | nickrubin wrote:
       | Yes, all apps are down
        
       | rileytg wrote:
       | status, dashboard and landing page are down for me. (socal.)
        
       | sharksauce wrote:
       | % heroku status
       | 
       | > Warning: Our terms of service have changed:
       | 
       | https://dashboard.heroku.com/terms-of-service Apps: Yellow Data:
       | No known issues at this time. Tools: No known issues at this
       | time.
       | 
       | === Availability of Common Runtime apps 2022-02-24T16:50:47.249Z
       | https://status.heroku.com/incidents/2402 investigating
       | 2022-02-24T16:50:47.249Z (1 minute ago) Engineers are looking
       | into reports of connectivity issues to Common Runtime apps in the
       | US and EU regions.
        
         | jrochkind1 wrote:
         | oh I didn't know that was a CLI command, nice!
        
       | Justsignedup wrote:
       | is it related to an AWS outage?
       | https://downdetector.com/status/aws-amazon-web-services/
        
         | steeef wrote:
         | "User reports indicate problems at Amazon Web Services" User
         | reported outage sounds like people confusing the Heroku outage
         | with AWS.
        
         | jmartens wrote:
         | Functional monitoring of AWS by metrist.io (I'm a co-founder)
         | shows no AWS problems. Downdetector is not a reliable source.
        
       | Gaussian wrote:
       | We're down
        
       | andycloke wrote:
       | All 6 of my apps & dashboard are down too.
        
       | pwned1 wrote:
       | It would be nice if they updated their status page. _sigh_
        
       | greset wrote:
       | Same here
        
       | ferajay wrote:
       | Same here
        
       | drstewart wrote:
       | Cypress is also down
        
       | teagee wrote:
       | HN is truly a market leader in status page technology
        
         | [deleted]
        
         | js4ever wrote:
         | HN was also down for a moment because too much traffic
        
           | Tainnor wrote:
           | I experienced that too
        
         | odiroot wrote:
         | Such a lightweight tool at that! Barely any JS loaded.
        
         | chasd00 wrote:
         | A couple of outages that have affected my client i learned
         | about first on HN. We were able to mobilize a team and get on
         | top of it faster than any monitoring team at the client ( a
         | state government ). I feel like HN should invoice us haha
        
           | castaway3000 wrote:
           | You might benefit from some better monitoring!
        
         | duxup wrote:
         | Except when it is a false alarm.
         | 
         | I've seen a few get to the front page.
         | 
         | But when we're right we're right.
        
           | dec0dedab0de wrote:
           | Even when it's a false alarm it's usually something else is
           | having a problem that is affecting many people and
           | manifesting itself as a particular service being down.
        
           | jcuenod wrote:
           | 10/9 times
        
             | for1nner wrote:
             | HN: All of our amps go to 11.
        
         | rileytg wrote:
         | *was
        
         | jonpon wrote:
         | I came to HN to make sure it was actually down.
        
       | typeofhuman wrote:
       | Yes all of my apps are down.
        
       | ab-dm wrote:
       | Well that was a fun one to wake up to. We were down for ~33
       | minutes overnight.
        
       | ericpauley wrote:
       | Anyone still seeing outages? Our services are all back up.
        
         | ezekg wrote:
         | All our applications are still down.
        
           | ezekg wrote:
           | We're back up. Total 47 mins of downtime.
        
         | cbonser wrote:
         | We are still down as well.
        
       ___________________________________________________________________
       (page generated 2022-02-24 23:02 UTC)