[HN Gopher] Heroku was Down
___________________________________________________________________
Heroku was Down
Incident: https://status.heroku.com/incidents/2402 Update: our
apps appear back up after 23 minutes total downtime. Others are
reporting applications still down. Update: it appears most or all
services have been restored.
Author : ericpauley
Score : 245 points
Date : 2022-02-24 16:33 UTC (6 hours ago)
| leetrout wrote:
| Heroku just sent an email to us confirming issues. Links to a
| status page but that isn't loading.
|
| https://status.heroku.com/incidents/2402
| neilobremski wrote:
| https://status.heroku.com/incidents/2402
|
| (It took a few minutes to get this link to work for me)
| jamil7 wrote:
| And we're back? I think.
| iamricks wrote:
| So... what are the chances this is related to a cyber attack from
| Russia? Also our apps are down
| adamrezich wrote:
| what would lead you to this conclusion?
| 2OEH8eoCRo0 wrote:
| https://www.cisa.gov/shields-up
| adamrezich wrote:
| what does this have to do with Heroku? why would it be
| targeted? especially given that they're running on AWS?
| 2OEH8eoCRo0 wrote:
| > Every organization in the United States is at risk from
| cyber threats
|
| Heroku and AWS are organizations no?
|
| > While there are not currently any specific credible
| threats to the U.S. homeland, we are mindful of the
| potential for the Russian government to consider
| escalating its destabilizing actions in ways that may
| impact others outside of Ukraine.
| adamrezich wrote:
| this morning at work we had a database connection issue
| for a few hours. I work for a school district, which is
| an organization. therefore, it was probably a Russian
| cyberattack.
|
| like what
| 2OEH8eoCRo0 wrote:
| > what would lead you to this conclusion?
|
| I'm just posting what would lead somebody to that
| conclusion. Nobody is definitively saying anything right
| now.
| adamrezich wrote:
| yes, and I'm just pointing out how ridiculous of a
| conclusion this is.
| driverdan wrote:
| Practically zero. Why would Russia care about Heroku?
| Trasmatta wrote:
| Heroku runs on AWS, so it would be an AWS attack of some
| sort. But AWS in general seems to be up from what I can tell,
| so I'm feeling like it's not a cyberattack, and just a Heroku
| specific problem.
| briandear wrote:
| https://www.cnn.com/2022/02/24/tech/russia-ukraine-us-
| sancti...
| adamnemecek wrote:
| Maybe not heroku per se but something hosted on heroku?
| dc-programmer wrote:
| Collateral damage (but probably DNS lol)
| ezekg wrote:
| I wasn't gonna say it...
| 2OEH8eoCRo0 wrote:
| brimble wrote:
| Now is a good time for everyone to double-check their backups
| and go over their disaster recovery plans. I know I am.
|
| I mean, yesterday was a good day. But today's the best you've
| got if you didn't do it then.
| jrochkind1 wrote:
| My heroku app was back up again as of about 100 seconds before
| this comment timestamp.
|
| My app was down about 35 minutes total, according to my
| monitoring.
| seannui wrote:
| No update to their own status or their Twitter (@herokustatus)
| after 10+ minutes. Pro.
| canadianwriter wrote:
| Gotta love asking in slack "is uh... literally everything down
| right now?"
| craigkerstiens wrote:
| It appears to only be affecting apps that have multiple dynos.
|
| While Heroku is busy fixing the issue if you scale your app to a
| single dyno (assuming it can handle the traffic) it should
| restore availability.
| jrochkind1 wrote:
| My single-dyno app was affected.
|
| However, while I could not connect to my app, nobody I know
| could connect to my app, and my monitoring service could not
| connect to my app for a ping... my app logs showed that _some_
| traffic continued to connect throughout the outage. So it was
| not entirely universal. And was clearly a routing problem of
| some kind.
| bisRepetita wrote:
| I had single dynos affected.
| ericpauley wrote:
| We also had single-dyno apps affected.
| craigkerstiens wrote:
| It may have been during an earlier part of the instance, but
| we just scaled multiple affected apps to a single dyno and
| all recovered. ps restart with multiple dynos and no effect.
| jeremyjh wrote:
| I did too. I'm curious if you also have pre-boot enabled? I
| think that may set the routing configuration the same as if
| multiple dynos are up.
| reubano wrote:
| My 6 apps are all single dyno. Not sure if I have pre-boot
| enabled or not though....
| ageitgey wrote:
| Same here
| ezekg wrote:
| I've been with Heroku for 7 years and this is the first time in
| recent memory that I've seen Heroku completely go down. Nothing
| at all works.
| digerata wrote:
| Ahh... were you sleeping in 2021? They went down in September,
| November, and December. Probably some other times earlier in
| the year as well. I stopped keeping track at how bad their
| service is because of how bad AWS's service is.
| ezekg wrote:
| I'm well aware of their other outages. But I don't remember
| another time when they were _completely_ down like this, even
| down to their status page.
| digerata wrote:
| Well, they aren't completely down now either. Heroku Shield
| is up. Dashboard is up. EU is reported as up.
|
| As an aside, I spoke with our AE in early January on where
| they are going in dealing with AWS unreliability. One would
| think they would have a good answer. They don't.
| jrochkind1 wrote:
| My app didn't go down during any of those times.
|
| I recall at least one of them I could not deploy new versions
| of my app though, which is pretty bad. I don't believe I had
| any downtime at all due to heroku platform outages in 2021. i
| believe you if you say you did though!
|
| This particular outage definitely seems to be of a rare level
| of severity.
| heyoni wrote:
| Their authentication system completely broke I think in
| January...for hours.
| jrochkind1 wrote:
| Oh right, I remember that now.
|
| Didn't bring my deployed app down though! No user-facing
| outage for me.
|
| This time my app was down for about 35 minutes.
| heyoni wrote:
| I think apps still worked but I'm not sure. I know I
| couldn't create a privatelink to our AWS environment
| because the CLI kept failing, so we were locked out of a
| dev database. Not too bad.
| driverdan wrote:
| You haven't been paying attention. We're very heavy Heroku
| users. Their dashboard/API has gone down a few times in the
| past year. Overall they've been reasonably good.
|
| It's much better than it was 5+ years ago. Back then they had
| almost weekly downtime.
| ezekg wrote:
| I am paying attention. My business runs on Heroku. I never
| said they don't go down... they go down quite a lot (much to
| my dismay), but never _completely_ down like this one. Even
| their status page is down, which is a first for me. Must be
| bad.
| bin_bash wrote:
| they had their big layoff in 2020 or 2021 though so I
| wouldn't expect it to improve
| moeadham wrote:
| US down for us, EU is up
| jlangenauer wrote:
| All our apps are down as well. And it seems (and I really have to
| almost laugh here) that https://status.heroku.com no longer
| loads.
| dspillett wrote:
| As always, a note to infrastructure providers: HOST YOUR LIVE
| STATUS DETAILS ON OTHER INFRASTRUCTURE.
|
| Of course, they won't. If they host is on someone else's then
| that might look bad (tacitly saying that a competitor is
| reliable and might be up when they are down) and if they hive
| off an extra copy of some of their infrastructure there will
| still be single points of failure either accidentally, by human
| error (someone somehow messing up both segments at once), or by
| design (possibly through management trying to save pennies when
| they noticed this extra bit of infrastructure on the balance
| sheet).
| wuputah wrote:
| I agree with your lead statement and argued as such, but was
| overruled. Last I knew and understood, Heroku Status is
| static pages pushed out to Fastly, with the internal admin
| site (that does that work) running in a Heroku Private Space.
| If you look at the DNS, it still appears to be served by
| Fastly, and Heroku Private Spaces are generally pretty
| isolated infra, so I would be curious what the failure mode
| was here. But ultimately this is the fire you play with when
| you self-host your status site...
| bin_bash wrote:
| I am pretty sure Heroku has a Linode box or something for
| their status page. I may be wrong though. They may also have
| moved it but that would seem like an incredibly dumb idea.
|
| I worked there and I vaguely remember something like this but
| it's been a long time.
|
| If I had to guess this is probably a DNS issue (it's always
| DNS).
| amotinga wrote:
| yes, site is down
| reubano wrote:
| Yep, all 6 of my Heroku apps are down.
| reubano wrote:
| Interestingly enough, my 2 Squarespace sites weren't affected.
| Anyone know who they host with?
| andygcook wrote:
| Our app is currently down as well and we host on Heroku. We
| received a down alert from our monitoring service at 11:32AM EST.
| smoldesu wrote:
| Rust's crate repository (crates.io) seems to be down too. I
| wonder if they're connected...
| tempnow987 wrote:
| Status Page was running a bit slowly for me connecting to
| developers.salesforce.com.
|
| Link here to latest snapshot:
|
| https://postimg.cc/McZHpCWg
|
| No issues noted there.
| jrochkind1 wrote:
| At first the heroku status page was still all green for me...
| although took 30 seconds to load. I guess when the status page
| takes 30 seconds to load, that's an indicator!
|
| I did figure, okay, a 30-second-to-load status page probably
| means my app outage is a heroku platform problem.
|
| (Also an indicator the status page is sharing too much platform
| with the platform it's supposed to be reporting on? Also in this
| case an indication that the platform problems are pretty deep?)
|
| Interestingly, my app logs (via papertrail, which is still up)
| show that _some_ traffic is getting through continually through
| this current outage, although I can 't (and my monitoring app
| can't either, which pinged me).
| perryraskin wrote:
| My Heroku app is completely down as well
| [deleted]
| reubano wrote:
| AWS is down too, this is likely the cause since Heroku runs on
| it. https://downdetector.com/status/aws-amazon-web-services/
| jrochkind1 wrote:
| AWS status page is all green.... Just kidding! I know the
| information content of the AWS status page is literally zero,
| it's always green!
| somenewaccount1 wrote:
| Nah, it's just on a 12 hour delay between updates.
| nickphx wrote:
| No, AWS is not down.
| jermaustin1 wrote:
| AWS status page is up, but going VERY slow, and says no issues.
| reubano wrote:
| I still don't get why they use images <td
| class="bb pad04 top center" style="width: 32px">
| <img src="/images/status0.gif"> </td>
|
| Instead of
| unicode:https://www.htmlsymbols.xyz/unicode/U%2b2705
| cure wrote:
| It's because they can host the images in S3 buckets, so
| that their status page goes down when S3 is down.
|
| /s
|
| But, this really happened in 2017.
| jermaustin1 wrote:
| if the image tag is drawn dynamically then it is actually
| probably less bandwidth than the unicode character since
| the image can be cached at multiple locations including the
| browser.
| cure wrote:
| .... no way.
|
| Cache freshness checks involve a lot of headers, which
| take up way more bandwidth than one unicode character.
| jermaustin1 wrote:
| if you care about cache freshness - you could set the
| cache header to 10 years and it never be an issue again.
| gjs278 wrote:
| jrochkind1 wrote:
| Less bandwidth than the 1-4 bytes of a codepoint in
| UTF-8? How do you figure?
| jermaustin1 wrote:
| Because if it is cached, then it is 0 bytes transferred.
| First request could be probably a few hundred bytes, but
| never needed again. And once it is at a CDN, there is
| never another request to the server.
| XzAeRosho wrote:
| As usual whenever there's an AWS outage
| mattwad wrote:
| AWS must be only partially down, all of it is running fine for
| me in us-east-1. Elastic beanstalk ftw! :)
| throwthere wrote:
| If us-east-1 is up I have a heard time believing ANYTHING aws
| is down
| [deleted]
| jmartens wrote:
| What is the proof that AWS is down? Functional monitoring of
| AWS by metrist.io (I'm a co-founder) shows no AWS problems.
| Downdetector is not a reliable source.
| reubano wrote:
| What makes Downdetector unreliable? It's showing a huge spike
| right now.
| djbusby wrote:
| Their (metrist) claim is that DD is human reported and
| therefore unreliable.
|
| Metrist monitors via bots
| jmartens wrote:
| It can be useful, but you have to take it with a grain of
| salt. A perfect example is the recent Facebook (Meta)
| outage. When that happened, Downdetector showed that ATT,
| Verizon, and T-Mobile all had issues. They didn't, it was
| just Facebook and users mentioned or otherwise claimed
| that it was their mobile carrier.
| zorpner wrote:
| Downdetector relies on user reports, so e.g. if a user's
| ISP is down and they can't get to Facebook, they might
| report Facebook being down (or vice versa). DD spikes are
| typically indicative of _something_, but it's not always
| the actual down service.
| reubano wrote:
| Gotcha. Although for a spike this large (over 1000) for a
| tech service (AWS vs Facebook) I'd give some credence to
| it. It _could_ be that everyone who reported AWS as down
| is running on Heroku. Definitely possible. For comparison
| Azure [0] and Google Cloud [1] have spikes under 30.
|
| [0] https://downdetector.com/status/windows-azure/ [1]
| https://downdetector.com/status/google-cloud/
| jaywalk wrote:
| It's solely based off social media and user reports. It's
| the "smoke" in the saying "where there's smoke, there's
| fire" with the caveat that in some cases there's actually
| no fire even if there's a decent amount of smoke.
| djbusby wrote:
| Neat! Wish you had a clickable demo rather than just
| screenshots.
| jmartens wrote:
| Thanks for the feedback, we can do that soon.
| nuggien wrote:
| maybe your service isn't as reliable as you think it is ;)
| mabsboza wrote:
| reporting from Nicaragua, our apps are down!
| rlopezc wrote:
| Down for us too
| yekurtal wrote:
| Mine is still up.
| mhale wrote:
| Heroku apps are down, but API access is available.
|
| Might be a good time to run: heroku pg:backups:download
| jpmoyn wrote:
| Down along with the status page right now
| VWWHFSfQ wrote:
| It is hard-down for me
| [deleted]
| emilsedgh wrote:
| Yes it appears to be down for us as well.
| mmarcant wrote:
| 2/3 of our apps have been timing out of synthetic checks since
| 8:30a PT
| leros wrote:
| I was able to get my apps up faster by restarting the dynos
| pwned1 wrote:
| Back up just now.
| pedroborges wrote:
| Our app is down too :(
| henryaj wrote:
| EU outage towards the end of last year was similarly bad, but
| lasted much longer. Asked Heroku for at least a refund for the
| dyno time when our apps were unavailable and was flatly told no.
| reubano wrote:
| Back up now!
| ericpauley wrote:
| Our app is currently 100% down at the router layer, and dashboard
| requests are mostly failing.
| VWWHFSfQ wrote:
| Routing and logs are down for me as well
| noodle wrote:
| Things just came back up for me
| nordec wrote:
| Same here
| dec0dedab0de wrote:
| I've only ever used heroku for free tier level personal projects,
| but as I understand it they use AWS to do the actual hosting. I
| can understand an outage affecting their deployment process, but
| what could cause running servers to go down?
|
| As I typed that out I remembered that they handle DNS, load
| balancing, and databases, so I guess any one of them.
| alpb wrote:
| All it takes is a load balancer misconfiguration on a core
| service to take a large scale service down.
| jrochkind1 wrote:
| it indeed behaved like a routing issue of some kind (my app
| was still UP, and was still logging, just no traffic could
| get to it), and a heroku incident status line said "Engineers
| are recovering affected routing components," so, yup.
| escot wrote:
| Mines back up, at least 20min down time
| vendiddy wrote:
| Down for us as well.
| btown wrote:
| Same here, hard down for us
| bisRepetita wrote:
| Some of my apps are back up. Others are not.
| digerata wrote:
| Ours are up. US East. Heroku Shield.
| forgingahead wrote:
| Our apps are all down as well, Heroku Status not loading either
| (or just incredibly sluggish).
| [deleted]
| bovage wrote:
| my apps are back up now.
| kennysmoothx wrote:
| Our Heroku applications are all down, it seems to be that it is a
| Heroku Router/DNS issue.
|
| We can access our logs and the application is running, just no
| incoming requests.
| jpmw wrote:
| Let me bet: it's DNS.
|
| It's always DNS.
| ne0n wrote:
| It was, but I don't know why. I'm curious to hear if Heroku
| releases any information about how this happened. Heroku's DNS
| was returning a single 100.64.x.x address which is in a
| reserved range.
| jpmw wrote:
| They typically publish post-mortems, I think? Not 100% sure.
| They definitely do in-house.
|
| We'll have to wait and see I guess.
| nickrubin wrote:
| Yes, all apps are down
| rileytg wrote:
| status, dashboard and landing page are down for me. (socal.)
| sharksauce wrote:
| % heroku status
|
| > Warning: Our terms of service have changed:
|
| https://dashboard.heroku.com/terms-of-service Apps: Yellow Data:
| No known issues at this time. Tools: No known issues at this
| time.
|
| === Availability of Common Runtime apps 2022-02-24T16:50:47.249Z
| https://status.heroku.com/incidents/2402 investigating
| 2022-02-24T16:50:47.249Z (1 minute ago) Engineers are looking
| into reports of connectivity issues to Common Runtime apps in the
| US and EU regions.
| jrochkind1 wrote:
| oh I didn't know that was a CLI command, nice!
| Justsignedup wrote:
| is it related to an AWS outage?
| https://downdetector.com/status/aws-amazon-web-services/
| steeef wrote:
| "User reports indicate problems at Amazon Web Services" User
| reported outage sounds like people confusing the Heroku outage
| with AWS.
| jmartens wrote:
| Functional monitoring of AWS by metrist.io (I'm a co-founder)
| shows no AWS problems. Downdetector is not a reliable source.
| Gaussian wrote:
| We're down
| andycloke wrote:
| All 6 of my apps & dashboard are down too.
| pwned1 wrote:
| It would be nice if they updated their status page. _sigh_
| greset wrote:
| Same here
| ferajay wrote:
| Same here
| drstewart wrote:
| Cypress is also down
| teagee wrote:
| HN is truly a market leader in status page technology
| [deleted]
| js4ever wrote:
| HN was also down for a moment because too much traffic
| Tainnor wrote:
| I experienced that too
| odiroot wrote:
| Such a lightweight tool at that! Barely any JS loaded.
| chasd00 wrote:
| A couple of outages that have affected my client i learned
| about first on HN. We were able to mobilize a team and get on
| top of it faster than any monitoring team at the client ( a
| state government ). I feel like HN should invoice us haha
| castaway3000 wrote:
| You might benefit from some better monitoring!
| duxup wrote:
| Except when it is a false alarm.
|
| I've seen a few get to the front page.
|
| But when we're right we're right.
| dec0dedab0de wrote:
| Even when it's a false alarm it's usually something else is
| having a problem that is affecting many people and
| manifesting itself as a particular service being down.
| jcuenod wrote:
| 10/9 times
| for1nner wrote:
| HN: All of our amps go to 11.
| rileytg wrote:
| *was
| jonpon wrote:
| I came to HN to make sure it was actually down.
| typeofhuman wrote:
| Yes all of my apps are down.
| ab-dm wrote:
| Well that was a fun one to wake up to. We were down for ~33
| minutes overnight.
| ericpauley wrote:
| Anyone still seeing outages? Our services are all back up.
| ezekg wrote:
| All our applications are still down.
| ezekg wrote:
| We're back up. Total 47 mins of downtime.
| cbonser wrote:
| We are still down as well.
___________________________________________________________________
(page generated 2022-02-24 23:02 UTC)