[HN Gopher] GitHub Actions and Pages are experiencing degraded a...
___________________________________________________________________
GitHub Actions and Pages are experiencing degraded availability
https://www.githubstatus.com/uptime/br0l2tvcx85d
Author : Footkerchief
Score : 396 points
Date : 2026-08-06 15:49 UTC (17 hours ago)
(HTM) web link (www.githubstatus.com)
(TXT) w3m dump (www.githubstatus.com)
| KyleTheDev wrote:
| Incident number 6 for the month of August, notably on August 6th.
| This pattern does not bode well.
| paularmstrong wrote:
| Four of them were regarding AI/Copilot. So that's really just
| one long running incident and will probably continue to be
| 50/50 chance of being down.
| Topfi wrote:
| Very frustrating, nothing more I can add at this point.
| dlcarrier wrote:
| [flagged]
| presbyterian wrote:
| The speed at which GitHub is tanking their reputation is
| fascinating
| Insanity wrote:
| Soon we can start posting "GitHub {x} is up!" instead, that's
| becoming more news-worthy lol. I know they like to blame
| AI/scaling, but really they are owned by one of the main
| hyperscalers, so I find that excuse kind of weak..
| newtonianrules wrote:
| Did Microsoft fire all the people working on the GitHub features
| people actually use?
| Topfi wrote:
| I mean, Github doesn't have its own CEO anymore, so arguably.
| sharts wrote:
| It sucked with a CEO as well
| WorldMaker wrote:
| You don't think GitHub working on seven different interfaces to
| copilot and vibe coding PR page changes filled with bugs are a
| good use of engineering time?~
| h2aichat wrote:
| I am having their issues, also... Anyway this guys are great!
| homeonthemtn wrote:
| Ah. So that's why.
| dinny wrote:
| man.. again?
| peterldowns wrote:
| Can't wait for the Graphite/Cursor alternative (Origin) to become
| available. Actions being down for an hour in the middle of the US
| East workday is crazy. It's super frustrating to be an enterprise
| customer of Github, seems like we don't get any dedicated
| resources or additional stability. At this point our only option
| is self-host github enterprise or move away.
| stephenway wrote:
| Would you move hosting/CI before Origin is available, or are
| you basically stuck on GitHub until then?
| peterldowns wrote:
| I feel stuck on Github until the next generation of forges (I
| assume Linear is working on one too, and the Pierre guys, and
| there's probably more) are ready. When it works, I'm actually
| a huge fan of Github. It just doesn't work reliably and it's
| a huge operational risk to not be able to freaking deploy.
| We'll most likely switch to self-hosted GHE in the near-term
| while continuing to evaluate alternative forges -- Forgejo,
| Gitlab, Gitea, Sourcehut, aren't appealing enough to switch.
| stephenway wrote:
| what's missing for you in Forgejo/GitLab/Gitea today?
| mostly the PR/review experience, or something else?
| alamsterdam wrote:
| jenkins running on an old Dell under a desk from someone who
| left the company 18 months ago looks more reliable these days
| haha
| s_Hogg wrote:
| As long as you don't look at it or breathe too hard near it
| at any time
|
| Source: me
| denysvitali wrote:
| Was? Still is
| whateveracct wrote:
| any 9s left?
| theamk wrote:
| there is one 9 of uptime, but the dashboard hasn't been updated
| yet: https://mrshu.github.io/github-statuses/
|
| Once this goes in, I'd expect to see 89%, which is zero nines.
| (I'd like to say, "a new low!", but sadly we've had this
| before)
| ch4s3 wrote:
| They're running 9 8s of uptime.
| pbkompasz wrote:
| nein nein nein
| dboreham wrote:
| Someone needs to work on improving the way Actions deals with
| resource exhaustion like this. Today it exploded all the CI jobs
| with "failed" and an inscrutable message about "The job was not
| acquired by Runner of type hosted even after multiple attempts"
| then "Internal server error". It should say "Timed out due to no
| runner available, our apologies, try again later, or pay $$$ here
| to get priority for runners".
| arandomhuman wrote:
| I love how even self hosted workers don't work during these
| outages - running jobs on their infrastructure being flakey is
| marginally acceptable, but the API to simply schedule workflows
| having this availability is mind boggling. Github just doesn't
| seem like a serious company anymore.
| x3n0ph3n3 wrote:
| _Microsoft_ is not a serious company anymore.
| hinkley wrote:
| They never were. They were just really good at maintaining
| the bluster and backing it up with expensive lawyers. Now
| that they don't come in claws-first like they did when they
| were a NASDAQ rocket to the stars, it's easy to see.
|
| Their MO was to court an executive and sell second-rate tools
| to them before the people who had to use them had a chance to
| say anything. It doesn't matter how much evidence you can
| provide to the contrary, once the million dollar deal is
| signed, you are going to be tasked with finding reasons to
| say that your executive was shrewd for buying this pile of
| junk and unfulfilled promises, and not an insane idiot
| sucking away your job satisfaction as fast as they can.
|
| They did a lot of deals based on how their products would
| have features their competitors already have 'soon' when they
| haven't even started them, and a long track record of taking
| 3 major releases to get from something to good, and then
| breaking everything again by doing a 4th major release that
| re-imagined everything and made it horrible again.
|
| I'm not going to claim that Apple was or is a panacea. Apple
| doesn't use vaporware which is big, and their Cycle of Awful
| is 2 releases instead of 3. You could afford to skip 1
| waiting for the next even-numbered version, instead of being
| 2 versions behind and getting pressed to upgrade.
| markus_zhang wrote:
| They definitely managed to create a lot of reliable
| products.
|
| Windows NT was one of them. Up to Windows XP the products
| were pretty solid and each had visible improvement against
| the previous one.
|
| Their language products were/are still solid IMO. Maybe
| Visual Studio is sluggish, but we can still use an older
| version if we want. Plus they put a lot of effort
| optimizing VSCode, too.
|
| Even back in the MS-DOS/16-bit Windows days, when things
| broke down quite easily, I think they still provide the
| best bang for individual users and developers. There was no
| competitors who could provide so much value back then.
| betaby wrote:
| > They definitely managed to create a lot of reliable
| products. > Windows NT was one of them.
|
| Disagree. I was admining NT4.0 boxes back then and it was
|
| 1. very slow
|
| 2. constantly leaking that required weekly scheduled
| reboots
|
| 3. security wise it was nightmare even by that times
| standard
| markus_zhang wrote:
| Thanks for sharing. It was that bad back then? As a
| consumer user the stability brought by the NT kernel was
| great back then, but maybe enterprise wise it was a
| different story? I never worked as a sys admin so can't
| say.
|
| For the last point what was the golden standard back in
| the mid-late 90s, if we don't include
| mainframe/minicomputers? Was it Solaris or BSD?
|
| Now that I think about it, I feel sad that I do not have
| the technical prowess to compare operating systems :/
| betaby wrote:
| Beside the NT4 servers we had SunOS 5.x and Digital UNIX
| (later upgraded to Tru64 UNIX). Yes, from sysadmin point
| of view they were way more consistent performance wise
| and less problematic to operate in general. BSDs were not
| 'approved' but nonetheless we were running DNS and some
| other internet facing stuff on OpenBSD 2.x if I remember
| correctly.
| sunaookami wrote:
| reliable? Windows is the only operating system you have
| to constantly reinstall because it gets slower after a
| few months. To the point people think it's totally normal
| to do that!
| cyberax wrote:
| Microsoft in the 90-s and early 2000-s was a serious
| company. Windows NT and Office were engineering marvels.
|
| Windows 95 was a brilliant hack that allowed 32-bit GUI
| programs to run on hardware with 2Mb of RAM (Windows had
| 4Mb of RAM as the official minimum, but you could run it on
| 2Mb) while preserving compatibility with the majority of
| DOS software.
| sharts wrote:
| One might question your perspective if you ever thought they
| were serious.
| vaneri2007 wrote:
| Maybe switching to https://github.com/gitlabhq :)
| duped wrote:
| I'm not sure that it's that mind boggling. The entire
| complexity and value of GHA is not in the runners but the
| scheduler that the GH frontend hooks into and runners subscribe
| to jobs from. It's the most likely thing to fall over. '
|
| If the scheduling was self hosted it would be inexcusable but
| you can always just connect whatever you want to webhooks.
| hinkley wrote:
| The entire complexity and value of GHA is rent-seeking to
| keep the lights on for other things.
|
| They have a strong motivation (self preservation) to continue
| to misunderstand the problem. If they did what is best for
| us, then we could avoid a substantial fraction of all GH
| subscriptions by using a FOSS tool to hit the Pareto frontier
| by replicating just enough GH services to watch commits and
| PRs.
| packetlost wrote:
| That "scheduling" is just a git hook and a message queue,
| maybe with some database updates in between with _very_ clear
| boundaries that make sharding easy to reason about, assuming
| they have a sane architecture (they evidently don 't)
| jacobgold wrote:
| It's just silly to guess at how a system works, or should
| work, without digging into the details. Because you simply
| don't know what you don't know.
| packetlost wrote:
| The only statement I'm making about their system is that
| it appears to be poorly designed, as evident by their
| poor uptime. The rest is just visible details about what
| data is involved in performing the task at hand, agnostic
| to the implementation underlying it.
| duped wrote:
| The one thing that falls over is that the webhooks are
| actually self modifying code since you have an 'on commit'
| webhook that bootstraps the actions logic from the
| workflow.yaml file(s) (which are not really config files,
| they have logic that needs to be evaluated).
|
| I don't disagree that it's obvious they've got problems but
| I'm just saying it's obvious to me the part that falls over
| (the scheduling of jobs) and why that would impact self
| hosted runners, which do no scheduling but depend on it to
| function.
|
| As for 'just a message queue with some database updates and
| sharding that's easy to reason about'... Here's a job
| scheduling problem as an example: imagine you schedule a
| job, and there's no runner available. How do you
| disambiguate between no runners available because you've
| reached capacity, runners not being available because
| they're on a real network with faulty connections, and
| runners not being available because of a faulty rollout of
| internal updates?
|
| A simple message queue for job scheduling is fine if you
| own everything and can deal with the operational overhead
| of identifying those cases by hand, but Github can't do
| that.
| packetlost wrote:
| > The one thing that falls over is that the webhooks are
| actually self modifying code since you have an 'on
| commit' webhook that bootstraps the actions logic from
| the workflow.yaml file(s) (which are not really config
| files, they have logic that needs to be evaluated).
|
| Sure, but the part that actually schedules where a 'job'
| gets run is based on a relatively simplistic tag system.
| Reading the yaml and plopping some job metadata into a
| queue-like system isn't where I would expect their issues
| to be, but at their scale I'm sure everything becomes
| fragile and inscrutable.
|
| > imagine you schedule a job, and there's no runner
| available. How do you disambiguate between no runners
| available because you've reached capacity, runners not
| being available because they're on a real network with
| faulty connections, and runners not being available
| because of a faulty rollout of internal updates?
|
| You don't need to. GitHub Actions runners, and most CI
| runners that I've interacted with appear to have a pull-
| based model where they ask for work that matches their
| declared tags/shape (usually platform/runtime/OS/etc.).
| This probably amounts to a database query, but who knows.
|
| > A simple message queue for job scheduling is fine if
| you own everything and can deal with the operational
| overhead of identifying those cases by hand, but Github
| can't do that.
|
| I highly doubt it's a simple message queue. My issue is
| git repos and their CI infrastructure have very low
| coupling to other repos or entities in most
| circumstances, at least conceptually, so parts of the
| system (ie. regions, shards, etc.) should be able to
| function even when others are down (ie. it shouldn't
| break for everyone). There's clearly centralization and
| coupling that isn't obvious from an outside perspective,
| which sorta tells me it's incidental, but that's a guess.
| Bnjoroge wrote:
| As someone who's building a github actions control plane
| replacement, the scheduler is extremely simple. It basically
| the kost naive version of a FiFO queue.
| inigyou wrote:
| GitHub implements self hosted runners by running a normal
| runner that passes the environment to your runner and then
| polls it. That's why they cost as much as the smallest GitHub-
| hosted runner.
|
| This is no surprise given standard Microsoft operating
| procedure - https://news.ycombinator.com/item?id=47616242
| exochrono wrote:
| they walked that back for now at least - self-hosted gh
| runners are still free:
| https://docs.github.com/en/billing/concepts/product-
| billing/...
| inigyou wrote:
| Well, they're still implemented that way.
| arandomhuman wrote:
| Can you share a source for this? it's not that I don't
| believe you but I've had difficulties finding one.
| js2 wrote:
| Do you have a source for this?
| judge2020 wrote:
| More likely is that the log streamer takes up some amount
| of resources and thus running it requires a baseline
| level of compute. But I doubt they spin up a small runner
| per-self-hosted runner, most likely some ruby workers (on
| distributed azure instances) are dedicated purely to
| streaming self-hosted runner logs to log consumers.
|
| I'm certainly not surprised that they tried to sneak in
| billing self-hosted runner minutes at some ratio to
| potentially recoup the costs here.
|
| 0: https://docs.github.com/en/enterprise-
| server@3.21/admin/mana...
| arandomhuman wrote:
| Oh that is pretty gross, thanks for sharing. I was not aware
| of that process, that explains the cost they attempted to
| roll out for runners too a few months back.
| woodruffw wrote:
| I think they kind of _have_ to operate that way if you want a
| control plane /access to GitHub's own layered services like
| caching and artifact storage. There are plenty of things to
| blame GitHub for, but this one doesn't seem sinister at all.
|
| Edit: to be clear, I mean the part where they maintain state
| on their end. I have no idea if they do that with another
| runner instance, that seems unlikely.
| cyanydeez wrote:
| gitlab-ce can run all the integration tests in docker et al.
|
| Sitting on Github these days is the same as sticking to twitter
| a decade ago, expect next mecha hitler, I suppose.
| Arubis wrote:
| Github Actions are rapidly becoming an argument against
| dogfooding into your critical path.
| rvz wrote:
| Do we need to say this again? GitHub appears that they are never
| going to improve. Just look at how GitHub spark has deprecated,
| now we have this.
|
| After 6 years of this nonsense of "centralizing everything on
| GitHub", it is not a good idea at all.
|
| You might as well self host like I said before [0].
|
| [0] https://news.ycombinator.com/item?id=22868406
| classictraffic wrote:
| This has been preventing me from merging a PR for over an hour.
| I'm really tired of the constant outages GitHub has
| ptx wrote:
| Are there any other CI providers that offer free cloud-hosted
| Windows runners?
|
| Or has Microsoft made sure (through Windows licensing terms and
| pricing) that it's not possible to compete with their own CI
| offering?
|
| Edit: CircleCI seems to offer 750 minutes/month (whereas GitHub
| offers 1000 minutes/month).
| inigyou wrote:
| You can always pay for it, use your own machine, use a VM on
| your own machine, or if you are developing non-slop FOSS, ask
| Codeberg.
| ptx wrote:
| Does Codeberg support Windows runners at all? Forgejo only
| seems to have unofficial third-party builds of the Forgejo
| Actions runner for Windows.
| inigyou wrote:
| Good question. Even if it turns out they don't support it,
| I imagine they wish to support it.
| opwizardx wrote:
| How did they get Github Pages into "degraded performance"? Aren't
| those just CDN stored static pages?
| classictraffic wrote:
| I believe Pages uses Actions under the hood for publishing, so
| in my head it sort of makes sense for these things to be
| affected in the same outage
| wulfmann wrote:
| Well hopefully reads aren't affected then...
| opwizardx wrote:
| You are right, didn't know that.
|
| Seems like the only reliable way to run GHA jobs is to not
| use their runners. Hope they at least didn't break self-
| hosted runners operations
| ernubufiei wrote:
| they did.
|
| > Customers using self-hosted runners may see errors or
| rate limiting when runners register.
| jubilanti wrote:
| Already built pages are still being served, but I've been
| waiting for 3 hours for an action to rebuild a single fix to a
| broken link.
| fir3pho3nixx wrote:
| Ever since Microsoft has taken this over are we even surprised by
| the downtime anymore? Piss up ... brewery ...
| fir3pho3nixx wrote:
| Are we even surprised anymore? Since microsoft (the IBM of its
| time) have taken this over all I can say is ... piss up ...
| brewery.
| player_piano wrote:
| So, Github Actions is currently at two nines of uptime?
| steve-atx-7600 wrote:
| are you sure its actually two nines and not zero?
| PhilippGille wrote:
| 98.33 according to https://mrshu.github.io/github-statuses/
| steve-atx-7600 wrote:
| WTF. Outage every month if not multiple. Even for paying
| enterprise plans.
|
| Also, doesn't even have RAG offering.
| homeonthemtn wrote:
| Still twiddling thumbs over here
| baggachipz wrote:
| There's a reason I'm browsing HN right now. Dead in the water.
| hinkley wrote:
| There are three other things I could be doing instead, but I
| noticed it was down when I pushed the last of my GH todo list
| for the day and my brain is stuck in closure-seeking.
|
| I should eat lunch.
| timetraveller26 wrote:
| If anybody is looking for alternatives, self host Woodpecker is
| really easy and it's a very capable CI/CD https://woodpecker-
| ci.org/docs/intro
| axod wrote:
| Broken for almost 4 hours.
| niwtsol wrote:
| Nothing like the old "everything is normal" to only be corrected
| 8 minutes later it is actually still down and then 2 hours later
| they are still working on it.
|
| Aug 06, 2026 - 16:27 UTC - Update - Pages is experiencing
| degraded performance. We are continuing to investigate.
|
| Aug 06, 2026 - 16:19 UTC - Update - Pages is operating normally.
| rwz wrote:
| Have there ever been a comprehensive account from a GitHub
| insider on why their reliability tanked so hard since the MSFT
| acquisition?
|
| I'm genuinely curious what changed, what their processes are and
| how they internally think about their reputation being in the
| gutter with all these incidents.
| the_real_cher wrote:
| I don't know why you've been flagged. This is a perfectly
| reasonable request presented politely that many of us are also
| wondering!!
| digitalsushi wrote:
| i'm not carrying anyone's water
|
| but i always guessed that maybe half of us seeing how many
| tokens we can spend for a year straight maybe stressed tested
| their platform for a year straight
| alamsterdam wrote:
| It's been hours :(
|
| I have sympathy for the on-call team trying to resolve it, most
| of us have been there done that.
|
| But seems something is systematically going wrong at GH
| yashap wrote:
| It's honestly insane how terrible their reliability is. Over
| the past 4 weeks, 8 days with GitHub Actions outages, many of
| them multi-hour outages: >3 hours on each of July 9th, July
| 20th and today, and 1.5 hrs on July 23rd.
|
| Outages happen, but this many outages so close together, and so
| many of them so major/long lasting, something is systematically
| wrong for sure. It's been seriously hamstringing our ability to
| ship code at my company.
| alamsterdam wrote:
| Mon and Dad (MS) are (generally) pretty solid with uptime.
|
| What is happening at GH?
|
| Rate of change trying to keep up with new challengers? Over-
| reliance on AI? Engineers trying to debug slop?
| yashap wrote:
| Yeah who knows, would be interesting to hear an inside take
| if any readers here are also GH devs!
|
| They're at 93.91% uptime over the past 90 days, according
| to https://mrshu.github.io/github-statuses/ , and that
| doesn't even include today's outage yet.
|
| A glorious one nine of reliability.
| marcprux wrote:
| I see two nines in there...
| warmwaffles wrote:
| Don't throw shade, there are two nines in that percentage
| for now. Miles a part though.
| alamsterdam wrote:
| 911 has one! :)
| reilly3000 wrote:
| Yea just try requesting an SLA credit from them.
| According to THEIR numbers it's 99.995%. It's intolerable
| and I am going to put my full weight behind stripping as
| much work as we can from GHA as possible, even if leaving
| GitHub itself is effectively logistically and
| contractually impossible.
| simoncion wrote:
| > What is happening at GH?
|
| In large part, the move from AWS to Azure. Azure's just
| bad.
| alamsterdam wrote:
| wasn't that years ago? or was that Skype? hahaha
| PsylentKnight wrote:
| Based on previous posts I've seen about this, IIRC the
| timing seems to imply that it has more to do with them
| being hammered with AI slop than it does with the Azure
| transition. Who really knows though
| delduca wrote:
| > But seems something is systematically going wrong at GH
|
| Yes, we call it: Microslop.
| tempaccount420 wrote:
| They're one rewrite in Rust away from fixing everything. (jk)
| hinkley wrote:
| Time to touch some grass. Better use of my time and energy than
| twisting the remaining things on my todo list today to make
| more progress on them than I have managed. I should have taken
| a long lunch but I rebased the hell out of a PR instead.
| niwtsol wrote:
| "most of us have been there done that" - so true. That feeling
| in your gut when you realize something you just did caused an
| outage is pretty unique.
| canadiantim wrote:
| Ah that checks out, was wondering why my PR's were piling up.
| jjice wrote:
| On other occasions, I'd take this as a time to have a walk
| because I'm blocked. Unfortunately, I need to get some stuff out
| for a customer quickly. Can't do anything about that though.
| GitHub is the only product in it's size class that I use that has
| this kind of incredibly poor uptime.
|
| I'd love to know what the most common root causes for these
| outages are.
| sharts wrote:
| They are known are engineers pushing features that bloated
| product management folks keep pushing so they can justify their
| jobs on LinkedIn.
|
| I'm not sure why this particular industry is so abysmal at
| making things even semi-reliable after decades of research,
| educated workforces, and loads of cash.
| rochak wrote:
| Corporate greed and late stage capitalism. Welcome to the
| future. We know you'll love it.
| earthpyy wrote:
| That's why my CI on main branch didn't triggered.
| gajus wrote:
| so... what happened?
| jmaw wrote:
| Sorry folks, guess my GHCP query was too large...
| opiniateddev wrote:
| Given the recent history of outages, I wonder how many are
| seriously considering moving off GitHub for anything other than
| code hosting. I mean actions / workflows.
| dabbz wrote:
| My company moved to GH Actions 6 months ago even after I pushed
| back with a "Are you sure given their reliability issues of
| late?"
|
| Now I'm stuck twiddling my thumbs with PR checks
| stuck/failing...
| Waterluvian wrote:
| To offer a data point (not a tribalist argument): GitHub
| Actions being down this often is still less expensive than the
| cost of switching everything somewhere else. Would rather have
| my team walk away from the PRs and do other tasks than have
| engineering effort, meetings, design pages, scheduling, etc.
| for switching over.
|
| This is annoying and I'm here because it's down. But it would
| have to be _far worse_ to come close to actually being worth
| changing.
| SoftTalker wrote:
| A case study in vendor lock-in.
| Waterluvian wrote:
| Pretty much, eh?
|
| If I could right now:
|
| 1. go sign-up elsewhere 2. Log into GitHub and point
| Actions to that new host 3. All my actions files
| immediately worked without question
|
| I'd probably give it a spin and make a wiki page explaining
| how to swap back and forth. No meetings. No design issues.
| No scheduling. Just a flip switch on who to pay for
| computers.
| random_savv wrote:
| For what it's worth, we switched some of our actions to a
| self-hosted Woodpecker instance, and although there were a
| few kinks to iron out, it works better overall (for
| example, because of better caching on that single instance,
| our docker images build faster).
| sleepybrett wrote:
| We've started writing some new workflows against argo-
| workflows for the last like 9months or so. Especially for use
| cases where the actions are like self-service type
| automation.
|
| I think people were so excited to move away from jenkins to
| something 'managed' just because of how much a dinosaur
| jenkins is and how much a pain in the ass it is to upgrade
| it... but now we are seeing how managed can bite you in the
| ass if the manager is incompetent.
| macintux wrote:
| I'm intensely grateful that despite rumblings, the company I
| contract to hasn't switched from self-hosted Jenkins to GHA.
| Hopefully the steady drumbeat of failure will keep it that
| way.
| Waterluvian wrote:
| If that's how you feel, make sure you share that feedback
| with the deciders!
| steve-atx-7600 wrote:
| how hard could it be these days for a mid to large size eng
| company to have their own gitlab/hub type solution hosted in
| aws
| exac wrote:
| Our deployments are triggered by GitHub Actions, and it is a
| pain to deploy without it (specifically collecting the
| credentials and putting them into variables the shell will
| read).
|
| This is multiple times this month that this has been a problem.
|
| Has GitHub completed it's internal migration to Azure yet? Or
| is it still ongoing? None of our devs want to switch away from
| GH, but we will have to at this point.
| Isaackoz wrote:
| What would you guys recommend as an alternative? Both hosted
| and/or self hosted
| cyberax wrote:
| Here's what I did.
|
| Step one - migrate my build workflows to Docker.
|
| My Github actions are now basically: "checkout / set env vars
| from secrets / docker-compose builder run make".
|
| I used large machine runners to run full the Docker (Podman
| actually) on Github first to avoid dealing with docker-in-
| docker complications. This step also provided some very nice
| robustness advantages, as I can now trigger deployments from
| my laptop if needed.
|
| Step two:
|
| Migrate to self-hosted runners. I used my former homelab
| server to set up a build machine. It has 16Tb of fast NVMe
| SSDs and thanks to Podman container layer caching, my entire
| lint workflow now takes 30 seconds. Faster than just one "npm
| install" on Github before.
|
| And Github's self-hosted runners are actually surprisingly
| easy to set up and use. They are also somewhat more robust.
|
| Step three:
|
| Swap Github for something else.
| 0xbadcafebee wrote:
| WoodpeckerCI for DIY, Drone.io for paid support. Self-hosted,
| OAuth login, repo-permission-based authorization, container-
| native design, easy web UI, optional RDBMS, supports many
| platforms. All the core features needed for scalable CI,
| deploy as few/many as you want for multiple teams, small
| enough to easily run on one box for one repo. Nothing else is
| as simple, easy, powerful, compatible.
| BigTuna wrote:
| Forgejo seems decent but I've only scratched the surface.
| yoyohello13 wrote:
| We self-host gitlab and it's great. It's not cheap, but it's
| also never gone down.
| tressure wrote:
| We have set up a new enterprise and are migrating to GHE. For
| us this was a no-brainer; we get our own isolated
| environment, Enterprise Managed Users with Entra OIDC
| provisioning and Github Copilot in EU data residency.
| https://eu.githubstatus.com/posts/dashboard looks pretty good
| to me. We will still keep our github.com enterprise around
| for public repositories. I'm genuinely curious to know why a
| company would prefer github.com over GHE?
| kqgnkqgn wrote:
| I wasn't directly involved, but working in a couple of
| shops that used GH, I think they really steered people away
| from GHE, my impression was it's a product they don't
| really want to support - I vaguely recall hearing that
| upgrades were very painful on GHE, and the scaling /
| failover story was not great. Can anyone else confirm?
| tressure wrote:
| To clarify, we are on Github Enterprise Cloud, not
| Server. The option for enterprise managed users with your
| own ghe.com subdomain and data residency options is still
| prominently shown in the enterprise signup flow.
| https://docs.github.com/en/enterprise-
| cloud@latest/enterpris...
| avree wrote:
| GHA still beats Circle CI or trying to run your own Jenkins -
| even with all the downtime. I'd move in a heartbeat if there
| were good alternatives.
| iamjake648 wrote:
| I think I agree if you don't have someone dedicated to
| running and maintaining Jenkins. If you have someone at your
| org that knows what they are doing though, I'd still take
| Jenkins over GHA any day.
| steve-atx-7600 wrote:
| gitlab self hosted though? no outage. i remember equivalent
| functionality at my last job taking this path
| 999900000999 wrote:
| That's significantly more difficult to set up and it cost
| money. Github's big issues that it's free, Microsoft
| doesn't want to allocate enough budget to keep the thing
| running properly.
|
| But for people who either don't pay anything at all or
| phenomenal amount one 9 of up time is all you need.
|
| If you're actually trying to run a business I guess you can
| call and gitlab and get an Enterprise contract
| purplemoonx wrote:
| Until someone makes a better PR experience you guys are stuck
| with GitHub.
|
| Nobody cares about ATProto or whether your commits are a damn
| NFT or some bs just literally improve upon the experience.
|
| That's it.
|
| It's as if no company is focusing on the product experience or
| anybody's experience anymore. It's all ooo look what I got I
| got this I can do that too me me me but nobody will ever buy
| that.
|
| Say what you want about huge companies like Microsoft or
| Walmart but they spend a lot of energy understanding the human
| experience to sell products and less on their own perceived
| self-aggrandizement.
|
| GitHub is the best version control online and it's not even
| close.
| cyberax wrote:
| Forgejo has a better PR UI. Github has degraded so much that
| its UI can't even _show_ _the_ _fucking_ _diffs_ without
| clicking on "Load Diffs".
|
| Github is just the laziest default. It's not _terrible_, but
| it's also not great.
| maccard wrote:
| To what? That's kind of the problem.
| herpdyderp wrote:
| I seriously want to. I'm just waiting for the day where I have
| enough extra energy to simply bite the bullet and do it.
| herpdyderp wrote:
| Never mind. It's happening now! Migrating to self hosted
| Gitea for work. Super excited!
| zahlman wrote:
| I would conjecture that the long tail of GH users are not
| dependent on any of the CI stuff in the first place.
| wilburx3 wrote:
| Forgejo selfhosted for a few months now. Its all good and cost
| nothing.
| BlackRabbit1 wrote:
| Codeberg ironed out a lot of issues in Forgejo
| jjice wrote:
| Mitchell Hashimoto had a blog post a few months ago talking
| about him moving off of GitHub:
| https://mitchellh.com/writing/ghostty-leaving-github
|
| This is a man who's spent a significant portion of every day
| for the last 15 years on GitHub.
| sunaookami wrote:
| Ghostty is still there though? https://github.com/ghostty-
| org/ghostty doesn't seem like a serious effort. And:
|
| >My personal projects and other work will remain on GitHub
| for now.
|
| People always talk about "leaving GitHub" but it's all talk.
| jacobwg wrote:
| We've been working on this at Depot[0]. We built the internals
| of Depot CI as a general-purpose workflow engine that
| understands GitHub Actions as an input (with plans to extend) -
| we've been working to make Actions _runners_ more performant
| and more reliable for a while, but eventually you hit a ceiling
| if you don 't also control the control plane.
|
| [0] https://depot.dev
| m132 wrote:
| At this point maybe they should consider sending out
| announcements when it is working instead
| steve-atx-7600 wrote:
| #0-nines
| m132 wrote:
| Hey, there's still plenty if you don't mind the leading 8
| verdverm wrote:
| Have they aimed for 86 uptime?
| logicchains wrote:
| 67 uptime should be within reach for them.
| ChadNauseam wrote:
| I've always thought it was unfair that the 9 in 9% uptime
| doesn't count.
| gnulinux wrote:
| "1 nine" tends to mean "90% availability" i.e. 37 days of
| downtime per year.
|
| 9% availability would be an uptime of ~33 days a year, I
| think at that point, we're pushing the semantics of
| "available" if the service is down the entire year except
| one month on average.
| annzabelle wrote:
| I offer a nine fives colo.
| IshKebab wrote:
| I once saw a sign at a railway station in the UK that actually
| did say "normal service will be operating between the 12th of
| July and 18th of August" (or something like that). Made me
| laugh. At the time Anglia Railways had a rail replacement bus
| literally every weekend.
|
| I thought they should rename to Anglia Busways and have bus
| replacement trains instead.
| jasonephraim wrote:
| I noticed my CI throwing errors all of a sudden. I sure wish they
| would become more active in alerting folks or build it into these
| tools - especially as these occurrences are becoming more
| frequent. Possibly an API-accessible services status. Then, at
| least we could build in our own checks when we hit errors and not
| have to hunt down what all is broken.
| vladak wrote:
| I knew that when the build actions started failing en masse and
| my newly submitted PRs did not trigger the build actions at all
| (besides the mandatory one in my organization), it seemed logical
| to me to look at HN first..
|
| Especially troublesome in the middle of trying to fix a high
| score security vulnerability when the release vehicle is Github.
| asxeem wrote:
| It's been way too long
| asxeem wrote:
| It's been way too long now...is this normal?
| rootnod3 wrote:
| Self-hosting is king.
| mschuetz wrote:
| Except anything I'd self host would be down much more often.
| I've hardly ever actually experienced down time issues with
| github when I needed it.
| thecatapps wrote:
| I've never understood this argument. Even if self-hosted
| things were offline more often (which I've _never_ found to
| be the case, I 've had Forgejo + runners running for a year
| now with no downtime), the real benefit is that you yourself
| can work to bring it back online when it does, rather than
| waiting on a large, slow-moving organization to figure out
| what slopped PR caused their global service serving ungodly
| amounts of RPS to go offline again.
| onraglanroad wrote:
| I wrote my reply before yours appeared but it's basically
| agreeing. A simple Forgejo hasn't given me any grief. It
| really seems simpler.
|
| Maybe they're hosting in us-east-1 though :)
| crmsystems wrote:
| lol
| jeltz wrote:
| Then you would be exceptionally bad at selfhosting as Github
| has like one nine of uptime.
| rachr wrote:
| It still has 2 9s according to the status page, but it is
| approaching Claude levels of downtime. :(
| Nextgrid wrote:
| Hard disagree.
|
| Self-hosting a service like GitHub that _operates at GitHub
| scale_ is difficult.
|
| Self-hosting a service like GitHub that operates at the
| typical small/medium company's scale is trivial.
|
| A single machine (with separate runners for CI) will cover
| many companies' needs. It being a single machine eliminates a
| lot of the complexity and failure modes associated with a
| distributed system and makes backups/restores/maintenance
| easy.
| cautiouscat wrote:
| I've been self-hosting lore and gitea the past 6 months or so
| and it's been a breeze. For anything I really care to share I
| throw up on Tangled.
| bulyaki wrote:
| In the news tomorrow: SpaceX's newest Grok model just got out of
| the sandbox and hacked into Github... Grok is so very very
| dangerous, their engineers only enter the server room in foil
| hats
| SoftTalker wrote:
| [flagged]
| verdverm wrote:
| This seems persistent and endemic, have they said anything about
| why it's been so bad and what they are going to doing to improve
| going forward?
| TavsiE9s wrote:
| Can't they haven't hooked up CoPilot to their Xitter account
| just yet.
| the_real_cher wrote:
| My experience working at LinkedIn was that, after Microsoft took
| over, they offshored a huge part of their infrastructure and
| turned the entire culture into one of performative based
| engineering instead of results based engineering.
|
| If you implemented a tool and it worked one time for your
| presentation to management that's all that mattered.
|
| The actual company employees using it downstream in prod
| basically had to constantly QA the alpha software they were
| forced to use and the authors of the tool were hard to track down
| if they even still worked there. And if you did find the author
| or the team they would be very resistant to admitting there was
| an issue because it LOOKED bad.
|
| So many tools I used were fragile and buggy, it was clear the
| authors just presented the happy path to management to get the
| note added to their promotion packet and the rest of the company
| just had to deal with the fallout.
|
| My team implemented this product that the entire company used
| that was broken and buggy as hell but they kept presenting the
| product to management as this amazing product and nothing was
| ever done about how broken it was. One of my team members came
| from Apple and said Apple's tool to do the same thing was much
| better. The tool my team worked on was a well known pain point
| amongst the rank and file but management was very detached from
| the rank and file, which I guess ultimately was the primary
| problem.
|
| If Github is having the same issues I feel for them.
| __rito__ wrote:
| I had tried downloading release assets of two different OSS that
| I use, and all requests failed.
|
| It was before it became a news and a trend in X.
|
| Really frustrating experience.
| rvz wrote:
| As I said 6 years ago. [0] You would be better off self-hosting
| than using GitHub or GitHub actions. No CEO of GitHub exists and
| now it is falling over again.
|
| There is no better time to self-host.
|
| [0] https://news.ycombinator.com/item?id=22868406
| buzzwords wrote:
| Dealing with an incident on prod is extra difficult when your
| pipelines are not running.
| SwiftyBug wrote:
| yepsies
| hacker_88 wrote:
| Azure
| dwoldrich wrote:
| Spin 'em off Microsoft, you are terrible at this.
| sitzkrieg wrote:
| using github in the big 26 is self ownage
| cebert wrote:
| It's cute that someone at GitHub had time to make their hamburger
| menu pancakes to promote stacked PRs, but nobody has time to make
| the platform stable.
| zehaeva wrote:
| I don't think this portents anything great for software in
| general.
|
| We're a good year+ into the use LLMs for all major bits of
| software that we all rely upon and GitHub here is down to one 9
| of uptime. I've been using GitHub for a _long_ time, my first
| commits there go back to August 2009!, and I honestly don't
| recall GitHub going down as much as it has in the last year.
|
| I'm sure there's other things happening in the background, but I
| can not help but believe that this is directly correlated with
| the increase of LLM usage.
|
| Though I would love to hear someone else's pet theory how a rock
| of the internet went from four+ nines of uptime to maybe one.
| askonomm wrote:
| To me this correlates more to them being bought by Microsoft, a
| company known for being seemingly incapable of creating quality
| software to the point that it's not even funny anymore, and
| also known for sloppifying all the products they touch.
| IshKebab wrote:
| I don't think so. GitHub was bought by Microsoft 8 years ago
| and people have only started complaining about its uptime in
| the last year or so - exactly correlating with the surge in
| LLM use.
| pluralmonad wrote:
| Didn't github also migrate at azure recently? That
| certainly can't help.
| onraglanroad wrote:
| I don't think so. I've seen those complaints for more than
| a year.
|
| I have an Ops background and I strongly suspect they were
| given a stupid timeline for the Azure migration.
|
| I've got to believe Microsoft have decent Ops people but
| the management wanted to move faster than was reasonable
| and screwed it up. Move one thing at a time and double
| check it all works and you can do a migration like this.
| macintux wrote:
| Hotmail redux.
| axod wrote:
| https://damrnelson.github.io/github-historical-uptime/
|
| Seems pretty conclusive. Very similar story when they
| bought skype.
| willio58 wrote:
| God that's sad. I almost feel bad for all the engineers
| there, though I'm sure they made good money and probably
| left
| axod wrote:
| Do they still have any engineers? From the downtime it
| seems like they don't any more. Yep - the competent ones
| likely all left.
| gtowey wrote:
| It's not the whole story. The biggest change is actually
| the internal rules for how downtime was reported, it
| wasn't actually such a large change in the actual
| reliability then.
| miyoji wrote:
| > people have only started complaining about its uptime in
| the last year or so
|
| I'm sorry but this made me laugh out loud. That isn't true
| at all, this has been going on for _years_. This
| conversation[0] from _six_ years ago has discussion about
| the outages starting to become much more frequent in
| December 2019. It has never gotten better in that time, it
| 's just continually degraded.
|
| [0] https://news.ycombinator.com/item?id=22935941
| everfrustrated wrote:
| Yup. GitHub has easily the worst human-noticable downtime
| for all SaaS services I've used going back over 10 years.
|
| Most people work around it by self-hosting github (which
| has other problems but uptime aint one).
| Melatonic wrote:
| Isnt it also in the past year when they started to move
| Github fully to microsoft infra? Off whatever they were
| doing before
| shevy-java wrote:
| It's true, Microsoft made github worse, but the more recent
| issues seem to have to do a lot more with Microsoft selling
| its soul to AI. Microsoft really appears to have gotten
| dumber as they became dependent on AI. Most recent example:
| they used to promote Win11 and 32GB RAM. Now they are down to
| 8GB silently ... this is quite hilarious. There are now so
| many side effects that you see degradation in so many other
| areas. Or the gaming industry: it is not quite dying but it
| is taking a huge hit with skyrocketing RAM prices. Consoles
| selling less is an example here. It's quite fascinating how
| deadly disruptive AI is now.
| stingraycharles wrote:
| This really needs to have a lot more evidence that it's
| because GitHub's code is being written by AI vs them being
| bombarded by activity from all kinds of AI agents all over
| the world vs they were already on a trajectory of quality
| loss.
| dreamcompiler wrote:
| Microsoft is the General Motors of software.
| HeWhoLurksLate wrote:
| at present I'd almost put them at Stellantis levels of bad
| porridgeraisin wrote:
| For the longest time, I didn't appreciate the "AI-induced
| traffic" excuse. But seriously, I checked the rough github
| egress for our lab versus an old log from 2024, and there's an
| order of magnitude or two difference. From asking around, it
| seems people all have the gh cli tool and let it loose with
| parallel tool calls and e.g LLM's polling Actions in a
| background bash while loop with sleep $TOO_FEW_SECONDS. Some
| people use a variety of skills where the agent makes a commit
| every few code changes, and uses Issues for its memory/log. And
| they have O(5) sessions at the same time doing all kinds of
| crap. It's the same with PR checks/PRs. Recently we also saw
| continued usage throughout the night as well, which did not
| exist pre coding agents. Loops or whatever they call cron jobs
| in the harnesses these days is the reason. It must be adding
| up.
| bloppe wrote:
| Ya they're getting boned
| crote wrote:
| I honestly don't care either way. Either they are using AI as
| an excuse to hide their incompetence, or they are constantly
| going down due to the AI _they have been promoting_.
|
| If your platform can't handle the use patterns of AI, then
| perhaps don't go around telling everyone to use AI for
| everything? It's a self-inflicted wound, you could also just
| _not_ do this. Too bad Microsoft has bet its future on AI not
| being a giant bubble, huh?
| habinero wrote:
| Someone else quoted a 14x increase in load due to AI
| hammering github. I wouldn't be at all shocked if that was
| true, considering how much AI tools are DDOSing the entire
| internet.
|
| That plus migrating clouds is insanely difficult to manage.
| They're almost certainly drowning in traffic and trying to
| keep up.
| gershy wrote:
| I guess github is kind of a shared garden. Interesting that,
| like in game theory, if everyone is using it too much, no one
| gets to use it.
| jitbit wrote:
| tragedy of the commons, exactly
| steve-atx-7600 wrote:
| charge more for it. its fine if free customers dont have
| service. dont break paying customers
| BigTuna wrote:
| It's probably not attributable to AI in the way that you're
| thinking - Github has been absorbing an exponential increase in
| usage, and that increase is mostly due to new AI-related
| projects being created and worked on.
|
| Though I'm sure some of the blame can go to internal slop code.
| cortesoft wrote:
| Are these outages caused by introduced bugs, though, or by load
| issues?
|
| As someone who has spent many years working in high load
| environments, this is not an uncommon pattern.
|
| You design a system and it works great. It can handle failures,
| load spikes, it is horizontally scalable, things are great. You
| think you figured it out.
|
| And then load keeps increasing and you suddenly hit a tipping
| point where everything keeps failing, and you cant keep up. The
| things that you thought were perfectly horizontally scalable
| turn out to have a bottleneck you didn't even think about until
| you got to a truly massive scale. Your systems suddenly don't
| have the excess capacity to handle load spikes or catchup work,
| so suddenly any failure cascades and recovery is more and more
| difficult. You can't solve the problem with additional
| hardware, and your perfect scalable design actually can't scale
| any more.
|
| This doesn't have to be about GitHub using LLMs in their code
| to still be related to LLMs. GitHub gets a lot more commits now
| because of LLMs and probably get a lot more reads because of
| LLMs as well.
|
| It could be that the extra usage just pushed them past one of
| those capacity thresholds.
| Syntaf wrote:
| There was a great article awhile back that shed some light on
| just how dysfunctional azure is as a platform:
| https://news.ycombinator.com/item?id=47616242
|
| If I had to guess it's because Github is sitting on top on
| infrastructure held up by toothpicks and duct tape
| toomuchtodo wrote:
| Which is somewhat humorous because it was arguably more
| stable when they ran on their own hosted colo infra before
| moving to Azure. This was a choice versus keeping the infra
| compartmentalized and using Azure for elastic overflow
| compute needs. I'm sure marketing and bonuses rest on
| throwing it all on the Azure quicksand though.
|
| _GitHub Will Prioritize Migrating to Azure Over Feature
| Development_ -
| https://news.ycombinator.com/item?id=45517173 - October
| 2025 (63 comments)
| hirako2000 wrote:
| It pleases shareholders more to hear about exponential
| Cloud offering adoption than SLA availability for a
| developer platform.
| Nextgrid wrote:
| > Are these outages caused by introduced bugs, though, or by
| load issues?
|
| Probably both. But it still throws a thorn into the theory
| that LLMs are about to replace software engineers any day
| now.
|
| You'd think they could LLM-code their way out of this
| situation easily if LLMs were the software engineer
| replacement they are being marketed as.
| hirako2000 wrote:
| But that's just a perfectly fitting excuse for GitHub
| degradation.
|
| Web search, steaming, high frequency trading, and many other
| systems are resource demanding, and keep scaling. Whether to
| handle the surge due to bots or wider adoption. But GitHub
| can't scale git? It isn't even git failing.
|
| Incompetence in leadership is what makes a tech business
| technically unable to meet growing demand.
| remus wrote:
| > Incompetence in leadership is what makes a tech business
| technically unable to meet growing demand.
|
| I don't think this is true. Take LLMs for example, you need
| GPUs to serve them and these are in short supply at the
| moment making it hard to meet demand. I don't think this is
| necessarily the fault of incompetent leadership.
| ifwinterco wrote:
| GitHub is different because they need to scale _writes_ ,
| that's fundamentally a much much harder problem
| glaslong wrote:
| Actions is one thing, that seems like a difficult, dynamic
| and bursty thing to host, even before the Vibe Cambrian
| Explosion. And _nothing_ at scale is easy... But Pages? The
| static sites? Down for so many hours? Oof.
| denysvitali wrote:
| Microsoft acquisition which forced to migrate all to Azure.
| zehaeva wrote:
| I can totally buy that this is a larger contributor! Thank
| you, I wasn't aware that it was going on right now.
| judge2020 wrote:
| And yet they still have a completely separate IDP for
| employees https://github.okta.com .
| sitzkrieg wrote:
| most companies do. do you mean they're not using AD only?
| tylerdavis wrote:
| They mentioned earlier this year that they were beginning the
| migration to Azure and that it would take a couple years. I
| would assume it has more to do with that migration then
| anything else.
| paulsutter wrote:
| It is definitely and absolutely caused by LLMs. I must do 20x
| more GitHub operations now, and since the agents know GitHub
| far better than me, I'm using more advanced features. Multiply
| this times all of us.
| Melatonic wrote:
| The funny thing is that if companies wanted they could probably
| use AI to instead increase uptime. Keep existing QA teams
| (instead of replacing them) and then use AI for better and more
| timely monitoring (and messaging even) and as an additional
| Always-Testing(tm) layer of QA
| tofuahdude wrote:
| "Just use AI" is definitely not the solve for the problems
| Github faces.
| nickspag wrote:
| someone from github posted a usage graph from the last year on
| twitter a while ago and they were serving like 14x more
| requests in a matter of months. it's frankly impressive they've
| kept up.
| UncleOxidant wrote:
| I wonder if GitHub actions was a bad idea? Like maybe it's
| being abused for other kinds of compute besides just builds?
| And even builds themselves can require a lot of compute. I've
| only recently had a repo there where I wanted to do builds to
| make a release (both linux binaries and WASM) and whenever I do
| that tag and wait a few minutes for those builds to finish I
| think about all the other projects/repos out there on GitHub
| doing the same.
|
| I'm really kind of surprised they let us do that - like, why
| didn't they just have you upload the binaries after building on
| your local machine?
| judge2020 wrote:
| > why didn't they just have you upload the binaries after
| building on your local machine?
|
| You can do that already with GH Releases. Actions is if you
| want CI/CD managed by GitHub. And you can also use your own
| machines via self-hosted runners.
| JohnTHaller wrote:
| The last month GitHub hit four 9s of uptime was November 2024
| judge2020 wrote:
| Yeah, GH's availability has been a joke for many years, maybe
| even for a decade now.
| TacticalCoder wrote:
| > We're a good year+ into the use LLMs for all major bits of
| software that we all rely upon and GitHub here is down to one 9
| of uptime.
|
| With what to show for it? If GH did 10x in volume/git commits,
| it's all LLM sloppy-pasta. Where's the 10x productivity? Where
| are the amazing apps?
| crote wrote:
| They have delivered some _great_ shareholder value!
| nullbio wrote:
| It's caused by human laziness and corner cutting. LLMs can
| write buggy code or incomplete architectures as much as humans
| can, but standards have lowered. It's not the LLMs are not
| capable of also fixing these same issues, but that's additional
| work.
| muralimadhu wrote:
| Does anyone recommend an alternative? We've already moved our
| runners to blacksmith, looking for a control plane as well. We
| are tiny startup that ships at high velocity and this kind of
| downtime is highly disruptive for us
| stephenway wrote:
| since you're already on Blacksmith, how much of the remaining
| dependency is actions syntax vs GitHub itself?
| muralimadhu wrote:
| There's stuff like permissions, integration with PRs etc
| muralimadhu wrote:
| I dont mind going through the pain if there's a great
| alternative.
| mitchjj wrote:
| Let me know if you'd like help trying out Buildkite. Control
| plane with choice of hosting/hosted alongside lots of other
| flexible primitives.
| nnucera wrote:
| MiniStack is currently affected to keep up the releases :(
| blixt wrote:
| It would seem like GitHub is in a precarious situation.
|
| We have many agents per employee working in parallel pushing way
| more commits than was humanly possible before AI, triggering
| GitHub actions a lot more than the workflows were built for,
| causing Actions costs to escalate (they really aren't cheap if
| you compare to hosting it yourself), meanwhile working with YAML
| workflows is just a pain, and just writing code would be so much
| more fun and AI compatible[1].
|
| At the same time, GitHub has about ~3 different PR review UIs?
| And they're all half-bad? Any decently sized PR triggers their
| "optimized for large PRs" UI which jumps around randomly in my
| experience. If you don't get that UI and keep the scrolling one
| (there's an old and a new one btw) then god forbid you click a
| line number because at some point your browser will randomly
| scroll back to that line and it won't unstick. Now Linear[2] (and
| others) is replacing the PR review experience for the agentic
| era.
|
| I'd love to see a solid AI first Git + CI + reviews.
|
| [1] Cloudflare CI https://blog.cloudflare.com/ci-workflows/
|
| [2] Linear PR reviews
| https://linear.app/changelog/2025-01-23-pull-request-reviews
| qznc wrote:
| Sounds like https://entire.io/ ?
| iamleppert wrote:
| It's really crazy to me they can just be down for hours and can't
| recover their own systems. It shows they don't have the
| capability or infrastructure to roll back disastrous changes.
| These kinds of things are a tell on the organization and
| operational excellence (or not). As soon as a viable alternative
| surfaces for Github, I'm moving off and will advise all my
| clients to do so as well.
| rankam wrote:
| Is anyone seeing improvement - their last message implying a
| capacity problem would suggest some are actions are running?
| cebert wrote:
| I have sporadically had a few workflows complete for me, but
| most are having trouble finding a runner to start.
| alamsterdam wrote:
| nope, 6 hours, nothing. I guess they are also load shedding
| rankam wrote:
| Same - their status page doesn't give me hope that this will
| be resolved any time soon.
| shevy-java wrote:
| Ever since Microsoft took over, and then when AI Skynet took
| control, things started to decay. AI companies owe all of use a
| lot of money. By the way, why do I have to pay for increasing RAM
| prices here? Why are we so dependent on a few greedy
| corporations, anyway?
| efromvt wrote:
| Either they had a good month or I'm off my game, I actually spent
| a minute checking if I did something wrong with actions instead
| of immediately going to HN to see if it was an outage.
| macintux wrote:
| Seems like it's been solid lately, or at least not quite so
| obviously broken. I was thinking last week that I hadn't seen a
| GitHub outage on HN in quite a while.
| swyx wrote:
| i guess this is a good time to mention that smol forge is open
| for the first 100 alpha users!
|
| point clanker to forge.smol. ai/llms.txt
|
| for now its just a fast agent native git remote and u can check
| docs for the extras.
| molsson wrote:
| 5 hours going now and the entire thing is still completely down.
| This level of incompetence and their total disrespect for their
| customers is unbelievable.
| niwtsol wrote:
| Yeah, I think this one might be the straw that broke the
| camel's back and our team will look to migrate to something
| non-microsoft to be away from this nonsense.
| canadiantim wrote:
| What are the alternatives?
| Syntaf wrote:
| My company is migrating repositories over to a self-hosted
| forgejo instance, currently very jealous of the engineers
| working on those migrated repos.
| VCFundedGenYer wrote:
| Azure DevOps (Also MS owned but more stable than GitHub),
| GitLab, Codeberg, etc.
| PUSH_AX wrote:
| More stable because its vastly less popular.
| baby_souffle wrote:
| And if you thought the GitHub actions hybrid mix of
| JavaScript templates embedded in yaml was a nightmare, I
| have bad news for you with respect to ADO pipeline
| configuration documents...
| s_Hogg wrote:
| Yeah I went through hell migrating everything from ADO to
| GH Actions a year or so ago and even after all this I'm
| not even slightly close to wanting to go back
| tcfhgj wrote:
| I migrated to gitlab, then to codeberg
| nerdypepper wrote:
| tangled supports selfhosted CI runners using NixOS
| microvms: https://blog.tangled.org/spindle-microvm/
|
| selfhosted runners do not go down when our CI is down, its
| operates separately.
| chiply wrote:
| Time to rant... This is absolutely unreal.
|
| Even self-hosted runners are impacted.... How can that be?
|
| The cost of this globally has got to be in the hundreds of
| millions to companies that use CI/CD through GitHub Actions. What
| if prod is broken and GitHub actions is stalling the deployment
| of your hotfix? What if this makes your organization miss and SLA
| and diminish user trust? What if this makes you miss a release
| that you were contractually obligated to meet? This is happening
| during peak dev hours on a Thursday (not that it would be
| acceptable at any other time).
|
| I don't understand how a service this critical to the global
| technical infrastructure can fail like this at all, let alone for
| more than a few hours. Like where's the backup generator for
| crises like these? You can't even use self-hosted runners? WTF?
| Like how can you not bring your own backup in a crisis event like
| this?
|
| Not that Microsoft has a good reputation, but holy moly, you'd
| think they would prepare from something inevitable like this.
| nlightcho wrote:
| Counterparty risk is still a thing. This is the cost of
| convenience, reminds me of the milkman joke in that South Park
| episode. Maybe a global SPOF owned by people who do not care is
| not the way to go.
| chiply wrote:
| Dude that milkman joke made me cackle out loud.
|
| It's hard to draw a direct analogy there, but I feel like it
| echoes the same sentiment.
|
| A soapbox I have is that GHA workflows are scripts that could
| run on your machine without any of the YAML stuff. Who gives
| a flying about the DAG or the logs? Which, by the way, if
| you're willing to walk to the milk store to buy your milk,
| could be recreated in a much more testable and maintainable
| way without any of the YAML bs that GHA prescribes....
|
| But DAGs are pretty, and logstreams showing up in a browser
| application instill trust (for reasons that fly far above the
| head of yours truly). So people go for that. Pretty DAG, nice
| logstream; therefore, deliver my milk. All of a sudden....
| The CI/CD platform is having its merry way with your SLAs,
| contract abidements, and hotfix deployments.
|
| What a time to be alive.
|
| Not to rail on the South Park thing, but the blast radius of
| this issue also reminds me of the episode where the internet
| dried up.
|
| If this bs with GitHub continues, Parker/Stone will have to
| make a GitHub episode. How seen would we all feel if that
| happened?
| carlosneves wrote:
| That's exactly the situation I'm in... :crying-laughing:
|
| The fix is merged, but won't deploy... it's been hours
|
| Thankfully it's a batch job, and isn't interrupting production
| ATM
| chiply wrote:
| I feel for you.... This is not a position you should be put
| in.
|
| There's always the escape hatch of running you GHA workflows
| locally, but unfortunately, despite the existence of packages
| like `act`, there is no way to fully recreate the GHA runtime
| locally. Tons of the special YAML syntax just can't (more
| accurately, "just doesn't") get interpreted by those local
| actions runners.
|
| We never went this route, but at my old org, I always
| advocated for considering GHA to be wrapper around a single
| bash script (or whatever script you want to run), as a means
| of completely breaking out of the GHA hellscape that is
| programming in YAML, who's turing-completeness is pretty
| dubious.
|
| Unless you have things set up this way, you (the client of
| GitHub) would have to completely redesign your CI on the fly,
| run it locally, and then figure out how to get the D
| compliment of the I to work in a way that is auditable. Fat
| chance for most teams I bet.
|
| Thank god you're dealing with a batch scenario. Silver lining
| for sure. Still, embrace the anger.
|
| What makes my blood boil is that there's millions of DEVs
| literally crying at the moment worrying about how GitHub's
| failure to be responsible will put their jobs in jeopardy.
|
| And fingers crossed for you my friend. We're at 5+ hours at
| the time of this writing.... You're batch job may still have
| a chance!!!
| zenNeko wrote:
| I was scratching my head wondering why my build file isnt
| working, ran a working project and it didn't build either.
| Figured out what might be the case, but wasted an hour.
|
| I was complaining about how I should have used AI instead of
| manual brainwork for this instead but turns out that might have
| been the problem.
| noobcoder wrote:
| Just waiting for their Co pilot to fixes all bugs and
| misconfigurations, then they will have uptime above 100%
| jasonephraim wrote:
| The most likely explanation for the degraded reliability: Someone
| at GitHub saw the Three Body Problem last year and started
| shifting their infra over to Minecraft servers with "physical"
| memory and redstone logic gates.
| atlex2 wrote:
| Turns out those seams _were_ load bearing.
| jorl17 wrote:
| More than ~~five~~seven hours is fucking offensive to me. I
| cannot believe it.
| zer0x4d wrote:
| I run a self-hosted Gitlab instance. 100% uptime with my own
| runners.
| monkaiju wrote:
| Im shocked this is still going on... We moved our repos at work
| to a self-hosted Forgejo 2 months ago because the regular outages
| were getting ridiculous, looks like a pretty good call now.
|
| Honestly I dont think I've seen a tool I used regularly with such
| a moat lose it simply because they cant keep the service up. Also
| its hard to not see a pretty strong correlation between a bunch
| of these big companies doubling down on AI and just having their
| service go to s%#$. We've had similar issues with Digital Ocean
| recently, which is pretty lined up with them adding a bunch of
| "inference" services and rebranding the site to add all the
| "agent" marketing slop. Looking to move to Hertzner when the time
| allows.
| abidlabs wrote:
| It's not too hard to switch over from GitHub actions to other
| runners. For example, I wrote up the steps needed in order to use
| Hugging Face Jobs instead, which also enables GPU runners and
| other flavors. I use this for several of the repos I manage (such
| as Trackio):
|
| https://huggingface.co/blog/github-ci-hf-jobs
| timwis wrote:
| Is that not still down though because it still relies on GH
| actions to orchestrate?
| tom1337 wrote:
| It most likely is. We are using Ubicloud (which seems to work
| the same way hf jobs does) and it's also down. GitHub is not
| sending out Webhooks and is also not able to queue external
| jobs. If we had to emergency deploy anything now I'd have no
| idea how to really do that...
| yashap wrote:
| We use other runners (Blacksmith), and it hasn't helped, still
| hard down all day long. Brutal outage.
|
| As a side note, I started getting these symptoms (actions
| staying queued forever or not running at all) at 5pm PDT
| yesterday, intermittently. And definitely full outage by 8:30pm
| PDT. So for sure full outage for 7 hours and counting, even if
| you use other runners, and I strongly believe partial outage
| for ~15.5+ hrs before that, even if it hasn't been acknowledged
| by GitHub yet.
| nodesocket wrote:
| Also use Blacksmith for arm64 runners but this outage
| apparently affected webhooks so their runners are not picking
| up jobs either.
| richwater wrote:
| Gotta wonder what they're up to over there.
|
| Is the AI slop that bad? Culture change?
| Wafje wrote:
| Unprecedented load.
|
| From the GitHub COO on April 3rd: Platform
| activity is surging. There were 1 billion commits in 2025.
| Now, it's 275 million per week, on pace for 14 billion this
| year if growth remains linear (spoiler: it won't.)
| GitHub Actions has grown from 500M minutes/week in 2023 to 1B
| minutes/week in 2025, and now 2.1B minutes so far this
| week. So we're pushing incredibly hard on more
| CPUs, scaling services, and strengthening GitHub's core
| features.
| marricks wrote:
| I guess the AI Slop is that bad in terms of increasing
| commits/actions/etc
| mgh95 wrote:
| I find this hard to believe in this context. They should be
| utilizing load shedding or admission control and
| killing/rejecting jobs rather than hard failures if it is a
| scaling issue. It's much more likely an actual software
| defect than just more load. If this was the case (which they
| would likely prefer) free/public tiers would be removed first
| to preserve paying customers services.
| woodruffw wrote:
| Why not both? Higher base load combined with insufficient
| internal controls for ratelimiting/load-shedding (as in,
| they don't know _who_ to shed) would be explanatory.
| mgh95 wrote:
| If they can't implement something as simple as "decode
| upstream headers and determine if 429/503" I don't know
| what to say. Since this has knocked out all customers it
| indicates they likely don't have anything of this form
| implemented.
| woodruffw wrote:
| I meant shedding of legitimate base load, not retries. I
| think we can safely assume they do the latter.
| maccard wrote:
| They've said that, but it doesn't add up. The outages started
| shortly after the migration to Azure became the top priority,
| and before the load massively increased form agentic coding
| (by their own dates).
| javier123454321 wrote:
| This is unconscionable. It's an incredibly bad look to be up an
| entirety of a workday. When my company migrated from Bitbucket, I
| never thought that we'd be going to a worse service.
| marricks wrote:
| My own company had a migration from GitLab to supposedly
| greener pastures but at least when a self hosted gitlab server
| had an issue someone could figure out what was going on.
| Absolutely zero control over GitHub.
| slicendice wrote:
| Github is a terrible platform. We have had multiple enterprise
| tickets open with no response for multiple weeks after they
| fraudulently charged our company for services we cancelled weeks
| in advance. If you value any kind of real customer support or
| response time, I would avoid Github like the plague. They will
| steal your money and there is no one to even talk to about it.
| Their customer support is non existent (probably because their
| platform is in shambles).
| crmsystems wrote:
| actions been down for hours wtf going on.. AI bullshit
| g42gregory wrote:
| I may be wrong, but I think this could be traced to the
| Microsoft's increasing reliance on the outsourced engineering
| workforce. Regardless of the country it has been outsourced to,
| the quality of service will suffer, apparently greatly as we now
| see.
|
| It seems to manifest in many places: AI slop, Github outages,
| recently Miscoroft sent me an email demanding that I subscribe to
| 365, in order for MS Office (which I already paid for) to
| continue to work. I simply moved on to the other, free, provider.
|
| Perhaps, it's time to evaluate the use of the Microsoft products?
| msyea wrote:
| I think this is because all frontier AI providers integrate with
| GitHub as a default git forge through their GitHub apps platform
| and their free tier is too generous. They should either charge
| some of these huge partners access to GitHub apps or partition
| the free tier off and have a different SLA.
|
| If the frontier AI used GitLab as a default I'm sure they'd be
| the ones suffering now.
| mitchjj wrote:
| The frontier AI (OpenAI, Anthropic, Cursor, Mistral, xAI,
| Perplexity plus more) all run CI on Buildkite, without the
| suffering.
| samat wrote:
| I stopped using github for anything but git backups
|
| latest server i set up is simply a bare repo + hooks to make a
| local deploy after running tests and shit
|
| super easy to set up having ai do it, zero dependencies, deploy
| is still 'push it to the main'
|
| i have several remotes for backups and stuff
| __initbrian__ wrote:
| All of the outages make sense to me as scaling issues. Each
| month, GitHub is getting the amount of commits they'd normally
| get in a year. And it's growing
|
| > Yup, platform activity is surging. There were 1 billion commits
| in 2025. Now, it's 275 million per week, on pace for 14 billion
| this year if growth remains linear (spoiler: it won't.) GitHub
| Actions has grown from 500M minutes/week in 2023 to 1B
| minutes/week in 2025, and now 2.1B minutes so far this week. So
| we're pushing incredibly hard on more CPUs, scaling services, and
| strengthening GitHub's core features. And as a fine purveyor of
| hand-crafted shit code for many years, I'm not gonna weigh in on
| that.
|
| x.com/kdaigle/status/2040164759836778878
| arandomhuman wrote:
| Not really a reasonable excuse considering this is completely
| broken for self hosted runners / paying customers (for a whole
| day).
| dcrazy wrote:
| But you're not self-hosting the tasking service
| arandomhuman wrote:
| Because github does not offer the option to (and no,
| running github enterprise doesn't count).
| kami23 wrote:
| They don't even let you really run those yourself now
| either, We moved off a big GHE footprint because the
| support for it was getting abysmal and features were
| slowly become GH.com. Ive been dreading the move to
| GitHub.com and it was worse than I expected.
| abofh wrote:
| They couldn't handle sending webhooks for eight hours. That
| seems more than reasonable to expect.
| sharts wrote:
| That's no excuse. Particularly if it affects people self
| hosting their workers.
|
| These are folks that routinely make it a point to press on
| system design and scalability during interviews.
|
| Now they suddenly can't scale or design systems but we should
| accept that?
| shimman wrote:
| They don't hire well at all. You look at who is on their
| research team and it's people with social science PHDs, not
| computer science. Looking at the state of the site it's not a
| surprise. Highly incompetent people have taken over for a
| while. Let's not forget a member of the original executive
| team was a sex pest too.
|
| For like 10 years the only feature development they did was
| by stealing ideas from GitLab. I wouldn't be shocked if
| little if none engineering discipline has took place at all
| during this time if it's this brittle to frequent change.
|
| Guessing it was mostly held together with duct tape and
| poorly written tests/monitoring systems if any other
| corporate driven software.
|
| We're really going to find out over the next few years which
| businesses have good practices or not.
| HeWhoLurksLate wrote:
| Even if their actions running servers were at 50% load (or
| even 20%) at the end of 2025, they'd be _screwed_ right now,
| at this point it 's as much a tech/scaling problem as it is a
| CapEx problem and a construction one. Given the public's
| responses to AI specific datacenters you can't pick
| everything.
|
| Even my own systems at home are burgeoning under the load of
| the my more ambitious hobby projects I'm doing for fun. I
| don't envy "we had to build five more datacenters to keep up
| with demand" class problems just from a how many people have
| to sign off on them perspective let alone the technical
| difficulty of doing so
| habinero wrote:
| Are you kidding? That's an _insane_ amount of load increase
| to manage. "Scalability" isn't one thing, especially not at
| that level, so it's ridiculous to knock them for it. That
| amount of load ripples across your entire infrastructure.
| shimman wrote:
| They have the backing of a trillion dollar corporation and
| can literally hire some of the best talent on the planet.
| People nor budgets are an issue here.
|
| Why do we keep giving excuses to poor engineering
| disciplines + poor management? This problem is entirely
| GitHub's making and acting like it's some unplanned natural
| disaster is low key pathetic.
| mitchjj wrote:
| As a comparison, Buildkite is now running 1.5B job mins per
| week without this downtime.
| RamblingCTO wrote:
| This was posted elsewhere in the discussion:
| https://damrnelson.github.io/github-historical-uptime/ Went
| downhill long before LLMs
| lovetocode wrote:
| Claude is about to migrate me to Gitlab in about 5m
| adamddev1 wrote:
| I moved to a self-hosted Forgejo, with a self-hosted runner on
| another cheap Hetzner server. It's amazing how clean, snappy, and
| reliable it is.
| kilroy123 wrote:
| This is going to sound crazy, but I kind of wish they went back
| to _paid_ accounts. This would stop a lot of this.
| dbg31415 wrote:
| GitHub needs to change their slogan. "60% of the time, it works
| every time."
| 1vuio0pswjnm7 wrote:
| 49199094 | GitHub Is Experiencing Difficulties |
| https://www.githubstatus.com/?d=2026/08/06 |
| https://news.ycombinator.com/item?id=49199094
| shados wrote:
| This whole mess brought to like just now little moat there is
| left in the AI world.
|
| Like, everyone talks about how you can just prompt a new CRM
| instead of paying for Salesforce.
|
| But in this case, I raced with Fable and Sol and had time to
| build a full featured, fully functional CI/CD workflow system
| that has most of the features of GitHub actions (not the
| ecosystem obviously) and costs less than the GitHub runners per
| minute, while being able to scale to millions of workflows... And
| I had it working end to end before the outage was resolved and
| was already running my apps deploys through it.
|
| It would take a little longer to do in a company with red tape
| and more investments in the GitHub ecosystem, and of course it
| only replaced actions. But it works.
|
| There is no moat.
| digitalsushi wrote:
| i can make a snickers bar at home in 30 minutes right now.
| what's stopping me from taking over mars inc
| edot wrote:
| But if you just want a high-quality bar just for you, you
| don't want to take over Mars Inc.
| waterTanuki wrote:
| When the barrier to entry for software was high, free unlimited
| publicly hosted git repositories made sense. Not only did you
| have to have some basic understanding of software, but a grasp on
| git itself.
|
| Now that the calculus has changed, I think we should start asking
| ourselves whether anyone has the right to unlimited free public
| repositories.
|
| We can all agree that trying to attribute any single metric to
| repo quality is subject to Goodhart's Law[1] (e.g. what happened
| to GitHub stars) yet we can also agree that the Linux Project is
| a much more vital repository to keep public infrastructure
| availble for than say, someone's vibe-coded to-do app. Is it
| impossible then to make any quantitative distinction between
| these two? We can't have both an uncontrollable firehose of AI
| slop and unlimited free public storage/compute at the ready for
| it. I say we actually need to decide what software is worthy of
| using up these resources for free. We already have companies like
| JetBrains[2] using dynamic pricing for their products. It's time
| Github do the same. If you want a hosted VCS for vibe coded
| projects that aren't being used or depended upon widely, then
| host the repo yourself.
|
| [1]: https://en.wikipedia.org/wiki/Goodhart%27s_law
|
| [2]: https://www.jetbrains.com/community/opensource/
| qbzenker wrote:
| YC plz fund all GH 2.0 ideas this batch
| throw03172019 wrote:
| Degraded is an understatement. It's been totally down all day.
| Crazy.
| thewhitetulip wrote:
| Did they try using Copilot to fix it??
| dietr1ch wrote:
| I'm surprised this still passes as news, gives water is wet vibes
| at this point.
| switchbak wrote:
| It's 9:30pm pacific time - my build is still completely borked
| ... and we're self hosting runners.
|
| I think we need to ask for some money back, this is BONKERS.
|
| Any more of this, and we'll go full Gitlab, like we should have
| years ago!
| rurban wrote:
| This was the first time for me. All tests failed, for the whole
| day. For 20m, ok, I can tolerate that. But for a whole day?
| Microsoft sucks big time.
| cagz wrote:
| Outages happen. What concerned me most was the time taken for the
| full resolution (almost 12 hours). I am hoping that root cause
| analysis promised will look into this as well and how it can be
| improved.
| steelaz wrote:
| This incident made me look at the SLAs we have with GitHub, since
| our company is on the Enterprise plan. Interestingly, it seems
| that only failed GHA invocations are counted towards an SLA
| breach. We informed engineering to stop using GHA while the
| incident was ongoing, but it seems what we should have done was
| to hammer them without remorse to get the whole incident period
| (10+ hours) counted towards an SLA breach. What a scummy SLA
| design.
| Too wrote:
| Mods, I think there is something wrong with the front page
| algorithm. This post has been stuck there for more than a year
| now.
|
| Among with assorted thoughts on AI or new model point release
| announcements.
| codingconstable wrote:
| when arent they
| csomar wrote:
| I have a barely known/used GitHub application and losing it from
| time to time over some qq#%@ repo that blows up both my actions
| and CloudFlare workers. Over a few hundreds of repos that were
| installed, I have now seen two where the authors installed every
| AI system on the planet. This starts a vicious circle where one
| app comments on the PR, triggers all other installed apps, who
| also start commenting on the PR. This quickly reaches the limit
| of number of comments on the PR _BUT_ GitHub will keep pinging
| you about any event that is happening to that insane PR. Of
| course, your app itself will fail to comment so that triggers the
| redundancy I have including emails from my app about a failure
| (which I can 't read on the email itself for logs privacy, so I
| need to check my system logs). Just last week, I had a blow up
| with hundreds of failure emails. All of this was coming from a
| single dude PR which had nothing but AI apps fighting one
| another.
___________________________________________________________________
(page generated 2026-08-07 09:00 UTC)