Breakpoint

Inside GitHub's August 17 outage and retry storm

On August 17, 2026, GitHub was down for 7 hours 47 minutes: github.com, authentication, Actions, APIs, pull requests, issues and Copilot, with roughly 20%…

github··PT2M23S

video loads only when you press play

On August 17, 2026, GitHub was down for 7 hours 47 minutes: github.com, authentication, Actions, APIs, pull requests, issues and Copilot, with roughly 20%…

On August 17, 2026, GitHub was down for 7 hours 47 minutes: github.com, authentication, Actions, APIs, pull requests, issues and Copilot, with roughly 20% of web/API requests failing at peak and 50% of archive downloads.

  • A traffic peak exhausted an Istio sidecar limit that the autoscaler was not watching.
  • Gateway degradation and optimistic retries amplified a capacity problem into a broad outage.
  • Retry budgets, concurrency-aware scaling and stronger isolation are separate parts of the repair.

GitHub was down for 8 hours last week, and now they've said why. Pull requests, issues, Actions, the API and Copilot all broke, 1 in 5 requests failing at the peak, and it was the second big outage in August. Here's what the post-mortem says. Every request into GitHub takes the same path: load balancers at the front door, a gateway that checks who you are, then the service you asked for. Each service runs with a small helper that handles its network traffic, and an autoscaler adds copies of anything that gets busy. On the seventeenth, traffic hit a new peak, and one of those helpers hit its limit on open connections. The autoscaler never reacted, because its policy was watching the app's limits, not the helper's. The app looked healthy, so nothing scaled. That one failure spread to its neighbours, and the load rolled back up to the front door, where 4 of the load balancers ran out of room. The gateway that checks who you are sat behind them, so logins failed, and everything behind a login failed with it. And every failed request came straight back, because clients retry, and GitHub's own services retried too, so the busier those balancers got, the more traffic they were sent. Recovery came in stages. Pausing those 4 balancers at once brought most of the site back about 3 hours in. Copilot took longer. A slow reply from one endpoint tripped a latent retry bug in VS Code, and editors everywhere kept asking, sending the token service 10x its normal traffic. GitHub had to block those requests outright and let them back in one site at a time, and full recovery came just short of 8 hours. GitHub's explanation for the peak is growth. Since April, commits per month have doubled, and pull requests, new repos and Actions runs are on the same curve. As a result of this outage they added retry limits, retry budgets and variable timeouts on every call between services, autoscaling fixed to watch the helpers too, and a review of the quieter alerts for the next thing that would fail under a spike. And critical systems are being isolated from each other, so one failing can't take the rest with it. Along with that they're continuing their migration to Azure, which now carries more than half of the platform's resources, up from about a tenth in May. Ultimately, both outages were capacity failures, and the growth explains the pressure, but it does not excuse them. And if you were trying to ship software that day, they let you down.

This explainer is based on The August 17 outage and the work ahead by GitHub ↗. The original reporting and technical work belong to its publisher.