1 week ago
GitHub CTO Apologises After Platform's Major 7.5-Hour Outage
GitHub stopped working for nearly eight hours on August 17.
Many tools, including the website, APIs, Actions, pull requests, issues, and Copilot, were affected.
GitHub said the outage was caused by unusually heavy traffic and infrastructure that did not scale properly.
An Istio component reached its limit, while the system watching it monitored the wrong thing.
A retry problem in Visual Studio Code sent much more traffic to one Copilot service.
Scraping attacks also made recovery harder.
Most services returned by 1636 UTC, but the Copilot Token Service recovered at 2102 UTC.
GitHub apologized and said it will improve how its systems handle traffic and retries.
GitHub was disrupted for 7 hours and 47 minutes on August 17.
An Istio sidecar reached its concurrency limit while autoscaling monitored the wrong service.
A Visual Studio Code retry bug amplified traffic to the Copilot Token Service.
Scraping attacks further complicated recovery across archive and raw repository downloads.
GitHub plans retry controls and architecture changes as platform usage continues growing.
- Who
- GitHub and its CTO, Vlad Fedorov, addressed the incident; GitHub users and services were affected.
- What
- A major platform outage lasted 7 hours and 47 minutes, disrupting GitHub.com, Actions, APIs, pull requests, issues, and Copilot.
- Where
- The failure began in infrastructure in GitHub’s Central US data centre and affected the GitHub platform.
- When
- The outage occurred on August 17; GitHub published its explanation on August 20.
- Why
- Unprecedented traffic overwhelmed infrastructure after an Istio sidecar reached its concurrency limit; incorrect autoscaling, amplified retries, and scraping attacks worsened the disruption.
GitHub’s Explanation
Reliability Concerns
What the outage means
GitHub’s Explanation
GitHub attributed the incident to unprecedented traffic, an infrastructure scaling failure, a Visual Studio Code retry bug, and scraping attacks, and apologized to users.
Reliability Concerns
CloudBees CEO Moritz Plassnig said GitHub may not remain the default choice, predicting a more fragmented ecosystem after the incident.
How to respond
GitHub’s Explanation
GitHub plans to standardize retry limits, add retry budgets and variable timeouts, revise alerts, and redesign read-capacity scaling for large monorepos.
Reliability Concerns
The outage, which followed another major Actions incident on August 6, raised concerns about whether GitHub can reliably support its continuing growth.
Key facts
- Outage duration
- 7 hours and 47 minutes
- Peak web and API errors
- Nearly 20%
- Peak archive and raw-download errors
- Close to 50%
- Most services recovered
- 1636 UTC
- Actions recovered
- 1803 UTC
- Copilot Token Service recovered
- 2102 UTC
- Monthly commits
- Nearly doubled from 1.4 billion in April to 2.9 billion
- Azure share of platform load
- Approximately 58%, up from 12% in May
Quotes
GitHub
The company, in its outage explanation attributed to CTO Vlad Fedorov’s blog post
“If you were trying to ship software that day, we let you down”
thehansindia.com








