1 week ago

GitHub CTO Apologises After Platform's Major 7.5-Hour Outage

GitHub CTO Apologises After Platform's Major 7.5-Hour Outage
GitHub CTO Apologises After Platform's Major 7.5-Hour Outage · thehansindia.com

GitHub stopped working for nearly eight hours on August 17.

Many tools, including the website, APIs, Actions, pull requests, issues, and Copilot, were affected.

GitHub said the outage was caused by unusually heavy traffic and infrastructure that did not scale properly.

An Istio component reached its limit, while the system watching it monitored the wrong thing.

A retry problem in Visual Studio Code sent much more traffic to one Copilot service.

Scraping attacks also made recovery harder.

Most services returned by 1636 UTC, but the Copilot Token Service recovered at 2102 UTC.

GitHub apologized and said it will improve how its systems handle traffic and retries.

Key facts

Outage duration
7 hours and 47 minutes
Peak web and API errors
Nearly 20%
Peak archive and raw-download errors
Close to 50%
Most services recovered
1636 UTC
Actions recovered
1803 UTC
Copilot Token Service recovered
2102 UTC
Monthly commits
Nearly doubled from 1.4 billion in April to 2.9 billion
Azure share of platform load
Approximately 58%, up from 12% in May

Quotes

GitHub

The company, in its outage explanation attributed to CTO Vlad Fedorov’s blog post

“If you were trying to ship software that day, we let you down”
thehansindia.com

Sources

Related news