1 week ago

GitHub Blames Seven-Hour Outage on Critical Infrastructure Failure

GitHub Blames Seven-Hour Outage on Critical Infrastructure Failure
GitHub traces 7-hour outage to critical infrastructure failure: Here’s what we know · indianexpress.com

GitHub, a service that helps people store and manage computer code, had a major outage.

The outage lasted seven hours and 47 minutes on August 17.

A key part of GitHub’s infrastructure could not handle a new peak in traffic.

This created pressure that spread to other parts of the platform.

Users had trouble signing in and using tools such as GitHub Actions, pull requests, issues, APIs, and Copilot.

GitHub said the problem was not caused by a code or configuration change.

The company is adding computer capacity and changing how its systems connect to reduce the chance of another large outage.

It is also moving more of its platform to Microsoft Azure.

Key facts

Outage duration
Seven hours and 47 minutes
Primary failure
A critical infrastructure component did not scale when traffic reached a new peak
Affected services
github.com, authentication, GitHub Actions, APIs, pull requests, issues, and Copilot
Commit growth
Monthly commits increased from 1.4 billion to 2.9 billion since April
Recent incidents
GitHub reported two major outages in August, on August 6 and August 17
Azure platform load
Microsoft Azure serves roughly 58% of GitHub’s platform load and half of all Git operations
Added capacity
GitHub said it added more than 3 million CPU cores and 120 petabytes of high-speed storage

Quotes

GitHub spokesperson

GitHub representative explaining outage cause.

“Our next milestone is an architecture that scales read capacity linearly with the number of readers, enabling unlimited read operations.”
indianexpress.com
“We failed to scale critical components before demand exceeded their capacity.”
indianexpress.com

Sources

Related news