Workloads depending on compute resources in this cage became degraded due to packet loss, and exhibited intermittent errors:
- Actions saw 10% of jobs fail during the impact window, and 5% of jobs succeeded but with delayed starts.
- 27% of GitHub issues interactions saw slow requests or timeouts.
- 4% of GitHub Copilot requests experienced errors, though most automatically retry.
- 4% of git push operations saw impacts during the affected window.
- Authentication requests saw increased latency during the affected window, but error rates, while elevated, were < 1% in all cases.
We were able to mitigate the outage by re-routing affected connections to available fiber paths that were allocated for future capacity upgrades. Sufficient network capacity to eliminate packet loss was restored at 17:07, with most services showing full recovery by 17:16. All paths were restored and services healthy at 17:36.
This incident affected 25% of available network interconnect capacity. Older cages utilize a 100Gbps network interface standard. To remove risk of reoccurrence, a planned upgrade to 400Gbps interfaces is being accelerated as much as possible, ensuring increased bandwidth available at all layers of the switch fabric for resiliency to path or device loss.
Jul 24, 17:36 UTC
Update -
We are seeing recovery across all services
Jul 24, 17:24 UTC
Update -
The degradation affecting API Requests, Actions, Copilot, Issues, Pages and Pull Requests has been mitigated. We are monitoring to ensure stability.
Jul 24, 17:16 UTC
Update -
Actions is experiencing degraded performance. We are continuing to investigate.
Jul 24, 16:41 UTC
Update -
We have applied a mitigation and are monitoring for recovery
Jul 24, 16:40 UTC
Update -
Actions is experiencing degraded availability. We are continuing to investigate.
Jul 24, 16:28 UTC
Update -
Pages is experiencing degraded performance. We are continuing to investigate.
Jul 24, 16:27 UTC
Update -
Copilot is experiencing degraded performance. We are continuing to investigate.
Jul 24, 16:26 UTC
Update -
We are investigating timeouts to some GitHub services
Jul 24, 16:22 UTC
Update -
Pull Requests is experiencing degraded performance. We are continuing to investigate.
Jul 24, 16:20 UTC
Update -
Actions is experiencing degraded performance. We are continuing to investigate.
Jul 24, 16:19 UTC
Investigating -
We are investigating reports of degraded performance for API Requests and Issues
Jul 24, 16:17 UTC