Back to System Design Index

Cloud InfrastructureSeptember 20266 min read

GCP us-central1 Outage: Fiber Isolation Across Two Zones

On September 1, 2026, a network-maintenance error isolated compute capacity in specific clusters in us-central1-b and us-central1-f. Google's incident report attributes the failure to incompatible optical transceivers plus a work sequence that omitted required verification. The reported incident window was 4 hours 11 minutes.

TL;DR: The primary failure was physical network isolation during maintenance, not a demonstrated regional control-plane queue. This reconstruction separates Google's published incident facts from the resilience analysis that follows.

Zones with affected clusters
us-central1-b and -f
Reported incident window
4h 11m
All paths disconnected
Within 13m
Published root cause
Maintenance error

By Mukul Kumar Mishra · Research-led architecture teardown · Published September 30, 2026 · Updated October 1, 2026

Summary of Google's GCP outage report: two affected zones, a 4 hour 11 minute incident window and fiber paths disconnected during maintenance
Figure 1. Google's report records fiber isolation in parts of two zones, with an incident window of 4 hours 11 minutes.

1. What Google's report establishes

On September 1, a scheduled capacity upgrade replaced optical transceivers on network routers supporting parts of us-central1-b and us-central1-f. The replacement transceivers were incompatible with equipment remaining in the network fabric. A manually orchestrated work list omitted the required one-router-at-a-time sequence plus the normal human and software verification steps. Fiber paths were disconnected across the affected devices within 13 minutes.

The loss isolated compute capacity from the network. Virtual machines in the affected clusters could not reach external resources, nor could outside clients reach those machines. This was a physical connectivity failure. Google's report does not attribute the event to a slow disk, a regional control-plane queue or a measured retry storm.

Failure boundary: physical work crossed device boundaries faster than safeguards could catch it. The design question is how maintenance sequencing, independent verification and traffic diversion should constrain that boundary.

2. Scope: affected clusters, not a whole region

The incident affected specific clusters in parts of two zones. Google says many multi-zonal deployments using capacity in unaffected zones largely continued serving. Regional products that depended on workloads in the affected clusters still saw symptoms, so a zone label alone did not describe every customer's blast radius.

The September 29 addendum adds important detail: traffic diversion reduced the impact to regional products, while Cloud Run and App Engine saw a later thundering herd as workloads initialized on replacement capacity. GKE control-plane requests also experienced errors and delays. Those are reported outcomes. They do not establish the invented per-minute queue and retry counts previously shown here.

Sequence in Google's incident report: incompatible transceiver replacement, fiber-path loss, then network isolation in affected clusters
Figure 2. The reported chain from incompatible maintenance hardware to network isolation. It does not depict an unreported retry multiplier.

3. Timeline from the primary report

  • 07:41 US/Pacific: Google's report marks the start of customer impact.
  • 07:45 and 08:00: traffic diversion actions began to reduce regional impact.
  • 08:50: most physical connections had been restored and traffic began recovering.
  • 09:19: diversion actions were removed as network capacity recovered.
  • 11:52: Google confirmed recovery across nearly all affected services and long-running operations.

The 4-hour-11-minute figure is Google's incident window, not a claim that every listed product was continuously unavailable for that entire period. The report describes different impact windows by product and workload.

4. Resilience lessons supported by the evidence

Sequence hazardous work in the tooling. The report says the procedure intended one router at a time, but the work list did not carry that instruction. Encode ordering constraints in the workflow rather than relying on an operator to infer them from a long list.

Require independent confirmation. The final report describes missing human and software checks plus a standing stop-work procedure that was not followed. A safe workflow should pause when a live signal remains on a cable after disconnection, require a second confirmation and make the next step unavailable until checks pass.

Automate traffic diversion with a bounded blast radius. Google says diversion took about 19 minutes in this event and describes a target near five minutes for its automated system. That is Google's stated remediation target, not a measured result. Applications should still tolerate partial network isolation with explicit timeouts, bounded retries and gradual traffic shifts.

Test recovery behavior, not only replica placement. Rehearse how services discover healthy capacity, how retries behave during packet loss and how cold caches respond when traffic shifts. The report documents a thundering herd after restoration. It does not publish request-rate measurements, so any load test should use your own service telemetry.

5. The verdict: make maintenance stoppable

This incident was not evidence that multi-zone design failed everywhere. Google's report says unaffected zones allowed many multi-zonal workloads to continue. It was evidence that routine maintenance can defeat physical redundancy when a procedure permits multiple paths to be removed before verification.

For an infrastructure review, ask for the ordered maintenance plan, the independent signal that blocks the next operation, the stop-work threshold, the traffic-diversion objective and a tested recovery path for workloads that move to cold capacity. Measure each against your own environment. Do not turn a provider's incident into a cost estimate without workload and rate-card data.

For comparison, see the distinct storage event in GCP us-west1. Its failure mechanism and reported scope differ from this network-maintenance incident.

Frequently Asked Questions

What caused the September 2026 GCP us-central1 outage?

Google's incident report says incompatible optical transceivers were installed during network maintenance. Missing sequencing and verification safeguards allowed fiber paths to be disconnected in parts of us-central1-b and us-central1-f.

How long did the Google Cloud incident last?

Google reports an incident window from 07:41 to 11:52 US/Pacific on September 1, 2026, a duration of 4 hours 11 minutes.

Did the outage affect an entire region?

No. Google describes network isolation in specific clusters in parts of two zones. Many multi-zonal deployments continued through unaffected capacity, though dependent regional products also saw symptoms.

What engineering lessons does Google's report support?

Use explicit maintenance sequencing, stop-work checks, automated traffic diversion and rehearsed application behavior during network isolation. Treat stated remediation targets as targets until measured.

Sources and Method

Incident facts and the timeline come from Google's incident report and its September 29 addendum. Recommendations are engineering analysis based on those reported failure mechanisms. This revision removes previously published queue, retry and cost figures that had no primary-source support.