AWS, Google Cloud, and Microsoft Cloud Outages Show How Anything Can, and Will, Go Wrong
Shared by John Foley from Cloud Database Report · October 8, 2026
Read the original on clouddb.substack.com
View this post on the web at https://clouddb.substack.com/p/aws-google-cloud-and-microsoft-cloud
Welcome to the Cloud Database Report. I’m John Foley, a long-time tech journalist who also worked in strategic comms at Oracle, IBM, and MongoDB. A big welcome to new subscribers and thank you to paid subscribers, your support is appreciated. Connect with me on LinkedIn [ https://substack.com/redirect/dd0132cf-945f-490a-8aff-a0a97d29c388?j=eyJ1IjoiMTR3eXB3In0.f-xldB9G03IXstk7BAMCqX62DEM2cFHI5JhVsvW8cAA ]. Beware the blast radius Drone strikes, power disruption, water damage, human error, cable cuts, a faulty “blast radius” analysis system. What’s next, a swarm of grasshoppers? The technical issues and breakdowns that have long plagued IT systems and cloud services continue unabated. Months back, I recapped a long list of miscues involving Microsoft, Oracle, Google Cloud, Snowflake, and others. Here’s that earlier blog post of unfortunate events. Now, nine months later, system, network, and cloud availability seem to be as bad as ever, if not worse. The causes go beyond software misconfigurations or overheated hardware. Three examples: In March, an AWS facility in the United Arab Emirates got hit by a drone strike that caused structural damage, power disruption, and water damage. Earlier this month, a Google Cloud network outage was blamed on “inadvertent physical disconnection” of network cables. Last week, a telecom switch failure at an air traffic control facility, compounded by a cut cable, forced ground stops that lasted up to eight hours at major airports including JFK, LGA, EWR, PHL, and BOS. Such incidents are bad enough when they affect online and in-person services, travel, e-commerce, smart homes, and health services. But in a world of growing digital dependencies, more can go wrong. Blast radius is the term being used to describe the expanding impact of these interruptions and breakdowns. Here’s a closer look at some recent high-profile incidents. AWS: ‘Unable to restore access’ AWS recently admitted to what is every CIO’s worst nightmare: It is unable to provide access to the data or cloud resources that were hosted in one of its Middle East availability zones that was knocked out by a drone strike. The update from AWS comes a full six months after the attack, and there’s no indication that it will be able to recover what’s been lost. The issue dates back to early March, when AWS’s Health Dashboard signaled connectivity and power issues in availability zone ME-CENTRAL-1 Region (mec1-az2). That AZ is physically located in the United Arab Emirates and was/is in the line of fire from Iran and its proxies in the conflict with the U.S. In subsequent days, AWS revealed more information. Drone strikes had impacted two AWS facilities in the region, resulting in structural damage, power disruption, and water damage due to fire suppression. Amazon’s dashboard listed 135 services that were affected by the outage, including S3, DynamoDB, AWS Lambda, Amazon Kinesis, Amazon CloudWatch, and Amazon RDS. At the time, AWS said, “We are working to restore full service availability as quickly as possible.” Now, the world’s largest hyperscaler seems to be acknowledging that may not happen at all. Here’s the statement from its latest update: “After a thorough assessment, we have determined that we are unable to restore access to the resources and data hosted exclusively in the mec1-az2 Availability Zone. We continue to work on recovering regional resources, as well as zonal resources hosted in the other affected Availability Zones (mec1-az1 and mec1-az3).” You can view AWS’s detailed update here [ https://substack.com/redirect/ca4a0614-6443-414e-9c16-eae36fc24c40?j=eyJ1IjoiMTR3eXB3In0.f-xldB9G03IXstk7BAMCqX62DEM2cFHI5JhVsvW8cAA ]. The idea that data could be permanently lost runs contrary to the conventional wisdom of well-planned disaster recovery and IT resiliency. We understand that, with appropriate distribution and replication, data can and should survive even if one cloud availability zone goes kaput. In this case, attempts at failover may have been complicated by the fact that the impact extended across AZs. Three are mentioned in the AWS statement above. Reuters reports [ https://substack.com/redirect/005624ef-8143-4050-884c-a71f720bb096?j=eyJ1IjoiMTR3eXB3In0.f-xldB9G03IXstk7BAMCqX62DEM2cFHI5JhVsvW8cAA ] that the AWS facility most directly affected is located in Bahrain. How did AWS customers fare? In its Sept. 15 update, AWS states that “most” have been able to restore their systems: “Since the disruption began in March, most customers have been able to re-establish their operations in other Regions by restoring backups or copying data that remained accessible.” Yet, saying most customers have recovered isn’t the same as a saying all have done so. One can only hope that the others had backup and recovery processes in place. Google Cloud: ‘inadvertent disconnection’ On Sept. 1, Google Cloud suffered a four-hour service outage in its us-central1 region. Google Cloud blamed the incident on what it called “a procedural error.” In this case, the faux pas was that a technician accidentally unplugged cables that brought the system down. “The event was triggered by the inadvertent physical disconnection of network fiber-optic cables during routine hardware maintenance,” explained Google Cloud on its Service Health site [ https://substack.com/redirect/0dfd4589-9d06-42df-a300-4000aead4bc4?j=eyJ1IjoiMTR3eXB3In0.f-xldB9G03IXstk7BAMCqX62DEM2cFHI5JhVsvW8cAA ]. The glitch affected more than a dozen Google Cloud services, including some of its most essential data management services: Cloud SQL, AlloyDB for PostgreSQL, Bigtable, Cloud Filestore, Cloud Spanner, and Cloud Firestore. The impact extended beyond Google Cloud. Cockroach Labs indicated that its CockroachDB Cloud service, which is available via Google Cloud, was affected. “Google Cloud is experiencing networking outages which is affecting Cockroach Cloud customers,” Cockroach noted in an incident report [ https://substack.com/redirect/f025d246-9769-4b91-8084-346265ea7784?j=eyJ1IjoiMTR3eXB3In0.f-xldB9G03IXstk7BAMCqX62DEM2cFHI5JhVsvW8cAA ]. “Cockroach Cloud customers may experience heightened latency, delays, or connection issues.” By day’s end, Cockroach reported that its services had been restored. Shattered.io characterized [ https://substack.com/redirect/5f91b62d-0cff-43ce-8258-61abbf65f506?j=eyJ1IjoiMTR3eXB3In0.f-xldB9G03IXstk7BAMCqX62DEM2cFHI5JhVsvW8cAA ] Google Cloud’s Sept. 1 mishap as part of “a pattern, not an anomaly.” The site listed six Google Cloud outages over the past 18 months with causes including loss of power, human error, and a data center cooling failure. Microsoft Azure: ‘defect in blast radius analysis system’ The Google Cloud outage was followed two days later, on Sept. 3, by what some blamed on a failure [ https://substack.com/redirect/9ea7e316-613e-4805-99d6-3712fefc53c3?j=eyJ1IjoiMTR3eXB3In0.f-xldB9G03IXstk7BAMCqX62DEM2cFHI5JhVsvW8cAA ]in Microsoft Azure’s East US region. The incident, which lasted for five hours, had a ripple effect on popular AI services ChatGPT, Claude, Grok, and Copilot. In fact, that widely reporting “AI outage” may not have been Microsoft’s fault, or at least not entirely. It’s hard to know exactly what went wrong. Despite finger pointing by some that Azure was involved, SpaceXAI apologized [ https://substack.com/redirect/f86b82a4-21fa-4184-9674-995f7ab8466a?j=eyJ1IjoiMTR3eXB3In0.f-xldB9G03IXstk7BAMCqX62DEM2cFHI5JhVsvW8cAA ] to Grok users for an outage at one of its data centers. This may have been a case where various cloud infrastructure and LLM providers encountered performance issues at or about the same time. And there may have been interdependencies, making it harder to pinpoint the root cause. Regardless, the fact that AI services were experiencing issues simultaneously is wakeup call. InfoWorld dubbed it “the cloud outage that should terrify the CIO.” [ https://substack.com/redirect/ef2b89df-47d3-4aad-8c5e-4847cad6e306?j=eyJ1IjoiMTR3eXB3In0.f-xldB9G03IXstk7BAMCqX62DEM2cFHI5JhVsvW8cAA ] Even if Azure wasn’t the culprit on Sept. 3, I’m not letting Microsoft off the hook in this analysis. There have been other notable Azure disruptions in recent memory, including a five-hour outage [ https://substack.com/redirect/308923c9-31d1-4eae-9be8-6fe51060680e?j=eyJ1IjoiMTR3eXB3In0.f-xldB9G03IXstk7BAMCqX62DEM2cFHI5JhVsvW8cAA ] in July. In its post-incident review [ https://substack.com/redirect/3339c1e7-1ef8-4b5a-b875-83b8292a9c09?j=eyJ1IjoiMTR3eXB3In0.f-xldB9G03IXstk7BAMCqX62DEM2cFHI5JhVsvW8cAA ], Microsoft said that problem was triggered by its “blast radius analysis system,” which had a cascading effect: “A defect in our blast radius analysis system incorrectly expanded the scope of the repair event to include all optical devices egressing a specific datacenter.” It should be noted some of Microsoft’s core database services were impacted, including Azure Cosmos DB, Azure Database for PostgreSQL, and Azure Databricks. Not getting any better I’ve been writing system, network, database, and cloud availability for years in the Cloud Database Report. See the links below for more examples of the many points of vulnerability. Unfortunately, there’s no reason to believe that hyperscalers and infrastructure providers are getting better at curtailing these all-too-frequent breakdowns. Tech teams have little choice but to continue prioritizing investment and planning for system and service availability, redundancy, and resilience.