← Back to Blog

Regional vs Zonal Failures: Which Outages You Can Actually Claim

August 29, 2026

The zone died and took your app with it, and the provider still owes you nothing. That outcome surprises people every time, because the outage was the provider's fault and the SLA mentions the very zone that failed. But cloud SLAs do not promise that infrastructure stays up. They promise that properly redundant deployments stay up, and the wording of each promise decides which outages convert into credits.

Two glass server towers on a dark reflective floor, one glowing cyan inside and the other dark, a thin electric arc between them, blurred amber rack lights behind

The commitment follows the deployment shape

Read the SLAs closely and the pattern is identical across AWS, Azure, and Google: the high uptime commitments only apply once you deploy with redundancy. A single instance, or instances stacked in one zone, sit under lower or no commitments.

ProviderDeployment shapeCommitmentCredit tiers
AWS EC2Instances in two or more Availability Zones (Compute SLA)99.99%10% below 99.99%, 25% below 99.0%, 100% below 95.0%
AWS EC2Single Instance (Instance Level SLA)99.5%10% below 99.5%, 25% below 99.0%, 100% below 95.0%
AWS RDSMulti-AZ DB Instance or Cluster99.95%10% below 99.95%, 25% below 99.0%, 100% below 95.0%
AWS RDSSingle-DB Instance99.5%Same tiers as single-instance EC2
AzureVMs across two or more Availability Zones99.99%10% below 99.99%, 25% below 99%, 100% below 95%
AzureVMs in an Availability Set or Dedicated Host Group99.95%10% below 99.95%, 25% below 99%, 100% below 95%
AzureSingle-instance VM (Premium SSD)99.9%10% below 99.9%, 25% below 99%, 100% below 95%
Google Compute EngineInstances in Multiple Zones (Premium Tier)99.99%10% at 99.0% to under 99.99%, 25% below 99.0%, 100% below 95.0%
Google Compute EngineSingle instance, most families99.9%10% below 99.9%, 25% below 95.0%, 100% below 90.0%

Google's single-instance tiers look unusual, and they are: a lone VM gets a real commitment and real credits, down to 100% below 90% uptime. AWS and Azure do not offer anything equivalent at the same bar; AWS's instance-level 99.5% is the closest cousin.

Why a zonal outage often fails the claim

The definitions do the work. Google's Compute Engine SLA counts Downtime "for virtual machine instances: loss of external connectivity or persistent disk access for the Single Instance or, with respect to Instances in Multiple Zones, all applicable running instances". All instances. If you run three VMs across three zones in one region and one zone burns, your other two keep serving, so the SLA's downtime clock never starts. No downtime period, no breach, no credit, even though a quarter of your fleet was dark for six hours.

Azure measures downtime across the same shape. For VMs in Availability Zones, Maximum Available Minutes begin when at least two VMs across two or more zones are running, and Downtime means minutes with no Virtual Machine Connectivity in the region. One dead zone with healthy VMs elsewhere does not register. The single-instance tiers behave differently because there is nothing else to fail over to: a zonal failure that kills your only VM is downtime for that VM, and if it runs long enough in the month, a claim exists.

AWS's Compute SLA follows the same logic for multi-AZ deployments, and the us-east-1 event of October 19-20, 2025 is the case study: DynamoDB API errors, NLB connection errors, and EC2 launch failures in one region. Customers whose workloads ran in that region and nowhere else had claimable downtime windows; customers whose fleets spanned regions often saw degraded performance but no regional SLA breach, because the commitment measured their whole regional deployment, not their worst hour.

The SLA measures your redundancy, not your pain. A zonal failure that your architecture absorbed is a non-event to the credit schedule, and a zonal failure that flattened a single-instance deployment can be worth more than the failure that made the news.

What actually earns a credit after a zonal or regional failure

Three configurations turn the same outage into three different outcomes.

Single instance, single zone. The provider's failure is your downtime. Check the instance-level commitment: 99.5% for AWS EC2 and RDS single instances, 99.9% for Azure single-instance VMs on Premium SSD, 99.9% for most Google single-instance families. Collect your logs, because these claims succeed on the provider's own incident records matching yours.

Multi-zone, same region, whole region down. This is the claim everyone pictures. If the regional failure kept your multi-zone deployment from reaching connectivity thresholds for long enough, the monthly uptime math crosses a tier and the credit schedule applies. Evidence matters more here than sympathy: AWS wants dates, times, and request logs per incident; Google wants log files showing the downtime periods.

Multi-zone, one zone down. Usually nothing, unless the remaining zones also failed to serve. The 10% credit tiers sit behind breaches measured across the whole month and the whole deployment, and a zone failure absorbed by design does not move that number.

The practical checklist

Before the next regional incident makes the decision for you, write down which deployment shapes you actually run, and read the matching SLA line for each. The mismatch between what teams assume ("the cloud is 99.99% up") and what their configuration entitles them to (a single-instance 99.5% commitment, or nothing at all outside instance level) is where unclaimed credits accumulate. The claim guide lists the filing route per provider, and the multi-region redundancy math covers when engineering around outages beats claiming them.

When a regional failure does cross your threshold, UptimeAudit flags the breach against your monitored services and regions and drafts the claim with the incident window attached. The monitoring, again, is free to imitate: an external check on each regional endpoint, capturing failures the provider's status page rounds away.