← Back to Blog

What Counts as a Covered Outage: The Exclusions Hidden in Cloud SLAs

August 23, 2026

When production falls over during a provider incident, most engineers assume a credit follows automatically. It doesn't. Every cloud SLA carries an exclusions section that strips entire categories of downtime out of coverage before the uptime math even starts: anything touching a beta or preview service, failures traced to your own configuration, outages caused by deploying into too few zones, slow-but-up degradation, scheduled patching. Denied claims usually lose on these clauses, not on arguments about whether the servers actually went down.

A brass magnifying glass hovering over an open contract dense with fine print, one clause glowing red under the lens

Here is what AWS, Azure, Google Cloud, and DigitalOcean exclude, so you can tell a payable outage from a dead end before spending an evening assembling evidence.

Four exclusions every provider shares

Providers decide claims by answering two questions. Did measured downtime cross a threshold, and does an exclusion apply? The first question is arithmetic. The second is where claims go to die, because the clauses are written broadly and applied literally.

Four families repeat across all four providers' contracts:

Exclusion familyAWSAzureGoogle CloudDigitalOcean
Events outside the provider's controlExcluded beyond the service's demarcation pointExcluded, including network failure between your site and their datacenterExcluded, factors outside reasonable controlExcluded, including third-party outages
Your actions, code, or configurationActions or inactions of you, down to ignoring resource health promptsFailure to follow required configurations or published guidanceErrors caused by your software or hardwareApplication code or configuration errors
Suspension, misuse, quotasSuspension or termination under the agreementUnauthorized action, exceeded quotas, throttled abusive accountsAbuses violating the agreement, system-applied quotasAccount restrictions, misuse, terms violations

None of this is sinister. A provider cannot refund an outage your own firewall caused. But read that middle row again: configuration errors are the escape hatch for almost any incident where their platform technically worked and your workload didn't.

The beta trap

The strangest denial category catches teams who did nothing wrong. AWS's Service Terms state that Service Level Agreements "do not apply to Beta Services or Beta Regions." Azure's SLA says previews and free tiers are "not included or eligible for SLA claims or credits." Google's Compute Engine SLA excludes features designated pre-general availability unless its documentation says otherwise.

Run production traffic through a preview endpoint, a trial tier, or a beta region and there is no uptime promise behind it at all, not merely a weaker one. The requests don't care what the console called the service, and neither will the support agent reading your claim.

An outage you can prove is not automatically an outage they will pay for. Assume the provider's first move on any claim is hunting for an exclusion, because it usually is.

Your deployment shape decides coverage too

Azure writes the bluntest version: its SLA excludes downtime that results from failures in a single Microsoft datacenter location when your connectivity depends on that location in a non-geo-resilient manner. Build on one datacenter and a failure there is contractually your problem.

AWS splits EC2 into two commitments. The headline 99.99% Region-Level SLA applies only when all running instances sit across two or more availability zones. A single-zone fleet falls back to the Instance-Level SLA at 99.5%, which pays smaller credit tiers and demands per-instance proof. Google prices its commitments the same way: instances spread over multiple zones get 99.99%, while other single instances get 99.9%. Its SLA also blocks double dipping, so a VM's downtime gets claimed either as a Single Instance or as Instances in Multiple Zones, never both.

DigitalOcean skips the architecture games: its Droplet SLA is instance-level at 99.99% per Droplet, so one badly behaved VM can breach it on its own. The exclusions still apply though, which is why code errors and config mistakes knock out so many small claims.

Azure adds two quieter carve-outs. Downtime from your own restart, stop, failover, or scale operations stays out of the uptime math, and so does the monthly patching window.

TrapProviderEffect on a claim
Fleet in one availability zoneAWSOnly the 99.5% instance-level commitment applies; 99.99% regional never triggers
Dependence on a single datacenterAzureDowntime excluded outright as non-geo-resilient
Restart, stop, failover, or scale operationsAzureDowntime removed from the uptime calculation
Monthly maintenance windowAzurePatching downtime removed from the uptime calculation
Claiming both single-instance and multi-zone credit for one VMGoogle CloudNot permitted; Google forces one classification
Errors triggered by hitting a quotaGoogle CloudExcluded from the SLA

Slow is not down

Every SLA defines downtime narrowly, and those definitions matter more than the marketing number on the product page. AWS counts a single instance as unavailable only when it has no external connectivity. Google counts loss of external connectivity or persistent disk access, and intermittent blips shorter than a minute don't stack toward a downtime period. Azure states it plainly: performance degradation or latency issues without actual service unavailability fall outside the SLA unless a service carries a performance-based commitment.

Response-time dashboards won't carry a claim on their own. What carries a claim is hard connection failure: timeouts, failed handshakes, 5xx storms with unreachable infrastructure behind them. If your monitoring records percentiles but not connectivity verdicts, fix that before the next incident rather than after.

One Azure sentence deserves taping above a desk. Outage communications exist to help customers take preventive actions and "are not a confirmation of missed Service Levels or Service Credits eligibility." A status page entry opens an investigation. It doesn't award the credit.

What survives every filter

Strip the clauses away and covered downtime looks identical at all four providers: loss of external connectivity, caused by the provider's own infrastructure, on a paid generally available service, running in a supported configuration, outside scheduled maintenance. Evidence should be shaped to match.

Five checks before filing:

  1. Is the affected component generally available? Preview, beta, trial, and free-tier pieces carry no promise.
  2. Does the deployment match the SLA tier being claimed? Two availability zones minimum for AWS's regional commitment, multiple zones for Google's top number.
  3. Did the downtime meet the definition? Connectivity lost beats latency degraded.
  4. Could the provider pin any of it on you? Recent config changes, quota limits, ignored health warnings, patch windows.
  5. Is the clock still running? AWS and Azure allow 60 days; every window is measured in weeks, so file early.

Ten minutes on that list beats a week of waiting for a rejection that was decided by clause three.

Do one thing today: open your production inventory and tag every component that runs on a preview, beta, or free-tier service. Anything tagged has no guarantee behind it at AWS, Azure, or Google Cloud alike. Move it to general availability or accept the risk knowingly, and the next outage becomes a recovery exercise instead of an argument you were always going to lose.