August 24, 2026
A pricing page tells you what a cloud region costs. An SLA page tells you what the provider pays when it fails. Neither tells you how often it actually fails, and that is the number a buyer wants. The failure record is public. For three of the big four providers it goes back five years or more, and almost nobody reads it before committing workloads. You can open the archive, count every major incident the provider has admitted to in your target region, time them, and check whether the region you are about to choose is a repeat offender. This post covers where each provider keeps its history, what the logs quietly leave out, and the counts worth taking before you sign.

The live dashboard gets all the attention during an outage. The archive is the part worth studying before one happens.
| Provider | Where the history lives | How far back | What gets published |
|---|---|---|---|
| AWS | Service History on the AWS Health Dashboard, plus the Post-Event Summaries page | 12 months on the dashboard; summaries kept a minimum of 5 years | Dashboard rows for service events; written summaries reserved for broad, severe events |
| Azure | Azure status history page | Post-incident reviews retained 5 years | Reviews for publicly communicated, broad incidents only |
| Google Cloud | status.cloud.google.com, View incident history and per product See more | 1 year on the main view, up to 5 years per product | Broad Severe incidents, with public analysis reports |
| DigitalOcean | status.digitalocean.com/history | Full backfill; filters only work for incidents after May 10, 2023 | Platform level incidents and scheduled maintenance |
Google's archive is the easiest to analyze. The whole incident stream downloads as a JSON file from status.cloud.google.com/incidents.json, with a published schema, stable product IDs, and structured affected locations attached to every update, so you can compute over years of history without scraping anything. AWS's Post-Event Summaries page is smaller but deeper: eighteen writeups spanning April 2011 to October 2025, each covering scope of impact, contributing factors, and the fixes that followed. Azure sits in between, one review per broad incident, filterable by region.
Every provider decides for itself what qualifies for publication, and every one publishes only the big ones. AWS commits to a public summary when an event causes failure of a significant percentage of control plane API calls, hits a significant percentage of a service's infrastructure, or stems from total power or significant network failure. Google's public dashboard carries Broad Severe incidents only, defined as global impact or trouble hitting a significant percentage of customer projects across multiple regions; anything narrower lives in Personalized Service Health inside the console, if it surfaces to you at all. Microsoft's history page holds reviews for broad Scenario 1 events, while smaller problems go out as targeted notifications to exactly the subscriptions that were hit. DigitalOcean publishes platform level events, so a dying disk on your Droplet will never appear anywhere.
Absence from the log therefore proves very little. It mostly means the blast radius stayed small. Once you are a customer, the account-level views fill part of the gap: AWS Health API retains events for about 90 days, Azure's Service Health keeps a Health history tab, and Google's Personalized Service Health lists issues touching your projects. Before you are a customer, independent trackers are the closest thing to a census. StatusGator alone has been recording status changes across thousands of services since 2015.
The incident log documents a provider's worst days, not its average day. Read it to compare regions, services and root causes between providers. Never read it as a prediction of the uptime you will personally see, because only your own measurements answer that question.
Pull the last 12 months for the services and region you plan to use, then take these counts.
That last one deserves an example, because the data is sitting in plain sight. Of the eighteen post-event summaries AWS has published since 2011, eleven name Northern Virginia or US-East, including DynamoDB in October 2025, Kinesis in November 2020 and again in July 2024, Lambda in June 2023, and the December 2021 event that took out many services at once. Us-east-1 is the region most customers default to and the one where AWS's published failures keep landing. These summaries only cover the largest events, which cuts both ways: the sample is biased, but it is biased toward exactly the disasters you need to plan for. If you intend to run single-region in us-east-1 anyway, you now have a documented reason to price that risk rather than discover it.
Azure's review of its May 29, 2026 cooling failure rewards a close read. Utility power sag and swells hit multiple datacenters, mechanical cooling dropped into a protective lockout, and thermal alerts fired within ten minutes. Infrastructure shut itself down to prevent heat damage. Cooling was restored in about ninety minutes, half the affected virtual machines were back within two hours, ninety-five percent within eight, and the storage tier dragged on for most of a day before Log Analytics and Application Insights cleared the following morning at 02:30 UTC.
Four things to extract from any review you read: the true cause, not the headline; the blast radius in regions and datacenters; the recovery curve, which here split into fast compute and slow storage; and whether the prevention list names specific engineering work or just promises better process. Comparing reviews across providers on those four axes tells you more than any benchmark slide. A vague review is a fact worth weighing on its own.
The log gives you probability. The SLA gives you payout. Together they give you expected value: if a region produced several claim-grade breaches in a year and the credit schedule pays 10% to 100% of the affected service's bill, you can put a rough dollar figure on the risk you are accepting, and decide whether multi-zone deployment buys more than it costs. Remember too that credits are claim-only with windows as short as 30 days, so the same log tells you which billing months deserve a re-check after you sign. Thresholds and deadlines for all four providers are in the claim guide.
Open the provider's incident archive today, filter to your target region and the three services you would bet the business on, and read the trailing 12 months. Count the incidents, note the durations, and open every review whose outage window would have crossed your SLA threshold. If the count embarrasses the uptime figure on the sales page, look at the second region instead. Let that number, not the logo, pick where your workloads land.