← Back to Blog

Cloud Outages at Night: An Alert Setup That Lets You Sleep

August 31, 2026

The October 19, 2025 us-east-1 outage started at 11:48 PM Pacific, which is 2:48 AM in New York and 7:48 AM in London. Whatever on-call arrangement you had that night decided how fast you knew, what evidence you captured, and whether the SLA claim got filed at all. Alert design for outages is mostly staging design: deciding in advance which signal wakes a human, which one waits for morning, and what gets recorded either way.

A smartphone glowing cyan on a dark bedside table at night, a ripple of alert light rising from its screen beside a warm alarm clock glow

Two alert classes, two different reactions

The mistake that ruins sleep is treating every provider signal as page-worthy. Most are not. Your monitoring should split what you hear about from the big four into two classes with different handling.

Signal classExamplesNight behaviorMorning behavior
Provider-side noiseStatus page updates, "investigating" posts, service health advisories that do not touch your regionSilence: log it, never pageDigest review with coffee
Your-stack impactExternal probe failing on your production endpoint, error rate crossing your own threshold, provider health event matching a region you deploy inPage the on-call humanPost-incident review and claim check

Provider-side noise deserves silence because it is usually not actionable at 3 AM. A status page that flips to "investigating" for a service you do not use, in a region you do not deploy in, is information, not an emergency. The October 2025 AWS event produced dozens of status updates across 141 affected services; an arrangement that paged on every one of them trained its humans to ignore pages within a week.

Your-stack impact deserves the pager because it is rare and it is exactly the case where minutes matter. The rule of thumb: page on what your users experience, log what your provider publishes.

Page on what your users experience, log what your provider publishes. Every alert that fires for a signal no one can act on at 3 AM spends credibility you will need for the one that matters.

What the 3 AM responder actually needs

When the page fires, the responder's job is containment, not documentation, so the evidence capture should already be running. Three setup pieces make that true.

First, an external probe per critical endpoint, from a different provider than the one it watches. The probe is your timestamp authority: it records UTC start and end of every failure window, which is precisely what AWS claim procedures ask for and what Google's SLA requires as log files showing downtime periods. Internal monitoring that shares the provider's network goes dark with the outage, which is how the June 2025 Google incident left the status page itself down for the first hour.

Second, log retention sized to outlive the claim window. AWS gives you until the end of the second billing cycle after the incident, Azure 60 days from the incident, Google 60 days (30 on several services). If your application logs rotate in 14 days and your error-rate dashboard holds 30, the claim will be built on memory. Thirty to ninety days of retention on the services that carry your bill is cheap insurance.

Third, a staging script or runbook that snapshots the moment: export probe history, capture the provider health event ID from Personal Health Dashboard or Service Health, and drop both in the incident folder. Two minutes of scripting beats reconstructing evidence at day 50 from chat logs.

The escalation ladder that preserves sleep

A workable ladder for a small team, tuned over years of on-call pain, looks like this.

StageTriggerWho is woken
0Provider status change, no impact on your endpointsNobody, logged
1First probe failure on a non-critical endpointOn-call phone notification, no siren
2Probe failing on a critical endpoint for 5+ consecutive minutesOn-call paged, acknowledged within 15 minutes or it escalates
3Critical endpoint down plus provider health event in your regionOn-call paged, secondary notified, evidence capture runs

Stage 3 is the shape the October 2025 event had: provider confirms a regional event while your probes fail. That combination is also the strongest predictor of a claimable SLA breach, which is why the evidence capture belongs in the same stage: the moment you know an outage is real and regional is the moment the claim clock, documentation included, starts.

After the wake-up

When the incident closes, three tasks, in order: write the internal timeline while it is fresh, file or stage the claim per the claim guide, and hold a 20-minute review asking the only question that improves next month: did we learn about this before or after our customers? If the answer is "after", the probe placement is wrong. The evidence checklist keeps the documentation half honest, and the deadline tracking options cover what to do when several claims are open at once.

UptimeAudit lives at stage 3: it watches the big four's health feeds, matches events to your monitored services and regions, and drafts the claim while the incident is fresh. The alert ladder above, though, is yours to build, and it costs nothing but an afternoon.