July 20, 2026
Your dashboards are red. Deploys are hanging. Support tickets are trickling in. Somewhere in the back of your mind a question forms: is this us, or is this them?

The next 60 minutes decide two things: how fast you recover, and whether you'll get the SLA credits you're owed afterward. Most teams handle the first and completely fumble the second. Here's the runbook for both.
The single most important takeaway: the outage is only half the hour. The credits are decided by what you capture while it is live and whether you file before the deadline. Most teams fumble exactly those two steps, and the provider will not chase you.
| Time window | What to do | Why it matters |
|---|---|---|
| Minutes 0-10 | Confirm it is really the provider: check the status page but treat it as lagging, corroborate externally, and scope the failure to service and region | Every later step, including the SLA claim, depends on getting the scope right |
| Minutes 10-30 | Protect the system: failover if you are multi-region, communicate if single-region, freeze deploys on control-plane issues | You recover faster and avoid making the incident worse |
| Minutes 30-45 | Capture the evidence: UTC timestamps, error screenshots, resource IDs, and the provider's incident ID | Claims are adjudicated weeks later and you must prove impact |
| Minutes 45-60 | Set the deadline trap: calendar reminder and ticket assigned to a named human | Claim windows run 30-60 days and are the most commonly missed step |
Before you failover anything or wake anyone up, establish where the failure lives.
With scope established, triage by blast radius:
This is the step everyone skips and everyone regrets. SLA credit claims are adjudicated weeks later, and the provider's support team will ask you to prove impact. Their status page acknowledgment is not, by itself, a claim.
While the incident is live, capture:
Five minutes of screenshotting during the fire saves an hour of archaeology afterward, and is frequently the difference between an approved claim and a stalled one.
The outage will end. The adrenaline will fade. And then the claim window starts quietly ticking:
| Provider | Claim deadline |
|---|---|
| AWS | within 60 days of the end of the billing cycle |
| Azure | within 60 days of the end of the month |
| Google Cloud | within 30 days of the end of the month, the shortest and most commonly missed |
| DigitalOcean | within 30 days of the end of the month |
Before you close the incident channel: create the calendar reminder, the Jira ticket, whatever your org actually looks at. Assign it to a named human. "We'll file it later" is how four-figure credits evaporate: the thresholds are low enough that even a sub-hour blip can qualify (at a 99.99% SLA, about 4½ minutes of downtime in a month is technically a breach).
Then, when things are calm, work out the credit tier. Our pillar guide has the per-provider thresholds and submission channels: How to Claim SLA Credits from AWS, Azure, GCP, and DigitalOcean.
The reason teams miss credits isn't laziness, it's that evidence capture and deadline tracking are boring chores bolted onto an already stressful hour. So we automated them. UptimeAudit watches the big four providers' status pages and health endpoints down to the service and region level, records the incident timeline as it happens, and when observed downtime crosses your provider's SLA threshold it drafts the credit claim for you, pre-filled with your account details, the evidence, and the deadline counted down in your dashboard. You review and submit; the provider decides the outcome, as always.
Next outage, spend your 60 minutes on your customers. Let the paperwork take care of itself.