← Back to Blog

Error Budgets and SLA Credits: Running Your Own SLO Next to the Provider Promise

September 5, 2026

Your provider promises 99.99%. Your users need 99.9% of the whole experience to work, which is a different sentence with different math. Error budgets are how engineering teams turn the second sentence into a number they can act on, and provider SLA credits are one of the inputs that number produces. Teams that conflate the two promises over-engineer in the wrong places and miss money they are owed, usually in the same quarter.

A wooden hourglass on a dark desk, its upper bulb filled with glowing golden coins instead of sand, a few coins scattered beside it, blurred server rack lights behind

Two promises, measured differently

The provider SLA measures a slice of the world: connectivity to a specific service in a specific region, error rates on specific request types, computed from their telemetry. Your SLO measures what your users experience: successful requests served by your whole stack, computed from your telemetry. The slice and the whole fail differently. Your stack can breach its SLO while every provider SLA holds (your own deploy breaks checkout), and a provider can breach its SLA while your SLO holds (a zone fails and your multi-zone deployment absorbs it).

QuestionProvider SLA answersYour SLO answers
What is measuredOne service, one region, provider's definition of DowntimeYour whole request path, your definition of success
Who measuresThe provider, from their telemetryYou, from your probes and application metrics
What it producesA credit percentage of your bill for the affected serviceA burn rate that decides feature work vs reliability work
What a breach meansMoney, if you claim in the windowUsers failed; fix the stack

Google's Site Reliability Engineering book, which introduced the error budget into mainstream practice, frames it simply: the SLO is the target, and the budget is the amount of unreliability you are allowed to spend while meeting it. At a 99.9% SLO, the budget is 0.1% of requests or minutes per window, which over a 30-day month is about 43 minutes. Spend it on launches, experiments, or the month's outages; when it is gone, reliability work takes priority until it refills.

The provider's promise is an input to your budget, not a substitute for it. A 99.99% SLA on one service says nothing about the 12 services your request path touches.

How the two numbers interact

The useful arrangement puts provider outages inside your budget as a category, not outside it. Concretely: your monthly error budget absorbs everything, provider-caused or self-caused. When a provider incident burns budget, two ledgers update at once: the budget consumption (reliability planning) and the SLA claim (recovery). The first decides whether you need to redesign; the second decides whether you get money back. They are connected but not interchangeable, and the difference shows up in real incidents.

The October 19-20, 2025 AWS us-east-1 event is the worked case. Teams with multi-zone deployments and healthy SLOs burned budget on degraded latency while never breaching a provider SLA, and correctly filed no claim. Teams whose single-region stacks went dark both burned budget and crossed provider SLA thresholds, and had a claim to file with evidence their own probe logs supplied. The budget ledger told both teams the truth about their users; only the second had money to recover.

Building the budget, step by step

Pick the SLO from user impact, not provider marketing. Start from what downtime costs you per hour, the calculation in the cost of downtime piece, and the availability your users actually need. A 99.9% SLO is a defensible default for internal tools and a bad one for checkout.

Measure from your side. External probes on the critical endpoints, per the monitoring stack, plus application metrics on success rates. The provider's dashboard cannot supply these numbers.

Track burn rate, not just totals. A budget that will exhaust in four hours at the current rate needs attention now, not at month end. Alert on burn rate; log the rest.

File claims in parallel. When budget burn traces to a provider incident, the claim is a separate workstream with its own deadline: end of the second billing cycle for AWS, 60 days for Azure, 60 days for Google (30 on several services), two billing cycles for DigitalOcean Droplets. The claim guide covers the routes.

Where credits fit in the money math

Credits do not make your budget whole; they discount the month's bill for the affected service. The schedules pay 10% at the first tier for most AWS, Azure, and Google services, 25% and 100% above it, and DigitalOcean pays 100% of a breached Droplet's charge outright. As a fraction of what a serious outage costs a real business in revenue and churn, the credit is small. Its value is in the incentives: filing claims makes outage costs visible in finance, which makes the reliability conversation concrete, and tracking claims month over month gives you a provider-side outage history no status page archive supplies.

UptimeAudit sits on the claim side of this arrangement: monitoring the big four's health feeds, detecting when your monitored services and regions cross a published SLA threshold, drafting the claim, and counting the window down. Your SLO tooling and your budget reviews stay yours. The two systems answer different questions, and running both is cheaper than either alone.