September 28, 2026
The monitor is a dependency like any other, and it fails like one. When the observability platform goes down, you learn about outages late, you lose the evidence you would need for other claims, and the dashboards that would have told you all this are, for the moment, dark. All three major providers now carry explicit SLAs for the watching services, mostly at 99.9% to 99.95%, and their definitions of downtime are observability problems in themselves: alerts that arrive too late, metrics that fail to publish, log requests that come back as server errors. Read them before you need them, because a claim on the watcher requires evidence the watcher may no longer be able to produce.

| Service | Commitment | Unavailability is measured as |
|---|---|---|
| AWS CloudWatch (per Function: Metrics API, Logs ingestion, Alarms, Network Monitoring) | 99.9% per region, per Function | API requests failing with 500 or 503 errors per five-minute interval; alarm rules that fail to process; metrics that fail to publish |
| Azure Monitor Alerts | 99.9% | alert rules that fail to analyse telemetry, or that miss their scheduled start by more than five minutes |
| Azure Monitor Notification Delivery (action groups) | 99.9% | minutes where sending alerts or managing notification registrations fails |
| Azure Monitor managed Prometheus | 99.9% | minutes where retrieving metric data fails |
| Application Insights | 99.9% | minutes where no query operation succeeds |
| Google Cloud Observability (Monitoring, Logging, Trace APIs and web interface) | 99.95% | more than a 5% error rate on requests to retrieve metrics or logs |
Two details stand out immediately. Google carries the highest baseline in the set, 99.95%, and the lowest ceiling: its ladder stops at 50%. And the Azure surfaces are split finely: alerts, notification delivery, Prometheus queries and Application Insights queries each have their own clock, so a month can contain a broken alerting rule and working metrics at the same time, and only one of them is a claim.
Each measurement carries a design decision worth understanding.
AWS counts a Function per region, and its unavailability is request-level: the percentage of Metrics API calls or log ingestion calls that fail with server errors, averaged over five-minute intervals. Alarms are measured differently again, as the percentage of rules that fail to be processed. The practical effect is that a partial degradation, a steady trickle of failed calls, still accumulates into a claimable month, and CloudWatch is one of the few services here with a 100% tier on its ladder.
Google's Cloud Observability family defines downtime as more than a 5% error rate, with a ten-minute consecutive period before anything counts; alert notifications get fifteen minutes and a stricter condition, requiring at least two notification channels to be configured and all of them to fail before a missed alert counts as downtime. A single channel that quietly stops working is not covered, which is a good argument for configuring two.
Azure's five-minute grace is the subtle one: an alert rule is only unavailable if it does not succeed within five minutes of its scheduled start. A chronically late alert, arriving at four minutes fifty-nine, never becomes downtime at all, even though it is late every time.
The deeper pattern is that observability SLAs measure the pipeline, not the outcome. CloudWatch counts failed API calls, not whether your alarm fired. Google counts HTTP 5XX responses, not whether the pager went off. Azure counts the alert rule's analysis attempt, not whether a human saw the result. Between a metric landing in the platform and a person being woken up sit several more failure points, and none of them appear on any of these ladders. That is not a scandal, but it is a boundary: the SLAs promise the pipe, and the responsibility for the outcome stays with you.
| Provider | First tier | Second tier | Top tier |
|---|---|---|---|
| AWS CloudWatch | 10% (below 99.9%) | 25% (below 99%) | 100% (below 95%) |
| Azure Monitor surfaces | 10% (below 99.9%) | 25% (below 99%) | none |
| Google Cloud Observability | 10% (below 99.95%) | 25% (below 99%) | 50% cap (below 95%) |
On $2,000 a month of observability spend, a month at 98.5% pays $500 on the AWS ladder, $500 on Azure's, and $500 on Google's. Push below 95% and the AWS bill is refunded entirely, Google caps at $1,000, and Azure still stops at $500. The windows are the familiar ones: AWS wants the case by the end of the second billing cycle after the incident, Google wants notification within 30 days with log files, and Azure wants the claim within 60 days of the incident.
The watcher is a dependency, and its failure compounds: you lose the alert first and the evidence second, and the claim you file about the watcher needs exactly the logs the watcher failed to produce. Keep one synthetic check outside the platform you are monitoring, because that check is both your early warning and your claim evidence.
Three habits close the loop. Keep an out-of-band check, hosted somewhere other than the platform whose availability you care about; it is the only evidence source that survives the incident it documents. Configure the second notification channel, since Google's alert condition explicitly requires multiple channels to fail before a missed alert counts, and the same redundancy argument applies everywhere else. And treat the observability SLA as the first claim of any incident, not the last: it is small, defined in request-level terms, and easy to prove while the timestamps are fresh.
UptimeAudit lives on exactly this boundary, tracking the big four's health feeds and drafting the credits these documents actually pay.