← Back to Blog

Messaging SLA Credits: What SQS, SNS, EventBridge, Service Bus, and Pub/Sub Pay When They Fail

September 17, 2026

Your queue can lose a full day of processing and the SLA pays you nothing, because every messaging SLA in cloud measures the API, not the messages. SQS, SNS, EventBridge, Service Bus, and Pub/Sub all define downtime as failed API calls. If your producers can publish and your consumers can poll, the service is "available", no matter how many messages sit undelivered. This post has the exact commitments, credit tiers, claim deadlines, and the measurement rules that decide whether an outage counts.

A dark conveyor belt carrying glowing envelopes, the middle one stalled and glowing hot orange while the rest move on, an illustration of a message queue that stays "available" while messages stop

The measurement rule that decides everything

Every one of these SLAs rests on the same idea: a request that succeeds is an available interval, a request that fails with an error is not. Nothing else counts. Delayed delivery is not downtime. A consumer that stops pulling is not downtime. A redrive re-sending poison messages for hours is not downtime. Only your own failed API calls produce credits.

ServiceUptime commitmentCredit tiersWhat counts as downtimeClaim window
SQS / SNS99.9%10% / 25% / 100%API calls returning 500 or 503End of 2nd billing cycle
EventBridge99.99%10% / 25% / 100%API calls returning 500 or 503End of 2nd billing cycle
Service Bus (Standard)99.9%10% / 25%All attempts fail for a full minute60 days from incident
Service Bus (Premium, AZ regions)99.99%10% / 25%All attempts fail for a full minute60 days from incident
Event Grid99.99%10% / 25%Publish requests fail60 days from incident
Pub/Sub99.95%10% / 25% / 50%No publish succeeds for 60+ seconds30 days from eligibility
Pub/Sub Lite (regional / zonal)99.95% / 99.5%10% / 25% / 50%No publish succeeds for 60+ seconds30 days from eligibility

The pattern behind every row: downtime is failed API calls, and nothing else. A queue that accepts sends but never delivers to consumers is 100% "available" under the SLA. Credits only exist when your own requests return errors.

AWS: the empty-interval loophole

AWS bundles SQS and SNS into one agreement, the Amazon Messaging SLA, last updated May 2022. It promises 99.9% monthly uptime per region and pays 10% of the affected service's bill below 99.9%, 25% below 99.0%, and 100% below 95.0%. Availability is measured per 5-minute interval: the percentage of requests that do not fail with an error, where an error is a 500 or 503 response. SQS counts Send, Receive, and Delete invocations. SNS counts Publish calls.

The clause that quietly kills claims: an interval with no requests is assumed 100% available. If your producers were down too, or traffic was light, the outage barely registers. A regional SQS outage at 2am with no traffic is invisible to the SLA. You need requests hitting the API during the incident to have a claim at all.

The evidence requirement is equally specific. AWS wants request logs documenting the errors, broken into dates, times, and availability for every 5-minute interval below 100%. Most teams do not keep per-interval error logs for their queues, so most cannot file this claim even when the outage was real.

EventBridge: the 99.99% outlier

EventBridge is the only AWS messaging service with a four-nines commitment: 99.99%, about 4 minutes of allowed downtime per month. The credit schedule runs 10% below 99.99%, 25% below 99.0%, 100% below 95.0%, with the same 500/503 definition and no-traffic-means-available rule. The tighter commitment matters. A 30-minute EventBridge failure in a month with normal traffic puts almost any account below 99.99% and into the 10% tier. This is the claim most likely to succeed if you have the logs.

Azure: the all-attempts minute

Service Bus measures downtime in minutes, and a minute counts only if every continuous attempt to send, receive, or operate in that minute returns an error or fails to succeed within five minutes. One successful operation in the minute resets it. The September 2026 consolidated Azure SLA puts Standard at 99.9% and Premium in Availability Zone regions at 99.99%, both paying 10% below the commitment and 25% below 99%. Relays sit at 99.9%. Event Grid is separate: 99.99%, measured only on publish requests that fail to succeed within one minute, paying 10% and 25%.

The five-minute success timeout is the trap. Operations that hang for four minutes and then succeed keep the minute green. Slow is not down under this SLA, only fully failed is.

Google: publish-only, with a 60-second floor

Pub/Sub's SLA is the narrowest of the five. Downtime means no valid publish request succeeds and no streaming publish connection can be established. Subscriber-side failures never count, and any downtime period under 60 consecutive seconds is ignored entirely. The commitment is 99.95%, paying 10% below that, 25% below 99%, and 50% below 95%, with a hard cap: total credits for a month never exceed 50% of the Pub/Sub bill. Claims go to Google support within 30 days of eligibility.

Pub/Sub Lite is the weakest messaging promise of the group: 99.5% for zonal topics, about 3.6 hours of allowed downtime per month. It is the only service here where a bad day is part of the deal rather than a breach.

The money math: small credits, real risk

Credits are a percentage of the affected service's bill, and messaging bills are small. EventBridge charges $1.00 per million custom events, SQS gives you the first million requests free each month, and AWS does not pay credits under $1. The table below is the honest picture.

Monthly messaging spend10% credit25% credit100% credit
$5 (small app, ~5M EventBridge events)$0.50, not paid, under the $1 floor$1.25$5
$50 (busy pipeline)$5$12.50$50
$500 (fan-out heavy, multi-account)$50$125$500

Google caps at 50% of the bill, so the worst Pub/Sub month pays at most half of what you spent. Azure caps credits at the monthly fees. For most teams a successful claim is a rounding error, which is why these SLAs go almost entirely unclaimed. That is not the real exposure though. A stalled queue holding up order processing or payment events costs far more than the 10% tier, and no row in these tables compensates for that.

What to do this month

Three things separate a team that collects from a team that finds out about the SLA in the claim denial.

First, log the errors. Set 5-minute metrics for SQS, SNS, and EventBridge API error counts, keep request logs for 90 days, and do the same for Service Bus and Pub/Sub. That is the exact evidence every claim above demands, and it is worthless to collect after the incident.

Second, know the deadlines. AWS claims arrive by the end of the second billing cycle after the incident. Azure wants claims within 60 days. Google gives 30 days from eligibility. Put all three on a calendar the day you see the status-page incident, not the day you remember one happened.

Third, engineer what the SLA will not cover. Dead-letter queues with redrive, replay, and a cross-region fallback are your real availability story. The SLA is a refund on API failures. Your messages still need protection from everything else.

One action makes this real: open your metrics console and add a widget for the 5-minute error rate of each messaging service you run in production. Ten minutes today. When the next incident hits, you will have the numbers the claim requires, and the choice to file will be yours instead of being made by missing evidence.