← Back to Blog

Streaming Data SLAs: What Kinesis, MSK, and Event Hubs Pay When the Stream Breaks

October 7, 2026

Your streaming pipeline has an SLA, and the provider will pay you when it fails, but the money is tied to a very narrow definition of failure. Kinesis Data Streams and MSK count only failed requests in five-minute windows, Event Hubs counts only whole minutes where every attempt fails, and MSK covers only Multi-AZ deployments. Miss those details and you either file a claim that gets denied or skip one that would have paid.

A long steel pipe system carrying glowing blue streams of data, one section burst at a joint with light leaking out, dark industrial setting

The comparison table first, then the traps that decide whether your claim survives.

ServiceCommitmentCredit tiersMeasurement windowClaim deadline
Kinesis Data Streams + Video Streams99.9% per region10% below 99.9%, 25% below 99.0%, 100% below 95.0%Per 5-minute interval, request error rateEnd of second billing cycle
Amazon MSK (Multi-AZ only)99.9% per region10% below 99.9%, 25% below 99.0%, 100% below 95.0%Per 5-minute interval, API + Kafka error codesEnd of second billing cycle
Event Hubs Basic/Standard99.95%10% below 99.95%, 25% below 99.0%Whole minutes where all attempts fail60 days from the incident
Event Hubs Premium/Dedicated99.99%10% below 99.99%, 25% below 99.0%Whole minutes where all attempts fail60 days from the incident

Kinesis: 99.9%, and the empty-interval rule

The Amazon Kinesis SLA (last updated February 9, 2024) covers Kinesis Data Streams and Kinesis Video Streams. The commitment is 99.9% in each region per billing cycle, and the tiers are the familiar AWS ladder: 10% below 99.9%, 25% below 99.0%, 100% below 95.0%, applied to what you paid for the service in the affected region that month.

Two details shape every Kinesis claim:

  • Availability is "the percentage of Requests processed that do not fail with Errors" in each 5-minute interval, where an Error is a 500 or 503 response. If you made no requests in an interval, that interval counts as 100% available. A silent stream that stops ingesting because the service is wedged earns nothing if you were not calling the API at that moment.
  • The claim itself must document the interval-by-interval availability: billing cycle, region, dates and times of each incident, plus request logs with sensitive data redacted. AWS asks for the specific intervals below 100%, not just an incident ID.

The money: Kinesis Data Streams on-demand pricing runs $0.08 per GB ingested and $0.040 per GB retrieved in US-East, plus $0.040 per stream-hour, which is $28.80 per month for one idle on-demand stream. A 60-minute outage in a 30-day month puts you at roughly 99.86%, under 99.9%, so a single stream earns 10% of $28.80: $2.88. Small, but the SIEM or clickstream pipeline running 20 streams gets $57.60, and the claim is one support case.

MSK: the Multi-AZ gate

Amazon MSK's SLA (May 19, 2022) looks like Kinesis's with one difference that matters more than the percentages: it covers only Multi-AZ Deployments. The definition of a Multi-AZ Deployment for a provisioned cluster is a topic with a replication factor of at least two and partition replicas in a different availability zone than the lead replica. Run a single-AZ cluster, or a topic with replication factor one, and the SLA does not apply at all. No breach, no credit, regardless of what broke.

The error definition is wider than AWS's usual 500/503 pair: MSK counts Amazon MSK API requests returning 500 or 503, plus Kafka protocol errors 2, 8, 9, 15, 56 and 72, plus 19 and 20 when they persist on retry. Claims must include the cluster ARN or connector ARN for every affected resource, and AWS explicitly excludes failures caused by not following operational guidance: overloading brokers, using an excessively large number of partitions, or running insufficient capacity all void the claim. The underlying Kafka and ZooKeeper engine software is also excluded, which is worth rereading: a bug in the managed engine itself can fall outside the commitment.

The single biggest streaming claim killer is not the threshold, it is the deployment shape. MSK pays nothing for single-AZ clusters, and Kinesis treats intervals with no API traffic as fully available. Check both before you spend an hour assembling logs.

Event Hubs: whole minutes, all attempts

Azure Event Hubs (October 2026 consolidated SLA) works differently from AWS on both ends. The commitment is higher: 99.95% for Basic and Standard, 99.99% for Premium and Dedicated. But a minute is "unavailable" only if all continuous attempts to send, receive, or operate throughout that minute either return an Error Code or fail to produce a Success Code within five minutes. One request succeeding in a minute resets the clock for that minute. Partial degradation, throttling that still lets some traffic through, and brief blips inside a minute all vanish from the calculation.

Credits cap at 25% on all four tiers (10% below the commitment, 25% below 99.0%), and the free tier is not covered. The claim clock is Microsoft's standard Azure rule: file within 60 days of the incident, with a description, times, affected resources, user counts, and the errors observed.

Pricing is per throughput unit: $0.015 per hour for Basic, $0.03 for Standard, with each unit granting 1 MB/s ingress and 2 MB/s egress. Ten Standard units for a month is about $216. A 60-minute outage there drops the month to 99.86%, below 99.95%, so 10% of $216: $21.60. The same outage in Kinesis terms paid $2.88 per stream. Streaming SLAs are real money, but the arithmetic rewards knowing which service you run.

What to do this week

  1. If you run MSK, confirm every cluster is Multi-AZ with replication factor two or higher. Single-AZ clusters are the cheapest way to forfeit your only claim channel.
  2. Put the deadlines on a calendar: end of the second billing cycle for AWS, 60 days from the incident for Azure. Both are shorter than they feel.
  3. Collect request logs for the outage window before you open the case, and send per-interval availability, not a status page screenshot.

The providers publish these numbers because they know most customers never file. Yours are owed either way.