← Back to Blog

NoSQL SLA Credits: What DynamoDB, Cosmos DB, Firestore, and Bigtable Pay When They Fail

September 29, 2026

The four biggest NoSQL databases hold the strongest uptime promises in the cloud. DynamoDB, Cosmos DB, Firestore, and Bigtable all commit to 99.999% availability for certain configurations, a number that allows about 26 seconds of downtime in a 30-day month. Breach it and the credit schedule starts at 10% of that service's monthly bill, reaches 100% on AWS, and stays capped at 25% on Azure and 50% on Google. What most teams never check is that the five nines are conditional on how you deployed. Single region instead of multi region, and the commitment drops a tier or two, sometimes to 99.9%. This post maps the commitments to the deployment shapes that earn them, the measurement rules that decide whether an outage counts, and the filing deadlines.

Hourglass with four colored sand streams merging into a single golden bowl, the uptime math behind four NoSQL databases

The five nines are a deployment choice, not a product feature

Each provider publishes one SLA with several commitments inside it, and which one applies to you depends on the account configuration. DynamoDB splits into Standard and Global Tables. Cosmos DB splits four ways: single region, with or without availability zones, multi-region reads, and multiple write locations. Firestore splits two ways and Bigtable four.

Deployment shapeCommitmentWhere it is written
DynamoDB, any account (Standard SLA)99.99%DynamoDB SLA, May 14 2025
DynamoDB, all tables in region part of Global Tables, with failover attempts99.999%Global Tables SLA, same document
Cosmos DB, single region99.99%Azure SLA, September 2026
Cosmos DB, single region with availability zones99.995%Azure SLA, September 2026
Cosmos DB, two or more regions, one writable99.999% readsAzure SLA, September 2026
Cosmos DB, two or more writable regions99.999%Azure SLA, September 2026
Firestore, multi-region location99.999%Firestore SLA, March 2025
Firestore, regional location99.99%Firestore SLA, March 2025
Bigtable, replicated, multi-cluster routing in 3+ regions99.999%Bigtable SLA, March 2025
Bigtable, replicated, multi-cluster routing in fewer than 3 regions99.99%Bigtable SLA, March 2025
Bigtable, replicated, single-cluster routing, or zonal (one cluster)99.9%Bigtable SLA, March 2025

The pattern is consistent: every five-nines promise requires geography. One cluster, one zone, or one write region costs you nines, sometimes two of them.

The deployment shape is the claim. Cosmos DB's 99.999% needs two or more regions, Firestore's needs multi-region mode, Bigtable's needs multi-cluster routing across three or more regions, and DynamoDB's needs every table in the region to be a Global Table with failover attempted. Run the cheaper shape and your outage is measured against a weaker promise.

What each nine allows in real minutes

The gap between tiers is where unclaimed money hides. In a 30-day month of 43,200 minutes:

CommitmentDowntime allowedWhat breaches it
99.999%~26 secondsthe first five-minute error window
99.995%~2 minutes 10 secondsthe first five-minute error window
99.99%~4 minutes 20 secondsroughly two bad intervals
99.95%~22 minutesfive bad intervals
99.9%~43 minutesa solid hour of failure

Firestore and Bigtable count errors only in five consecutive minutes, so a four-minute spike earns nothing even at the five-nines tier. DynamoDB and Cosmos DB count partial windows. Bigtable single-cluster teams watch error panels for half an hour and collect zero; an operator with multi-region Firestore files a claim over a blip that woke them at 3 a.m.

The measurement rules decide the claim

Each provider defines an outage differently, and none of them count the same thing you see on a status page.

ProviderHow downtime is measuredErrors that count
DynamoDBper-5-minute interval, share of requests; zero-traffic intervals count as 100%HTTP 500 and 503 only; throttling (400/429) never counts
Cosmos DBhourly error rate per subscription, averaged over the Applicable Period; failed = error code or no success code within the latency bounds (5 seconds for resource operations)500-class failures; 429 rate limiting is excluded from availability and handled by a separate throughput SLA
Firestoremore than 5% server-side error rate for 5 consecutive minutes; shorter windows never countHTTP 500 with Internal Error, Unknown, or Unavailable codes
Bigtablemore than 5% error rate for 5 consecutive minutes, minimum 60 requests per minute in the windowHTTP 50x with Internal Error, Unknown, or Unavailable codes

Two consequences follow. Your traffic defines the record: on all four services, an interval with no requests counts as healthy. And retry behavior can erase it: both Google services discard repeated identical requests unless your client backs off, 1 second doubling to 32. Hand-rolled retry loops without backoff can wipe out the error data your claim depends on; the SDKs handle it by default.

The credit schedules

All four pay a percentage of what you actually spent, applied to future bills. The tiers:

ServiceCredit tiersCap and notes
DynamoDB10% below the commitment, 25% below 99.0%, 100% below 95.0%$1 minimum; AWS may refund to your card at its discretion; credits stay in the account
Cosmos DB10% below each commitment, 25% below 99.0%no tier above 25%; if one incident misses several service levels (availability, throughput, consistency, latency) you choose exactly one to claim under
Firestore10% below the commitment, 25% below 99.0%, 50% below 95.0%50% of the monthly bill, per project per region
Bigtable10% below the commitment, 25% below 99.0%, 50% below 95.0%50% cap, per instance

Azure adds one filter: the credit applies only to the affected resource or tier, and only to actual downtime, not the overall incident. Google applies approved credits within 60 days of the request.

Money math: what an outage actually pays

Same spend in every row: $2,000 per month, 30-day month.

Outage lengthDynamoDB Standard (99.99%)DynamoDB Global Tables (99.999%)Firestore multi-region (99.999%)Bigtable single cluster (99.9%)
4 minutesnothing (99.991%)$200, 10%nothing, under the 5-minute windownothing
20 minutes$200, 10%$200, 10%$200, 10%nothing (99.954%, still above 99.9%)
~15 hours$500, 25%$500, 25%$500, 25%$500, 25%

The 15-hour row is real. On October 20, 2025, an internal DNS incident in AWS us-east-1 took down the DynamoDB regional endpoint for roughly 15 hours, pushing the month to about 97.98% availability there. A $2,000 DynamoDB bill at the 25% tier returned $500. Only DynamoDB has a 100% tier, and it needs availability below 95%, more than 36 hours of downtime in a month.

Filing inside the windows

The claim routes follow each provider's standard channel, with NoSQL-specific evidence:

  • AWS DynamoDB: open a support case with "SLA Credit Request" in the subject, and include the billing cycle, region, computed monthly uptime percentage, the dates, times, and availability of every 5-minute interval below 100%, and request logs documenting the errors. Deadline: end of the second billing cycle after the incident month.
  • Azure Cosmos DB: support request within 60 days of the incident, with the incident description, time and duration, affected resource names, number and location of affected users, and the errors observed. For pay-as-you-go the Applicable Period is the 30 days before and including the incident day; the credit applies to those fees.
  • Google Firestore and Bigtable: contact Google Cloud support within 30 days of eligibility. Firestore claims are per project per region; Bigtable claims are per instance. Total credits for the month are capped at half the bill.

The AWS and Azure evidence checklists are the same regardless of service: timestamps in a stated time zone, the provider's incident ID, your resource identifiers, and error logs. The claim evidence guide covers what each provider accepts, and the nines explainer does the annual math.

Here is the concrete move for this month. List every NoSQL database you run and write down the commitment its deployment shape actually earns: region count, write locations, routing policy, Global Tables status. That list is your claim map. When a provider health feed shows errors in your region, check the map, then the clock, then the window. If the commitment was breached and the deadline is open, file. The credit already sits in the provider's budget; it waits for someone to ask. The mapping and clock-watching is what UptimeAudit automates: it watches big four health feeds by service and region, computes the interval math, drafts the claim, and tracks the deadline. Nothing is submitted for you; the provider makes the final call.