← Back to Blog

Key Management SLAs: What KMS, Key Vault, Cloud KMS and Secret Manager Pay When Keys Fail

October 10, 2026

Key management services carry the strictest SLAs in the cloud, and nearly nobody claims them. AWS KMS promises 99.999% uptime, which is about 26 seconds of downtime allowed per month. Azure Key Vault promises 99.99%. The catch: credits are a percentage of what you paid for the key service itself, and those bills are often single digits. The tightest contract in your account protects the cheapest line item, so the claim game is different here. You are not chasing dollars. You are chasing the documented breach, and the credit is a paper trail that costs ten minutes to file.

A heavy brass keyring holding large ornate keys, one key snapped in two with a glowing red crack, dark stone wall behind

ServiceCommitmentCredit tiersWhat counts as downClaim window
AWS KMS99.999% per region10% below 99.999%, 25% below 99.0%, 100% below 95.0%500/503 errors per 5-min intervalEnd of second billing cycle
AWS Secrets Manager99.99%10% below 99.99%, 25% below 99.0%, 100% below 95.0%500/503 errors per 5-min intervalEnd of second billing cycle
Azure Key Vault99.99%10% below 99.99%, 25% below 99.0%Whole minutes, all attempts fail or no Success Code within 5 seconds60 days from incident
Azure Key Vault Managed HSM99.9% (99.99% if replicated across regions)10%/25%Whole minutes, all attempts fail or no Success Code within 5 seconds60 days from incident
Cloud KMS / Cloud HSM (direct app ops)99.95%10%/25%/50% cap>10% error rate, 5+ consecutive minutes30 days from eligibility
Cloud KMS / Cloud HSM (service account ops)99.99%10%/25%/50% cap>10% error rate, 5+ consecutive minutes30 days from eligibility
Secret Manager99.95%10%/25%/50% cap>10% error rate, 5+ consecutive minutes30 days from eligibility

AWS KMS: 99.999% and a $1 floor

The AWS KMS SLA (last updated November 29, 2022) is the strictest AWS publishes. The commitment is 99.999% per region per billing cycle, the top tier is 10% for any month below 99.999% but at or above 99.0%, and the ladder runs 25% and 100% below that. Errors are 500 and 503 responses per five-minute interval, and intervals with no requests count as fully available.

Because the budget is 26 seconds a month, an outage of a few minutes almost always breaches. An hour of KMS failure in a 30-day month puts the month at roughly 99.86%, squarely in the 10% tier. The problem is the base. KMS costs $1 per key per month, plus $0.03 per 10,000 requests after a 20,000-request free tier. An account with 10 customer managed keys and light traffic pays about $10, so a 10% credit is $1, and AWS does not pay out credits under $1. You need roughly ten keys plus real request volume before a KMS breach produces a check that clears the floor.

AWS also writes out exclusions that matter: custom key stores, both CloudHSM key stores and external key stores, are excluded from the SLA entirely. If your keys live in a CloudHSM custom key store and the store fails, the KMS SLA does not apply. Disabled keys and misconfigurations on your side are excluded too.

Secrets Manager's SLA (December 5, 2023) sits one rung lower at 99.99%, with the same 10/25/100 ladder, the same $1 floor, and the same end-of-second-billing-cycle deadline. Errors, again, are 500/503 only. Its exclusion clause references the documented best practices, so a rotation setup that ignores the user guide is a weaker claim.

Azure Key Vault: the 5-second, whole-minute rule

Key Vault's October 2026 consolidated SLA commits to 99.99% across your subscription's vaults. A minute is unavailable if all continuous attempts to perform transactions throughout that minute return an Error Code or do not produce a Success Code within five seconds. Note the exclusions baked into the definition: transactions that create, update, or delete vaults, keys, or secrets do not count at all. The SLA covers read and use operations on keys and secrets, not the management plane. A morning where vault creation fails for two hours is not a covered minute.

Managed HSM starts at 99.9% for a single region, and rises to 99.99% when the pool is replicated across regions. Credits top out at 25% on every tier. The claim window is 60 days from the incident. Support wants a description, incident times, affected resource names, affected user counts, and the errors observed.

Key Vault pricing is per operation: $0.03 per 10,000 transactions on standard vaults, with HSM keys at $1 per key per month and advanced key types climbing from $5. A subscription doing a million secret reads a month pays about $3, so a 25% credit on that is $0.75. The credit exists, the invoice line barely does.

The strictest SLA in your account protects the smallest bill. KMS at 99.999% pays 10% for any real outage, but after the $1 floor and a $1/key/month base, most accounts get nothing they could spend. File anyway: the credit is not the point, the documented breach is.

Google Cloud: one service, two promises

Cloud KMS and Cloud HSM run two different SLAs in the same document. Operations that originate directly from your application or an end user carry a 99.95% commitment. Operations requested by a service account, which is how CMEK-enabled services like Compute Engine, GKE, and BigQuery call KMS, carry 99.99%. Most of your production traffic lands on the higher number.

Downtime means more than a 10% error rate for five or more consecutive minutes with at least ten valid requests. Errors are 500-range responses or no valid response within ten seconds. The credit ladder is 10/25/50 with a hard 50% cap on monthly credits, and the window is the Google standard: notify technical support within 30 days of eligibility, with log files showing the downtime. Secret Manager sits at 99.95% with the same ladder, cap, and window.

Google pricing is the smallest of the three: software-protected key versions run about $0.06 per key per month ($0.000082192 per hour), HSM keys about $1, and crypto operations $0.03 per 10,000. A hundred software keys cost about $6 a month. The 10% tier on a bad month is $0.60, and there is no floor, so Google will process it. What a claim buys you is the audit trail at month end, not a refund that moves your budget.

The money math nobody prints

SetupMonthly bill1-hour outage puts month atCredit
AWS, 10 KMS keys, light traffic~$1099.86% (10% tier)$1.00, but AWS floor means likely $0
AWS, 100 KMS keys, 1M requests~$10399.86% (10% tier)~$10.30
Azure, 1M vault reads~$399.86% (25% tier needs below 99%)~$0.30
Azure, 2M reads + HSM keys~$1199.86% (10% tier)~$1.10
Google, 100 software keys~$699.86% (10% tier)~$0.60

The pattern is consistent: the tighter the SLA, the smaller the base, and the less the credit matters as money. That is not an argument to skip claims. Key management failures take down every service pinned to the key, so a KMS or Key Vault breach is usually the root cause of a much larger incident. When the postmortem blames the key service, the SLA credit is the only compensation the contract provides, and the evidence you file with it is exactly what the incident review needs.

Do this now

Check which tier your workloads actually sit in: service-account CMEK calls on Google (99.99%), standard key reads on Azure (99.99%), KMS with a CloudHSM custom key store (excluded). Then put each window on your claim calendar: end of the second billing cycle for AWS, 60 days from the incident for Microsoft, 30 days from eligibility for Google. When the next key-service incident hits, file the claim with interval-level logs within the window, whatever the dollar figure says.