← Back to Blog

Do Spot Instances Have an SLA? What AWS, Azure, and GCP Actually Guarantee

September 12, 2026

Spot capacity is the cheapest compute in the cloud and the only compute sold without an uptime promise. That saves real money, up to 90 percent off on-demand rates, and it changes what you can ever claim back from your provider. Before you move more of the fleet onto spare capacity, it pays to know exactly where AWS, Azure, and Google stand on paper, what a reclaim looks like in practice, and whether the discount is worth more than the credit it replaces. It almost always is, and the numbers below show why.

A robotic claw lifting one glowing amber cube out of a wall of identical dark cyan cubes, leaving a gap where it was taken

Where each provider stands on paper

There is no ambiguity hiding in the contracts. Microsoft's docs say it outright: "there's no SLA for these VMs." Google's is just as explicit: "Spot VMs are not covered by any Service Level Agreement and are excluded from the Compute Engine SLA." AWS publishes nothing that dramatic, but the pieces line up the same way: spot instances run "whenever capacity is available," get reclaimed when "Amazon EC2 needs the capacity back," and the compute SLA's credit process has no procedure for a reclaim.

ProviderSpot productListed discountReclaim warningWhat the SLA says
AWSSpot Instancesup to 90% off On-Demand2-minute interruption notice, plus earlier rebalance recommendationsInterruptions are documented behavior; the credit process covers outages, not reclaims
AzureSpot VMsup to 90% cheaper than pay-as-you-gominimum 30 seconds before eviction"No SLA for these VMs"
Google CloudSpot VMsup to 91% offmetadata signal, 0 or 120 seconds notice, shutdown period up to 30 seconds"Excluded from the Compute Engine SLA"
DigitalOceannonenot offerednot applicableNo spot tier exists, so Droplets keep their standard coverage

What a reclaim actually looks like

For a workload built for it, an interruption is boring: the instance drains, the queue picks the job up somewhere else, nobody notices. For a workload that is not built for it, the same event is a page at 3am and a half-finished job. The mechanics differ by provider, and the differences matter if you are writing handlers.

AWS issues a two-minute interruption notice as an EventBridge event and in instance metadata, and recommends checking every five seconds. Delivery is best effort, and rebalance recommendations often arrive earlier when an instance is at higher risk. Azure gives a minimum of 30 seconds before an eviction, through opt-in Scheduled Events, also best effort. Google signals preemption through instance metadata, with a notice duration you choose at creation: 0 seconds by default, 120 seconds if your workload needs the runway. None of these promise that you will finish your work.

The billing rules are friendlier than most people assume:

  • AWS: an instance interrupted by EC2 in its first hour is not charged at all, and after that you pay only for the seconds you used. Attached EBS volumes keep billing while the instance is stopped.
  • Azure: the default deallocate policy keeps your disks and keeps charging for them. The delete policy removes the VM and its disks together.
  • Google: a preempted VM that stops is not charged for VM hours while it sits terminated, and a VM preempted within its first minute is not billed at all. Disks keep billing either way.

A reclaim is not an outage. It is the product working as documented, and you were paid for it up front, in the discount. The only questions that matter are whether the discount outweighs the disruption and whether the workload can absorb it.

The money math

Put real numbers on it. Take a $10,000 monthly compute bill, move a share onto spot at a 70 percent average discount (the ceilings in the table are higher, but averages are what show up on invoices), and compare the monthly saving against the credit you could claim if things went badly.

Compute mixMonthly billMonthly savingCredit at a 10 percent tier
All on-demand$10,000n/a$1,000
Half on spot$6,500$3,500$650
80% on spot$4,400$5,600$440

Every SLA in play calculates credits as a percentage of what you actually paid for the affected service, so a smaller bill means smaller credits. In the 80 percent mix, one month of spot savings ($5,600) is worth roughly 13 months of a possible credit ($440), and that credit only arrives after a region-level failure, a claim, and a filing before the deadline. The savings land every month regardless. That asymmetry is the whole decision.

One extra detail for AWS accounts holding a Compute Savings Plan: spot spend does not apply toward the commitments, so a fleet that leans too far onto spot can leave a shortfall behind it. Balance the mix with that in mind.

Keep a covered core

Spot is a trade, and it is only fair when you decide, workload by workload, who can lose an instance mid-run. Batch jobs, queue workers, CI runners, and stateless services behind a load balancer are spot material. Databases, single-instance stateful services, and anything your own contracts promise at a specific uptime stay on on-demand or reserved capacity. That group is your covered core: your uptime floor, and the spend any future credit calculation draws on.

It is smaller than most teams assume, too. AWS measures instance-level unavailability as loss of external connectivity and automatically waives the charge for any single instance unavailable more than six minutes in a clock hour, no claim required. That kind of detail is the difference between an SLA you can use and an SLA you only quote.

Then watch the supply side of the trade. Azure publishes per-SKU eviction rates for the last 28 days through Resource Graph, Google's managed instance groups can pick machine types with the lowest observed preemption rates, and AWS rebalance recommendations are the early warning that a pool is heating up.

Running spot without surprises

Handlers first. Every provider gives you a signal, and all of them treat delivery as best effort, so build as if the notice might not arrive:

  • AWS: an EventBridge rule or a metadata check every five seconds, then drain the instance and let the fleet replace it. Rebalance recommendations often land before the two-minute notice.
  • Azure: opt in to Scheduled Events and react to the Preempt signal. Choose deallocate if the workload should wait for capacity in place, or delete with ephemeral OS disks if it should redeploy elsewhere quickly.
  • Google: read the preempted flag from metadata and run a shutdown script. If your workload needs more than a few seconds, set the notice duration to 120 seconds when you create the VM.

Then spread the risk across instance types and zones, so a single hot pool cannot take the whole fleet at once, and checkpoint on a schedule rather than only when a signal arrives. Spot capacity is finite and shared, and the moments you most want it are the moments everyone else wants it too. The handlers are best effort; your job progress should not be.

The action worth taking this week

Open last month's compute bill and split it into two columns: workloads that can lose an instance mid-run, and workloads that cannot. Move the first column to spot, keep the second on capacity that carries a promise, and run the arithmetic above against your own bill. That covered core is the only part of your infrastructure that ever converts a provider failure into money back, and it is the part uptimeaudit.io monitors and, when a breach crosses a threshold, drafts the claim for. The discount pays out every month. The credit, if it ever arrives, arrives once and after a fight. Price both accordingly.