← Back to Blog

No Capacity, No Claim: Why Out of Stock Is Not an Outage

September 27, 2026

An outage you cannot claim is still an outage. When a scale-up dies at the quota ceiling, or a launch request fails with an insufficient capacity error, your dashboards go red and the SLA maths stays green. The reason is in the general terms: uptime is measured over resources that already exist. Azure says it outright, excluding operations "such as restart, stop, start, failover, scale compute, and storage (which by their nature include capacity constraints) that incur downtime". Google, in SLA after SLA, excludes errors "that resulted from quotas applied by the system or listed in the Admin Console". The capacity you cannot get is not an outage, and the two Azure products that sell capacity guarantees show what it costs to change that.

A night harbour with container ships docked and a crane holding a glowing container suspended over an empty berth

The uptime maths counts what already runs

Where it is writtenWhat it excludes
Azure general terms"Your initiated operations such as restart, stop, start, failover, scale compute, and storage (which by their nature include capacity constraints) that incur downtime"
Azure general termsattempts to "perform operations that exceed reasonable use or prescribed quotas", plus provider throttling
Azure Kubernetes Automatic Clusteran "Excluded Window" is any five-minute interval when large-scale operations of 100 or more nodes are initiated, or quota errors appear
Google, across Compute, Cloud Run, Kubernetes Engine, Cloud Armor, Cloud Storage and moreerrors "that resulted from quotas applied by the system or listed in the Admin Console"
Google Cloud SQLalso excludes downtime from "Customer's restart of an Instance"
AWSthe EC2 SLA measures the availability of the instances you run; AWS documentation treats "insufficient instance capacity" as a standard launch failure to troubleshoot, not a credit event

The pattern is consistent: the provider's clock runs on running resources. If your platform failed to add capacity, the failure is real, the cause is real, and the uptime percentage simply does not see it.

What this looks like in an incident

A traffic spike pushes an autoscaling group against its vCPU quota. The extra instances never launch, the surviving fleet saturates, and the incident channel fills up. The SLA number does not move, because the capacity that failed to materialise was never part of the measurement, and the scale-up itself is an excluded operation. The same logic catches a failed deployment into a region that is short of a particular instance type: the error is a standard one, documented in the troubleshooting guides, and no commitment attaches to the launch.

What remains covered is the stewardship of what is running. If the provider's own infrastructure loses your instances, or availability for running resources drops below the commitment, the clock runs and the ladder applies. The exclusions carve out creation and control-plane actions, not provider-side failure of the resources you already hold. That distinction is worth writing into incident reviews, because it decides which parts of a bad afternoon were claimable and which were never going to be.

The exception: capacity as a product

Two Azure reservation products carry their own SLA sections, and they read unlike anything else on the big four's shelves.

Capacity Blocks reads its commitment in capacity units rather than minutes: Capacity Availability is the percentage of the reserved capacity that is available for provisioning during the period, with a ladder of 10% below 99.9%, 25% below 99%, and an unusual cliff at the bottom, 100% below 50%. The SLA explicitly covers only the failure to provision or make available the reserved capacity, and hands over to the Virtual Machine SLA once deployed resources are running.

On Demand Capacity Reservations covers a different failure: a Supported Deployment that consumes an unused reservation receiving a "lack of capacity" error. Minutes not available accumulate per reserved unit until a deployment succeeds, another failure occurs, or fifteen minutes elapse. The ladder is 10% below 99.9%, 25% below 99%, and 100% below 95%, and the credit attaches to the cost of each reserved unit rather than the reservation as a whole.

The arithmetic shows what is being bought. On a $10,000 a month Capacity Block, a month that offers 60% of the reserved capacity pays $2,500; a month below half pays the full $10,000. For comparison, the standard compute ladders pay nothing at all for a month of missing capacity, because the failure sits outside their measurement entirely. AWS sells the same idea as a product: Capacity Blocks for ML let you reserve GPU capacity starting on a future date, for durations up to six months.

The trade is the whole story: providers sell uptime on the resources you have, not on the resources you want. Quotas, scale-ups and launch failures sit outside the uptime maths by design, and the only way to move them inside is to buy capacity as a product. Everywhere else the fix for a capacity failure is architecture, not a claim.

What to do with this

Three habits cover almost every capacity incident. First, raise quotas before they bite: every major provider treats quota increases as a self-serve request, and an afternoon spent raising ceilings is cheaper than a night spent explaining a saturation. Second, if a specific workload cannot tolerate a failed launch, buy the reservation product and read its ladder, because it is the only capacity guarantee that pays. Third, when a big incident unwinds, separate "we could not create" from "what we had fell over": the second bucket still produces credits, the first almost never does, and mixing them weakens the claim that matters.

None of this is a loophole in anyone's favour; it is the line providers draw between selling you capacity and promising it. UptimeAudit works the other side of that line, tracking the big four's health feeds and drafting the credits the documents actually pay.