← Back to Blog

When Your Model API Fails: What Bedrock, Azure OpenAI, and Vertex AI SLAs Pay

September 15, 2026

Model APIs are production dependencies now. When Bedrock, Azure OpenAI, or Vertex AI slows down or fails requests, the feature on top of it breaks in front of paying customers, and the contract underneath is a 99.9% monthly uptime headline with credits measured in percentages of the AI bill. That promise is much narrower than the compute SLAs most teams already know. Bedrock counts only one HTTP error code. Google pays nothing for a single-node deployment. Microsoft's model endpoint table stops at a 25% credit. Here is what each one measures, excludes, and pays.

A brass hourglass with glowing blue shards falling through it, pooling around a single gold coin in the lower chamber

The three promises, side by side

All three lead with the same number for their flagship tiers. The differences start in how that number is measured.

PlatformCommitmentCredit tiersDowntime measurement
AWS Bedrock99.9% per region10% / 25% / 100% of regional Bedrock chargesRequests that return HTTP 500, averaged across 5-minute intervals
Azure model endpoints (Foundry Models, including Azure OpenAI)99.9%10% / 25% of the affected charges, with no third tierRequests that return an error code, averaged per minute
Vertex AI99.9% for training, deployment, batch prediction and AutoML online prediction; 99.5% for custom online prediction and pipelines; 99% for the training cluster control plane10% / 25% / 50% of the covered service's billMinutes where the server-side error rate exceeds 5%, based on HTTP 500 and 503 responses

Two details decide more claims than the percentages: a credit applies to the affected service's charges, not your whole cloud bill, and it lands on future bills rather than as a refund.

What each platform counts as an outage

Bedrock's definition of an Error is one sentence long: any request that returns HTTP 500. Throttling responses, auth failures and client-side timeouts never move the number, and 5-minute intervals with no requests count as 100% available. A failure severe enough to empty your traffic, because users gave up or your app stopped calling, makes those quiet minutes look perfect.

Azure averages an error rate per minute across the month: requests returning an error code divided by total requests in that minute. Zero-request minutes count as 0% error rate, the same loophole in different clothes.

Google measures server-side error rate above a 5% threshold per minute, using only HTTP 500 and 503 responses to valid requests. It also discards repeated identical failing requests unless they follow the documented back-off requirements, starting at one second and doubling up to 32. A retry loop that hammers the same call without back-off shrinks the incident on paper.

The gates: nodes, tiers and preview status

Three structural gates decide claims before a single error is counted.

Two nodes or nothing. Vertex AI's online prediction commitments attach to deployments on two or more nodes, for custom models at 99.5% and AutoML tabular and image models at 99.9%. A single-node deployment carries no uptime commitment, so a pilot promoted straight to production cannot earn a credit no matter how long it was down.

The Developer tier is excluded. Microsoft's credit terms cover models sold by Azure, and the Foundry Models Developer tier sits outside the SLA entirely.

Preview models are excluded. Google excludes pre-general-availability features by name, Microsoft excludes previews and free tiers from claims, and AWS offers SLAs for paid, generally available services only. If your model is still marked preview, its promise is zero on all three.

Slow is not down, unless you reserved capacity

Every uptime number above counts failed requests, never slow ones. A model answering at a crawl for a full day records 100% availability and pays nothing.

Microsoft is the exception, and only on reserved capacity. Provisioned deployments and priority processing carry a separate latency service level: generation speed in tokens per second, measured as a 5-minute median or 1-minute average depending on model, against a per-model target published in the product docs, written like 99% of requests above 50 tokens per second. Miss 99% attainment and the compensation is a 10% credit. Standard pay-as-you-go deployments have no latency commitment and emit no metrics to evidence one.

If one incident breaches both metrics, you pick a single service level to claim under, so run both calculations and file the larger number.

The pattern underneath all three SLAs: the provider measures its endpoint, not your experience. Slow responses, quiet minutes, preview models, single-node deployments and retry loops without back-off all record as healthy, and the payout is always a fraction of the model bill. The only minutes that pay are the minutes you can prove.

What an outage pays

The arithmetic: $2,500 of monthly spend, a 30-day month of 43,200 minutes. The 99.9% tolerance is 43.2 minutes. The 25% tier opens after 7.2 hours of failed requests. The top tiers require more than 36 hours.

ScenarioRecorded uptimeBedrock paysAzure paysVertex AI pays
6 hours down99.17%$250 (10%)$250 (10%)$250 (10%)
8 hours down98.89%$625 (25%)$625 (25%)$625 (25%)
40 hours down94.44%$2,500 (100%)$625 (25% ceiling)$1,250 (50% ceiling)

A six-hour model outage is a bad day that pays one first-tier credit. Bedrock is the only schedule that reaches 100%, and it takes more than 36 hours of failed requests in a month to get there. Microsoft's model endpoint table has no third tier, so 25% is the ceiling however bad the month gets. Small bills can earn nothing at all: AWS will not issue a Bedrock credit under $1.

None of that makes the claims pointless. A $625 credit is $625 and providers pay when the evidence lines up. It covers a slice of one bill, not the revenue the feature was supposed to earn.

How to file, and what to capture

All three run a claim process with a hard window and a fixed evidence list.

ProviderWindowWhere to fileEvidence required
AWS BedrockEnd of the second billing cycle after the incidentSupport case with "SLA Credit Request" in the subjectBilling cycle, region, monthly uptime percentage, and the dates, times and availability of every 5-minute interval under 100%, plus request logs of the errors
Azure model endpoints60 days from the incidentSupport request with issue type Billing and problem type Refund RequestDescription of the incident, its time and duration, affected resource names, number and location of affected users, and the errors you observed
Vertex AI30 days from the moment you become eligibleGoogle Cloud support contact formProject ID, job IDs, and the dates and times of the errors

Small print worth knowing before you file:

  • Google applies the credit within 60 days of the request and caps a single month at 50% of the service's bill.
  • Microsoft typically processes claims within 45 days, and CSP customers must file through their partner rather than the Azure portal.
  • For metered services, Microsoft reviews the 30 days up to and including the incident day, not the calendar month.
  • AWS issues the credit within one billing cycle after the month you filed.

Before the next model incident

Three habits cover almost everything.

  1. Capture evidence while the incident is happening. Export the status code distribution by timestamp (5-minute buckets on Bedrock, since that is the unit AWS audits), request IDs, region and deployment names. If you use Google, keep retry logs that show back-off, because that is what makes your failed requests count. For building your own downtime record, see measuring an outage yourself.

  2. Know which promises you hold. A preview model, a single-node deployment or a Developer-tier resource carries no claimable commitment. If the promise matters, put online prediction on two nodes and stick to GA models sold directly by the platform.

  3. Put the windows somewhere visible: 30 days from eligibility for Google, 60 from the incident for Microsoft, the second billing cycle for AWS. Miss one and the credit is gone.

One action for today: add the log export command for your model endpoints to your incident runbook. Status codes, timestamps, region, request IDs. Every one of these SLAs pays on minute-level proof, and that proof is easiest to grab while the incident is still open.