# Error Budgets: How to Calculate and Alert on Burn Rate

> An error budget is the unreliability an SLO allows. Learn to calculate it, track burn rate, set multiwindow alerts in PromQL and write an error budget policy.

Source: https://last9.io/blog/error-budget/

An error budget is the amount of failure a service is allowed under its service level objective (SLO). It equals one minus the SLO target: a 99.9% availability SLO over 30 days gives an error budget of 0.1%, which is about 43 minutes of full downtime or one failed request in every thousand.

Teams spend an error budget on releases and risk, measure how fast they spend it with burn rate, and slow down changes when it runs out.

This guide covers how to calculate an error budget, how burn rate works, the multiwindow alert setup Google's SRE teams recommend, the PromQL to run it, and what an error budget policy should say. The definitions and alert thresholds come from Google's [Site Reliability Engineering book](https://sre.google/sre-book/embracing-risk/) and [SRE workbook](https://sre.google/workbook/alerting-on-slos/).

## What is an error budget?

An error budget turns an SLO into a quantity a team can spend. The SRE book describes it as the difference between the SLO and actual measured uptime: "the 'budget' of how much 'unreliability' is remaining."

The idea changes how teams argue about releases. Without a budget, product teams push for speed and operations teams push for stability, and the argument never ends. With a budget, the data decides. The SRE book states the rule: "as long as the system's SLOs are met, releases can continue." When the budget is spent, releases pause and effort moves to reliability.

An error budget needs three things defined first:

- **An SLI (service level indicator):** the measurement, such as the ratio of successful HTTP requests to all requests.
- **An SLO target:** the goal for that SLI, such as 99.9%.
- **A window:** the period the target covers, usually a rolling 28 or 30 days.

If those terms are new, our guides to [SLA vs SLO vs SLI](https://last9.io/blog/sla-vs-slo-vs-sli/) and [implementing SLOs](https://last9.io/blog/a-practical-guide-to-implementing-slos/) cover the groundwork.

## How do you calculate an error budget?

The formula is the same for every SLO:

**Error budget = (1 minus SLO target) x total events or time in the window**

There are two ways to count, and they give different kinds of budget.

**Time-based budget.** Count the minutes in the window. A 30-day window has 43,200 minutes, so a 99.9% SLO allows 0.001 x 43,200 = 43.2 minutes of downtime.

**Request-based budget.** Count the requests in the window. If a service handles 10 million requests in 30 days, a 99.9% SLO allows 0.001 x 10,000,000 = 10,000 failed requests. Request-based budgets are more accurate for most services, because a partial outage that fails 5% of requests for an hour costs less than a full outage of the same length.

Here is what each common SLO target allows over 30 days:

| SLO target | Error budget | Downtime allowed per 30 days | Failed requests allowed per 10M |
| ---------- | ------------ | ---------------------------- | ------------------------------- |
| 99%        | 1%           | 432 minutes (7.2 hours)      | 100,000                         |
| 99.5%      | 0.5%         | 216 minutes (3.6 hours)      | 50,000                          |
| 99.9%      | 0.1%         | 43.2 minutes                 | 10,000                          |
| 99.95%     | 0.05%        | 21.6 minutes                 | 5,000                           |
| 99.99%     | 0.01%        | 4.32 minutes                 | 1,000                           |

The jump from 99.9% to 99.99% leaves about four minutes a month. At 99.99%, a single slow rollback can use the whole month's budget, so a higher target needs faster detection and safer deploys to go with it.

## What is burn rate?

Burn rate measures how fast a service is using its error budget. Google's SRE workbook defines it as "how fast, relative to the SLO, the service consumes the error budget."

**Burn rate = observed error rate / (1 minus SLO target)**

A burn rate of 1 means the service fails at exactly the rate the SLO allows, so the budget runs out right at the end of the window. For a 99.9% SLO, that is a steady 0.1% error rate. Higher burn rates use the budget up faster:

| Burn rate | Error rate at a 99.9% SLO | Time to use the full 30-day budget |
| --------- | ------------------------- | ---------------------------------- |
| 1         | 0.1%                      | 30 days                            |
| 2         | 0.2%                      | 15 days                            |
| 6         | 0.6%                      | 5 days                             |
| 14.4      | 1.44%                     | 50 hours                           |
| 1,000     | 100%                      | 43.2 minutes                       |

Burn rate matters because the error rate alone does not say how urgent a problem is. A 0.2% error rate sounds small, but on a 99.9% SLO it doubles the planned spend and will miss the SLO halfway through the month.

## How do you alert on error budget burn rate?

The simplest SLO alert fires when the error rate crosses the SLO threshold. That alert is noisy, because short spikes trip it, and slow, because a steady low burn never trips it at all. Google's SRE workbook recommends **multiwindow, multi-burn-rate alerts** instead.

Each alert checks the burn rate over two windows, a long one and a short one, and fires only when both are above the threshold. The long window confirms that a meaningful part of the budget has been used. The short window confirms the problem is still happening, so the alert stops firing soon after a fix. The workbook's rule of thumb is to "make the short window 1/12 the duration of the long window."

The workbook's recommended starting point for a 99.9% SLO:

| Severity | Long window | Short window | Burn rate | Budget used when it fires |
| -------- | ----------- | ------------ | --------- | ------------------------- |
| Page     | 1 hour      | 5 minutes    | 14.4      | 2%                        |
| Page     | 6 hours     | 30 minutes   | 6         | 5%                        |
| Ticket   | 3 days      | 6 hours      | 1         | 10%                       |

The numbers fit together. A 30-day window has 720 hours, so 1 hour at a burn rate of 14.4 uses 14.4 / 720, or 2%, of the budget. Six hours at 6 uses 36 / 720, or 5%, and 3 days at 1 uses 72 / 720, or 10%. The pages catch fast burns while there is still budget left to protect, and the ticket catches the slow leak that would otherwise go unnoticed until the month ends.

## How do you write burn rate alerts in PromQL?

Start with recording rules for the error ratio over each window. This example assumes an `http_requests_total` counter with a `code` label and a 99.9% SLO on the `api` job:

```yaml
groups:
  - name: slo-api-recording
    rules:
      - record: job:slo_errors_per_request:ratio_rate5m
        expr: |
          sum(rate(http_requests_total{job="api",code=~"5.."}[5m]))
          /
          sum(rate(http_requests_total{job="api"}[5m]))
      - record: job:slo_errors_per_request:ratio_rate30m
        expr: |
          sum(rate(http_requests_total{job="api",code=~"5.."}[30m]))
          /
          sum(rate(http_requests_total{job="api"}[30m]))
      - record: job:slo_errors_per_request:ratio_rate1h
        expr: |
          sum(rate(http_requests_total{job="api",code=~"5.."}[1h]))
          /
          sum(rate(http_requests_total{job="api"}[1h]))
      - record: job:slo_errors_per_request:ratio_rate6h
        expr: |
          sum(rate(http_requests_total{job="api",code=~"5.."}[6h]))
          /
          sum(rate(http_requests_total{job="api"}[6h]))
```

Then alert when both windows in a pair are above the burn rate threshold. The factor `0.001` is the error budget for a 99.9% SLO:

```yaml
groups:
  - name: slo-api-alerts
    rules:
      - alert: ErrorBudgetBurnFast
        expr: |
          (
            job:slo_errors_per_request:ratio_rate1h > (14.4 * 0.001)
            and
            job:slo_errors_per_request:ratio_rate5m > (14.4 * 0.001)
          )
          or
          (
            job:slo_errors_per_request:ratio_rate6h > (6 * 0.001)
            and
            job:slo_errors_per_request:ratio_rate30m > (6 * 0.001)
          )
        labels:
          severity: page
        annotations:
          summary: "api is burning its 30-day error budget too fast"
```

Add a third pair with 3-day and 6-hour recording rules and a burn rate of 1 for the ticket-level alert. Route the page and ticket severities to different receivers in [Alertmanager](https://last9.io/blog/prometheus-alertmanager/), so slow burns open a ticket instead of waking someone up.

To show how much budget is left on a dashboard, divide the error ratio over the full window by the budget and subtract from one:

```promql
1 - (
  sum(increase(http_requests_total{job="api",code=~"5.."}[30d]))
  /
  sum(increase(http_requests_total{job="api"}[30d]))
) / 0.001
```

A result of 0.4 means 40% of the budget is left. A negative result means the SLO is already missed for this window. A 30-day range query over raw counters is expensive, so on large services compute it from a recording rule or use a backend that handles long-range queries well.

## What should an error budget policy include?

An error budget only changes behavior if the team agrees in advance what happens when it runs low. That agreement is the error budget policy. Google's SRE workbook publishes an example policy with rules like these:

- **Release freeze:** "If the service has exceeded its error budget for the preceding four-week window, we will halt all changes and releases other than P0 issues or security fixes until the service is back within its SLO."
- **Postmortem threshold:** "If a single incident consumes more than 20% of error budget over four weeks, then the team must conduct a postmortem."
- **Planning commitment:** if a single class of outage consumes more than 20% of the budget over a quarter, the team adds a P0 item to next quarter's plan to fix it.
- **Escalation:** disagreements about the calculation or the actions go to the CTO.

A useful policy is short, signed off by both engineering and product leadership, and written before the first freeze, not during it. Read the full [example error budget policy](https://sre.google/workbook/error-budget-policy/) for the complete wording.

## What are common error budget mistakes?

Error budgets fail in practice for a handful of repeated reasons:

**Setting the SLO higher than users need.** A 99.99% target on an internal tool burns budget on failures nobody notices, and the team spends its time chasing noise.

**Counting the wrong failures.** If the SLI counts 4xx client errors as failures, a single misbehaving client can drain the budget. Most availability SLIs count only server-side failures such as 5xx responses and timeouts.

**Alerting on the error rate instead of the burn rate.** A static threshold on error rate either pages on every spike or misses slow burns. Burn rate alerts tie the page to budget impact.

**No policy behind the number.** A budget that runs out with no consequence is a dashboard, not a decision tool.

**Measuring from one place only.** Server-side metrics miss failures that happen before a request reaches the service, such as DNS or load balancer problems. External probes, like those from the [Prometheus Blackbox Exporter](https://github.com/prometheus/blackbox_exporter), help cover that gap.

## How does Last9 help with error budgets?

Last9 supports SLOs natively. In Last9, you define an SLI as a query, set a target such as 99.9% and a compliance window such as 1, 7 or 30 days, and choose a request-based or window-based SLO expression. Last9 then raises two alert severities: **Threat**, an early warning when the SLO is at risk, and **Breach**, when the target has been missed. Our [SLO documentation](https://last9.io/docs/slos/) covers the setup.

Last9 alert rules take PromQL, so you can write the burn rate conditions above as Last9 alert rules. An existing Prometheus alert rules file can be converted with our [Alertmanager migration endpoint](https://last9.io/docs/migrating-from-alertmanager/). [Last9 alerting](https://last9.io/alerting/) is built for high-cardinality use cases and keeps high-cardinality labels with no sampling, so SLIs can be broken down per customer or per endpoint. [Change Events](https://last9.io/docs/change-events/) record deployments and configuration changes as markers on your charts, so you can line up a burn rate spike with the release that caused it.

## Spend the error budget on purpose

An error budget turns a reliability target into something a team can plan with. Calculate it from the SLO, track how fast it is being spent with burn rate, page on fast burns over paired long and short windows, and open tickets for slow ones. Then write down, before you need it, what the team will do when the budget runs out.

Start with one user-facing service, a 99.9% request-based SLO over 30 days, and the two paging alerts from the table above. Tune from there using real incidents. If you want SLOs, burn alerts and the traces behind them in one place, [Last9](https://last9.io/) supports SLOs natively and keeps the high-cardinality detail you need to find which customer or endpoint is burning the budget.

## FAQ

### What is an error budget?

An error budget is the amount of unreliability a service is allowed within its service level objective (SLO) over a set window. It equals one minus the SLO target. A service with a 99.9% availability SLO over 30 days has an error budget of 0.1%, which is about 43 minutes of full downtime or one failed request in every thousand. Teams spend the budget on releases and experiments, and slow down when it runs out.

### How do you calculate an error budget?

An error budget is calculated as (1 minus the SLO target) multiplied by the total events or time in the SLO window. For a 99.9% SLO over 30 days, the time-based budget is 0.001 x 43,200 minutes, or 43.2 minutes. For a service that handles 10 million requests in that window, the request-based budget is 0.001 x 10,000,000, or 10,000 failed requests.

### What is burn rate in SRE?

Burn rate is how fast a service consumes its error budget relative to the SLO, as defined in Google's SRE workbook. At a burn rate of 1, the error budget runs out exactly at the end of the SLO window. On a 30-day window, a burn rate of 14.4 empties the whole budget in about 50 hours, and a burn rate of 6 empties it in about 5 days.

### What is a multiwindow, multi-burn-rate alert?

A multiwindow, multi-burn-rate alert fires only when the error budget burn rate is above a threshold over both a long window and a short window. Google's SRE workbook recommends paging at a burn rate of 14.4 over 1 hour and 5 minutes, paging at 6 over 6 hours and 30 minutes, and opening a ticket at 1 over 3 days and 6 hours. The short window lets the alert clear soon after a fix.

### What is the difference between an SLO and an error budget?

An SLO is the reliability target, such as 99.9% of requests succeeding over 30 days. The error budget is the remaining room for failure under that target, which is 0.1% of requests in that example. The SLO says how reliable a service should be, and the error budget measures how much unreliability is still available before the SLO is missed.

### What happens when an error budget runs out?

What happens when an error budget runs out is set by the team's error budget policy. In the example policy in Google's SRE workbook, a service that exceeds its error budget over the preceding four weeks halts all changes and releases except P0 issues and security fixes until it is back within its SLO. The same policy requires a postmortem when a single incident consumes more than 20% of the error budget over four weeks.
