ConceptsSLOs and error budgets

SLOs and error budgets

How service level objectives translate metric status into a contractual signal.

A Service Level Objective (SLO) is a commitment that a metric will remain healthy for a defined fraction of a rolling window. SLOs turn the binary "is this healthy right now" question into a running balance: the error budget, which is the remaining allowance of unhealthy time.

Definition

An SLO has three core fields:

  • Metric: which metric the SLO observes.
  • Target percentage: the fraction of the window the metric must be healthy. Common values: 99, 99.5, 99.9, 99.95, 99.99.
  • Window in days: the rolling period the target applies to. Common values: 7, 30, 90.

A fourth field, the unknown-data policy, controls how periods without samples are counted. It is optional and defaults to the strictest reading; see Unknown-data policy below.

The window is rolling: at any instant, the SLO looks back N days and computes the fraction of that time the metric was healthy. There is no calendar boundary that resets the budget.

Window start

The rolling window never reaches back before the SLO was created. A new SLO on a metric with months of prior history measures only from its own creation time: the older samples remain visible on the metric, but they neither earn nor burn budget. Until the SLO has been observing for a full window length, the calculation covers the period since creation; after that, the window rolls normally.

Because no state survives deletion, deleting an SLO and recreating it starts a clean epoch. This is the supported way to reset an SLO after a definitional change (for example, a corrected threshold that invalidates the old verdicts).

For the first 24 hours of observation the budget reads collecting instead of a number. A percentage computed from a few hours of samples swings too widely to act on.

The budget also reads collecting whenever the metric is stale. A metric that is currently blind makes no budget claim in either direction: it neither burns nor recovers while the agent is silent. It resumes reporting a number on the first evaluation after fresh samples arrive. See Thresholds and dwell for how staleness is decided.

Error budget

Given a 99.9% target over 30 days, the budget allowance is:

allowance = 30 days * (1 - 99.9 / 100)
         = 30 days * 0.001
         = 43.2 minutes per 30-day window

The budget burns on any non-healthy sample. A sample counts as good only when it resolves to healthy; degraded and no_data samples both burn budget (the threshold operators reference covers each status). "Resolves to" matters for no_data: the SLO reads the same folded status the status page paints, so a no_data sample on a metric whose no-data treatment is healthy counts as good rather than burning. The contract number and the published page never disagree.

Periods with no samples at all are governed by the SLO's unknown-data policy, described next. Under the default policy they also burn.

The reported percentage is clamped to the range -9999.9999 to 9999.9999. An SLO that has burned many multiples of its allowance reads as a large negative number rather than an unbounded one.

Unknown-data policy

Every SLO carries an unknown-data policy that answers one question: what does an unobserved period mean? When the agent stops pushing (a network partition, a stopped agent, a decommissioned exporter), the SLO has no verdict for that time. The policy decides how that silence is counted.

PolicySilence counts asNotes
downAgainst the budget and the SLAThe strictest reading. Default.
excludeRemoved from the calculationMonitoring coverage is disclosed alongside the number, so the smaller denominator is visible. Matches the common vendor-SLA reading that downtime is a failed check, not an absent check.
upUpOnly for contracts that specify it. Not recommended.

The policy applies consistently to every surface that reports the SLO: the live error budget, the SLO page, SLA reports, and compliance evidence packs. There is one policy and one number; a report never uses a different reading than the dashboard.

In config-as-code, the slos entry accepts an optional field:

unknown_policy: down | exclude | up

Omitting the field means down. Exports include the field only when it is set to a non-default value.

Monitoring coverage

The SLO page and compliance evidence packs show monitoring coverage: the share of expected samples that carried a verdict. The expected sample count derives from the metric's push cadence.

Coverage below the expectation signals monitoring gaps rather than service downtime. Under the exclude policy, coverage is the disclosure that keeps the availability number honest: a 100% availability figure computed over 60% coverage reads very differently than the same figure over full coverage.

Burn events

A burn event opens when the error budget goes negative — the metric has spent more unhealthy time than the objective allowed. It is not triggered by an individual unhealthy sample, and not by the budget merely dropping below 100% remaining: a budget at 40% remaining is being consumed as designed and opens nothing. The event closes on the first evaluation where the budget is back above zero. Each burn event records its start, end, and the percent of the budget it consumed.

Two grace conditions gate the opening, so a young or thinly-sampled SLO cannot page on noise:

  • The SLO must be at least one hour old.
  • The window must hold at least 10 samples.

Until both hold, a negative budget is still reported on the SLO page, but no burn event opens and no slo.burn_started fires.

Webhook subscribers receive slo.burn_started when an event opens and slo.burn_resolved when it closes. Pair the two by their burn_event_id.

Recovery projection

When the budget is negative, the SLO page shows the projected date the budget returns to positive. The projection assumes no further breaches: as the window rolls forward, old unhealthy samples age out and the healthy fraction recovers on its own. The date answers the operational question "when does this SLO stop being red if nothing else goes wrong?"

Exclusion windows

An operator can exclude a specific period from SLO and SLA math. Typical cases: a misconfigured threshold that produced false unhealthy verdicts, or a telemetry outage that the unknown-data policy would otherwise count against the service.

Exclusions are deliberately constrained:

  • A written reason is required for every window.
  • Each window is capped at 30 days.
  • Creating or deleting a window is recorded in the audit log.
  • Every exclusion is listed, with its reason, in compliance evidence packs.

The constraints exist for one property: numbers never improve silently. An excluded period changes the result, and the change is visible to anyone reading the evidence.

Picking a target

The right SLO target reflects the system's actual achieved availability over the prior 90 days, plus a margin for the behaviour you want to drive. Three common starting points:

  • 99.5% for a new service or unknown baseline. Loose enough that noise does not drive false alerts.
  • 99.9% for a service with a stable history and a reasonable remediation pipeline.
  • 99.99% for systems where customers feel every minute of unhealthy time. Requires investment in error-handling and rapid remediation; otherwise the target produces churn rather than signal.

Tighter is not always better

A target tighter than the system's achieved availability burns budget on noise and trains the on-call team to ignore alerts. Start at the 90-day baseline and only tighten as the underlying system improves.

Per-customer targets

Different customers can sign different SLO targets against the same underlying metric. The model and configuration steps live in Customer scopes.