Acceldata
AIO

Last updated: Oct 06, 2026 15:10 UTC

Latency regression

This page explains what the Latency regression rule checks, when it raises a finding, and how to set its parameters.

Definition

Latency regression (latency_regression) is an aggregate rule. It checks whether the 95th percentile (p95) latency of your traces has gone up compared with their own recent history.

Property

Value

Unit

window (an aggregate across many traces over a time window)

Check kind

Aggregate

Category

Operational thresholds

Score type

Numeric

Default severity

HIGH

The rule doesn't judge traces one at a time. Instead, it measures p95 trace latency over a time window and compares it with a baseline period. It makes this comparison separately for each cohort. A cohort is a project, request model, and prompt version taken together.

Each time the rule runs, it works like this:

  1. It takes the most recently closed evaluation window. Windows sit on a fixed time grid.
  2. It measures p95 latency for each cohort in that window.
  3. It compares each cohort with its baseline, which is the period immediately before the window.
  4. It raises one finding for each cohort whose p95 is more than your threshold percentage above its baseline.

The following table shows what each outcome records.

Outcome

What's recorded

A cohort rises above the threshold

One failed verdict for that cohort

A cohort stays within the threshold

Nothing. No passed verdict is written.

A cohort has no baseline, or its baseline is zero

Nothing. This is never a finding.

The rule can't be evaluated

One error verdict for the rule, with the failure message

Each finding includes these details:

  • Score: the measured percent change.
  • Observed value: the cohort's p95 latency in the window.
  • Excerpt: the cohort label.
  • Reasoning: a sentence that names the cohort and gives the window value, the baseline value, the percent change, the threshold, and how many rows were measured.

Why it matters

Latency often gets worse slowly, or only for one model or one prompt version. A fixed latency budget can miss that. This rule compares each cohort with its own past, so it catches a slowdown even when overall latency still looks normal. It also catches a slowdown after you switch models or release a new prompt version.

What it requires

  • Permission: You need permission to modify rules on the project. For details, see Project permissions.
  • Traces: The project needs to send traces. For setup, see Instrument your code.
  • History: A cohort needs traffic in the baseline period. A cohort with no baseline never raises a finding, so a new model or prompt version is only checked after it has some history.

This rule has no Sampling setting. It always measures all traffic in the window.

Configuration examples

To find this template, go to the New rule page. Select the Aggregate chip, and then select Configure › on the Latency regression card.

The New rule template catalog, showing template cards that you can filter by kind and category

On the Configure rule page, set Name, Project, and Severity as described in Create a rule. Severity starts at HIGH for this template. The read-only "How this rule evaluates" panel shows the metric, how it's grouped, and the comparison the rule uses.

The Configure rule form for Latency regression, with fields for Evaluation window, Baseline window, and Percent increase to alert on

Parameters

All three parameters are required.

Parameter

Type

Default

Allowed values

What it controls

Evaluation window

Text

1h

A whole number followed by m, h, or d, for example 15m, 1h, or 6h

How much recent traffic each evaluation covers

Baseline window

Text

7d

Same format, for example 24h, 7d, or 30d

The period before the window that the window is compared with

Percent increase to alert on

Number

25

1 to 1000

How far above the baseline p95 can rise, in percent, before the rule raises a finding

Examples

The following table shows some ways to combine the parameters.

Goal

Evaluation window

Baseline window

Percent increase to alert on

Use the template defaults

1h

7d

25

Catch large, sudden slowdowns quickly

15m

24h

50

Catch smaller, sustained slowdowns

6h

30d

10

When you're done, select Save & enable. The rule starts working right away.

Troubleshooting

When a parameter value isn't valid, the form shows a message after you try to save. For example:

  • "Evaluation window is required."
  • "Percent increase to alert on must be a number."
  • "Percent increase to alert on must be at least 1."
  • "Percent increase to alert on cannot exceed 1000."

If a cohort you expect to see never produces a finding, check whether it had traffic during the baseline period. A cohort with no baseline is never flagged.

Related

Next steps