Latency regression
This page explains what the Latency regression rule checks, when it raises a finding, and how to set its parameters.
Definition
Latency regression (latency_regression) is an aggregate rule. It checks whether the 95th percentile (p95) latency of your traces has gone up compared with their own recent history.
Property | Value |
Unit |
|
Check kind | Aggregate |
Category | Operational thresholds |
Score type | Numeric |
Default severity | HIGH |
The rule doesn't judge traces one at a time. Instead, it measures p95 trace latency over a time window and compares it with a baseline period. It makes this comparison separately for each cohort. A cohort is a project, request model, and prompt version taken together.
Each time the rule runs, it works like this:
- It takes the most recently closed evaluation window. Windows sit on a fixed time grid.
- It measures p95 latency for each cohort in that window.
- It compares each cohort with its baseline, which is the period immediately before the window.
- It raises one finding for each cohort whose p95 is more than your threshold percentage above its baseline.
The following table shows what each outcome records.
Outcome | What's recorded |
A cohort rises above the threshold | One failed verdict for that cohort |
A cohort stays within the threshold | Nothing. No passed verdict is written. |
A cohort has no baseline, or its baseline is zero | Nothing. This is never a finding. |
The rule can't be evaluated | One error verdict for the rule, with the failure message |
Each finding includes these details:
- Score: the measured percent change.
- Observed value: the cohort's p95 latency in the window.
- Excerpt: the cohort label.
- Reasoning: a sentence that names the cohort and gives the window value, the baseline value, the percent change, the threshold, and how many rows were measured.
Why it matters
Latency often gets worse slowly, or only for one model or one prompt version. A fixed latency budget can miss that. This rule compares each cohort with its own past, so it catches a slowdown even when overall latency still looks normal. It also catches a slowdown after you switch models or release a new prompt version.
What it requires
- Permission: You need permission to modify rules on the project. For details, see Project permissions.
- Traces: The project needs to send traces. For setup, see Instrument your code.
- History: A cohort needs traffic in the baseline period. A cohort with no baseline never raises a finding, so a new model or prompt version is only checked after it has some history.
This rule has no Sampling setting. It always measures all traffic in the window.
Configuration examples
To find this template, go to the New rule page. Select the Aggregate chip, and then select Configure › on the Latency regression card.
On the Configure rule page, set Name, Project, and Severity as described in Create a rule. Severity starts at HIGH for this template. The read-only "How this rule evaluates" panel shows the metric, how it's grouped, and the comparison the rule uses.
Parameters
All three parameters are required.
Parameter | Type | Default | Allowed values | What it controls |
Evaluation window | Text |
| A whole number followed by | How much recent traffic each evaluation covers |
Baseline window | Text |
| Same format, for example | The period before the window that the window is compared with |
Percent increase to alert on | Number |
| 1 to 1000 | How far above the baseline p95 can rise, in percent, before the rule raises a finding |
Examples
The following table shows some ways to combine the parameters.
Goal | Evaluation window | Baseline window | Percent increase to alert on |
Use the template defaults |
|
|
|
Catch large, sudden slowdowns quickly |
|
|
|
Catch smaller, sustained slowdowns |
|
|
|
When you're done, select Save & enable. The rule starts working right away.
Troubleshooting
When a parameter value isn't valid, the form shows a message after you try to save. For example:
- "Evaluation window is required."
- "Percent increase to alert on must be a number."
- "Percent increase to alert on must be at least 1."
- "Percent increase to alert on cannot exceed 1000."
If a cohort you expect to see never produces a finding, check whether it had traffic during the baseline period. A cohort with no baseline is never flagged.
Related
- Trace slower than budget checks each trace against a fixed latency budget. Latency regression compares p95 latency with a cohort's own baseline.
- Request over token budget is another operational threshold that's checked once per request.

Have a suggestion?