Overview
This page explains what rules check in your AI telemetry, when they run, and how to read the verdict that each check produces.
Rules run continuous checks on live traffic from your application. For example, a rule can catch a blank completion, a banned term, a refusal, or a rise in latency. Each rule watches one project. When a check finds a problem, the failure links back to the trace that caused it. You see problems as they happen, without reading traces one by one.
What a rule checks
Each rule starts from a built-in template in the rule catalog. The template sets what the rule checks and how it decides. You choose the template's parameters, the project, and a severity of LOW, MEDIUM, HIGH, or CRITICAL.
What a rule evaluates
Each template evaluates one kind of item, called its unit. The unit always comes from the template, and you can't change it.
Unit | What it evaluates |
| One model or tool call |
| One request |
| One conversation |
| Many traces together, over a time window |
How a rule decides
Each template also has a check kind, which sets how the rule reaches a decision.
Check kind | How it decides |
Deterministic | Tests the item's recorded values against a condition. No model call. |
Judge | A model gives the item a score. |
Classifier | A model assigns the item a label. |
Aggregate | Measures a value across a time window and compares it with a baseline. |
Deterministic, judge, and classifier rules make their decision in two parts:
- Precondition: decides whether the rule applies to the item at all. For example, many rules apply only to LLM calls. If the rule doesn't apply, no result is recorded. It doesn't count as a pass or as a skip.
- Predicate: the condition that marks a problem. When it's true, the rule records a failure, also called a finding. When it's false, the rule records a pass.
Aggregate rules work differently. To learn how they compare a window with its baseline, see Latency regression.
For what each template checks and which parameters it takes, see the rule catalog, starting with Empty output.
What a verdict is
A verdict is the result of one rule on one span, trace, or session. It's one of the following:
Verdict | Meaning |
Pass | The check ran and found no problem. |
Fail | The check found a problem. A failure is a finding. |
Error | The check couldn't decide. For example, a judge model that can't be reached gives an error. An error doesn't count as a pass or as a failure. |
Skipped | The check didn't run on this item. A skipped result always includes the reason. |
A failure can include evidence to help you understand it:
- Observed value: the measured value, such as the length of a completion.
- Excerpt: a piece of the text that caused the failure, up to 500 characters.
- Summary: a sentence that explains the finding, such as "completion is empty".
Rules that produce a numeric score, such as Groundedness, also show that score with each result.
When rules run
Rules run in the background on new telemetry. You don't trigger them yourself.
- Start: saving a rule enables it right away. AIO starts evaluating it within 30 seconds.
- New telemetry: AIO looks for new spans, traces, and sessions every 30 seconds.
- Existing traffic: when a project has no rule evaluation yet, evaluation starts 24 hours back. So you see findings on recent traffic right away, without waiting for new traffic.
- Settle time: AIO waits a short time after an item arrives before it checks the item. This gives a trace or conversation time to finish arriving.
- Updated results: as more of a trace or conversation arrives, a deterministic rule can check it again. The new result replaces the earlier one.
- Judge and classifier rules: these rules need a model to run. If no model API key is set up for rule evaluation, they don't run, and their results stay "pending".
- Disabled rules: a disabled rule isn't evaluated. Its findings are kept.
How it works
- Create a rule. Pick a template from the catalog, set its parameters, and choose a project and a severity. Then save the rule. For the steps, see Create a rule.
- AIO checks new traffic. For every new span, trace, or session the rule applies to, AIO records a verdict.
- Review rules on the Rules page. The rules list shows each rule's template, unit, and severity. Failed 24h shows how many times the rule failed in the last 24 hours. Last failed shows how long ago it last failed.
- Read verdicts on traces, conversations, and steps. The trace, conversation, and step lists each have a Rules column. Select a cell to open Check results, which lists each rule's verdict, score, reasoning, and severity. A cell shows one of the following:
Cell | Meaning |
"2 failed · 3 passed" | At least one check failed on this row. |
"couldn't run" | Checks gave an error. |
"✓ 4 passed" | Every check that ran passed. "so far" after it means some checks haven't reached this row yet. |
"pending" | Checks haven't reached this row yet, and no results exist. |
"—" | No checks of this kind have run in this project. |
- Open the details. On a detail page, the Checks panel lists the latest result of each rule. Failures appear first, ordered by severity, then errors, then passes. The panel opens automatically when a check failed or gave an error.
- Find failing traces. On a list, use the Rules filter and select Any failed rule, or select one failing rule. The list then shows only traces with failing checks, most recent failure first. To filter and read traces in other ways, see Explore traces.
To see rules and their verdicts, you need permission to view the project. To create or change rules, you need permission to modify it. For details, see Project permissions.

Have a suggestion?