Rules
Rules tell AIO what a good trace should look like. After telemetry is ingested, AIO evaluates each enabled rule and records a verdict. Use verdicts to find empty answers, refusals, budget overruns, and missing tool calls. A rule belongs to one project. Open a project and go to the Rules tab.
How a rule works
- Click + New rule and pick a template.
- Configure name, severity, optional environment, sample rate, and template fields.
- Save enabled. Disabled rules are not evaluated.
- The AIO rules service evaluates enabled rules on ingested traces, spans, and sessions.
- Verdicts appear on the matching entity. On Traces and Steps, Checks opens the per-rule result. If a rule does not apply to an entity, AIO records no verdict.
What you can set
- Template: Chosen at creation (cannot change after create).
- Name and description: Identifies the rule purpose and configuration.
- Enabled toggle: Controls whether the rule is evaluated during ingestion.
- Environment: Target execution environment (empty = all).
- Severity: Assigned impact level (Low, Medium, High, Critical).
- Sample rate: Proportion of traffic evaluated, from 0 to 1 (judge rules cost a model call per sampled entity).
- Template parameters: Specific variables defined for the chosen template.
The rules list view shows Fired 24h and Last fired tracking indicators.
Kinds of checks
- Deterministic: Simple yes/no validation with no model call required.
- Aggregate: Windowed metric measurement with no sample rate required.
- Judge: LLM scores the entity and needs a judge model; groundedness also needs retrieval documents.
Units
- Span: Represents one individual step.
- Trace: Represents one full request lifecycle.
- Session: Represents a multi-turn conversation context.
- Window: Represents a designated time window for aggregate rules.
Template catalog
Output integrity
(Applies to Span level unless noted)
Template Name | Evaluation Level | Type / Details | Severity |
|---|---|---|---|
Empty output | Span | Deterministic | Medium |
Output length out of bounds | Span | Deterministic | Low |
Truncated output | Span | Deterministic | Medium |
Stopped by content filter | Span | Deterministic | High |
Refusal detected | Span | Deterministic | Low |
Banned term in output | Span | RE2 pattern match | High |
Missing citation | Span | Deterministic | Medium |
Expected format absent | Span | Deterministic | Medium |
Groundedness | Trace | Judge (LLM as Judge) | High |
Operational thresholds
Template Name | Evaluation Level | Type / Details | Severity |
|---|---|---|---|
Latency regression | Window | Aggregate | High |
Trace contains failed spans | Trace | Deterministic | Medium |
Trace slower than budget | Trace | Deterministic | Low |
Conversation had failed turns | Session | Deterministic | Medium |
Conversation over token budget | Session | Deterministic | Medium |
LLM call over token budget | Span | Deterministic | Medium |
Request over token budget | Trace | Deterministic | Medium |
Tool and agent structure
Template Name | Evaluation Level | Type / Details | Severity |
|---|---|---|---|
Agent answered without using a tool | Trace | Deterministic | Medium |
Agent behaviour
Template Name | Evaluation Level | Type / Details | Severity |
|---|---|---|---|
Conversation ran long | Session | Deterministic | Low |
Hygiene and governance
Template Name | Evaluation Level | Type / Details | Severity |
|---|---|---|---|
Personal data in completion | Span | Classifier (not regex) | High |
What's next
Send traces first, add one or two deterministic rules, then add a judge rule after retrieval content and a judge model are available.

Have a suggestion?