LLM call over token budget
This page explains what the LLM call over token budget rule checks, when it raises a finding, and how to set its token budget.
Definition
The LLM call over token budget rule (key span_token_budget) flags a single model call that uses more tokens than you budgeted. It counts input and output tokens together for one call, not for a whole request or conversation.
Property | Value |
Unit | span (one model or tool call) |
Check kind | deterministic |
Category | Operational thresholds |
Score type | BOOL |
Default severity | MEDIUM |
The rule applies to LLM spans only. Tool calls and other spans aren't checked, and no result is recorded for them.
For each LLM span, the rule fails when the call's total tokens are greater than the Token budget. Otherwise, it passes. On the configure and edit pages, the read-only How this rule evaluates panel shows this logic:
- Precondition:
span.span_kind == 'llm' - Predicate:
span.total_tokens > params.max_total_tokens
When a call fails, the result's reasoning gives the call's token use and the budget, for example: "call used 18450 tokens (16200 in, 2250 out), budget 16000".
Why it matters
One expensive call is easy to miss. A stuffed context or a runaway generation can hide inside a conversation that looks normal overall, because a per-conversation budget averages it away. This rule catches that one call on its own.
It's a deterministic check, so it makes no model calls of its own to reach a verdict.
What it requires
- Traces that contain LLM spans with token usage. To learn how token usage is recorded and shown, see Cost tracking.
- Permission to change rules in the project the rule watches. For details, see Project permissions.
The rule has one parameter.
Parameter | Required | Default | Allowed values |
Token budget | Yes |
| A whole number, at least |
The call fails when its input and output tokens together exceed this number. Set it close to your model's practical context size rather than to a spending figure.
If the value isn't valid, you see one of these messages after you select Save & enable:
- "Token budget is required."
- "Token budget must be a whole number."
- "Token budget must be at least 1."
Configuration examples
The template suggests these budgets. Select an example under the field to insert it.
Token budget | Use it to |
| Flag calls in a flow that should stay short |
| Keep the default budget |
| Flag only calls that come close to a large context |
To set up the rule:
- On the Rules page, select + New rule.
- On the New rule page, enter
span_token_budgetin Search templates, and then select Configure › on the LLM call over token budget card.
- For Token budget, enter the budget, for example
16000. - For Project, select the project to watch. Then set Name and Severity if you don't want the defaults.
- Select Save & enable.
For the full procedure, including all fields and error messages, see Create a rule.
Related
- Request over token budget checks all the model calls in one request together.
- Conversation over token budget checks the total tokens across a whole conversation.
- Output length out of bounds checks the length of a completion in characters.

Have a suggestion?