Request over token budget
This page explains how the Request over token budget rule works: it flags a request whose model calls together used more tokens than you allow, and you set that limit yourself.
Definition
Request over token budget (trace_token_budget) is a built-in rule template. It adds up the tokens from every model call in one request (one trace). The rule fails the trace when that total is greater than the budget you set.
Property | Value |
Unit |
|
Check kind | Deterministic: no model call, so the check doesn't cost anything to run |
Category | Operational thresholds |
Score type | Pass or fail |
Default severity |
|
Applies to | Traces with at least one LLM call |
The rule checks only traces that contain at least one LLM call. A trace with no LLM calls gets no result from this rule, so it doesn't count as a pass or a failure.
When a trace goes over the budget, its reasoning gives the token count, the number of calls, and the budget. For example: "request used 62400 tokens across 7 call(s), budget 50000". This text appears in the Check results popover and the Checks panel for the trace. For more about where results appear, see Explore traces.
Why it matters
Agent loops often show up here first. Each model call can look reasonable on its own, but together they make the request too expensive. Neither a per-call budget nor a per-conversation budget catches this case well:
- LLM call over token budget checks one model call at a time. It misses a request made up of many average-sized calls.
- Conversation over token budget checks a whole conversation. It's the right place to catch slow, steady growth over many requests.
Set this budget higher than a single call's budget and lower than a conversation's budget. A request is made up of several calls, and a conversation is made up of several requests.
What it requires
- Traces with LLM calls and token usage. The rule reads the total tokens of each request. If a trace has no token counts, its total is read as
0, so the trace never goes over the budget. To learn how token usage gets into traces, see Instrument your code and Cost tracking. - Permission to change rules in the project. To create this rule, you need permission to change rules in the project. For details, see Project permissions.
- No model API key. This deterministic rule doesn't call a model.
Parameters
The template has one parameter.
Parameter | Required | Default | Allowed values | What it does |
Token budget | Yes |
| Whole number, at least | Fails the request when its calls together use more tokens than this |
The rule fails only when the total is greater than the budget. A request that uses exactly the budget passes.
If the value isn't valid, the form shows one of these messages when you select Save & enable:
- "Token budget is required."
- "Token budget must be a whole number."
- "Token budget must be at least 1."
Configuration examples
To add this rule, go to the New rule page. Select Deterministic, or search for trace_token_budget, and then select Configure › on the Request over token budget card.
On the Configure rule page, enter a value for Token budget, choose a Project and a Severity, and then select Save & enable. For the full procedure, see Create a rule.
The following table shows example budgets for different kinds of workloads.
Token budget | Use it when |
| Requests are short, with one or two calls and small prompts. You want to catch any request that grows past that. |
| You want the default: requests make several calls, and you're looking for agent loops. |
| Requests are expected to be large, for example agent workflows with long context or many tool rounds. You want to catch only extreme outliers. |
Related
- Rules overview: what rules check, when they run, and what a verdict is
- LLM call over token budget: the same kind of check on a single model call
- Conversation over token budget: the same kind of check on a whole conversation
- Trace slower than budget: a time budget for the whole request

Have a suggestion?