Truncated output
Find out what the Truncated output rule checks, what it needs, and how to set it for the provider your project uses.
Definition
Truncated output flags a model call that stopped because it hit its token limit. When that happens, the call succeeds, but the answer is cut off partway through.
The rule checks each LLM span, that is, each model call. It reads the finish reason that the provider reported for the call. If the finish reason matches the value you set in the rule, the call fails the check.
Property | Value |
Key |
|
Evaluates | Span (one model call). Only LLM spans are checked. |
Check kind | Deterministic |
Category | Output integrity |
Default severity | MEDIUM |
Result | Pass or fail |
When a call fails, the check result shows:
- The observed value, which is the length of the output.
- An excerpt of the output, up to 500 characters.
- A reason such as "provider stopped at the token limit, finish_reason length".
Why it matters
A truncated answer looks like a success. The call doesn't fail, and it isn't a refusal. The user still gets an incomplete reply. This rule separates these cut-off answers from errors and refusals, so you can see how often your token limits are too tight.
What it requires
- LLM spans with a finish reason. The rule checks only LLM spans. It compares the finish reason that your application's traces report for each call.
- No model calls. This is a deterministic check. It doesn't call a model to decide, so it doesn't need a judge model.
- Permission to change rules. To create the rule, you need permission to change rules in the project. For more information, see Project permissions.
- One parameter. The rule has one required parameter:
Parameter | Type | Default | What it does |
Finish reason meaning "hit the limit" | Text, at least 1 character |
| The finish reason that means the provider stopped at the token limit |
Providers use different words for this finish reason. For example, OpenAI reports length and Anthropic reports max_tokens. Set the value your provider uses. The rule fires only when the finish reason is exactly the same as this value, so max_tokens and MAX_TOKENS count as different values.
If you leave the parameter blank, you see "Finish reason meaning "hit the limit" is required." when you try to save.
Configuration examples
To add this rule to a project:
- On the Rules page, select + New rule.
- On the New rule page, enter
truncatedin Search templates. Then on the Truncated output card, select Configure ›.
- On the Configure rule page, for Finish reason meaning "hit the limit", enter the finish reason your provider uses. To use one of the examples shown under the field, select it.
- Select Save & enable.
For the Name, Project, and Severity fields, see Create a rule.
The following table shows values to enter for Finish reason meaning "hit the limit".
If your project uses | Enter |
OpenAI |
|
Anthropic |
|
A provider that reports the reason in capitals |
|
A rule matches one value. If your project calls providers that report different finish reasons, create one rule for each value. Give each rule its own name, because rule names must be unique.
Related
- Output length out of bounds catches answers that are shorter or longer than expected, based on character count.
- Empty output catches calls that return a blank answer.
- Stopped by content filter catches calls the provider stopped on its own safety filter.
- Refusal detected catches answers where the model declined to respond.

Have a suggestion?