Acceldata
AIO

Last updated: Oct 06, 2026 15:42 UTC

Cost tracking

This page explains how AIO estimates what each LLM call costs, and where you'll find that cost for spans, traces, and sessions.

AIO estimates the cost of every large language model (LLM) call in US dollars. It multiplies the token counts your instrumentation reports by a price for the model. It then adds those costs into totals for traces, sessions, and filtered lists, so you can see which requests and conversations cost the most.

A call that AIO couldn't price shows no cost. That's different from a cost of $0. A blank cost means "not priced", not "free".

Model prices

Each model price has four rates, in US dollars per million tokens:

  • Input
  • Cache read
  • Cache write
  • Output

AIO gets prices from two places:

  • Built-in catalog: a read-only price list built from LiteLLM's model price data. Everyone shares it.
  • Custom prices: your tenant can set its own price for any model. A custom price replaces the built-in price for that model in every project in your tenant. Other tenants keep the built-in price.

You set custom prices through the AIO API. The UI has no page for viewing or editing prices.

How it works

  1. Your application sends an LLM span with token counts. The span reports input, output, cache read, and cache write tokens. The input count includes cached tokens. To start sending spans, see Instrument your code.
  2. AIO picks the model to price. It uses the response model, which is the model that actually served the call. If the span has no response model, AIO uses the request model. The name must exactly match a model that has a price.
  3. AIO picks the price. Your tenant's custom price wins over the built-in catalog. All four rates always come from the same source. AIO never mixes rates from a custom price and the catalog.
  4. AIO calculates the cost. It multiplies each token class by its rate, rounds each result to a whole millionth of a dollar, and adds the four parts:
  • Uncached input (input minus cache read minus cache write) × input rate
  • Cache read × cache read rate
  • Cache write × cache write rate
  • Output × output rate
  1. AIO adds up totals. A trace's cost is the sum of its priced LLM calls. A session's cost is the sum of the priced LLM calls across all its traces.

Which calls get a cost

An LLM call gets a cost only when all of the following are true:

  • It's an LLM span.
  • It reports all four token counts: input, output, cache read, and cache write.
  • Cache read plus cache write isn't more than input.
  • A price exists for the model, either a custom price or a built-in catalog entry.

Totals for traces, sessions, and lists count only priced calls. Unpriced calls add nothing, and the UI doesn't mark a total as partial. AIO calculates totals when you view them, so the total for a trace or session that's still receiving spans can change.

Where cost appears

Trace, session, and span lists

The Traces, Sessions, and Spans lists each have a Cost column that's visible by default. The cell is blank when the row isn't priced.

Each list's stats strip also has a Cost tile. It totals the whole filtered result, not just the rows on the current page. It shows "—" when nothing in the result is priced.

Traces list with the Cost column and the Cost tile in the stats strip

Sessions list with the Cost column and the Cost tile in the stats strip

Spans list with the Cost column and the Cost tile in the stats strip

To compare cost with token usage, add token columns from the column picker: Tokens, In, Out, and Cache read. On the Traces and Sessions lists, Tokens is visible by default and the others are hidden. On the Spans list, all four are hidden by default. To filter and read these lists, see Explore traces.

Trace detail

The stats strip on a trace's page has a Cost figure. It's the sum of the trace's priced LLM calls, and it shows "—" when none were priced.

Trace page with the Cost figure in the stats strip

Session detail

The stats strip on a session's page has a Cost figure. It's the sum of the priced LLM calls across all traces in the session, and it shows "—" when none were priced. The turns rail shows each turn's tokens but not its cost.

Session page with the Cost figure in the stats strip

To group traces into sessions, see Track users and sessions.

A single LLM call

Select a span in the waterfall on a trace or session page to open Span detail. On the Input / Output tab, the Cost row shows the call's cost to four decimal places, for example 0.0123 USD. The row appears only when the call has a cost. The Tokens row on the same tab shows the call's token counts, for example 1200 in · 350 out.

How cost values are shown

Lists and stats strips format cost as shown in the following table.

Cost

Shown as

Not priced

Blank in table cells; "—" in stats

Exactly 0

$0

Below $0.0001

<$0.0001

Below $1

Four decimals, for example $0.0123

Below $1,000

Two decimals, for example $12.34

$1,000 and above

Whole dollars with thousands separators, for example $1,235

The Span detail pane uses its own format: the value to four decimals followed by USD.

Next steps