Skip to content

Cost & retention — price every span, keep every trace

The thesis: you should know what every agent run costs — and still be able to afford to keep it. Spancache prices each span itself, from its token counts and the model’s list price, so you get accurate per-call, per-trace, per-model and per-user cost even when the source reports no dollar figure at all. Then it stores those traces on a losslessly compressed, never-sampled store, so keeping every run for the long term costs a fraction of typical observability retention.

1. Spancache prices the span

Most traces carry token counts but not money. A Claude Code session on a Pro or Max subscription, for example, reports input/output/cache tokens on every step — and no cost, because the seat is flat-rate. Dashboards that only read a cost attribute show $0.00 for that whole run.

Spancache doesn’t wait for the source to tell it the price. For every span it recognises as an LLM call, it reads the token usage off the OpenTelemetry attributes, looks up the model’s list price, and computes the cost itself. That gives you a real number per call, which rolls up exactly to per-trace, per-model and per-user spend — whether or not anything upstream reported a dollar.

2. Cache-aware pricing (the detail that matters)

Modern agents are cache-heavy. A long-running coding or research agent replays a large, stable prompt prefix on every step, so prompt-cache reads usually dominate the bill — and a pricing model that lumps all input tokens together gets the number badly wrong.

Spancache prices the four token streams separately, at their real list multipliers:

Token streamPriced at (Anthropic-style)
Input (uncached)1× input list rate
Outputoutput list rate
Cache read~0.1× input rate
Cache creation (write)~1.25× input rate

Worked example. An Opus agent run with roughly 2 fresh input tokens, 750 output tokens, 5.4M cache-read tokens and 20k cache-creation tokens prices out to about $9.58 in list-price terms — and the cache reads alone are ~$8.10 of that. Price the same run by its handful of fresh input tokens, as a naive meter would, and you’d miss almost the entire cost.

This is an estimate, not an invoice. The figure is computed at public list price. Each span carries a cost_source field that flags whether the cost was estimated by Spancache or taken from a value the source reported, so you always know which is which. On a subscription seat (like Claude Code on Pro/Max) the estimate is a shadow cost — a showback figure for what the run would cost at list — not what you were billed.

3. The Cost view

The console (console) rolls this up into per-model and per-version views of your fleet’s spend. For each one it shows spend, call volume, error rate, p95 latency and token counts side by side:

  • Per model — compare models on cost and behaviour in the same row. A switch from one model to another shows up as a step in spend or latency immediately, not three weeks later on an invoice.
  • Per version — group by prompt or agent release. A new prompt template or agent build that quietly got more expensive, slower, or more error-prone is visible the moment traffic hits it — before your users feel it or finance asks about the bill.

Because every value derives from the same per-span cost, the rollups reconcile exactly: the sum of the model rows equals the sum of the version rows equals your total.

4. Showback and chargeback

The same per-span attribution answers “who spent what.” Group spend by model, by user, or by trace to allocate cost back to the team, customer, or feature that drove it — internal showback, or chargeback when another team foots the bill. Slice it in the query surface the same way you’d slice latency or errors; cost is just another dimension of the span.

5. Retention economics

Pricing every span is only half the story. The other half is being able to keep the traces that prove where the money went.

Typical trace tooling makes you choose: short retention, head/tail sampling, or a punishing bill. Every one of those throws away exactly the history you later wish you had — the week the cache-hit rate collapsed, the slow regression that crept in over a month, the one expensive run a customer is asking about.

Spancache removes the choice. The store is losslessly compressed and never sampled, so:

  • Keep every trace, for the long term, at a fraction of typical observability cost. Strong compression means the stored bytes behind a month (or a year) of traces are small, so retention is a storage decision, not a budget fight.
  • No sampling means the cost rollups are complete. Per-model and per-user spend is computed over all your runs, not an extrapolation from a sampled slice — the number is the number.
  • History you’d otherwise drop stays queryable. Answering “what did this agent cost us last quarter, and why” doesn’t require a restore or a re-ingest; the traces are still there and still searchable.

6. What Spancache itself costs

Spancache bills in flat capacity tiers, set by your sustained ingest rate and the volume of data you keep — not per trace, per query, or per seat. Within a tier, queries, retention, users and dashboards are unlimited; you pick the smallest tier that covers your throughput and storage. Tiers run free → pro → team → BYOC (bring-your-own-cloud), so retention scales with storage cost rather than with how many questions you ask of your data.


See also: console (where the Cost view lives), query (slice spend by model, user, or trace), sending-traces (get your runs into the store).