Query
Everything in Spancache is a span — one LLM call, tool call, or agent step — and every span belongs to a trace, one agent run end to end. You explore by searching spans and aggregating them. Both the Traces view and the Cost view in the console are built on this same engine; there is no separate index to maintain.
Searches run directly over the losslessly compressed store — no content is decompressed to filter,
and results merge exactly across the fleet. There is no sampling and no approximation: a count
is a true count, a sum(cost) is every dollar, a p95 is the real 95th percentile.
Queryable fields
Every span carries these fields (populated from the OpenTelemetry gen_ai.* attributes you
send). All are directly searchable and groupable:
| Field | Meaning |
|---|---|
kind | span — the only record type |
trace_id | the agent run this span belongs to |
service.name | the app or agent that emitted the span |
service.version | release/build identifier |
model | model id, e.g. claude-opus-4-8 |
provider | anthropic, openai, … |
op / gen_ai.operation.name | operation, e.g. chat, embeddings, tool |
tool | tool name, for tool-call spans |
status | ok or error |
error.type | exception/class name when status=error |
cost | span cost in USD |
cost_micros | span cost in millionths of a USD (integer) |
input_tokens | prompt tokens |
output_tokens | completion tokens |
cache_read_tokens | tokens served from prompt cache |
cache_creation_tokens | tokens written to prompt cache |
total_tokens | all tokens for the span |
duration_ms | wall-clock span duration |
ttft_ms | time to first token |
attempt | retry attempt number |
stop_reason | why generation stopped (end_turn, max_tokens, …) |
Cost fields are cache-aware and computed from tokens × list price on every span — see cost.
Search syntax
- Equality:
model=claude-opus-4-8,status=error,provider=anthropic. - Comparisons:
cost>1,duration_ms>30000,cache_read_tokens>500000,input_tokens>=10000—>,<,>=,<=on any numeric field. - Boolean: combine terms with
AND/OR; multiple bare termsANDtogether. - Free text: a bare word matches anywhere in the span (names, attributes, captured content when opted in).
- Time range is chosen in the UI (last 15m, 24h, custom) — you don’t put it in the query string.
The pipeline
Chain a search into aggregations and shaping with |:
<search> | stats <aggs> by <field> | sort <field> | head NAggregations: count, sum, avg, min, max, p50, p95, p99, dc (distinct count).
sort -<field> sorts descending; head N keeps the top N rows. Every aggregation is computed
exactly across the whole fleet.
Examples
kind=span | stats sum(cost) as cost, count as calls by model | sort -costSpend and call volume by model, most expensive first.
kind=span status=error | stats count by toolWhich tool fails most often.
kind=span | stats p95(duration_ms) by service.versionLatency regression across releases — watch a new deploy move the tail.
kind=span | stats sum(cost) by trace_id | sort -cost | head 20The 20 most expensive agent runs.
error.type=ToolExecutionErrorJump straight to failing tool calls.
See sending-traces to get spans flowing, cost for how spend is priced, and the console for the views built on this engine.