Getting started
Spancache shows the cost, latency, and reliability of every AI-agent run on a losslessly compressed store you can keep forever. You point any OpenTelemetry exporter at our endpoint — there is no SDK to install — and your agent’s traces show up in the console within seconds. This page takes you from zero to seeing your first traces and cost in a few minutes.
1. Sign up and get your ingest token
- Create a workspace at spancache.ai.
- Open Overview in your workspace and copy the Ingest token — a per-tenant bearer token that authorizes your traces. Keep it secret; treat it like an API key.
At onboarding you choose a capture level. Spancache is metadata-first by default: it records span timing, model, token counts, and cost — but no prompt or response content. Capturing content is an explicit opt-in, and when you turn it on, Spancache scrubs secrets at ingest before anything is stored. You can change this later in your workspace settings. See console for where to find it.
2. Connect your agent
Point an OpenTelemetry exporter at https://ingest.spancache.ai/v1/traces (OTLP/HTTP) and send your token as a bearer header. No code changes, no SDK.
Claude Code
Set these environment variables, then run Claude Code as usual:
export CLAUDE_CODE_ENABLE_TELEMETRY=1export CLAUDE_CODE_ENHANCED_TELEMETRY_BETA=1export OTEL_TRACES_EXPORTER=otlpexport OTEL_LOGS_EXPORTER=otlpexport OTEL_EXPORTER_OTLP_PROTOCOL=http/jsonexport OTEL_EXPORTER_OTLP_ENDPOINT=https://ingest.spancache.aiexport OTEL_EXPORTER_OTLP_HEADERS="Authorization=Bearer <YOUR_INGEST_TOKEN>"Each Claude Code run becomes one trace, with a span per model call and tool step.
Any OpenTelemetry app
If your service already emits OTLP traces, point the traces exporter at the endpoint:
export OTEL_TRACES_EXPORTER=otlpexport OTEL_EXPORTER_OTLP_TRACES_ENDPOINT=https://ingest.spancache.ai/v1/tracesexport OTEL_EXPORTER_OTLP_TRACES_HEADERS="Authorization=Bearer <YOUR_INGEST_TOKEN>"For Python (OpenLLMetry), Node, the Vercel AI SDK, LangChain, LlamaIndex, and a one-line cURL smoke test, see sending-traces.
3. See your traces
Open the console at app.spancache.ai and run your agent once.
- Traces — one row per agent run, with its total cost and duration. Click a row to open the span waterfall: every LLM, tool, and agent step with its own per-span cost and latency.
- Cost — spend rolled up per model and per version. Spancache computes cost from tokens × list price, cache-aware, for every span — so even sources that report no cost (like Claude Code on a subscription) get accurate per-call spend. See cost.
Want to slice it yourself? Use the query bar: <search> | stats <aggs> by <field>, e.g.
kind=span | stats sum(cost) by modelMore in query.
4. Add a guardrail
Catch regressions before they cost you. In the console:
- Go to Detections → Templates.
- Enable a template in one click — e.g. Runaway LLM spend — to get alerted when a run’s cost crosses a threshold.
The Anomalies view surfaces cost outliers automatically.
Next steps
| Go here | For |
|---|---|
| sending-traces | Every exporter recipe — Python/OpenLLMetry, Node, cURL, no-SDK setups. |
| cost | How cache-aware cost is computed, and the per-model / per-version rollups. |
| console | Traces, Cost, Detections, and Anomalies in depth. |
| query | The `search |