Skip to content

Getting started

Spancache shows the cost, latency, and reliability of every AI-agent run on a losslessly compressed store you can keep forever. You point any OpenTelemetry exporter at our endpoint — there is no SDK to install — and your agent’s traces show up in the console within seconds. This page takes you from zero to seeing your first traces and cost in a few minutes.

1. Sign up and get your ingest token

  1. Create a workspace at spancache.ai.
  2. Open Overview in your workspace and copy the Ingest token — a per-tenant bearer token that authorizes your traces. Keep it secret; treat it like an API key.

At onboarding you choose a capture level. Spancache is metadata-first by default: it records span timing, model, token counts, and cost — but no prompt or response content. Capturing content is an explicit opt-in, and when you turn it on, Spancache scrubs secrets at ingest before anything is stored. You can change this later in your workspace settings. See console for where to find it.

2. Connect your agent

Point an OpenTelemetry exporter at https://ingest.spancache.ai/v1/traces (OTLP/HTTP) and send your token as a bearer header. No code changes, no SDK.

Claude Code

Set these environment variables, then run Claude Code as usual:

Terminal window
export CLAUDE_CODE_ENABLE_TELEMETRY=1
export CLAUDE_CODE_ENHANCED_TELEMETRY_BETA=1
export OTEL_TRACES_EXPORTER=otlp
export OTEL_LOGS_EXPORTER=otlp
export OTEL_EXPORTER_OTLP_PROTOCOL=http/json
export OTEL_EXPORTER_OTLP_ENDPOINT=https://ingest.spancache.ai
export OTEL_EXPORTER_OTLP_HEADERS="Authorization=Bearer <YOUR_INGEST_TOKEN>"

Each Claude Code run becomes one trace, with a span per model call and tool step.

Any OpenTelemetry app

If your service already emits OTLP traces, point the traces exporter at the endpoint:

Terminal window
export OTEL_TRACES_EXPORTER=otlp
export OTEL_EXPORTER_OTLP_TRACES_ENDPOINT=https://ingest.spancache.ai/v1/traces
export OTEL_EXPORTER_OTLP_TRACES_HEADERS="Authorization=Bearer <YOUR_INGEST_TOKEN>"

For Python (OpenLLMetry), Node, the Vercel AI SDK, LangChain, LlamaIndex, and a one-line cURL smoke test, see sending-traces.

3. See your traces

Open the console at app.spancache.ai and run your agent once.

  • Traces — one row per agent run, with its total cost and duration. Click a row to open the span waterfall: every LLM, tool, and agent step with its own per-span cost and latency.
  • Cost — spend rolled up per model and per version. Spancache computes cost from tokens × list price, cache-aware, for every span — so even sources that report no cost (like Claude Code on a subscription) get accurate per-call spend. See cost.

Want to slice it yourself? Use the query bar: <search> | stats <aggs> by <field>, e.g.

kind=span | stats sum(cost) by model

More in query.

4. Add a guardrail

Catch regressions before they cost you. In the console:

  1. Go to Detections → Templates.
  2. Enable a template in one click — e.g. Runaway LLM spend — to get alerted when a run’s cost crosses a threshold.

The Anomalies view surfaces cost outliers automatically.

Next steps

Go hereFor
sending-tracesEvery exporter recipe — Python/OpenLLMetry, Node, cURL, no-SDK setups.
costHow cache-aware cost is computed, and the per-model / per-version rollups.
consoleTraces, Cost, Detections, and Anomalies in depth.
queryThe `search