Skip to content

Spancache overview

AI agents are expensive, slow, and unreliable in ways you can’t see. A single run fans out into dozens of model calls, tool calls, and retries — and when it costs 40× what you expected, stalls for nine seconds, or quietly returns truncated output, the run is a black box. Worse, most observability tools sample traces or expire them after a few weeks, so by the time you notice a regression the history that would explain it is already gone.

Spancache is agent and LLM trace observability on a losslessly compressed store you keep forever. Send your traces once; see the cost, latency, and reliability of every run; and keep the full history at a fraction of typical retention cost — no sampling, no 30-day cliff.

The data model: spans and traces

  • A span is one unit of work: a single LLM call, a tool call, or an agent step.
  • A trace is one agent run — a tree of spans from the top-level request down through every model call and tool it made.

Every span carries the fields that matter for cost, performance, and reliability:

FieldWhat it tells you
cost, cost_sourcespend for the call, and whether it was source-reported or estimated
tokens (input / output / cache-read / cache-write)token usage, including prompt-cache hits
latency, ttfttotal duration and time to first token
modelthe model that served the call
toolthe tool invoked, for tool-call spans
status, errorsuccess or failure, with the error detail
stop_reasonwhy generation ended (e.g. end_turn, max_tokens)
service.versionthe release that produced the run — group cost and errors by deploy

How it works, end to end

No SDK to install. Spancache speaks OpenTelemetry, so any OTLP exporter you already have works.

  1. Send OTLP traces. Sign up at spancache.ai and copy your per-tenant ingest token from the workspace. Point any OpenTelemetry exporter at https://ingest.spancache.ai/v1/traces (OTLP/HTTP) with the header Authorization: Bearer <token>. Works with Claude Code, the OpenAI and Anthropic SDKs (via OpenLLMetry), LangChain, LlamaIndex, and the Vercel AI SDK.
  2. They’re priced and stored. On ingest, every span is priced (cache-aware, see below), scrubbed for secrets if content is captured, and written to the compressed store — kept in full, not sampled.
  3. Explore in the console. Open app.spancache.ai to browse runs, roll up cost, arm detections, and triage anomalies.

See getting-started.md for the five-minute setup and sending-traces.md for exporter recipes.

What you get in the console

  • Traces — one row per agent run; click into the span waterfall for per-span cost, latency, tool calls, and (if you opted in) the conversation.
  • Cost — per-model and per-release-version rollups of spend, latency, error rate, and tokens, so you can spot a regression right after a prompt or model change.
  • Detections — one-click guardrail templates for agent failure modes: runaway cost, tool-error loop, error-rate spike, output truncation, retry storm, slow first token, context burn, slow agent step.
  • Anomalies — cost-outlier traces (runaway loops, cache misses), novel agent behavior, and incompressible output that can signal prompt injection or exfiltration.

More in console.md.

Query language

Beyond the surfaces above, every field is queryable with a search plus a pipeline:

<search> | stats <aggs> by <field>

For example, spend by model across all runs:

kind=span | stats sum(cost) by model

What makes Spancache different

  • Accurate cost on every call. Spancache computes cost from tokens × list price, cache-aware (Anthropic cache reads bill at ~0.1× input, cache writes at ~1.25×), for every span. So even sources that report no cost — like Claude Code on a subscription — get accurate per-call spend, and a cost_source field flags estimated versus source-reported values. See cost.md.
  • Keep everything, cheaply. The losslessly compressed store makes long retention affordable, so you keep full-fidelity history instead of sampling it away or letting it expire.
  • Metadata-first governance. By default Spancache captures only metadata — cost, tokens, latency, models, tools, outcomes — and no prompt or response content. Content capture is an explicit opt-in, and anything captured is scrubbed for secrets at ingest (API keys, tokens, JWTs, and private keys are replaced with [REDACTED]). Strict per-tenant isolation; TLS everywhere.
  • OTLP-native, no lock-in. Standards-based ingest means no proprietary SDK and no agent to run — point your existing exporter and go.

Who it’s for

Developers and LLM engineers who ship agents and need to understand what each run costs, why it was slow, and why it failed — and the security teams who need those traces governed and retained. Spancache runs as a SaaS product across capacity tiers (free, pro, team, and BYOC enterprise); within a tier, nodes, seats, queries, and retention are unlimited.

Next steps