Spancache overview
AI agents are expensive, slow, and unreliable in ways you can’t see. A single run fans out into dozens of model calls, tool calls, and retries — and when it costs 40× what you expected, stalls for nine seconds, or quietly returns truncated output, the run is a black box. Worse, most observability tools sample traces or expire them after a few weeks, so by the time you notice a regression the history that would explain it is already gone.
Spancache is agent and LLM trace observability on a losslessly compressed store you keep forever. Send your traces once; see the cost, latency, and reliability of every run; and keep the full history at a fraction of typical retention cost — no sampling, no 30-day cliff.
The data model: spans and traces
- A span is one unit of work: a single LLM call, a tool call, or an agent step.
- A trace is one agent run — a tree of spans from the top-level request down through every model call and tool it made.
Every span carries the fields that matter for cost, performance, and reliability:
| Field | What it tells you |
|---|---|
cost, cost_source | spend for the call, and whether it was source-reported or estimated |
tokens (input / output / cache-read / cache-write) | token usage, including prompt-cache hits |
latency, ttft | total duration and time to first token |
model | the model that served the call |
tool | the tool invoked, for tool-call spans |
status, error | success or failure, with the error detail |
stop_reason | why generation ended (e.g. end_turn, max_tokens) |
service.version | the release that produced the run — group cost and errors by deploy |
How it works, end to end
No SDK to install. Spancache speaks OpenTelemetry, so any OTLP exporter you already have works.
- Send OTLP traces. Sign up at spancache.ai and copy your per-tenant
ingest token from the workspace. Point any OpenTelemetry exporter at
https://ingest.spancache.ai/v1/traces(OTLP/HTTP) with the headerAuthorization: Bearer <token>. Works with Claude Code, the OpenAI and Anthropic SDKs (via OpenLLMetry), LangChain, LlamaIndex, and the Vercel AI SDK. - They’re priced and stored. On ingest, every span is priced (cache-aware, see below), scrubbed for secrets if content is captured, and written to the compressed store — kept in full, not sampled.
- Explore in the console. Open app.spancache.ai to browse runs, roll up cost, arm detections, and triage anomalies.
See getting-started.md for the five-minute setup and sending-traces.md for exporter recipes.
What you get in the console
- Traces — one row per agent run; click into the span waterfall for per-span cost, latency, tool calls, and (if you opted in) the conversation.
- Cost — per-model and per-release-version rollups of spend, latency, error rate, and tokens, so you can spot a regression right after a prompt or model change.
- Detections — one-click guardrail templates for agent failure modes: runaway cost, tool-error loop, error-rate spike, output truncation, retry storm, slow first token, context burn, slow agent step.
- Anomalies — cost-outlier traces (runaway loops, cache misses), novel agent behavior, and incompressible output that can signal prompt injection or exfiltration.
More in console.md.
Query language
Beyond the surfaces above, every field is queryable with a search plus a pipeline:
<search> | stats <aggs> by <field>For example, spend by model across all runs:
kind=span | stats sum(cost) by modelWhat makes Spancache different
- Accurate cost on every call. Spancache computes cost from
tokens × list price, cache-aware (Anthropic cache reads bill at ~0.1× input, cache writes at ~1.25×), for every span. So even sources that report no cost — like Claude Code on a subscription — get accurate per-call spend, and acost_sourcefield flags estimated versus source-reported values. See cost.md. - Keep everything, cheaply. The losslessly compressed store makes long retention affordable, so you keep full-fidelity history instead of sampling it away or letting it expire.
- Metadata-first governance. By default Spancache captures only metadata — cost, tokens, latency,
models, tools, outcomes — and no prompt or response content. Content capture is an explicit
opt-in, and anything captured is scrubbed for secrets at ingest (API keys, tokens, JWTs, and
private keys are replaced with
[REDACTED]). Strict per-tenant isolation; TLS everywhere. - OTLP-native, no lock-in. Standards-based ingest means no proprietary SDK and no agent to run — point your existing exporter and go.
Who it’s for
Developers and LLM engineers who ship agents and need to understand what each run costs, why it was slow, and why it failed — and the security teams who need those traces governed and retained. Spancache runs as a SaaS product across capacity tiers (free, pro, team, and BYOC enterprise); within a tier, nodes, seats, queries, and retention are unlimited.
Next steps
- getting-started.md — create a workspace and send your first trace.
- sending-traces.md — OTLP exporter recipes for common stacks.
- cost.md — how cache-aware cost is computed.
- console.md — a tour of Traces, Cost, Detections, and Anomalies.