Scaling

At small volume Spanlens runs at default settings without thought. Past a few requests per second you start hitting trade-offs between log fidelity, latency, and cost. This page is the explicit map of those trade-offs and the levers available.

Latency budget

StepTypicalp95Notes
DNS + TLS handshake (first call)~80 ms~250 msAmortized to ~0 ms for keep-alive connections.
Auth (API key lookup)~5 ms~15 msHash + DB lookup, cached in-process per warm container.
Provider key decrypt~2 ms~5 msAES-256-GCM via Web Crypto.
Upstream provider callvariesvariesThis is your model latency. Spanlens does not add to it.
Stream pump (per chunk)< 1 ms~3 mstee() into our log buffer, passthrough to client.
Log write (async, off critical path)~30 ms~150 msRuns after response leaves; does not delay your user.

Bottom line: for warm connections the user-visible overhead is ~10 ms typical, ~50 ms p95. That number is recorded on every Request row in theproxy_overhead_ms column so you can audit it directly in /requests.

Lever 1: log body mode

The biggest log size driver is the request and response body. By default we keep them fully. For high-throughput apps where bodies are large but you only need cost and trace structure, drop to meta or none.

Moderequest_body / response_bodyuser_id / session_idWhen to use
full (default)keptkeptMost teams. Bodies are essential for debugging.
metaempty stringkeptYou only need cost / latency / token / user-level analytics.
noneempty stringnullStrict PII zones. You still get cost and structure, no identifying data.

Set per-call via header or SDK helper:

ts
import { createOpenAI, withLogBody } from '@spanlens/sdk/openai'

const openai = createOpenAI()

const res = await openai.chat.completions.create(
  { ... },
  { headers: withLogBody('meta').headers },
)

Or set process-wide via the SDK:

ts
import { observeOpenAI } from '@spanlens/sdk/openai'

await observeOpenAI(openai, { logBody: 'meta' }, async () => { ... })

Storage impact: full mode adds ~2 KB per call once the bodies are compressed. meta mode is ~150 bytes per call. At 10M calls/month the difference is roughly 20 GB against 1.5 GB of table space.

Lever 2: sampling

For very high volumes (10k+ rps), even compressed bodies add up. Sample at the SDK level:

ts
import { createOpenAI } from '@spanlens/sdk/openai'

const openai = createOpenAI({
  sampling: {
    rate: 0.1,           // log 10% of calls
    alwaysLogErrors: true, // override sampling for status >= 400
  },
})

Sampled-out calls still flow through the proxy normally; they are just not persisted. Errors are always logged regardless of sample rate so debugging stays intact. Aggregate metrics (cost, latency) are scaled up by the inverse of the sample rate when displayed.

Sampling is a per-call decision made before the request is sent. Use alwaysLogErrors: true so you never miss a 5xx.

Lever 3: trace sampling

Traces are usually orders of magnitude lighter than request bodies, so most teams log every trace. If you need to sample, do it at the trace-creation site:

ts
const shouldTrace = Math.random() < 0.2  // 20% sample
const trace = shouldTrace
  ? client.startTrace({ name: 'agent' })
  : null

// ... wrap observe() calls only if trace is non-null

Sampled-out traces simply do not exist. There is no per-span sampling; if a trace is on, all its spans are kept.

Lever 4: streaming for long generations

Anything that might exceed the 290s stream deadline (Vercel Pro cap) should use stream: true. The proxy streams chunks straight through; first byte arrives in ~200 ms regardless of total duration. If the stream does hit the deadline, the Request is logged with truncated: true and the partial response body is kept.

For non-streaming requests, the upstream fetch is gated at UPSTREAM_TIMEOUT_MS = 35000 for initial headers. Bigger jobs should use streaming, period.

Connection reuse

Most provider SDKs maintain an HTTPS keep-alive pool. Make sure yours does:

  • OpenAI Node SDK: keep-alive on by default.
  • Anthropic Node SDK: keep-alive on by default.
  • Raw fetch: no keep-alive by default in some runtimes. In Node, use undici's default Agent.

TLS handshake cost dominates first-call latency. With keep-alive, the per-call overhead drops to the <15 ms numbers in the table above.

Concurrency on the proxy

Spanlens cloud runs on Vercel Pro with a per-region invocation pool. There is no per-account concurrency limit we enforce at the proxy layer; the upstream provider's rate limit is the real ceiling.

If you are bursting hard enough to saturate Vercel's pool, you will see 503 from the proxy. The provider SDKs retry these. For sustained high traffic, talk to us about a dedicated deployment, or self-host on infra you control.

Self-host tuning

When you self-host, the bottleneck is your Postgres project. Defaults carry a deployment a long way. Past that, the levers below are the ones that matter, roughly in the order you will reach for them.

The request log

requests is the only table that grows with traffic. It is partitioned by month on created_at, with a primary key of (created_at, id).

  • Keep partitions ahead of the clock. Call SELECT ensure_requests_partitions(3); monthly. Miss it and rows land in the requests_default catch-all, which is recoverable but tedious.
  • Monthly granularity is right for anything up to a few hundred million rows a month. Weekly partitions make the planner do more work on every dashboard query and buy nothing until the monthly tables get unwieldy.
  • Bodies use lz4 compression, which only applies once a value crosses the TOAST threshold. That is the right trade for a table read by dashboards: worse ratio than zstd, noticeably faster to decompress.
  • Every index is on (organization_id, created_at DESC) or a narrower version of it. Tenant-scoped time ranges are the only shape the dashboard queries in. If you add an index, keep organization_id first.
  • Dropping a partition is how retention happens. DROP TABLE requests_2025_08 is instant and leaves no bloat, unlike a bulk DELETE that would leave the table needing a vacuum.

The connection pool

Request-log reads go through the transaction pooler on port 6543, not PostgREST. Two environment variables govern it.

  • PG_POOL_MAX (default 2) is connections per server instance. Serverless platforms scale out, so raising this multiplies across instances and exhausts the pooler rather than adding throughput. Raise it only when you run one server of a fixed size.
  • PG_STATEMENT_TIMEOUT_MS(default 60000) bounds a runaway dashboard query. Worth keeping low, because the proxy's auth path shares this database and a query with no ceiling can starve it.

Traces, prompts, evals

These tables are not append-only at request volume. Even at 1M traces/month one Supabase project handles them comfortably.

  • RLS adds ~2 ms per query. Worth it for the multi-tenant isolation.
  • spans_refresh_trace_aggregates trigger fires on every span INSERT / UPDATE. Heavy span churn on a single trace amplifies. If you measure this as a hotspot, switch to a periodic recompute.

Replay queue

The requests_fallback queue drains 50 rows per 5-minute cron tick by default (~10 rows/sec). For higher recovery throughput, change the cron schedule in vercel.json or run the replay handler as a long-lived worker instead.

What to monitor

MetricWhereAlert threshold
proxy_overhead_ms p95aggregate /requests column> 80 ms for 10 minutes
5xx rategroup requests by status_code> 1% for 5 minutes
Fallback queue sizeGET /health/deep> 1000 sustained for 30 min
Truncated streamsrequests where truncated=true> 5% of streaming requests

Cost optimization tactics

  • Switch to meta for chatty internal services. A customer support bot that gets the same 50 messages over and over does not need bodies stored 50 times.
  • Use sampling for high-volume embeddings. Embedding calls are often 100x more frequent than completions and contain less debugging value. Sample at 10% and you keep all the signal at 1/10 the storage.
  • Self-host if you do millions of calls per day. Cloud pricing crosses over with self-host TCO somewhere around 10M calls/month for most teams.

Next: reliability for failure modes and recovery, or self-hosting for full control.