mirror of
https://github.com/BerriAI/litellm.git
synced 2026-10-06 02:48:13 +00:00
docs(tracing): document the trace contract, payloads and correlation
This commit is contained in:
parent
fa00072a58
commit
b5636a8519
1 changed files with 245 additions and 0 deletions
245
litellm/tracing/README.md
Normal file
245
litellm/tracing/README.md
Normal file
|
|
@ -0,0 +1,245 @@
|
|||
# Agent tracing
|
||||
|
||||
LiteLLM accepts OpenTelemetry traces from agent frameworks (LangChain, LangGraph, Deep Agents, anything
|
||||
that speaks OTEL GenAI semconv or OpenInference). It stores them in ClickHouse and joins every agent LLM
|
||||
call to the LiteLLM request that served it, so one trace shows the agent tree along with the key, team,
|
||||
model and cost of each call.
|
||||
|
||||
```
|
||||
agent app (LangChain / Deep Agents)
|
||||
│ OTLP/HTTP POST /v1/traces (Authorization: Bearer sk-...)
|
||||
▼
|
||||
TraceReceiver.ingest ── decode + normalize ── stamp tenant from auth ──► ClickHouse otel_traces
|
||||
│ │ (MV)
|
||||
│ ▼
|
||||
│ agent_traces
|
||||
agent LLM calls ─► LiteLLM /chat/completions ─► `clickhouse` callback ─► spend_logs
|
||||
│
|
||||
GET /v1/traces[/{id}] ◄── ClickHouseTraceStore ── join LiteLLMRequestId = response_id
|
||||
```
|
||||
|
||||
Code: `receiver.py` (entry point), `decode.py` (OTLP → `SpanRow`), `store.py` (writes + SQL reads),
|
||||
`types.py` (every shape in this doc), `litellm/proxy/tracing_endpoints.py` (HTTP),
|
||||
`litellm/integrations/clickhouse/` (client, DDL, batch writer, spend-log callback).
|
||||
|
||||
## Setup
|
||||
|
||||
Proxy:
|
||||
|
||||
```yaml
|
||||
general_settings:
|
||||
tracing:
|
||||
store: clickhouse # enables POST/GET /v1/traces
|
||||
litellm_settings:
|
||||
callbacks: ["clickhouse"] # optional: writes spend_logs so LLM spans get key/team/cost
|
||||
```
|
||||
|
||||
```bash
|
||||
CLICKHOUSE_URL=http://localhost:8123 CLICKHOUSE_USER=default CLICKHOUSE_PASSWORD=... CLICKHOUSE_DATABASE=litellm
|
||||
```
|
||||
|
||||
Tables are created on startup if they don't exist. Agent side (LangSmith's built-in OTEL exporter, no LangSmith account needed):
|
||||
|
||||
```bash
|
||||
LANGSMITH_TRACING=true
|
||||
LANGSMITH_TRACING_MODE=otel
|
||||
OTEL_EXPORTER_OTLP_ENDPOINT=http://localhost:4000 # exporter appends /v1/traces
|
||||
OTEL_EXPORTER_OTLP_HEADERS="Authorization=Bearer sk-..." # a LiteLLM key; its team owns the traces
|
||||
OTEL_SERVICE_NAME=research-agent
|
||||
```
|
||||
|
||||
Point the agent's model client at the same proxy (`base_url=http://localhost:4000`). That makes LLM spans
|
||||
carry LiteLLM response ids.
|
||||
|
||||
## Inbound payload
|
||||
|
||||
OTLP/HTTP, protobuf or JSON, gzip optional. This is one real LangSmith LLM span, trimmed (`…` marks cuts):
|
||||
|
||||
```json
|
||||
{"resourceSpans": [{
|
||||
"resource": {"attributes": [
|
||||
{"key": "service.name", "value": {"stringValue": "agent-demo"}},
|
||||
{"key": "telemetry.sdk.language", "value": {"stringValue": "python"}}]},
|
||||
"scopeSpans": [{"scope": {"name": "langsmith"}, "spans": [{
|
||||
"traceId": "hFUOSZOtjJ0Z08vMDtomjA==", "spanId": "+6NV/7MYH4g=", "parentSpanId": "XUYYPvnrVnU=",
|
||||
"name": "ChatOpenAI", "kind": "SPAN_KIND_INTERNAL",
|
||||
"startTimeUnixNano": "1790742982590574080", "endTimeUnixNano": "1790742984795666944",
|
||||
"status": {"code": "STATUS_CODE_OK"},
|
||||
"attributes": [
|
||||
{"key": "langsmith.span.kind", "value": {"stringValue": "llm"}},
|
||||
{"key": "gen_ai.operation.name", "value": {"stringValue": "chat"}},
|
||||
{"key": "gen_ai.request.model", "value": {"stringValue": "claude-sonnet-4-5"}},
|
||||
{"key": "langsmith.metadata.lc_agent_name", "value": {"stringValue": "support_triage_agent"}},
|
||||
{"key": "langsmith.metadata.langgraph_node", "value": {"stringValue": "model"}},
|
||||
{"key": "gen_ai.usage.input_tokens", "value": {"intValue": "675"}},
|
||||
{"key": "gen_ai.usage.output_tokens", "value": {"intValue": "103"}},
|
||||
{"key": "gen_ai.tool.definitions", "value": {"stringValue": "[{\"type\":\"function\",…}]"}},
|
||||
{"key": "gen_ai.prompt", "value": {"bytesValue": "eyJtZXNzYWdlcyI6W1t7ImxjIjoxLC…"}},
|
||||
{"key": "gen_ai.completion", "value": {"bytesValue": "eyJnZW5lcmF0aW9ucyI6W1t7InRleH…"}}
|
||||
]}]}]}]}
|
||||
```
|
||||
|
||||
`gen_ai.prompt` decodes to `{"messages": [[{"lc":1, "id":[…,"SystemMessage"], "kwargs":{"content":"You are a LiteLLM support agent…","type":"system"}}, …]]}`.
|
||||
`gen_ai.completion` decodes to `{"generations": [[{"message": {"kwargs": {"content": "", "tool_calls": […], "response_metadata": {"id": "chatcmpl-3d66069a-…", …}}}}]]}`.
|
||||
|
||||
What we read (`decode.py`):
|
||||
|
||||
| Attribute | Becomes |
|
||||
|---|---|
|
||||
| `service.name` (resource) | `ServiceName` |
|
||||
| scope `langsmith` or `langsmith.span.kind` | LangSmith normalizer (else OpenInference if `openinference.span.kind`, else GenAI semconv) |
|
||||
| `langsmith.span.kind` = `llm` / `tool` | `ObservationType` llm / tool; root span or name == `lc_agent_name` → `agent`; `*.wrap_model_call`, `*.before_agent`, … → `framework`; rest → `chain` |
|
||||
| `langsmith.metadata.lc_agent_name` | `AgentName` (enclosing agent) |
|
||||
| `gen_ai.request.model` | `Model` |
|
||||
| `gen_ai.usage.input_tokens` / `output_tokens` | `InputTokens` / `OutputTokens` |
|
||||
| `gen_ai.completion` → `…message.kwargs.response_metadata.id` | `LiteLLMRequestId` |
|
||||
| `gen_ai.prompt` / `gen_ai.completion` | `Input` / `Output`: llm → `[{role, content, tool_calls?}]`; tool → args / result text (LangGraph `Command` → last message content); agent → first input messages / last output message |
|
||||
| GenAI semconv: `gen_ai.operation.name`, `gen_ai.agent.name`, `gen_ai.response.id`, `gen_ai.input/output.messages` | same columns |
|
||||
| OpenInference: `openinference.span.kind`, `agent.name`, `llm.model_name`, `llm.token_count.*`, `input/output.value` | same columns |
|
||||
| `exception` event (`exception.message`) | `StatusMessage` when `status.message` is empty |
|
||||
|
||||
The heavy attributes (`gen_ai.prompt`, `gen_ai.completion`, `gen_ai.tool.definitions`, `*.messages`,
|
||||
`input/output.value`) move into `Input`/`Output` and are dropped from `SpanAttributes`. Everything else
|
||||
is kept as a string map.
|
||||
|
||||
## Correlation
|
||||
|
||||
**Agents and subagents inside one trace.** Spans form a tree through `parent_span_id`. A span is an
|
||||
`agent` when it is the root, or a chain whose name equals its `lc_agent_name`. A subagent is an agent span
|
||||
nested under another agent. In Deep Agents that looks like `research_lead` → tool `task` → chain
|
||||
`researcher`. Every span's `agent` field names the agent it runs inside. `AgentNode.parent_agent` walks
|
||||
up from each agent span to the nearest agent span with a *different* name. All invocations of one name
|
||||
collapse into one node: 200 calls to `researcher` give one `AgentNode` with `invocations: 200` and
|
||||
summed `llm_calls`, `tool_calls`, `spend` and `duration_ms`.
|
||||
|
||||
**Agent LLM call ↔ LiteLLM request.**
|
||||
1. Response id (the default). `otel_traces.LiteLLMRequestId = spend_logs.response_id`. LiteLLM returns
|
||||
its id as the chat completion `id`, and LangChain records that as `response_metadata.id`. On cache
|
||||
hits LiteLLM appends `_cache_hit<ts>` to the request id, so the callback strips that suffix before
|
||||
writing `response_id`. When several spend-log rows share a response id, the read keeps the row
|
||||
closest in time to the span.
|
||||
2. W3C `traceparent` (optional). If the agent forwards `traceparent` on its LLM calls, the callback
|
||||
stores it as `spend_logs.trace_id` / `span_id`. That also covers failed calls, which have no response id.
|
||||
3. `x-litellm-session-id` as a fallback: stored as `spend_logs.session_id`.
|
||||
|
||||
**Across services (agent-to-agent over HTTP / A2A).** Propagate W3C `traceparent` on the outbound call.
|
||||
The remote service then continues the same `trace_id`, and its spans land in the same trace under the
|
||||
calling span. A different `service.name` on those spans marks the hop. For fire-and-forget handoffs
|
||||
that start a new trace, use OTEL span Links. Links are stored (`Links.*` columns) but the read API does
|
||||
not return them yet.
|
||||
|
||||
## Stored row
|
||||
|
||||
`SpanRow` = one `otel_traces` row. The standard columns match the OTel Collector `clickhouseexporter`, so a
|
||||
collector can write to the same table.
|
||||
|
||||
| Column | Notes |
|
||||
|---|---|
|
||||
| `Timestamp`, `Duration` | span start (ns), duration (ns) |
|
||||
| `TraceId`, `SpanId`, `ParentSpanId`, `TraceState` | hex ids; `ParentSpanId` = `''` for roots |
|
||||
| `SpanName`, `SpanKind`, `ServiceName`, `ScopeName`, `ScopeVersion` | from OTLP |
|
||||
| `ResourceAttributes`, `SpanAttributes` | `Map(String, String)`, values capped at `OTLP_MAX_ATTRIBUTE_VALUE_BYTES` |
|
||||
| `StatusCode`, `StatusMessage` | `STATUS_CODE_OK/ERROR/UNSET`; message falls back to the exception event |
|
||||
| `TeamId`, `ApiKeyHash` | from the auth'd key, never from the payload |
|
||||
| `ObservationType`, `AgentName` | `agent / llm / tool / chain / framework`, enclosing agent |
|
||||
| `LiteLLMRequestId`, `Model`, `InputTokens`, `OutputTokens` | join key + usage |
|
||||
| `Input`, `Output` | normalized I/O (ZSTD(3)); `InputPreview` = first 240 chars (column DEFAULT) |
|
||||
|
||||
Tables (`litellm/integrations/clickhouse/schema.py`):
|
||||
|
||||
- `otel_traces`: one row per span. MergeTree, `PARTITION BY toDate(Timestamp)`, `ORDER BY (TeamId, ServiceName, toDateTime(Timestamp), TraceId)`, bloom filters on `TraceId` and `LiteLLMRequestId`, TTL `AGENT_TRACING_RETENTION_DAYS` (30).
|
||||
- `agent_traces` + `agent_traces_mv`: one row per trace (counts, tokens, models, request ids, root name). AggregatingMergeTree, `ORDER BY (TeamId, TraceId)`. It backs the list endpoint. The MV writes one partial row per insert, and reads merge them.
|
||||
- `spend_logs`: one row per LiteLLM request, written by the `clickhouse` callback. ReplacingMergeTree(end_time), `PARTITION BY toYYYYMM(start_time)`, `ORDER BY (team_id, toDateTime(start_time), request_id)`, bloom filters on `response_id` and `trace_id`, TTL `AGENT_TRACING_SPEND_LOG_RETENTION_DAYS` (90).
|
||||
|
||||
## Read API
|
||||
|
||||
`GET /v1/traces/{trace_id}` → `Trace`. This is real output from a Deep Agents run (`research_lead` fans out
|
||||
to 4 `researcher` subagents and a `critic`). 216 spans, trimmed to five:
|
||||
|
||||
```json
|
||||
{
|
||||
"summary": {
|
||||
"trace_id": "e309a123963901e74c29cd2d3c86ff9e", "name": "research_lead", "service": "research-agent",
|
||||
"input_preview": "[{\"role\": \"user\", \"content\": \"Should we store OTEL agent spans in ClickHouse or Postgres at 50k spans/sec?\"}]",
|
||||
"start_time": "2026-09-30T06:43:54.291000+00:00", "duration_ms": 40198.10688, "status": "ok",
|
||||
"span_count": 216, "agent_count": 3, "llm_calls": 21, "tool_calls": 25, "error_count": 0,
|
||||
"input_tokens": 69506, "output_tokens": 2960, "spend": 0.12833025, "models": ["claude-sonnet-4-5"]
|
||||
},
|
||||
"agents": [
|
||||
{"name": "research_lead", "parent_agent": null, "invocations": 1, "llm_calls": 3, "tool_calls": 5, "spend": 0.025758, "duration_ms": 40198.10688},
|
||||
{"name": "researcher", "parent_agent": "research_lead", "invocations": 4, "llm_calls": 17, "tool_calls": 20, "spend": 0.089172, "duration_ms": 41634.618112},
|
||||
{"name": "critic", "parent_agent": "research_lead", "invocations": 1, "llm_calls": 1, "tool_calls": 0, "spend": 0.01340025, "duration_ms": 6897.236224}
|
||||
],
|
||||
"spans": [
|
||||
{"span_id": "3586edf49d446541", "parent_span_id": null, "name": "research_lead", "type": "agent",
|
||||
"agent": "research_lead", "start_offset_ms": 0.0, "duration_ms": 40198.10688, "status": "ok",
|
||||
"input_preview": "[{\"role\": \"user\", \"content\": \"Should we store OTEL agent spans in ClickHouse or Postgres at 50k spans/sec?\"}]",
|
||||
"model": null, "input_tokens": 0, "output_tokens": 0, "litellm": null},
|
||||
{"span_id": "2c51e6ddab97ed21", "parent_span_id": "ae58781d8e2e70eb", "name": "ChatOpenAI", "type": "llm",
|
||||
"agent": "research_lead", "start_offset_ms": 5.357056, "duration_ms": 8283.9168, "status": "ok",
|
||||
"input_preview": "[{\"role\": \"system\", \"content\": \"You are a research lead. Split the question into exactly 4 narrow sub-questions…",
|
||||
"model": "claude-sonnet-4-5", "input_tokens": 3378, "output_tokens": 514,
|
||||
"litellm": {"request_id": "chatcmpl-028008eb-a34b-4132-97c2-526b707e961e", "model": "openai/claude-sonnet-4-5",
|
||||
"model_group": "claude-sonnet-4-5", "provider": "openai", "key_alias": "research-bot",
|
||||
"team_alias": "research-agents", "spend": 0.0087315, "prompt_tokens": 3378, "completion_tokens": 514,
|
||||
"cache_read_tokens": 3375, "cache_write_tokens": 0, "latency_ms": 8278, "ttft_ms": 8278, "status": "success"}},
|
||||
{"span_id": "53e84f1d468bc2a6", "parent_span_id": "dcf2045cfb34c569", "name": "task", "type": "tool",
|
||||
"agent": "research_lead", "start_offset_ms": 8290.994944, "duration_ms": 6624.634112, "status": "ok",
|
||||
"input_preview": "{\"subagent_type\":\"researcher\",\"description\":\"Research and answer this narrow question: What are the write performance…",
|
||||
"model": null, "input_tokens": 0, "output_tokens": 0, "litellm": null},
|
||||
{"span_id": "8febe34dc2549d84", "parent_span_id": "53e84f1d468bc2a6", "name": "researcher", "type": "agent",
|
||||
"agent": "researcher", "start_offset_ms": 8291.454976, "duration_ms": 6624.08704, "status": "ok",
|
||||
"input_preview": "[{\"role\": \"user\", \"content\": \"Research and answer this narrow question: What are the write performance…",
|
||||
"model": null, "input_tokens": 0, "output_tokens": 0, "litellm": null},
|
||||
{"span_id": "ca8488012a54b761", "parent_span_id": "96abed5e22e54c83", "name": "search_docs", "type": "tool",
|
||||
"agent": "researcher", "start_offset_ms": 10082.61504, "duration_ms": 0.475904, "status": "ok",
|
||||
"input_preview": "{\"query\":\"Postgres write performance time-series data ingestion 50k inserts per second high-volume\"}",
|
||||
"model": null, "input_tokens": 0, "output_tokens": 0, "litellm": null}
|
||||
]
|
||||
}
|
||||
```
|
||||
|
||||
Spans are flat and ordered by start time. Build the tree from `parent_span_id`. `litellm` is set only on
|
||||
llm spans that matched a spend-log row. `status` is the root span's status, while `error_count` counts
|
||||
every span with an error status (a tool that raised shows up there even when the agent recovered).
|
||||
|
||||
`GET /v1/traces?start_ms=&end_ms=&cursor=` → `TracePage` (newest first, default window 24h, page size
|
||||
`AGENT_TRACING_LIST_PAGE_SIZE`=50):
|
||||
|
||||
```json
|
||||
{"data": [
|
||||
{"trace_id": "f78f6df35480060fafadac887e234241", "name": "support_triage_agent", "service": "research-agent",
|
||||
"input_preview": "[{\"role\": \"user\", \"content\": \"Customer acme-404 says billing is wrong. What plan are they on?\"}]",
|
||||
"start_time": "2026-09-30T06:43:52.928000+00:00", "duration_ms": 1315.0, "status": "ok",
|
||||
"span_count": 5, "agent_count": 1, "llm_calls": 1, "tool_calls": 1, "error_count": 2,
|
||||
"input_tokens": 659, "output_tokens": 60, "spend": 0.002877, "models": ["claude-sonnet-4-5"]}
|
||||
],
|
||||
"next_cursor": null}
|
||||
```
|
||||
|
||||
`next_cursor` is set when the page is full. Pass it back as-is.
|
||||
`GET /v1/traces/{trace_id}/spans/{span_id}` → `SpanDetail`: `{span_id, input, output, attributes}`, with the full
|
||||
`Input`/`Output` and `SpanAttributes` for the span drawer.
|
||||
|
||||
## Scoping and privacy
|
||||
|
||||
- Writes: `TeamId`, `ApiKeyHash` and `litellm.team_id` / `litellm.api_key_hash` / `litellm.org_id` in
|
||||
`ResourceAttributes` always come from the authenticated key. Values the client sends for these are overwritten.
|
||||
- Reads (`scope_for`): proxy admins (incl. view-only) see every team. A key with a team sees that team's
|
||||
traces. A team-less key sees only traces it sent (`TeamId = ''` and `ApiKeyHash` = its hash). A request
|
||||
with neither a team nor a key gets 403.
|
||||
- Spend-log rows are scoped the same way, so a join never exposes another team's cost.
|
||||
- `litellm.turn_off_message_logging = True` blanks `messages` / `response` in `spend_logs`. Span `Input` /
|
||||
`Output` come from the agent's own exporter; to keep prompts out, turn off content capture on the agent
|
||||
side (e.g. `LANGSMITH_HIDE_INPUTS=true` / `LANGSMITH_HIDE_OUTPUTS=true`).
|
||||
|
||||
## Limits
|
||||
|
||||
| Setting | Default | Behaviour |
|
||||
|---|---|---|
|
||||
| `OTLP_MAX_BODY_BYTES` | 8 MiB | larger bodies → 413 |
|
||||
| `OTLP_MAX_ATTRIBUTE_VALUE_BYTES` | 64 KiB | attribute values, `Input`, `Output` truncated with `…[truncated N bytes]` |
|
||||
| `OTLP_OFFLOAD_DECODE_BYTES` | 256 KiB | bodies above this are decoded in a worker thread |
|
||||
| `CLICKHOUSE_MAX_BUFFERED_ROWS` | 200,000 | buffer full → 429 + `Retry-After: OTLP_RETRY_AFTER_SECONDS` (2); OTLP exporters retry |
|
||||
| `CLICKHOUSE_BATCH_SIZE`, `CLICKHOUSE_FLUSH_INTERVAL_SECONDS` | 10,000 rows, 1s | spans are written in batches; `POST` never waits on ClickHouse |
|
||||
| `CLICKHOUSE_MAX_RETRIES` | 3 | after this many failed inserts a batch is dropped and logged |
|
||||
Loading…
Add table
Reference in a new issue