mirror of
https://github.com/BerriAI/litellm.git
synced 2026-09-30 01:52:18 +00:00
feat(otel): typed semconv-aligned OpenTelemetry instrumentation
Some checks are pending
Unit Tests: Proxy DB Operations / auth-checks (push) Blocked by required conditions
Unit Tests: Proxy DB Operations / assert-shard-coverage (push) Waiting to run
Unit Tests: Proxy DB Operations / budgets (push) Blocked by required conditions
Unit Tests: Proxy DB Operations / custom-logging (push) Blocked by required conditions
Unit Tests: Proxy DB Operations / db-and-spend (push) Blocked by required conditions
Unit Tests: Proxy DB Operations / endpoints-and-responses (push) Blocked by required conditions
Unit Tests: Proxy DB Operations / guardrails-hooks (push) Blocked by required conditions
Unit Tests: Proxy DB Operations / jwt-and-keys (push) Blocked by required conditions
Unit Tests: Proxy DB Operations / key-generation (push) Blocked by required conditions
Unit Tests: Proxy DB Operations / logging-misc (push) Blocked by required conditions
Unit Tests: Proxy DB Operations / proxy-runtime (push) Blocked by required conditions
Unit Tests: Proxy DB Operations / proxy-server-core (push) Blocked by required conditions
Unit Tests: Proxy DB Operations / schema-migration (push) Blocked by required conditions
Unit Tests: Proxy DB Operations / proxy-utils (push) Blocked by required conditions
Unit Tests: Security / security (push) Waiting to run
Some checks are pending
Unit Tests: Proxy DB Operations / auth-checks (push) Blocked by required conditions
Unit Tests: Proxy DB Operations / assert-shard-coverage (push) Waiting to run
Unit Tests: Proxy DB Operations / budgets (push) Blocked by required conditions
Unit Tests: Proxy DB Operations / custom-logging (push) Blocked by required conditions
Unit Tests: Proxy DB Operations / db-and-spend (push) Blocked by required conditions
Unit Tests: Proxy DB Operations / endpoints-and-responses (push) Blocked by required conditions
Unit Tests: Proxy DB Operations / guardrails-hooks (push) Blocked by required conditions
Unit Tests: Proxy DB Operations / jwt-and-keys (push) Blocked by required conditions
Unit Tests: Proxy DB Operations / key-generation (push) Blocked by required conditions
Unit Tests: Proxy DB Operations / logging-misc (push) Blocked by required conditions
Unit Tests: Proxy DB Operations / proxy-runtime (push) Blocked by required conditions
Unit Tests: Proxy DB Operations / proxy-server-core (push) Blocked by required conditions
Unit Tests: Proxy DB Operations / schema-migration (push) Blocked by required conditions
Unit Tests: Proxy DB Operations / proxy-utils (push) Blocked by required conditions
Unit Tests: Security / security (push) Waiting to run
This commit is contained in:
parent
76f56c3283
commit
ab489eb28d
47 changed files with 6106 additions and 1 deletions
155
litellm/integrations/otel/README.md
Normal file
155
litellm/integrations/otel/README.md
Normal file
|
|
@ -0,0 +1,155 @@
|
|||
# OpenTelemetry instrumentation
|
||||
|
||||
This package produces OpenTelemetry traces for LiteLLM. It is enabled by the
|
||||
`LITELLM_OTEL_V2` environment variable (`is_otel_v2_enabled()` in
|
||||
[`config.py`](./config.py)); when unset, nothing in this package runs.
|
||||
|
||||
## What gets traced
|
||||
|
||||
A traced proxy request produces one trace with two kinds of spans:
|
||||
|
||||
```
|
||||
SERVER span "POST /v1/chat/completions" ← FastAPI instrumentation
|
||||
├── CLIENT span "chat gpt-4o" ← LLM call ┐
|
||||
├── INTERNAL span "execute_guardrail …" ← guardrail │ this package
|
||||
└── INTERNAL span "redis" … ← service call ┘
|
||||
```
|
||||
|
||||
The gen-ai spans are siblings under the server span. In particular the guardrail
|
||||
span is a sibling of the LLM call, not a child of it: pre/during/post-call
|
||||
guardrail hooks are part of the request lifecycle (a pre-call guardrail runs
|
||||
before the LLM call even starts), so they parent to the server span via the
|
||||
ambient OpenTelemetry context, alongside the LLM call.
|
||||
|
||||
- **Server spans** (one per HTTP route) are created by the
|
||||
`opentelemetry-instrumentation-fastapi` package. It stamps `http.*` attributes
|
||||
and extracts inbound `traceparent` headers. This package does **not** create
|
||||
or modify server spans — request routes never touch spans.
|
||||
- **Gen-AI spans** (LLM calls, guardrails, internal service calls) are created
|
||||
by this package from LiteLLM's logging callbacks and parent to the active
|
||||
server span via ambient OpenTelemetry context.
|
||||
|
||||
Both kinds share a single `TracerProvider`, so they belong to the same trace
|
||||
and export through the same configured exporters. FastAPI middleware can only be
|
||||
added before the app starts serving, so the app is instrumented at
|
||||
import time **without** a provider — it binds to the OTel global
|
||||
`ProxyTracerProvider`. Once config (and the callbacks) is loaded, the proxy
|
||||
publishes the chosen logger's `TracerProvider` as the global via
|
||||
`trace.set_tracer_provider(...)`, and the server spans delegate to it. When a
|
||||
preset callback (`arize`, `langfuse_otel`, …) is configured, its provider
|
||||
becomes the global, so server spans export to that backend too.
|
||||
|
||||
## How a request flows
|
||||
|
||||
1. **App creation** (`proxy_server` import): when the gate is on,
|
||||
`FastAPIInstrumentor.instrument_app(app)` is called with no provider (the
|
||||
middleware stack is frozen once the app serves, so this can't wait for
|
||||
startup). It binds to the OTel global `ProxyTracerProvider`. Health-check
|
||||
routes (`/health*`) are excluded by default so load-balancer polling doesn't
|
||||
flood traces; set `OTEL_PYTHON_FASTAPI_EXCLUDED_URLS` to override (e.g. `""`
|
||||
to trace everything, or your own comma-separated path list).
|
||||
2. **Startup** (`proxy_server.proxy_startup_event`): after the config (and
|
||||
callbacks) is loaded, the already-registered preset `OpenTelemetryV2` logger
|
||||
is reused — or a generic one reading `OTEL_*` envs is built when no preset is
|
||||
configured — and its `TracerProvider` is published as the OTel global with
|
||||
`trace.set_tracer_provider(...)`. The proxy tracer then delegates to it, so
|
||||
server spans and gen-ai spans share one provider and the same trace.
|
||||
3. **Request**: the FastAPI instrumentation starts the server span and makes it
|
||||
the active context for the request task.
|
||||
4. **LLM call logging**: LiteLLM's async logging worker copies the request's
|
||||
context when it enqueues the success/failure callback, so
|
||||
`OpenTelemetryV2.async_log_success_event` runs with the server span as the
|
||||
ambient parent. It builds an `LLMCallSpanData` from the request's
|
||||
`standard_logging_object` and hands it to the engine, which creates the LLM
|
||||
span as a child of the server span. Emission is **async-only**; the
|
||||
synchronous callback runs in a worker thread without the request context and
|
||||
is a no-op.
|
||||
5. **Guardrails / services**: the post-call and service hooks emit guardrail and
|
||||
service spans the same way — typed data → engine → span.
|
||||
6. **Export**: each span ends and is handed to the provider's span processors,
|
||||
which export to the configured backends (OTLP, console, in-memory, …).
|
||||
|
||||
## Components
|
||||
|
||||
### Sources of truth (no OpenTelemetry import)
|
||||
|
||||
These define the shape of a span without depending on the OTel SDK, so they can
|
||||
be imported anywhere:
|
||||
|
||||
- [`semconv.py`](./semconv.py) — attribute-key constants (`gen_ai.*`, `http.*`,
|
||||
`litellm.*`), the GenAI operation/provider enums, and the functions that map
|
||||
LiteLLM provider/call-type strings onto convention values.
|
||||
- [`spans.py`](./spans.py) — the span registry: every span role, its OTel span
|
||||
kind, its place in the hierarchy, and its name builder.
|
||||
- [`payloads.py`](./payloads.py) — frozen dataclasses (`LLMCallSpanData`,
|
||||
`GuardrailSpanData`, `ServiceSpanData`, …) built from heterogeneous logging
|
||||
payloads via `from_*` classmethods.
|
||||
- [`config.py`](./config.py) — `OpenTelemetryV2Config`, a pydantic-settings
|
||||
model that reads `OTEL_*` / `LITELLM_OTEL_*` env vars, plus the feature gate.
|
||||
`capture_span_content` gates whether prompt/response bodies may be written as
|
||||
span attributes; it defaults **off** (`no_content`).
|
||||
|
||||
### Engine
|
||||
|
||||
- [`emitter.py`](./emitter.py) — `SpanEmitter.emit(role, data)`: dedupe → start
|
||||
the span → run the mapper chain to stamp attributes → set status → end. It
|
||||
owns no attribute keys. The dedupe set (which coalesces the sync+async firing
|
||||
of one request) is a bounded LRU so it can't grow without limit.
|
||||
- [`mappers/`](./mappers) — each mapper turns typed span data into a flat
|
||||
`{attribute key: value}` dict. They compose: listing several mapper names in
|
||||
the config layers multiple attribute vocabularies onto the same span.
|
||||
- `genai` — the canonical OpenTelemetry GenAI vocabulary, always present.
|
||||
- `legacy` — an additional vocabulary using the older semconv-ai / Traceloop
|
||||
attribute key names, for backends that read those.
|
||||
- `openinference`, `langfuse`, `weave`, `langtrace` — vendor vocabularies.
|
||||
- `resolve_mappers(names)` turns config names into mapper instances.
|
||||
|
||||
### Plumbing
|
||||
|
||||
- [`providers.py`](./providers.py) — builds the `TracerProvider`, its exporters
|
||||
(from `ExporterSpec`s), and the span processor that copies allowlisted Baggage
|
||||
entries onto every span. `register_exporter_factory(kind, factory)` lets a
|
||||
preset contribute a custom exporter `kind` (e.g. one that fetches an auth
|
||||
token lazily) without coupling this module to any vendor.
|
||||
- [`context.py`](./context.py) — trace-context and Baggage read/write helpers.
|
||||
- [`baggage.py`](./baggage.py) — the single definition of which request-identity
|
||||
values are promoted into Baggage (so child spans inherit them) and under which
|
||||
attribute keys.
|
||||
- [`routing.py`](./routing.py) — `TenantTracerCache`: when a request carries
|
||||
team/key-scoped vendor credentials, route its spans through a credential-keyed
|
||||
`TracerProvider` so one logger serves many tenants. The cache is a bounded LRU
|
||||
that flushes + shuts down evicted providers, since the key derives from
|
||||
request-supplied credentials and must not grow (or leak threads) without limit.
|
||||
- [`utils.py`](./utils.py) — value coercion, JSON serialization, and
|
||||
extractor-table application, shared across the package.
|
||||
- [`metrics.py`](./metrics.py) — GenAI client metric instruments.
|
||||
|
||||
### Adapter
|
||||
|
||||
- [`logger.py`](./logger.py) — `OpenTelemetryV2`, a `CustomLogger` that
|
||||
translates LiteLLM's logging callbacks into typed span data and hands them to
|
||||
the engine.
|
||||
|
||||
### Presets
|
||||
|
||||
- [`presets/`](./presets) — each preset reads one integration's env vars and
|
||||
returns an `OpenTelemetryV2Config` (exporter destination + mapper vocabularies
|
||||
+ resource attributes). `PRESET_BY_CALLBACK` maps a callback name (`"arize"`,
|
||||
`"langfuse_otel"`, …) to its preset. Integrations that support team/key-scoped
|
||||
credentials also provide a per-request OTLP header builder
|
||||
(`DYNAMIC_HEADERS_BY_CALLBACK`). Presets do **no** network I/O at build time:
|
||||
AgentOps, for example, mints its JWT lazily inside a custom exporter on the
|
||||
first export (in the `BatchSpanProcessor` worker thread), never on the event
|
||||
loop.
|
||||
|
||||
## Extending
|
||||
|
||||
- **A new attribute vocabulary for a backend**: add a mapper in `mappers/`
|
||||
(a class with a `map(data) -> AttributeMap` method, typically built from
|
||||
`key -> extractor` tables) and register it in `mappers/__init__._MAPPER_BY_NAME`.
|
||||
- **A new integration**: add a preset in `presets/` that returns an
|
||||
`OpenTelemetryV2Config`, and register it in `presets/__init__.PRESET_BY_CALLBACK`.
|
||||
If it supports dynamic credentials, add a header builder to
|
||||
`DYNAMIC_HEADERS_BY_CALLBACK`.
|
||||
- **A new span kind**: add a role to `spans.py` (registry entry + name builder),
|
||||
a payload dataclass in `payloads.py`, and a branch in the relevant mapper(s).
|
||||
92
litellm/integrations/otel/__init__.py
Normal file
92
litellm/integrations/otel/__init__.py
Normal file
|
|
@ -0,0 +1,92 @@
|
|||
"""Typed, semconv-aligned OpenTelemetry instrumentation for LiteLLM.
|
||||
|
||||
The three sources of truth — attribute keys (:mod:`semconv`), the span and
|
||||
hierarchy registry (:mod:`spans`), and the typed span-data inputs
|
||||
(:mod:`payloads`) — plus :mod:`config` are exported here and are free of any
|
||||
``opentelemetry`` import. The engine layer (``emitter``, ``providers``,
|
||||
``context``, ``metrics``) and the ``CustomLogger`` adapter (``logger``) are
|
||||
reached via their submodule paths so that importing this package never
|
||||
requires the OTel SDK.
|
||||
|
||||
The ``LITELLM_OTEL_V2`` env var gates whether the factory in
|
||||
``litellm_core_utils.litellm_logging`` constructs the ``OpenTelemetryV2``
|
||||
class (from :mod:`logger`).
|
||||
"""
|
||||
|
||||
from litellm.integrations.otel.config import (
|
||||
OTEL_V2_ENV,
|
||||
OpenTelemetryV2Config,
|
||||
is_otel_v2_enabled,
|
||||
)
|
||||
from litellm.integrations.otel.baggage import (
|
||||
BAGGAGE_PROMOTED_KEYS,
|
||||
DEFAULT_BAGGAGE_METADATA_KEYS,
|
||||
promoted_baggage,
|
||||
)
|
||||
from litellm.integrations.otel.payloads import (
|
||||
GuardrailSpanData,
|
||||
LLMCallSpanData,
|
||||
LLMRequestParams,
|
||||
LLMUsage,
|
||||
ProxyRequestSpanData,
|
||||
RequestIdentity,
|
||||
ServerInfo,
|
||||
ServiceSpanData,
|
||||
SpanError,
|
||||
)
|
||||
from litellm.integrations.otel.semconv import (
|
||||
Error,
|
||||
GenAI,
|
||||
GenAIOperation,
|
||||
GenAIProvider,
|
||||
HTTP,
|
||||
LiteLLM,
|
||||
Metric,
|
||||
Server,
|
||||
resolve_operation,
|
||||
resolve_provider,
|
||||
)
|
||||
from litellm.integrations.otel.spans import (
|
||||
SPAN_REGISTRY,
|
||||
LiteLLMSpanKind,
|
||||
SpanRole,
|
||||
SpanSpec,
|
||||
validate_registry,
|
||||
)
|
||||
|
||||
__all__ = [
|
||||
# config
|
||||
"OTEL_V2_ENV",
|
||||
"OpenTelemetryV2Config",
|
||||
"is_otel_v2_enabled",
|
||||
# semconv
|
||||
"BAGGAGE_PROMOTED_KEYS",
|
||||
"DEFAULT_BAGGAGE_METADATA_KEYS",
|
||||
"Error",
|
||||
"GenAI",
|
||||
"GenAIOperation",
|
||||
"GenAIProvider",
|
||||
"HTTP",
|
||||
"LiteLLM",
|
||||
"Metric",
|
||||
"Server",
|
||||
"resolve_operation",
|
||||
"resolve_provider",
|
||||
# spans
|
||||
"SPAN_REGISTRY",
|
||||
"LiteLLMSpanKind",
|
||||
"SpanRole",
|
||||
"SpanSpec",
|
||||
"validate_registry",
|
||||
# payloads
|
||||
"GuardrailSpanData",
|
||||
"LLMCallSpanData",
|
||||
"LLMRequestParams",
|
||||
"LLMUsage",
|
||||
"ProxyRequestSpanData",
|
||||
"RequestIdentity",
|
||||
"ServerInfo",
|
||||
"ServiceSpanData",
|
||||
"SpanError",
|
||||
"promoted_baggage",
|
||||
]
|
||||
72
litellm/integrations/otel/baggage.py
Normal file
72
litellm/integrations/otel/baggage.py
Normal file
|
|
@ -0,0 +1,72 @@
|
|||
"""Baggage promotion: request-identity values carried across child spans.
|
||||
|
||||
A bounded set of identity values is written into OpenTelemetry Baggage on the
|
||||
LLM-call span so that child spans (guardrail, service) inherit them.
|
||||
``providers.LiteLLMBaggageSpanProcessor`` reads Baggage at span start and stamps
|
||||
the allowlisted keys onto every span.
|
||||
|
||||
This module is the single place baggage is defined: ``_PROMOTABLE`` maps each
|
||||
promotable attribute key to how its value is read, and the two ``*_KEYS``
|
||||
defaults select what is promoted unless the config overrides them.
|
||||
"""
|
||||
|
||||
from collections.abc import Callable
|
||||
from typing import Final
|
||||
|
||||
from litellm.integrations.otel.payloads import RequestIdentity
|
||||
from litellm.integrations.otel.semconv import GenAI, LiteLLM
|
||||
|
||||
# Attribute key -> value extractor over (identity, request_model). The single
|
||||
# definition of what may be promoted and under which key.
|
||||
_PROMOTABLE: Final[dict[str, Callable[[RequestIdentity, str | None], str | None]]] = {
|
||||
LiteLLM.TEAM_ID: lambda identity, model: identity.team_id,
|
||||
LiteLLM.TEAM_ALIAS: lambda identity, model: identity.team_alias,
|
||||
LiteLLM.KEY_HASH: lambda identity, model: identity.key_hash,
|
||||
LiteLLM.END_USER: lambda identity, model: identity.end_user,
|
||||
GenAI.REQUEST_MODEL: lambda identity, model: model,
|
||||
}
|
||||
|
||||
# Keys promoted by default (a subset of ``_PROMOTABLE``). ``END_USER`` is
|
||||
# promotable but off by default — it identifies an individual user, so stamping
|
||||
# it onto every span is opt-in via ``config.baggage_promoted_keys``.
|
||||
BAGGAGE_PROMOTED_KEYS: Final[tuple[str, ...]] = (
|
||||
LiteLLM.TEAM_ID,
|
||||
LiteLLM.TEAM_ALIAS,
|
||||
LiteLLM.KEY_HASH,
|
||||
GenAI.REQUEST_MODEL,
|
||||
)
|
||||
|
||||
# Metadata sub-keys eligible for promotion under the ``litellm.metadata.*``
|
||||
# namespace. The full metadata blob is never promoted; only this allowlist is.
|
||||
DEFAULT_BAGGAGE_METADATA_KEYS: Final[tuple[str, ...]] = (
|
||||
"user_api_key_org_id",
|
||||
"user_api_key_user_id",
|
||||
"user_api_key_alias",
|
||||
"user_api_key_end_user_id",
|
||||
"requester_ip_address",
|
||||
)
|
||||
|
||||
|
||||
def promoted_baggage(
|
||||
identity: RequestIdentity,
|
||||
request_model: str | None,
|
||||
promoted_keys: tuple[str, ...],
|
||||
metadata_keys: tuple[str, ...] = DEFAULT_BAGGAGE_METADATA_KEYS,
|
||||
) -> dict[str, str]:
|
||||
"""Identity values to write into Baggage, filtered to ``promoted_keys``.
|
||||
|
||||
``promoted_keys`` selects from ``_PROMOTABLE``; ``metadata_keys`` selects
|
||||
sub-keys of ``identity.metadata`` to promote under ``litellm.metadata.*``.
|
||||
Empty values are dropped.
|
||||
"""
|
||||
out: dict[str, str] = {}
|
||||
for key, extract in _PROMOTABLE.items():
|
||||
if key in promoted_keys:
|
||||
value = extract(identity, request_model)
|
||||
if value:
|
||||
out[key] = value
|
||||
for meta_key in metadata_keys:
|
||||
value = identity.metadata.get(meta_key)
|
||||
if value:
|
||||
out[f"{LiteLLM.METADATA_PREFIX}{meta_key}"] = value
|
||||
return out
|
||||
194
litellm/integrations/otel/config.py
Normal file
194
litellm/integrations/otel/config.py
Normal file
|
|
@ -0,0 +1,194 @@
|
|||
"""Typed configuration for the OpenTelemetry instrumentation."""
|
||||
|
||||
from pydantic import AliasChoices, BaseModel, Field, model_validator
|
||||
from pydantic_settings import BaseSettings, SettingsConfigDict
|
||||
|
||||
from litellm.integrations.otel.baggage import (
|
||||
BAGGAGE_PROMOTED_KEYS,
|
||||
DEFAULT_BAGGAGE_METADATA_KEYS,
|
||||
)
|
||||
|
||||
#: Master feature-flag env var. The logger is inert until this is truthy.
|
||||
OTEL_V2_ENV = "LITELLM_OTEL_V2"
|
||||
|
||||
|
||||
class CaptureMessageContent(str):
|
||||
NO_CONTENT = "no_content"
|
||||
SPAN_ONLY = "span_only"
|
||||
EVENT_ONLY = "event_only"
|
||||
SPAN_AND_EVENT = "span_and_event"
|
||||
|
||||
|
||||
class _OTelV2Flag(BaseSettings):
|
||||
model_config = SettingsConfigDict(extra="ignore")
|
||||
|
||||
enabled: bool = Field(default=False, validation_alias=AliasChoices(OTEL_V2_ENV))
|
||||
|
||||
|
||||
def is_otel_v2_enabled() -> bool:
|
||||
return _OTelV2Flag().enabled
|
||||
|
||||
|
||||
class ExporterSpec(BaseModel):
|
||||
"""One span-export destination.
|
||||
|
||||
The shared ``TracerProvider`` attaches one ``SpanProcessor`` per spec, so
|
||||
listing several specs sends every span to all of them at once (e.g. Arize +
|
||||
Phoenix + your own Honeycomb).
|
||||
"""
|
||||
|
||||
model_config = {"extra": "forbid"}
|
||||
|
||||
kind: str = Field(
|
||||
default="console",
|
||||
description="console | in_memory | otlp_http | otlp_grpc | <factory kind>",
|
||||
)
|
||||
endpoint: str | None = None
|
||||
headers: str | None = None
|
||||
options: dict[str, str] | None = Field(
|
||||
default=None,
|
||||
description=(
|
||||
"Factory-specific configuration for a custom exporter ``kind`` "
|
||||
"registered via ``providers.register_exporter_factory`` (e.g. an "
|
||||
"API key a lazy-auth exporter fetches a token with). Ignored by the "
|
||||
"built-in console/in_memory/otlp exporters."
|
||||
),
|
||||
)
|
||||
use_simple_processor: bool | None = Field(
|
||||
default=None,
|
||||
description=(
|
||||
"Force SimpleSpanProcessor regardless of exporter kind. Default: "
|
||||
"auto (Simple for console/in_memory, Batch otherwise)."
|
||||
),
|
||||
)
|
||||
|
||||
|
||||
class OpenTelemetryV2Config(BaseSettings):
|
||||
model_config = SettingsConfigDict(populate_by_name=True, extra="ignore")
|
||||
|
||||
# ----- single-destination shorthand, read from standard OTEL_* envs ----- #
|
||||
exporter: str = Field(
|
||||
default="console",
|
||||
validation_alias=AliasChoices("OTEL_EXPORTER", "OTEL_EXPORTER_OTLP_PROTOCOL"),
|
||||
description=(
|
||||
"Exporter kind for the single-destination shorthand. The model "
|
||||
"validator folds this (with ``endpoint`` / ``headers``) into a "
|
||||
"one-entry ``exporters`` list when ``exporters`` is empty; set "
|
||||
"``exporters`` directly for multiple destinations."
|
||||
),
|
||||
)
|
||||
endpoint: str | None = Field(
|
||||
default=None,
|
||||
validation_alias=AliasChoices("OTEL_ENDPOINT", "OTEL_EXPORTER_OTLP_ENDPOINT"),
|
||||
)
|
||||
headers: str | None = Field(
|
||||
default=None,
|
||||
validation_alias=AliasChoices("OTEL_HEADERS", "OTEL_EXPORTER_OTLP_HEADERS"),
|
||||
)
|
||||
service_name: str = Field(
|
||||
default="litellm", validation_alias=AliasChoices("OTEL_SERVICE_NAME")
|
||||
)
|
||||
deployment_environment: str | None = Field(
|
||||
default=None, validation_alias=AliasChoices("OTEL_ENVIRONMENT_NAME")
|
||||
)
|
||||
|
||||
enable_metrics: bool = Field(
|
||||
default=False,
|
||||
validation_alias=AliasChoices("LITELLM_OTEL_INTEGRATION_ENABLE_METRICS"),
|
||||
)
|
||||
enable_events: bool = Field(
|
||||
default=False,
|
||||
validation_alias=AliasChoices("LITELLM_OTEL_INTEGRATION_ENABLE_EVENTS"),
|
||||
)
|
||||
capture_message_content: str = Field(
|
||||
default=CaptureMessageContent.NO_CONTENT,
|
||||
validation_alias=AliasChoices(
|
||||
"OTEL_INSTRUMENTATION_GENAI_CAPTURE_MESSAGE_CONTENT"
|
||||
),
|
||||
)
|
||||
legacy_compat: bool = Field(
|
||||
default=True, validation_alias=AliasChoices("LITELLM_OTEL_LEGACY_COMPAT")
|
||||
)
|
||||
|
||||
# ----- explicit multi-destination / vocabulary configuration ------------ #
|
||||
|
||||
exporters: list[ExporterSpec] = Field(
|
||||
default_factory=list,
|
||||
description=(
|
||||
"One destination per spec. The shared TracerProvider attaches a "
|
||||
"SpanProcessor per entry. When empty, the model validator folds "
|
||||
"the ``exporter`` / ``endpoint`` / ``headers`` shorthand into a "
|
||||
"single spec so there is always at least one destination."
|
||||
),
|
||||
)
|
||||
|
||||
mapper_names: list[str] = Field(
|
||||
default_factory=lambda: ["genai"],
|
||||
description=(
|
||||
"Ordered attribute vocabularies to emit. ``genai`` is the "
|
||||
"canonical OTel GenAI vocabulary and is always placed first. "
|
||||
"Vendor names: ``openinference`` (Arize + Phoenix), ``langfuse``, "
|
||||
"``weave``, ``langtrace``."
|
||||
),
|
||||
)
|
||||
|
||||
resource_attributes: dict[str, str] = Field(
|
||||
default_factory=dict,
|
||||
description=(
|
||||
"Extra Resource attributes beyond ``service.name`` and "
|
||||
"``deployment.environment`` (e.g. integration-specific markers)."
|
||||
),
|
||||
)
|
||||
|
||||
baggage_promoted_keys: list[str] = Field(
|
||||
default_factory=lambda: list(BAGGAGE_PROMOTED_KEYS)
|
||||
)
|
||||
baggage_metadata_keys: list[str] = Field(
|
||||
default_factory=lambda: list(DEFAULT_BAGGAGE_METADATA_KEYS)
|
||||
)
|
||||
|
||||
@model_validator(mode="after")
|
||||
def _normalize(self) -> "OpenTelemetryV2Config":
|
||||
# An endpoint with the default exporter kind implies OTLP/HTTP.
|
||||
if self.endpoint and self.exporter == "console":
|
||||
self.exporter = "otlp_http"
|
||||
# When no explicit destinations are given, fold the single-destination
|
||||
# shorthand into one spec so the provider always has a destination.
|
||||
if not self.exporters:
|
||||
self.exporters = [
|
||||
ExporterSpec(
|
||||
kind=self.exporter,
|
||||
endpoint=self.endpoint,
|
||||
headers=self.headers,
|
||||
)
|
||||
]
|
||||
# Ensure ``genai`` is always present and first.
|
||||
names = list(self.mapper_names)
|
||||
if "genai" in names:
|
||||
names = ["genai"] + [n for n in names if n != "genai"]
|
||||
else:
|
||||
names = ["genai"] + names
|
||||
# When enabled, also emit attribute keys under their semconv-ai /
|
||||
# Traceloop names via the ``legacy`` mapper. Append it at the tail so
|
||||
# the canonical ``genai`` keys win on any conflict.
|
||||
if self.legacy_compat and "legacy" not in names:
|
||||
names.append("legacy")
|
||||
self.mapper_names = names
|
||||
return self
|
||||
|
||||
@property
|
||||
def capture_span_content(self) -> bool:
|
||||
"""Whether prompt/response content may be stamped as span attributes.
|
||||
|
||||
Defaults off (``no_content``): an operator must opt in before message
|
||||
bodies leave the process, so a user request can never force its prompt
|
||||
or completion into the configured backend while capture is disabled.
|
||||
"""
|
||||
return self.capture_message_content in (
|
||||
CaptureMessageContent.SPAN_ONLY,
|
||||
CaptureMessageContent.SPAN_AND_EVENT,
|
||||
)
|
||||
|
||||
@classmethod
|
||||
def from_env(cls) -> "OpenTelemetryV2Config":
|
||||
return cls()
|
||||
51
litellm/integrations/otel/context.py
Normal file
51
litellm/integrations/otel/context.py
Normal file
|
|
@ -0,0 +1,51 @@
|
|||
"""Trace-context + Baggage helpers."""
|
||||
|
||||
from typing import Mapping
|
||||
|
||||
from opentelemetry import baggage
|
||||
from opentelemetry.context import Context, get_current
|
||||
from opentelemetry.trace import Span, set_span_in_context
|
||||
from opentelemetry.trace.propagation.tracecontext import (
|
||||
TraceContextTextMapPropagator,
|
||||
)
|
||||
|
||||
_PROPAGATOR = TraceContextTextMapPropagator()
|
||||
|
||||
|
||||
def set_request_baggage(
|
||||
values: Mapping[str, str], context: Context | None = None
|
||||
) -> Context:
|
||||
"""Return a context with ``values`` written into Baggage."""
|
||||
ctx = context
|
||||
for key, value in values.items():
|
||||
ctx = baggage.set_baggage(key, value, context=ctx)
|
||||
return ctx if ctx is not None else (context or get_current())
|
||||
|
||||
|
||||
def get_baggage_attributes(context: Context | None = None) -> dict[str, str]:
|
||||
"""All Baggage entries on ``context`` as strings."""
|
||||
return {key: str(value) for key, value in baggage.get_all(context).items()}
|
||||
|
||||
|
||||
def context_from_span(span: Span, context: Context | None = None) -> Context:
|
||||
"""A context with ``span`` as the active span (for explicit parenting)."""
|
||||
return set_span_in_context(span, context=context)
|
||||
|
||||
|
||||
def is_recordable_span(obj: object) -> bool:
|
||||
"""True if ``obj`` is a live span with a valid context (safe to parent under)."""
|
||||
if not isinstance(obj, Span):
|
||||
return False
|
||||
try:
|
||||
ctx = obj.get_span_context()
|
||||
except Exception:
|
||||
return False
|
||||
return ctx is not None and ctx.is_valid
|
||||
|
||||
|
||||
def extract_traceparent(headers: Mapping[str, str]) -> Context | None:
|
||||
"""Extract a remote parent context from incoming HTTP headers, if present."""
|
||||
if not any(key.lower() == "traceparent" for key in headers):
|
||||
return None
|
||||
carrier = {str(key).lower(): value for key, value in headers.items()}
|
||||
return _PROPAGATOR.extract(carrier)
|
||||
150
litellm/integrations/otel/emitter.py
Normal file
150
litellm/integrations/otel/emitter.py
Normal file
|
|
@ -0,0 +1,150 @@
|
|||
"""The span engine: dedup, start, run the mapper chain, set status, end."""
|
||||
|
||||
from collections import OrderedDict
|
||||
from typing import Callable, Sequence
|
||||
|
||||
from opentelemetry.context import Context
|
||||
from opentelemetry.trace import Span, Tracer
|
||||
from opentelemetry.trace.status import Status, StatusCode
|
||||
|
||||
from litellm.integrations.otel.config import OpenTelemetryV2Config
|
||||
from litellm.integrations.otel.mappers import resolve_mappers
|
||||
from litellm.integrations.otel.mappers.base import AttributeMapper, SpanData
|
||||
from litellm.integrations.otel.payloads import (
|
||||
GuardrailSpanData,
|
||||
LLMCallSpanData,
|
||||
ServiceSpanData,
|
||||
)
|
||||
from litellm.integrations.otel.providers import to_otel_span_kind
|
||||
from litellm.integrations.otel.semconv import Error
|
||||
from litellm.integrations.otel.spans import (
|
||||
SPAN_REGISTRY,
|
||||
SpanRole,
|
||||
guardrail_span_name,
|
||||
llm_call_span_name,
|
||||
service_span_name,
|
||||
)
|
||||
|
||||
# Roles emit() knows how to name and emit. PROXY_REQUEST and the management
|
||||
# routes are SERVER spans owned by the mounted FastAPI instrumentor, so they
|
||||
# have no builder here.
|
||||
_NAME_BUILDERS: dict[SpanRole, Callable[..., str]] = {
|
||||
SpanRole.LLM_CALL: llm_call_span_name,
|
||||
SpanRole.GUARDRAIL: guardrail_span_name,
|
||||
SpanRole.SERVICE: service_span_name,
|
||||
}
|
||||
|
||||
# Cap on the dedup cache. It only needs to coalesce the sync+async firing window
|
||||
# of a single in-flight request, so a bounded LRU keeps memory flat on a
|
||||
# long-running proxy while still covering every concurrently-open call.
|
||||
_DEDUP_CACHE_MAX = 10_000
|
||||
|
||||
|
||||
class SpanEmitter:
|
||||
def __init__(
|
||||
self,
|
||||
tracer: Tracer,
|
||||
config: OpenTelemetryV2Config,
|
||||
mappers: Sequence[AttributeMapper] | None = None,
|
||||
) -> None:
|
||||
self._tracer = tracer
|
||||
self._config = config
|
||||
# The mapper chain is the sole source of span attributes. When not
|
||||
# passed in, resolve it from the config so there's one source of truth.
|
||||
self._mappers: list[AttributeMapper] = (
|
||||
list(mappers)
|
||||
if mappers is not None
|
||||
else resolve_mappers(config.mapper_names)
|
||||
)
|
||||
# Bounded LRU (ordered by insertion / most-recent touch). Storing keys
|
||||
# only — the value is unused — so it behaves like a capped set.
|
||||
self._emitted: "OrderedDict[tuple[str, SpanRole], None]" = OrderedDict()
|
||||
|
||||
# -- low-level helpers --------------------------------------------------- #
|
||||
|
||||
def start_span(
|
||||
self,
|
||||
role: SpanRole,
|
||||
name: str,
|
||||
parent_context: Context | None = None,
|
||||
start_time_ns: int | None = None,
|
||||
*,
|
||||
tracer: Tracer | None = None,
|
||||
) -> Span:
|
||||
"""Start a span for ``role`` without dedup or attribute mapping.
|
||||
|
||||
For callers that own and manage their own span lifecycle. ``tracer``
|
||||
overrides the bound tracer for this span only, used for per-request
|
||||
multi-tenant credential routing.
|
||||
"""
|
||||
return (tracer or self._tracer).start_span(
|
||||
name,
|
||||
context=parent_context,
|
||||
kind=to_otel_span_kind(SPAN_REGISTRY[role].kind),
|
||||
start_time=start_time_ns,
|
||||
)
|
||||
|
||||
def _seen(self, dedup_key: str | None, role: SpanRole) -> bool:
|
||||
"""Return True once a ``(dedup_key, role)`` pair has been emitted.
|
||||
|
||||
Guards against emitting the same span twice when a streaming call
|
||||
fires both a sync and an async logging callback.
|
||||
"""
|
||||
if not dedup_key:
|
||||
return False
|
||||
marker = (dedup_key, role)
|
||||
if marker in self._emitted:
|
||||
self._emitted.move_to_end(marker)
|
||||
return True
|
||||
self._emitted[marker] = None
|
||||
if len(self._emitted) > _DEDUP_CACHE_MAX:
|
||||
self._emitted.popitem(last=False) # evict least-recently-used
|
||||
return False
|
||||
|
||||
# -- the engine ---------------------------------------------------------- #
|
||||
|
||||
def emit(
|
||||
self,
|
||||
role: SpanRole,
|
||||
data: SpanData,
|
||||
parent_context: Context | None = None,
|
||||
*,
|
||||
start_time_ns: int | None = None,
|
||||
end_time_ns: int | None = None,
|
||||
tracer: Tracer | None = None,
|
||||
) -> Span | None:
|
||||
"""Emit one complete span: dedup, start, map attributes, status, end.
|
||||
|
||||
Return the span, or ``None`` if it was deduplicated away. ``tracer``
|
||||
overrides the bound tracer for this span, used for per-request routing.
|
||||
"""
|
||||
# Only LLM-call spans carry a dedup key; LLM-call and service spans
|
||||
# carry an ``error`` field. ``isinstance`` narrows the type for mypy and
|
||||
# keeps the engine free of duck-typed attribute reads.
|
||||
dedup_key = data.identity.call_id if isinstance(data, LLMCallSpanData) else None
|
||||
if self._seen(dedup_key, role):
|
||||
return None
|
||||
span = self.start_span(
|
||||
role,
|
||||
_NAME_BUILDERS[role](data),
|
||||
parent_context=parent_context,
|
||||
start_time_ns=start_time_ns,
|
||||
tracer=tracer,
|
||||
)
|
||||
for mapper in self._mappers:
|
||||
for key, value in mapper.map(data).items():
|
||||
span.set_attribute(key, value)
|
||||
error = (
|
||||
data.error
|
||||
if isinstance(data, (LLMCallSpanData, ServiceSpanData, GuardrailSpanData))
|
||||
else None
|
||||
)
|
||||
if error and (error.error_type or error.message):
|
||||
span.set_attribute(Error.TYPE, error.error_type or "error")
|
||||
span.set_status(
|
||||
Status(StatusCode.ERROR, error.message or error.error_type or "error")
|
||||
)
|
||||
else:
|
||||
span.set_status(Status(StatusCode.OK))
|
||||
span.end(end_time=end_time_ns)
|
||||
return span
|
||||
425
litellm/integrations/otel/logger.py
Normal file
425
litellm/integrations/otel/logger.py
Normal file
|
|
@ -0,0 +1,425 @@
|
|||
"""``CustomLogger`` adapter on the OpenTelemetry span engine.
|
||||
|
||||
Thin adapter: it translates litellm's logging callbacks into typed ``*SpanData``
|
||||
and hands them to the engine (:mod:`emitter`), with multi-tenant tracer routing
|
||||
in :mod:`routing`. It emits the gen-ai spans (LLM call, guardrail, service).
|
||||
|
||||
The proxy server span is NOT owned here. It is created by the FastAPI
|
||||
instrumentation mounted in ``proxy_server``'s startup event, which stamps the
|
||||
``http.*`` attributes and handles inbound context propagation. The proxy-span
|
||||
methods below are therefore no-ops: routes never modify spans.
|
||||
|
||||
Gen-ai spans parent to that server span via the ambient OTel context rather
|
||||
than an explicitly threaded span. litellm's async logging worker copies the
|
||||
request's context at enqueue time, so ``async_log_success_event`` runs with the
|
||||
server span active. Emission is therefore async-only — the sync callback runs
|
||||
in an out-of-context thread, where there is no parent span, so it is a no-op.
|
||||
"""
|
||||
|
||||
from datetime import datetime
|
||||
from typing import Any, Mapping, cast
|
||||
|
||||
from opentelemetry.context import attach, get_current
|
||||
from opentelemetry.sdk.trace import TracerProvider
|
||||
from opentelemetry.trace import Span, Tracer, get_current_span
|
||||
|
||||
import litellm
|
||||
from litellm.integrations.custom_logger import CustomLogger
|
||||
from litellm.integrations.otel.baggage import promoted_baggage
|
||||
from litellm.integrations.otel.config import OpenTelemetryV2Config
|
||||
from litellm.integrations.otel.context import (
|
||||
context_from_span,
|
||||
is_recordable_span,
|
||||
set_request_baggage,
|
||||
)
|
||||
from litellm.integrations.otel.emitter import SpanEmitter
|
||||
from litellm.integrations.otel.mappers import resolve_mappers
|
||||
from litellm.integrations.otel.payloads import (
|
||||
GuardrailSpanData,
|
||||
LLMCallSpanData,
|
||||
RequestIdentity,
|
||||
ServiceSpanData,
|
||||
SpanError,
|
||||
)
|
||||
from litellm.integrations.otel.providers import build_tracer_provider, get_tracer
|
||||
from litellm.integrations.otel.routing import TenantTracerCache
|
||||
from litellm.integrations.otel.spans import SpanRole
|
||||
from litellm.integrations.otel.utils import to_ns
|
||||
|
||||
LITELLM_TRACER_NAME = "litellm"
|
||||
LITELLM_PROXY_REQUEST_SPAN_NAME = "Received Proxy Server Request"
|
||||
|
||||
# Any callback whose class belongs to one of these modules is "the OTel
|
||||
# callback" for proxy-global-registration purposes.
|
||||
_OTEL_MODULES = (
|
||||
"litellm.integrations.otel",
|
||||
"litellm.integrations.opentelemetry",
|
||||
)
|
||||
|
||||
|
||||
def _pre_call_guardrail_blocked(payload: Mapping[str, Any]) -> bool:
|
||||
"""True when a pre-call guardrail blocked the request (no LLM call happened).
|
||||
|
||||
A blocked pre-call guardrail raises before the upstream call, yet litellm
|
||||
still emits a failure log — which would otherwise produce a phantom CLIENT
|
||||
span for a call that never occurred. We detect the case (request failed AND a
|
||||
``pre_call`` guardrail intervened) so the caller can skip that span. A
|
||||
pre-call guardrail that merely *masks* lets the call proceed, so the request
|
||||
succeeds and this returns False — only genuine blocks fail the request.
|
||||
"""
|
||||
if payload.get("status") != "failure":
|
||||
return False
|
||||
info = payload.get("guardrail_information")
|
||||
if not isinstance(info, list):
|
||||
return False
|
||||
for entry in info:
|
||||
if not isinstance(entry, dict):
|
||||
continue
|
||||
mode = entry.get("guardrail_mode")
|
||||
is_pre_call = mode == "pre_call" or (
|
||||
isinstance(mode, (list, tuple)) and "pre_call" in mode
|
||||
)
|
||||
if is_pre_call and entry.get("guardrail_status") == "guardrail_intervened":
|
||||
return True
|
||||
return False
|
||||
|
||||
|
||||
class OpenTelemetryV2(CustomLogger):
|
||||
"""The ``CustomLogger`` for OpenTelemetry.
|
||||
|
||||
The constructor accepts an optional config, callback name, and pre-built
|
||||
OTel providers; when a provider is omitted it is built from the config.
|
||||
``logger_provider`` and ``meter_provider`` are accepted but reserved for
|
||||
future OTel logs and metrics support.
|
||||
"""
|
||||
|
||||
def __init__(
|
||||
self,
|
||||
config: OpenTelemetryV2Config | None = None,
|
||||
callback_name: str | None = None,
|
||||
tracer_provider: TracerProvider | None = None,
|
||||
logger_provider: Any | None = None, # reserved for OTel logs
|
||||
meter_provider: Any | None = None, # reserved for metrics
|
||||
**kwargs: Any,
|
||||
) -> None:
|
||||
super().__init__(**kwargs)
|
||||
self.config: OpenTelemetryV2Config = config or OpenTelemetryV2Config()
|
||||
self.callback_name = callback_name
|
||||
self._tracer_provider: TracerProvider = (
|
||||
tracer_provider
|
||||
if tracer_provider is not None
|
||||
else build_tracer_provider(self.config)
|
||||
)
|
||||
self.tracer: Tracer = get_tracer(self._tracer_provider, LITELLM_TRACER_NAME)
|
||||
self._emitter = SpanEmitter(
|
||||
self.tracer, self.config, mappers=resolve_mappers(self.config.mapper_names)
|
||||
)
|
||||
self._tenant_tracers = TenantTracerCache(
|
||||
self.config, callback_name, LITELLM_TRACER_NAME
|
||||
)
|
||||
self._init_otel_logger_on_litellm_proxy()
|
||||
|
||||
# ====================================================================== #
|
||||
# Proxy global registration
|
||||
# ====================================================================== #
|
||||
|
||||
def _init_otel_logger_on_litellm_proxy(self) -> None:
|
||||
"""Claim ``proxy_server.open_telemetry_logger`` if no one else has."""
|
||||
try:
|
||||
from litellm.proxy import proxy_server
|
||||
except Exception:
|
||||
return
|
||||
try:
|
||||
# Mutate ``litellm.service_callback`` in place. ``getattr(..) or []``
|
||||
# would bind a throwaway local when the list is empty (an empty list
|
||||
# is falsy), so the append would never reach the global and service
|
||||
# spans (Redis, Postgres, …) would be silently dropped from traces.
|
||||
service_callback = litellm.service_callback
|
||||
already_otel = any(
|
||||
cb.__class__.__module__.startswith(_OTEL_MODULES)
|
||||
for cb in service_callback
|
||||
if hasattr(cb, "__class__")
|
||||
)
|
||||
if not already_otel:
|
||||
service_callback.append(self)
|
||||
except Exception:
|
||||
pass
|
||||
if getattr(proxy_server, "open_telemetry_logger", None) is None:
|
||||
setattr(proxy_server, "open_telemetry_logger", self)
|
||||
|
||||
# ====================================================================== #
|
||||
# LLM-call callbacks
|
||||
# ====================================================================== #
|
||||
|
||||
# Async-only: the async path runs inside the request's restored OTel context
|
||||
# (the logging worker copies it at enqueue), so the span parents to the
|
||||
# instrumentor's server span via ambient context. The sync path runs in an
|
||||
# out-of-context thread with no parent span, so it is a no-op.
|
||||
|
||||
def log_success_event(self, kwargs, response_obj, start_time, end_time):
|
||||
return None
|
||||
|
||||
def log_failure_event(self, kwargs, response_obj, start_time, end_time):
|
||||
return None
|
||||
|
||||
async def async_log_success_event(self, kwargs, response_obj, start_time, end_time):
|
||||
self._emit_llm_call(kwargs, start_time, end_time)
|
||||
|
||||
async def async_log_failure_event(self, kwargs, response_obj, start_time, end_time):
|
||||
self._emit_llm_call(kwargs, start_time, end_time)
|
||||
|
||||
def _emit_llm_call(
|
||||
self,
|
||||
kwargs: Mapping[str, Any],
|
||||
start_time: datetime | float | None,
|
||||
end_time: datetime | float | None,
|
||||
) -> Span | None:
|
||||
payload = kwargs.get("standard_logging_object")
|
||||
if not payload:
|
||||
return None
|
||||
if _pre_call_guardrail_blocked(cast("Mapping[str, Any]", payload)):
|
||||
# A pre-call guardrail blocked the request, so the upstream LLM was
|
||||
# never called — litellm still emits a failure log, but a CLIENT
|
||||
# "chat …" span for a call that didn't happen is misleading. Skip it;
|
||||
# the guardrail span (ERROR, with the verdict) is the real outcome.
|
||||
return None
|
||||
data = LLMCallSpanData.from_standard_logging_payload(
|
||||
cast("Any", payload), capture_content=self.config.capture_span_content
|
||||
)
|
||||
# Parent is the ambient context (the instrumentor's server span,
|
||||
# restored by the logging worker); no span is threaded through metadata.
|
||||
parent_ctx = get_current()
|
||||
# Write identity into Baggage so child spans (guardrails, services)
|
||||
# inherit it.
|
||||
bag = promoted_baggage(
|
||||
data.identity,
|
||||
data.request_model,
|
||||
promoted_keys=tuple(self.config.baggage_promoted_keys),
|
||||
metadata_keys=tuple(self.config.baggage_metadata_keys),
|
||||
)
|
||||
if bag:
|
||||
parent_ctx = set_request_baggage(bag, context=parent_ctx)
|
||||
return self._emitter.emit(
|
||||
SpanRole.LLM_CALL,
|
||||
data,
|
||||
parent_context=parent_ctx,
|
||||
start_time_ns=to_ns(start_time),
|
||||
end_time_ns=to_ns(end_time),
|
||||
tracer=self._tenant_tracers.tracer_for(
|
||||
self.tracer, kwargs.get("standard_callback_dynamic_params")
|
||||
),
|
||||
)
|
||||
|
||||
# ====================================================================== #
|
||||
# Service hooks
|
||||
# ====================================================================== #
|
||||
|
||||
async def async_service_success_hook(
|
||||
self,
|
||||
payload: Any,
|
||||
parent_otel_span: Span | None = None,
|
||||
start_time: datetime | float | None = None,
|
||||
end_time: datetime | float | None = None,
|
||||
event_metadata: dict | None = None,
|
||||
) -> None:
|
||||
self._emit_service(
|
||||
payload,
|
||||
parent_otel_span=parent_otel_span,
|
||||
start_time=start_time,
|
||||
end_time=end_time,
|
||||
event_metadata=event_metadata,
|
||||
error_override=None,
|
||||
)
|
||||
|
||||
async def async_service_failure_hook(
|
||||
self,
|
||||
payload: Any,
|
||||
error: str | None = "",
|
||||
parent_otel_span: Span | None = None,
|
||||
start_time: datetime | float | None = None,
|
||||
end_time: datetime | float | None = None,
|
||||
event_metadata: dict | None = None,
|
||||
) -> None:
|
||||
self._emit_service(
|
||||
payload,
|
||||
parent_otel_span=parent_otel_span,
|
||||
start_time=start_time,
|
||||
end_time=end_time,
|
||||
event_metadata=event_metadata,
|
||||
error_override=error or "error",
|
||||
)
|
||||
|
||||
def _emit_service(
|
||||
self,
|
||||
payload: Any,
|
||||
*,
|
||||
parent_otel_span: Span | None,
|
||||
start_time: datetime | float | None,
|
||||
end_time: datetime | float | None,
|
||||
event_metadata: dict | None,
|
||||
error_override: str | None,
|
||||
) -> Span | None:
|
||||
if not is_recordable_span(parent_otel_span):
|
||||
return None
|
||||
data = ServiceSpanData.from_payload(payload, event_metadata=event_metadata)
|
||||
if error_override is not None and data.error is None:
|
||||
data = ServiceSpanData(
|
||||
service_name=data.service_name,
|
||||
call_type=data.call_type,
|
||||
error=SpanError(message=error_override),
|
||||
event_metadata=data.event_metadata,
|
||||
)
|
||||
# Parent to the server span, but layer it over the ambient context so the
|
||||
# identity Baggage seeded in ``async_pre_call_hook`` rides along and the
|
||||
# service span gets the same identity attributes as the LLM-call span.
|
||||
parent_context = context_from_span(
|
||||
cast(Span, parent_otel_span), context=get_current()
|
||||
)
|
||||
return self._emitter.emit(
|
||||
SpanRole.SERVICE,
|
||||
data,
|
||||
parent_context=parent_context,
|
||||
start_time_ns=to_ns(start_time),
|
||||
end_time_ns=to_ns(end_time),
|
||||
)
|
||||
|
||||
# ====================================================================== #
|
||||
# async_post_call_* hooks — emit guardrail spans. The server span's status
|
||||
# / errors are the FastAPI instrumentor's job, so we don't touch it here.
|
||||
# ====================================================================== #
|
||||
|
||||
async def async_pre_call_hook(
|
||||
self,
|
||||
user_api_key_dict: Any,
|
||||
cache: Any,
|
||||
data: dict,
|
||||
call_type: Any,
|
||||
) -> dict:
|
||||
"""Seed request identity into Baggage at the start of the request.
|
||||
|
||||
This runs in the request task (the server span is the ambient context),
|
||||
so attaching the identity Baggage here makes **every** span emitted for
|
||||
the request — LLM call, guardrail, and service — inherit it via
|
||||
``LiteLLMBaggageSpanProcessor``. Without this, only the LLM-call span got
|
||||
identity (it promoted Baggage locally) and the guardrail/service spans,
|
||||
which parent to the server span, had none. The async logging worker
|
||||
copies this context at enqueue time, so the LLM-call span inherits it too.
|
||||
"""
|
||||
try:
|
||||
identity = RequestIdentity.from_user_api_key_auth(user_api_key_dict)
|
||||
bag = promoted_baggage(
|
||||
identity,
|
||||
data.get("model") if isinstance(data, dict) else None,
|
||||
promoted_keys=tuple(self.config.baggage_promoted_keys),
|
||||
metadata_keys=tuple(self.config.baggage_metadata_keys),
|
||||
)
|
||||
if bag:
|
||||
# Attach (no detach): the contextvar is scoped to this request's
|
||||
# asyncio task and is reclaimed when the task ends.
|
||||
attach(set_request_baggage(bag, context=get_current()))
|
||||
# The server span was started by the instrumentor before this
|
||||
# hook ran, so the Baggage processor (which only fires at span
|
||||
# start) won't backfill it — stamp identity on it directly.
|
||||
server_span = get_current_span()
|
||||
if is_recordable_span(server_span):
|
||||
for key, value in bag.items():
|
||||
server_span.set_attribute(key, value)
|
||||
except Exception:
|
||||
pass
|
||||
return data
|
||||
|
||||
async def async_post_call_success_hook(
|
||||
self,
|
||||
data: Mapping[str, Any],
|
||||
user_api_key_dict: Any,
|
||||
response: Any,
|
||||
) -> Any:
|
||||
self._emit_guardrail_spans(data)
|
||||
return response
|
||||
|
||||
async def async_post_call_failure_hook(
|
||||
self,
|
||||
request_data: Mapping[str, Any],
|
||||
original_exception: BaseException | None,
|
||||
user_api_key_dict: Any,
|
||||
traceback_str: str | None = None,
|
||||
) -> None:
|
||||
self._emit_guardrail_spans(request_data)
|
||||
|
||||
def _emit_guardrail_spans(self, request_data: Mapping[str, Any]) -> None:
|
||||
# Post-call hooks run in the request task, so the ambient context is the
|
||||
# server span; guardrail spans parent to it implicitly via that context.
|
||||
metadata = request_data.get("metadata")
|
||||
guardrails: list[Any] = []
|
||||
if isinstance(metadata, dict):
|
||||
info = metadata.get("standard_logging_guardrail_information")
|
||||
if isinstance(info, list):
|
||||
guardrails = info
|
||||
elif isinstance(info, dict):
|
||||
guardrails = [info]
|
||||
for entry in guardrails:
|
||||
if not isinstance(entry, dict):
|
||||
continue
|
||||
self._emitter.emit(
|
||||
SpanRole.GUARDRAIL, GuardrailSpanData.from_logging_entry(entry)
|
||||
)
|
||||
|
||||
# ====================================================================== #
|
||||
# Management endpoint hooks — no-ops. Management endpoints are ordinary
|
||||
# FastAPI routes, so the mounted instrumentor already spans them.
|
||||
# ====================================================================== #
|
||||
|
||||
async def async_management_endpoint_success_hook(
|
||||
self,
|
||||
logging_payload: Any,
|
||||
parent_otel_span: Span | None = None,
|
||||
) -> None:
|
||||
return None
|
||||
|
||||
async def async_management_endpoint_failure_hook(
|
||||
self,
|
||||
logging_payload: Any,
|
||||
parent_otel_span: Span | None = None,
|
||||
) -> None:
|
||||
return None
|
||||
|
||||
# ====================================================================== #
|
||||
# Proxy SERVER-span API — no-ops. The FastAPI instrumentor owns the server
|
||||
# span (creation, http.* attributes, inbound propagation) and gen-ai spans
|
||||
# parent to it via ambient context. These methods are the surface the
|
||||
# proxy and auth call sites invoke; they intentionally do nothing.
|
||||
# ====================================================================== #
|
||||
|
||||
def create_litellm_proxy_request_started_span(
|
||||
self, start_time: datetime, headers: Mapping[str, str] | None
|
||||
) -> Span | None:
|
||||
"""Return the active server span instead of creating one.
|
||||
|
||||
The FastAPI instrumentor owns the server span, so V2 creates nothing
|
||||
here. But the proxy threads this return value as ``litellm_parent_otel_span``
|
||||
— and service logging (Redis, Postgres, …) only invokes the OTel service
|
||||
hook when that parent is non-None. Returning the ambient server span lets
|
||||
service spans nest under it. The proxy must NOT ``.end()`` this span (the
|
||||
instrumentor does); ``_close_dangling_otel_server_span`` skips it under V2.
|
||||
"""
|
||||
span = get_current_span()
|
||||
return span if is_recordable_span(span) else None
|
||||
|
||||
@staticmethod
|
||||
def set_proxy_request_route_attributes(
|
||||
span: Span | None,
|
||||
*,
|
||||
url_path: str | None = None,
|
||||
http_route: str | None = None,
|
||||
) -> None:
|
||||
"""No-op: the FastAPI instrumentor stamps ``http.route`` / ``url.path``."""
|
||||
|
||||
@staticmethod
|
||||
def set_response_status_code_attribute(
|
||||
span: Span | None, status_code: int | None
|
||||
) -> None:
|
||||
"""No-op: the FastAPI instrumentor stamps ``http.response.status_code``."""
|
||||
|
||||
@staticmethod
|
||||
def set_preprocessing_duration_attribute(span: Span | None, container: Any) -> None:
|
||||
"""No-op: the server span belongs to the FastAPI instrumentor."""
|
||||
58
litellm/integrations/otel/mappers/__init__.py
Normal file
58
litellm/integrations/otel/mappers/__init__.py
Normal file
|
|
@ -0,0 +1,58 @@
|
|||
"""Attribute mappers: pure ``LLMCallSpanData -> {attribute key: value}`` functions.
|
||||
|
||||
Composition over inheritance: vocabularies layer onto the same span. Listing
|
||||
``["genai", "openinference"]`` in ``config.mapper_names`` makes every span
|
||||
carry both the canonical ``gen_ai.*`` keys and the OpenInference (Arize +
|
||||
Phoenix) keys. Add ``"langfuse"`` and it works for all three backends at once.
|
||||
"""
|
||||
|
||||
from typing import Callable, Iterable
|
||||
|
||||
from litellm.integrations.otel.mappers.base import (
|
||||
AttributeMap,
|
||||
AttributeMapper,
|
||||
AttrValue,
|
||||
)
|
||||
from litellm.integrations.otel.mappers.genai import GenAIMapper
|
||||
from litellm.integrations.otel.mappers.langfuse import LangfuseMapper
|
||||
from litellm.integrations.otel.mappers.langtrace import LangtraceMapper
|
||||
from litellm.integrations.otel.mappers.legacy import LegacyMapper
|
||||
from litellm.integrations.otel.mappers.openinference import OpenInferenceMapper
|
||||
from litellm.integrations.otel.mappers.weave import WeaveMapper
|
||||
|
||||
# Registry keyed by ``config.mapper_names`` entries.
|
||||
_MAPPER_BY_NAME: dict[str, Callable[[], AttributeMapper]] = {
|
||||
"genai": GenAIMapper,
|
||||
"legacy": LegacyMapper,
|
||||
"openinference": OpenInferenceMapper,
|
||||
"langfuse": LangfuseMapper,
|
||||
"weave": WeaveMapper,
|
||||
"langtrace": LangtraceMapper,
|
||||
}
|
||||
|
||||
|
||||
def resolve_mappers(names: Iterable[str]) -> list[AttributeMapper]:
|
||||
"""Resolve mapper names to instances. Unknown names raise ``ValueError``."""
|
||||
out: list[AttributeMapper] = []
|
||||
for name in names:
|
||||
factory = _MAPPER_BY_NAME.get(name)
|
||||
if factory is None:
|
||||
raise ValueError(
|
||||
f"unknown mapper name {name!r}; known: " f"{sorted(_MAPPER_BY_NAME)}"
|
||||
)
|
||||
out.append(factory())
|
||||
return out
|
||||
|
||||
|
||||
__all__ = [
|
||||
"AttributeMap",
|
||||
"AttributeMapper",
|
||||
"AttrValue",
|
||||
"GenAIMapper",
|
||||
"LangfuseMapper",
|
||||
"LangtraceMapper",
|
||||
"LegacyMapper",
|
||||
"OpenInferenceMapper",
|
||||
"WeaveMapper",
|
||||
"resolve_mappers",
|
||||
]
|
||||
36
litellm/integrations/otel/mappers/base.py
Normal file
36
litellm/integrations/otel/mappers/base.py
Normal file
|
|
@ -0,0 +1,36 @@
|
|||
"""Mapper protocol and attribute value types."""
|
||||
|
||||
from typing import Sequence
|
||||
|
||||
from typing_extensions import Protocol, runtime_checkable
|
||||
|
||||
from litellm.integrations.otel.payloads import (
|
||||
GuardrailSpanData,
|
||||
LLMCallSpanData,
|
||||
ServiceSpanData,
|
||||
)
|
||||
|
||||
AttrScalar = str | bool | int | float
|
||||
# Mirrors ``opentelemetry.util.types.AttributeValue`` (homogeneous sequences)
|
||||
# without importing the SDK, so mappers stay OTel-free.
|
||||
AttrValue = (
|
||||
AttrScalar | Sequence[str] | Sequence[bool] | Sequence[int] | Sequence[float]
|
||||
)
|
||||
AttributeMap = dict[str, AttrValue]
|
||||
|
||||
# The closed set of span-data types the engine routes through the mapper chain.
|
||||
# Server spans (PROXY_REQUEST + management routes) belong to the mounted FastAPI
|
||||
# instrumentor, not the mapper chain.
|
||||
SpanData = LLMCallSpanData | GuardrailSpanData | ServiceSpanData
|
||||
|
||||
|
||||
@runtime_checkable
|
||||
class AttributeMapper(Protocol):
|
||||
"""Maps a typed span input to a flat dict of OTel span attributes.
|
||||
|
||||
One method per mapper, dispatched internally on the ``data`` type. The
|
||||
engine calls this uniformly for every span kind — mappers that don't speak
|
||||
a given type return ``{}``. This is why the engine contains no attribute keys.
|
||||
"""
|
||||
|
||||
def map(self, data: SpanData) -> AttributeMap: ...
|
||||
121
litellm/integrations/otel/mappers/genai.py
Normal file
121
litellm/integrations/otel/mappers/genai.py
Normal file
|
|
@ -0,0 +1,121 @@
|
|||
"""Canonical OpenTelemetry GenAI semantic-convention mapper (always active).
|
||||
|
||||
Owns the attribute schema for every span kind the engine emits — LLM call,
|
||||
guardrail, and service — so the engine itself never references attribute keys.
|
||||
|
||||
Each span kind declares its schema as a flat ``attribute key -> extractor``
|
||||
table: one lambda per mapping operation, applied against the typed span data.
|
||||
"""
|
||||
|
||||
from typing import Callable
|
||||
|
||||
from litellm.integrations.otel.mappers.base import AttributeMap, AttrValue, SpanData
|
||||
from litellm.integrations.otel.mappers.utils import collect, drop_none
|
||||
from litellm.integrations.otel.payloads import (
|
||||
GuardrailSpanData,
|
||||
LLMCallSpanData,
|
||||
ServiceSpanData,
|
||||
ToolDefinition,
|
||||
)
|
||||
from litellm.integrations.otel.semconv import Error, GenAI, LiteLLM, Server
|
||||
|
||||
|
||||
class GenAIMapper:
|
||||
|
||||
_LLM_CALL_ATTRS: dict[str, Callable[[LLMCallSpanData], AttrValue | None]] = {
|
||||
GenAI.OPERATION_NAME: lambda d: d.operation.value,
|
||||
GenAI.PROVIDER_NAME: lambda d: d.provider or None,
|
||||
GenAI.REQUEST_MODEL: lambda d: d.request_model or None,
|
||||
GenAI.REQUEST_TEMPERATURE: lambda d: d.request_params.temperature,
|
||||
GenAI.REQUEST_TOP_P: lambda d: d.request_params.top_p,
|
||||
GenAI.REQUEST_TOP_K: lambda d: d.request_params.top_k,
|
||||
GenAI.REQUEST_MAX_TOKENS: lambda d: d.request_params.max_tokens,
|
||||
GenAI.REQUEST_FREQUENCY_PENALTY: lambda d: d.request_params.frequency_penalty,
|
||||
GenAI.REQUEST_PRESENCE_PENALTY: lambda d: d.request_params.presence_penalty,
|
||||
GenAI.REQUEST_STOP_SEQUENCES: lambda d: (
|
||||
list(d.request_params.stop_sequences)
|
||||
if d.request_params.stop_sequences
|
||||
else None
|
||||
),
|
||||
GenAI.REQUEST_SEED: lambda d: d.request_params.seed,
|
||||
GenAI.RESPONSE_MODEL: lambda d: d.response_model,
|
||||
GenAI.RESPONSE_ID: lambda d: d.response_id,
|
||||
GenAI.RESPONSE_FINISH_REASONS: lambda d: (
|
||||
list(d.finish_reasons) if d.finish_reasons else None
|
||||
),
|
||||
GenAI.USAGE_INPUT_TOKENS: lambda d: d.usage.input_tokens,
|
||||
GenAI.USAGE_OUTPUT_TOKENS: lambda d: d.usage.output_tokens,
|
||||
Error.TYPE: lambda d: d.error.error_type if d.error else None,
|
||||
Server.ADDRESS: lambda d: d.server.address if d.server else None,
|
||||
Server.PORT: lambda d: d.server.port if d.server else None,
|
||||
LiteLLM.CALL_ID: lambda d: d.identity.call_id or None,
|
||||
f"{LiteLLM.COST_PREFIX}total": lambda d: d.response_cost,
|
||||
LiteLLM.REQUEST_STREAMING: lambda d: d.is_streaming,
|
||||
}
|
||||
|
||||
_TOOL_ATTRS: dict[str, Callable[[ToolDefinition], AttrValue | None]] = {
|
||||
"name": lambda t: t.name,
|
||||
"description": lambda t: t.description or None,
|
||||
"parameters": lambda t: t.parameters_json or None,
|
||||
}
|
||||
|
||||
_GUARDRAIL_ATTRS: dict[str, Callable[[GuardrailSpanData], AttrValue | None]] = {
|
||||
LiteLLM.GUARDRAIL_NAME: lambda d: d.guardrail_name,
|
||||
LiteLLM.GUARDRAIL_MODE: lambda d: d.mode,
|
||||
LiteLLM.GUARDRAIL_STATUS: lambda d: d.status,
|
||||
LiteLLM.GUARDRAIL_PROVIDER: lambda d: d.provider,
|
||||
LiteLLM.GUARDRAIL_ACTION: lambda d: d.action,
|
||||
LiteLLM.GUARDRAIL_RESPONSE: lambda d: d.response_json,
|
||||
LiteLLM.GUARDRAIL_VIOLATION_CATEGORIES: lambda d: (
|
||||
list(d.violation_categories) if d.violation_categories else None
|
||||
),
|
||||
LiteLLM.GUARDRAIL_CONFIDENCE_SCORE: lambda d: d.confidence_score,
|
||||
LiteLLM.GUARDRAIL_RISK_SCORE: lambda d: d.risk_score,
|
||||
LiteLLM.GUARDRAIL_MASKED_ENTITY_COUNT: lambda d: d.masked_entity_count,
|
||||
LiteLLM.GUARDRAIL_DURATION: lambda d: d.duration,
|
||||
}
|
||||
|
||||
_SERVICE_ATTRS: dict[str, Callable[[ServiceSpanData], AttrValue | None]] = {
|
||||
LiteLLM.SERVICE_NAME: lambda d: d.service_name,
|
||||
LiteLLM.SERVICE_CALL_TYPE: lambda d: d.call_type,
|
||||
}
|
||||
|
||||
def map(self, data: SpanData) -> AttributeMap:
|
||||
match data:
|
||||
case LLMCallSpanData():
|
||||
return self._llm_call(data)
|
||||
case GuardrailSpanData():
|
||||
return self._guardrail(data)
|
||||
case ServiceSpanData():
|
||||
return self._service(data)
|
||||
case _:
|
||||
return {}
|
||||
|
||||
@classmethod
|
||||
def _llm_call(cls, data: LLMCallSpanData) -> AttributeMap:
|
||||
attrs = collect(cls._LLM_CALL_ATTRS, data)
|
||||
attrs.update(
|
||||
drop_none(
|
||||
{
|
||||
f"gen_ai.tool.{idx}.{suffix}": extract(tool)
|
||||
for idx, tool in enumerate(data.tools)
|
||||
for suffix, extract in cls._TOOL_ATTRS.items()
|
||||
}
|
||||
)
|
||||
)
|
||||
return attrs
|
||||
|
||||
@classmethod
|
||||
def _guardrail(cls, data: GuardrailSpanData) -> AttributeMap:
|
||||
return collect(cls._GUARDRAIL_ATTRS, data)
|
||||
|
||||
@classmethod
|
||||
def _service(cls, data: ServiceSpanData) -> AttributeMap:
|
||||
attrs = collect(cls._SERVICE_ATTRS, data)
|
||||
attrs.update(
|
||||
{
|
||||
f"{LiteLLM.METADATA_PREFIX}{key}": value
|
||||
for key, value in data.event_metadata.items()
|
||||
}
|
||||
)
|
||||
return attrs
|
||||
84
litellm/integrations/otel/mappers/langfuse.py
Normal file
84
litellm/integrations/otel/mappers/langfuse.py
Normal file
|
|
@ -0,0 +1,84 @@
|
|||
"""Langfuse OTLP attribute mapper.
|
||||
|
||||
Langfuse ingests OTLP spans and reads from its own vendor namespace
|
||||
(``langfuse.observation.*``, ``langfuse.trace.*``). Compose this mapper after
|
||||
``GenAIMapper`` to send canonical + Langfuse-flavored spans simultaneously.
|
||||
|
||||
Every attribute is declared as a ``key -> extractor`` table entry (one callable
|
||||
per mapping operation): ``_LLM_CALL_ATTRS`` for scalars and ``_BLOB_ATTRS`` for
|
||||
the JSON-serialized payloads. ``_llm_call`` just applies both tables.
|
||||
"""
|
||||
|
||||
import json
|
||||
from typing import Callable
|
||||
|
||||
from litellm.integrations.otel.mappers.base import AttributeMap, AttrValue, SpanData
|
||||
from litellm.integrations.otel.mappers.utils import (
|
||||
collect,
|
||||
json_if,
|
||||
output_messages,
|
||||
serialize_messages,
|
||||
)
|
||||
from litellm.integrations.otel.payloads import (
|
||||
LLMCallSpanData,
|
||||
LLMRequestParams,
|
||||
LLMUsage,
|
||||
)
|
||||
|
||||
|
||||
class LangfuseMapper:
|
||||
|
||||
_LLM_CALL_ATTRS: dict[str, Callable[[LLMCallSpanData], AttrValue | None]] = {
|
||||
"langfuse.observation.type": lambda d: "generation",
|
||||
"langfuse.observation.model.name": lambda d: d.request_model or None,
|
||||
"langfuse.observation.metadata.provider": lambda d: d.provider or None,
|
||||
"langfuse.observation.id": lambda d: d.identity.call_id or None,
|
||||
"langfuse.trace.metadata.team_id": lambda d: d.identity.team_id or None,
|
||||
"langfuse.trace.metadata.team_alias": lambda d: d.identity.team_alias or None,
|
||||
}
|
||||
|
||||
# Sub-tables folded into their respective JSON blobs.
|
||||
_MODEL_PARAMS: dict[str, Callable[[LLMRequestParams], AttrValue | None]] = {
|
||||
"temperature": lambda rp: rp.temperature,
|
||||
"top_p": lambda rp: rp.top_p,
|
||||
"max_tokens": lambda rp: rp.max_tokens,
|
||||
"frequency_penalty": lambda rp: rp.frequency_penalty,
|
||||
"presence_penalty": lambda rp: rp.presence_penalty,
|
||||
"seed": lambda rp: rp.seed,
|
||||
}
|
||||
_USAGE_FIELDS: dict[str, Callable[[LLMUsage], AttrValue | None]] = {
|
||||
"input": lambda u: u.input_tokens,
|
||||
"output": lambda u: u.output_tokens,
|
||||
"total": lambda u: u.total_tokens,
|
||||
}
|
||||
|
||||
# JSON-payload attributes: each builder returns the serialized blob or None.
|
||||
_BLOB_ATTRS: dict[str, Callable[[LLMCallSpanData], AttrValue | None]] = {
|
||||
"langfuse.observation.model.parameters": lambda d: json_if(
|
||||
collect(LangfuseMapper._MODEL_PARAMS, d.request_params)
|
||||
),
|
||||
"langfuse.observation.input": lambda d: serialize_messages(d.messages_in),
|
||||
"langfuse.observation.output": lambda d: serialize_messages(output_messages(d)),
|
||||
"langfuse.observation.usage_details": lambda d: json_if(
|
||||
collect(LangfuseMapper._USAGE_FIELDS, d.usage)
|
||||
),
|
||||
"langfuse.observation.cost_details": lambda d: (
|
||||
json.dumps({"total": d.response_cost})
|
||||
if d.response_cost is not None
|
||||
else None
|
||||
),
|
||||
}
|
||||
|
||||
def map(self, data: SpanData) -> AttributeMap:
|
||||
match data:
|
||||
case LLMCallSpanData():
|
||||
return self._llm_call(data)
|
||||
case _:
|
||||
return {}
|
||||
|
||||
@classmethod
|
||||
def _llm_call(cls, data: LLMCallSpanData) -> AttributeMap:
|
||||
return {
|
||||
**collect(cls._LLM_CALL_ATTRS, data),
|
||||
**collect(cls._BLOB_ATTRS, data),
|
||||
}
|
||||
64
litellm/integrations/otel/mappers/langtrace.py
Normal file
64
litellm/integrations/otel/mappers/langtrace.py
Normal file
|
|
@ -0,0 +1,64 @@
|
|||
"""Langtrace attribute mapper.
|
||||
|
||||
Produces Langtrace's attribute vocabulary so a span can be ingested by a
|
||||
Langtrace backend. Compose it alongside other mappers like any other
|
||||
vocabulary.
|
||||
|
||||
Scalar attributes are declared as a flat ``key -> extractor`` table (one lambda
|
||||
per mapping operation); the prompt/completion blobs are serialized as a tail.
|
||||
"""
|
||||
|
||||
from typing import Callable
|
||||
|
||||
from litellm.integrations.otel.mappers.base import AttributeMap, AttrValue, SpanData
|
||||
from litellm.integrations.otel.mappers.utils import (
|
||||
collect,
|
||||
json_or_none,
|
||||
output_messages,
|
||||
)
|
||||
from litellm.integrations.otel.payloads import LLMCallSpanData
|
||||
|
||||
|
||||
class LangtraceMapper:
|
||||
|
||||
_LLM_CALL_ATTRS: dict[str, Callable[[LLMCallSpanData], AttrValue | None]] = {
|
||||
"gen_ai.operation.name": lambda d: "chat",
|
||||
"langtrace.service.name": lambda d: d.provider or None,
|
||||
"llm.model": lambda d: d.request_model or None,
|
||||
"gen_ai.response.model": lambda d: d.response_model or None,
|
||||
"gen_ai.response_id": lambda d: d.response_id or None,
|
||||
"gen_ai.system_fingerprint": lambda d: d.system_fingerprint or None,
|
||||
"llm.temperature": lambda d: d.request_params.temperature,
|
||||
"llm.top_p": lambda d: d.request_params.top_p,
|
||||
"llm.top_k": lambda d: d.request_params.top_k,
|
||||
"llm.max_tokens": lambda d: d.request_params.max_tokens,
|
||||
"llm.frequency_penalty": lambda d: d.request_params.frequency_penalty,
|
||||
"llm.presence_penalty": lambda d: d.request_params.presence_penalty,
|
||||
"llm.stream": lambda d: d.is_streaming,
|
||||
"llm.token.counts.prompt": lambda d: d.usage.input_tokens,
|
||||
"llm.token.counts.completion": lambda d: d.usage.output_tokens,
|
||||
"llm.token.counts.total": lambda d: d.usage.total_tokens,
|
||||
}
|
||||
|
||||
_BLOB_ATTRS: dict[str, Callable[[LLMCallSpanData], AttrValue | None]] = {
|
||||
"llm.prompts": lambda d: (
|
||||
json_or_none(list(d.messages_in)) if d.messages_in else None
|
||||
),
|
||||
"llm.completions": lambda d: (
|
||||
json_or_none(output_messages(d)) if d.choices_out else None
|
||||
),
|
||||
}
|
||||
|
||||
def map(self, data: SpanData) -> AttributeMap:
|
||||
match data:
|
||||
case LLMCallSpanData():
|
||||
return self._llm_call(data)
|
||||
case _:
|
||||
return {}
|
||||
|
||||
@classmethod
|
||||
def _llm_call(cls, data: LLMCallSpanData) -> AttributeMap:
|
||||
return {
|
||||
**collect(cls._LLM_CALL_ATTRS, data),
|
||||
**collect(cls._BLOB_ATTRS, data),
|
||||
}
|
||||
97
litellm/integrations/otel/mappers/legacy.py
Normal file
97
litellm/integrations/otel/mappers/legacy.py
Normal file
|
|
@ -0,0 +1,97 @@
|
|||
"""Mapper for the older semantic-convention attribute vocabulary.
|
||||
|
||||
Emits attributes under the semconv-ai / Traceloop key names (e.g.
|
||||
``gen_ai.system``, ``gen_ai.usage.prompt_tokens``, ``llm.is_streaming``) plus a
|
||||
few bare, unprefixed service keys (``service``, ``call_type``, ``error``), for
|
||||
backends that consume those names.
|
||||
|
||||
Like ``GenAIMapper``, each span kind declares its schema as a flat
|
||||
``attribute key -> extractor`` table: one lambda per mapping operation.
|
||||
"""
|
||||
|
||||
from typing import Callable, Final
|
||||
|
||||
from litellm.integrations.otel.mappers.base import AttributeMap, AttrValue, SpanData
|
||||
from litellm.integrations.otel.mappers.utils import collect, drop_none
|
||||
from litellm.integrations.otel.payloads import (
|
||||
LLMCallSpanData,
|
||||
ServiceSpanData,
|
||||
ToolDefinition,
|
||||
)
|
||||
|
||||
# Attribute keys in the semconv-ai / Traceloop vocabulary.
|
||||
_LEGACY_SYSTEM: Final = "gen_ai.system"
|
||||
_LEGACY_PROMPT_TOKENS: Final = "gen_ai.usage.prompt_tokens"
|
||||
_LEGACY_COMPLETION_TOKENS: Final = "gen_ai.usage.completion_tokens"
|
||||
_LEGACY_TOTAL_TOKENS: Final = "gen_ai.usage.total_tokens"
|
||||
_LEGACY_IS_STREAMING: Final = "llm.is_streaming"
|
||||
_LEGACY_TOP_K: Final = "llm.top_k"
|
||||
_LEGACY_FREQUENCY_PENALTY: Final = "llm.frequency_penalty"
|
||||
_LEGACY_PRESENCE_PENALTY: Final = "llm.presence_penalty"
|
||||
_LEGACY_STOP_SEQUENCES: Final = "llm.chat.stop_sequences"
|
||||
_LEGACY_SERVICE: Final = "service"
|
||||
_LEGACY_CALL_TYPE: Final = "call_type"
|
||||
_LEGACY_ERROR: Final = "error"
|
||||
|
||||
|
||||
class LegacyMapper:
|
||||
"""Emits LLM-call and service attributes under the older key names."""
|
||||
|
||||
_LLM_CALL_ATTRS: dict[str, Callable[[LLMCallSpanData], AttrValue | None]] = {
|
||||
_LEGACY_SYSTEM: lambda d: d.provider or None,
|
||||
_LEGACY_PROMPT_TOKENS: lambda d: d.usage.input_tokens,
|
||||
_LEGACY_COMPLETION_TOKENS: lambda d: d.usage.output_tokens,
|
||||
_LEGACY_TOTAL_TOKENS: lambda d: d.usage.total_tokens,
|
||||
_LEGACY_IS_STREAMING: lambda d: d.is_streaming,
|
||||
_LEGACY_TOP_K: lambda d: d.request_params.top_k,
|
||||
_LEGACY_FREQUENCY_PENALTY: lambda d: d.request_params.frequency_penalty,
|
||||
_LEGACY_PRESENCE_PENALTY: lambda d: d.request_params.presence_penalty,
|
||||
_LEGACY_STOP_SEQUENCES: lambda d: (
|
||||
list(d.request_params.stop_sequences)
|
||||
if d.request_params.stop_sequences
|
||||
else None
|
||||
),
|
||||
}
|
||||
|
||||
_TOOL_ATTRS: dict[str, Callable[[ToolDefinition], AttrValue | None]] = {
|
||||
"name": lambda t: t.name,
|
||||
"description": lambda t: t.description or None,
|
||||
"parameters": lambda t: t.parameters_json or None,
|
||||
}
|
||||
|
||||
_SERVICE_ATTRS: dict[str, Callable[[ServiceSpanData], AttrValue | None]] = {
|
||||
_LEGACY_SERVICE: lambda d: d.service_name,
|
||||
_LEGACY_CALL_TYPE: lambda d: d.call_type,
|
||||
_LEGACY_ERROR: lambda d: (
|
||||
d.error.message if d.error is not None and d.error.message else None
|
||||
),
|
||||
}
|
||||
|
||||
def map(self, data: SpanData) -> AttributeMap:
|
||||
match data:
|
||||
case LLMCallSpanData():
|
||||
return self._llm_call(data)
|
||||
case ServiceSpanData():
|
||||
return self._service(data)
|
||||
case _:
|
||||
return {}
|
||||
|
||||
@classmethod
|
||||
def _llm_call(cls, data: LLMCallSpanData) -> AttributeMap:
|
||||
attrs = collect(cls._LLM_CALL_ATTRS, data)
|
||||
attrs.update(
|
||||
drop_none(
|
||||
{
|
||||
f"llm.request.functions.{idx}.{suffix}": extract(tool)
|
||||
for idx, tool in enumerate(data.tools)
|
||||
for suffix, extract in cls._TOOL_ATTRS.items()
|
||||
}
|
||||
)
|
||||
)
|
||||
return attrs
|
||||
|
||||
@classmethod
|
||||
def _service(cls, data: ServiceSpanData) -> AttributeMap:
|
||||
attrs = collect(cls._SERVICE_ATTRS, data)
|
||||
attrs.update(dict(data.event_metadata))
|
||||
return attrs
|
||||
128
litellm/integrations/otel/mappers/openinference.py
Normal file
128
litellm/integrations/otel/mappers/openinference.py
Normal file
|
|
@ -0,0 +1,128 @@
|
|||
"""OpenInference attribute mapper (Arize + Arize-Phoenix shared vocabulary).
|
||||
|
||||
Spec: https://github.com/Arize-ai/openinference/tree/main/spec — the standard
|
||||
both Arize and Phoenix consume. Composing this mapper after ``GenAIMapper``
|
||||
gives the same span both vocabularies, so a single trace lights up Arize +
|
||||
Phoenix + any other OpenInference-aware backend simultaneously.
|
||||
"""
|
||||
|
||||
import json
|
||||
from typing import Callable, Sequence
|
||||
|
||||
from litellm.integrations.otel.mappers.base import AttributeMap, AttrValue, SpanData
|
||||
from litellm.integrations.otel.mappers.utils import (
|
||||
collect,
|
||||
drop_none,
|
||||
json_if,
|
||||
message_content,
|
||||
output_messages,
|
||||
)
|
||||
from litellm.integrations.otel.payloads import (
|
||||
LLMCallSpanData,
|
||||
LLMRequestParams,
|
||||
ToolDefinition,
|
||||
)
|
||||
|
||||
|
||||
class OpenInferenceMapper:
|
||||
"""Emits OpenInference attributes for LLM_CALL spans.
|
||||
|
||||
Key families (per the OpenInference spec):
|
||||
- ``openinference.span.kind`` — discriminator (``"LLM"`` here)
|
||||
- ``llm.model_name`` / ``llm.provider`` / ``llm.invocation_parameters``
|
||||
- ``llm.input_messages.{i}.message.role`` / ``...content``
|
||||
- ``llm.output_messages.{i}.message.role`` / ``...content``
|
||||
- ``llm.token_count.prompt`` / ``...completion`` / ``...total``
|
||||
- ``input.value`` / ``output.value`` — JSON-serialized request / response
|
||||
"""
|
||||
|
||||
_LLM_CALL_ATTRS: dict[str, Callable[[LLMCallSpanData], AttrValue | None]] = {
|
||||
"openinference.span.kind": lambda d: "LLM",
|
||||
"llm.model_name": lambda d: d.request_model or None,
|
||||
"llm.provider": lambda d: d.provider or None,
|
||||
"llm.token_count.prompt": lambda d: d.usage.input_tokens,
|
||||
"llm.token_count.completion": lambda d: d.usage.output_tokens,
|
||||
"llm.token_count.total": lambda d: d.usage.total_tokens,
|
||||
}
|
||||
|
||||
# Folded into the ``llm.invocation_parameters`` JSON blob.
|
||||
_INVOCATION_PARAMS: dict[str, Callable[[LLMRequestParams], AttrValue | None]] = {
|
||||
"temperature": lambda rp: rp.temperature,
|
||||
"top_p": lambda rp: rp.top_p,
|
||||
"top_k": lambda rp: rp.top_k,
|
||||
"max_tokens": lambda rp: rp.max_tokens,
|
||||
"frequency_penalty": lambda rp: rp.frequency_penalty,
|
||||
"presence_penalty": lambda rp: rp.presence_penalty,
|
||||
"seed": lambda rp: rp.seed,
|
||||
}
|
||||
|
||||
# Per-tool extractors, keyed by the ``llm.tools.{idx}.*`` suffix.
|
||||
_TOOL_ATTRS: dict[str, Callable[[ToolDefinition], AttrValue | None]] = {
|
||||
"tool.name": lambda t: t.name,
|
||||
"tool.description": lambda t: t.description or None,
|
||||
"tool.json_schema": lambda t: t.parameters_json or None,
|
||||
}
|
||||
|
||||
# JSON-payload attributes: each builder returns the serialized blob or None.
|
||||
_BLOB_ATTRS: dict[str, Callable[[LLMCallSpanData], AttrValue | None]] = {
|
||||
"llm.invocation_parameters": lambda d: json_if(
|
||||
collect(OpenInferenceMapper._INVOCATION_PARAMS, d.request_params)
|
||||
),
|
||||
}
|
||||
|
||||
def map(self, data: SpanData) -> AttributeMap:
|
||||
match data:
|
||||
case LLMCallSpanData():
|
||||
return self._llm_call(data)
|
||||
case _:
|
||||
return {}
|
||||
|
||||
@classmethod
|
||||
def _llm_call(cls, data: LLMCallSpanData) -> AttributeMap:
|
||||
return {
|
||||
**collect(cls._LLM_CALL_ATTRS, data),
|
||||
**collect(cls._BLOB_ATTRS, data),
|
||||
**cls._messages("llm.input_messages", "input.value", data.messages_in),
|
||||
**cls._messages(
|
||||
"llm.output_messages", "output.value", output_messages(data)
|
||||
),
|
||||
**cls._tools(data),
|
||||
}
|
||||
|
||||
@staticmethod
|
||||
def _messages(
|
||||
prefix: str, value_key: str, messages: Sequence[object]
|
||||
) -> AttributeMap:
|
||||
"""Per-message ``{prefix}.{idx}.message.*`` keys + the ``value_key`` blob."""
|
||||
parsed = [
|
||||
(m.get("role") if isinstance(m, dict) else None, message_content(m))
|
||||
for m in messages
|
||||
]
|
||||
attrs = drop_none(
|
||||
{
|
||||
key: value
|
||||
for idx, (role, content) in enumerate(parsed)
|
||||
for key, value in (
|
||||
(
|
||||
f"{prefix}.{idx}.message.role",
|
||||
role if isinstance(role, str) else None,
|
||||
),
|
||||
(f"{prefix}.{idx}.message.content", content),
|
||||
)
|
||||
}
|
||||
)
|
||||
if parsed:
|
||||
attrs[value_key] = json.dumps(
|
||||
[{"role": role, "content": content} for role, content in parsed]
|
||||
)
|
||||
return attrs
|
||||
|
||||
@classmethod
|
||||
def _tools(cls, data: LLMCallSpanData) -> AttributeMap:
|
||||
return drop_none(
|
||||
{
|
||||
f"llm.tools.{idx}.{suffix}": extract(tool)
|
||||
for idx, tool in enumerate(data.tools)
|
||||
for suffix, extract in cls._TOOL_ATTRS.items()
|
||||
}
|
||||
)
|
||||
76
litellm/integrations/otel/mappers/utils.py
Normal file
76
litellm/integrations/otel/mappers/utils.py
Normal file
|
|
@ -0,0 +1,76 @@
|
|||
"""Shared helpers for the attribute mappers.
|
||||
|
||||
Small, mapper-agnostic utilities — JSON serialization, message extraction, and
|
||||
extractor-table application — pulled out of the individual mapper modules so
|
||||
they live in one place.
|
||||
"""
|
||||
|
||||
import json
|
||||
from typing import Callable, Mapping, Sequence
|
||||
|
||||
from litellm.integrations.otel.mappers.base import AttributeMap, AttrValue
|
||||
from litellm.integrations.otel.payloads import LLMCallSpanData
|
||||
|
||||
|
||||
def drop_none(values: Mapping[str, AttrValue | None]) -> AttributeMap:
|
||||
"""Return ``values`` with ``None``-valued entries removed."""
|
||||
return {k: v for k, v in values.items() if v is not None}
|
||||
|
||||
|
||||
def collect(table: Mapping[str, Callable], source: object) -> AttributeMap:
|
||||
"""Apply an extractor table to ``source``, dropping ``None`` results."""
|
||||
return drop_none({key: extract(source) for key, extract in table.items()})
|
||||
|
||||
|
||||
def json_if(payload: Mapping[str, object]) -> str | None:
|
||||
"""JSON-serialize ``payload`` only when it's non-empty; else ``None``."""
|
||||
return json.dumps(payload) if payload else None
|
||||
|
||||
|
||||
def json_or_none(value: object) -> str | None:
|
||||
"""JSON-serialize ``value`` (falling back to ``str``); ``None`` on failure."""
|
||||
try:
|
||||
return json.dumps(value, default=str)
|
||||
except Exception:
|
||||
return None
|
||||
|
||||
|
||||
def stringify_message(message: object) -> str | None:
|
||||
"""JSON-serialize a chat message dict; ``None`` if not a dict or on failure."""
|
||||
if not isinstance(message, dict):
|
||||
return None
|
||||
try:
|
||||
return json.dumps(message, default=str)
|
||||
except Exception:
|
||||
return None
|
||||
|
||||
|
||||
def serialize_messages(messages: Sequence[object]) -> str | None:
|
||||
"""Round-trip a sequence of message dicts through ``stringify_message``."""
|
||||
serialized = [
|
||||
json.loads(s) for s in (stringify_message(m) for m in messages) if s is not None
|
||||
]
|
||||
return json.dumps(serialized) if serialized else None
|
||||
|
||||
|
||||
def message_content(message: object) -> str | None:
|
||||
"""Extract the textual ``content`` from a chat message dict."""
|
||||
if not isinstance(message, dict):
|
||||
return None
|
||||
content = message.get("content")
|
||||
if isinstance(content, str):
|
||||
return content
|
||||
if isinstance(content, list):
|
||||
# multimodal: concatenate text parts only
|
||||
parts = [
|
||||
part.get("text", "")
|
||||
for part in content
|
||||
if isinstance(part, dict) and part.get("type") == "text"
|
||||
]
|
||||
return "".join(p for p in parts if isinstance(p, str)) or None
|
||||
return None
|
||||
|
||||
|
||||
def output_messages(data: LLMCallSpanData) -> list:
|
||||
"""The ``message`` payload of each response choice."""
|
||||
return [c.get("message") for c in data.choices_out if isinstance(c, dict)]
|
||||
48
litellm/integrations/otel/mappers/weave.py
Normal file
48
litellm/integrations/otel/mappers/weave.py
Normal file
|
|
@ -0,0 +1,48 @@
|
|||
"""Weave (W&B) attribute mapper.
|
||||
|
||||
Weave consumes OpenInference + a small set of Weave-specific keys (display
|
||||
name, thread id, output value). This mapper layers the latter on top of
|
||||
OpenInference's vocabulary — compose ``["genai", "openinference", "weave"]``
|
||||
to feed a Weave backend.
|
||||
"""
|
||||
|
||||
from typing import Callable
|
||||
|
||||
from litellm.integrations.otel.mappers.base import AttributeMap, AttrValue, SpanData
|
||||
from litellm.integrations.otel.mappers.utils import collect, json_or_none
|
||||
from litellm.integrations.otel.payloads import LLMCallSpanData
|
||||
|
||||
|
||||
class WeaveMapper:
|
||||
"""Maps ``LLMCallSpanData`` to Weave's vendor attributes."""
|
||||
|
||||
_LLM_CALL_ATTRS: dict[str, Callable[[LLMCallSpanData], AttrValue | None]] = {
|
||||
# ``display_name`` has the form ``"{operation} {model}"``. The span
|
||||
# name already covers that, but Weave reads this attribute too.
|
||||
"weave.display_name": lambda d: (
|
||||
f"{d.operation.value} {d.request_model}" if d.request_model else None
|
||||
),
|
||||
"weave.call_id": lambda d: d.identity.call_id or None,
|
||||
}
|
||||
|
||||
# JSON-payload attributes: each builder returns the serialized blob or None.
|
||||
_BLOB_ATTRS: dict[str, Callable[[LLMCallSpanData], AttrValue | None]] = {
|
||||
# Weave treats the response choices as the "output" payload.
|
||||
"weave.output": lambda d: (
|
||||
json_or_none(list(d.choices_out)) if d.choices_out else None
|
||||
),
|
||||
}
|
||||
|
||||
def map(self, data: SpanData) -> AttributeMap:
|
||||
match data:
|
||||
case LLMCallSpanData():
|
||||
return self._llm_call(data)
|
||||
case _:
|
||||
return {}
|
||||
|
||||
@classmethod
|
||||
def _llm_call(cls, data: LLMCallSpanData) -> AttributeMap:
|
||||
return {
|
||||
**collect(cls._LLM_CALL_ATTRS, data),
|
||||
**collect(cls._BLOB_ATTRS, data),
|
||||
}
|
||||
28
litellm/integrations/otel/metrics.py
Normal file
28
litellm/integrations/otel/metrics.py
Normal file
|
|
@ -0,0 +1,28 @@
|
|||
"""GenAI client metrics (token usage + operation duration histograms)."""
|
||||
|
||||
from dataclasses import dataclass
|
||||
|
||||
from opentelemetry.metrics import Histogram, Meter
|
||||
|
||||
from litellm.integrations.otel.semconv import Metric
|
||||
|
||||
|
||||
@dataclass(frozen=True)
|
||||
class GenAIMetrics:
|
||||
token_usage: Histogram
|
||||
operation_duration: Histogram
|
||||
|
||||
|
||||
def create_genai_metrics(meter: Meter) -> GenAIMetrics:
|
||||
return GenAIMetrics(
|
||||
token_usage=meter.create_histogram(
|
||||
name=Metric.TOKEN_USAGE,
|
||||
unit="{token}",
|
||||
description="Number of tokens used per GenAI request.",
|
||||
),
|
||||
operation_duration=meter.create_histogram(
|
||||
name=Metric.OPERATION_DURATION,
|
||||
unit="s",
|
||||
description="GenAI operation duration.",
|
||||
),
|
||||
)
|
||||
409
litellm/integrations/otel/payloads.py
Normal file
409
litellm/integrations/otel/payloads.py
Normal file
|
|
@ -0,0 +1,409 @@
|
|||
"""Typed span-data inputs: frozen dataclasses the engine and mappers consume."""
|
||||
|
||||
from __future__ import annotations
|
||||
|
||||
import json
|
||||
from dataclasses import dataclass, field
|
||||
from typing import TYPE_CHECKING, ClassVar, Mapping, cast
|
||||
from urllib.parse import urlsplit
|
||||
|
||||
from litellm.integrations.otel.semconv import (
|
||||
GenAIOperation,
|
||||
resolve_operation,
|
||||
resolve_provider,
|
||||
)
|
||||
from litellm.integrations.otel.utils import (
|
||||
as_bool,
|
||||
as_float,
|
||||
as_int,
|
||||
as_str,
|
||||
as_str_tuple,
|
||||
)
|
||||
|
||||
if TYPE_CHECKING:
|
||||
from litellm.types.services import ServiceLoggerPayload
|
||||
from litellm.types.utils import StandardLoggingPayload
|
||||
|
||||
|
||||
# --- typed sub-structures ---------------------------------------------------- #
|
||||
|
||||
|
||||
@dataclass(frozen=True)
|
||||
class LLMRequestParams:
|
||||
temperature: float | None = None
|
||||
top_p: float | None = None
|
||||
top_k: int | None = None
|
||||
max_tokens: int | None = None
|
||||
frequency_penalty: float | None = None
|
||||
presence_penalty: float | None = None
|
||||
stop_sequences: tuple[str, ...] | None = None
|
||||
seed: int | None = None
|
||||
|
||||
@classmethod
|
||||
def from_model_parameters(cls, params: Mapping[str, object]) -> "LLMRequestParams":
|
||||
max_tokens = as_int(params.get("max_tokens"))
|
||||
if max_tokens is None:
|
||||
max_tokens = as_int(params.get("max_completion_tokens"))
|
||||
return cls(
|
||||
temperature=as_float(params.get("temperature")),
|
||||
top_p=as_float(params.get("top_p")),
|
||||
top_k=as_int(params.get("top_k")),
|
||||
max_tokens=max_tokens,
|
||||
frequency_penalty=as_float(params.get("frequency_penalty")),
|
||||
presence_penalty=as_float(params.get("presence_penalty")),
|
||||
stop_sequences=as_str_tuple(params.get("stop")),
|
||||
seed=as_int(params.get("seed")),
|
||||
)
|
||||
|
||||
|
||||
@dataclass(frozen=True)
|
||||
class LLMUsage:
|
||||
input_tokens: int | None = None
|
||||
output_tokens: int | None = None
|
||||
total_tokens: int | None = None
|
||||
|
||||
|
||||
@dataclass(frozen=True)
|
||||
class SpanError:
|
||||
error_type: str | None = None
|
||||
message: str | None = None
|
||||
|
||||
|
||||
@dataclass(frozen=True)
|
||||
class ServerInfo:
|
||||
address: str | None = None
|
||||
port: int | None = None
|
||||
|
||||
@classmethod
|
||||
def from_api_base(cls, api_base: str | None) -> ServerInfo | None:
|
||||
if not api_base:
|
||||
return None
|
||||
parsed = urlsplit(api_base if "://" in api_base else f"//{api_base}")
|
||||
if not parsed.hostname:
|
||||
return None
|
||||
return cls(address=parsed.hostname, port=parsed.port)
|
||||
|
||||
|
||||
@dataclass(frozen=True)
|
||||
class RequestIdentity:
|
||||
call_id: str | None = None
|
||||
team_id: str | None = None
|
||||
team_alias: str | None = None
|
||||
key_hash: str | None = None
|
||||
end_user: str | None = None
|
||||
metadata: Mapping[str, str] = field(default_factory=dict)
|
||||
|
||||
@classmethod
|
||||
def from_payload(cls, payload: "StandardLoggingPayload") -> "RequestIdentity":
|
||||
raw_meta = cast(Mapping[str, object], payload.get("metadata") or {})
|
||||
metadata = {
|
||||
key: str(value)
|
||||
for key, value in raw_meta.items()
|
||||
if isinstance(value, (str, bool, int, float))
|
||||
}
|
||||
return cls(
|
||||
call_id=as_str(payload.get("litellm_call_id")) or as_str(payload.get("id")),
|
||||
# StandardLoggingMetadata's canonical key is ``user_api_key_team_id``;
|
||||
# the bare ``team_id`` is a legacy alias and is often empty, so prefer
|
||||
# the canonical key and fall back to the alias.
|
||||
team_id=as_str(raw_meta.get("user_api_key_team_id"))
|
||||
or as_str(raw_meta.get("team_id")),
|
||||
team_alias=as_str(raw_meta.get("user_api_key_team_alias"))
|
||||
or as_str(raw_meta.get("team_alias")),
|
||||
key_hash=as_str(raw_meta.get("user_api_key_hash")),
|
||||
end_user=as_str(payload.get("end_user"))
|
||||
or as_str(raw_meta.get("user_api_key_end_user_id")),
|
||||
metadata=metadata,
|
||||
)
|
||||
|
||||
@classmethod
|
||||
def from_user_api_key_auth(cls, auth: object) -> "RequestIdentity":
|
||||
"""Identity from a ``UserAPIKeyAuth`` (duck-typed to keep this module
|
||||
free of a proxy import).
|
||||
|
||||
Used in the pre-call hook to seed Baggage early — before any LLM,
|
||||
guardrail, or service span is created — so the whole request's spans
|
||||
inherit identity, not just the LLM-call span. Metadata sub-keys use the
|
||||
``user_api_key_*`` names that ``baggage.DEFAULT_BAGGAGE_METADATA_KEYS``
|
||||
promotes.
|
||||
"""
|
||||
get = lambda name: getattr(auth, name, None) # noqa: E731
|
||||
metadata = {
|
||||
meta_key: str(value)
|
||||
for meta_key, attr in (
|
||||
("user_api_key_user_id", "user_id"),
|
||||
("user_api_key_org_id", "org_id"),
|
||||
("user_api_key_alias", "key_alias"),
|
||||
("user_api_key_end_user_id", "end_user_id"),
|
||||
)
|
||||
if (value := get(attr))
|
||||
}
|
||||
return cls(
|
||||
team_id=as_str(get("team_id")),
|
||||
team_alias=as_str(get("team_alias")),
|
||||
key_hash=as_str(get("api_key")),
|
||||
end_user=as_str(get("end_user_id")),
|
||||
metadata=metadata,
|
||||
)
|
||||
|
||||
|
||||
@dataclass(frozen=True)
|
||||
class GuardrailSpanData:
|
||||
guardrail_name: str
|
||||
mode: str | None = None
|
||||
status: str | None = None
|
||||
masked_entity_count: int | None = None
|
||||
provider: str | None = None
|
||||
action: str | None = None
|
||||
# The guardrail verdict / provider response (e.g. the moderation result),
|
||||
# JSON-serialized. This is the detail that belongs on the guardrail span.
|
||||
response_json: str | None = None
|
||||
violation_categories: tuple[str, ...] = ()
|
||||
confidence_score: float | None = None
|
||||
risk_score: float | None = None
|
||||
duration: float | None = None
|
||||
# Set when the guardrail intervened/blocked or failed, so the emitter marks
|
||||
# the span ERROR — a blocking guardrail is an error outcome for that span.
|
||||
error: SpanError | None = None
|
||||
|
||||
# Guardrail statuses that mean the guardrail did not pass the request through.
|
||||
_ERROR_STATUSES: ClassVar[frozenset[str]] = frozenset(
|
||||
{"guardrail_intervened", "guardrail_failed_to_respond"}
|
||||
)
|
||||
|
||||
@classmethod
|
||||
def from_logging_entry(cls, entry: Mapping[str, object]) -> "GuardrailSpanData":
|
||||
"""Build from one ``standard_logging_guardrail_information`` entry."""
|
||||
name = (
|
||||
as_str(entry.get("guardrail_name"))
|
||||
or as_str(entry.get("name"))
|
||||
or "guardrail"
|
||||
)
|
||||
status = as_str(entry.get("guardrail_status")) or as_str(entry.get("status"))
|
||||
response = entry.get("guardrail_response")
|
||||
error = (
|
||||
SpanError(error_type=status, message=as_str(entry.get("guardrail_action")))
|
||||
if status in cls._ERROR_STATUSES
|
||||
else None
|
||||
)
|
||||
return cls(
|
||||
guardrail_name=name,
|
||||
mode=as_str(entry.get("guardrail_mode")) or as_str(entry.get("mode")),
|
||||
status=status,
|
||||
masked_entity_count=_total_masked_entities(
|
||||
entry.get("masked_entity_count")
|
||||
),
|
||||
provider=as_str(entry.get("guardrail_provider")),
|
||||
action=as_str(entry.get("guardrail_action")),
|
||||
response_json=_json_or_none(response) if response is not None else None,
|
||||
violation_categories=as_str_tuple(entry.get("violation_categories")) or (),
|
||||
confidence_score=as_float(entry.get("confidence_score")),
|
||||
risk_score=as_float(entry.get("risk_score")),
|
||||
duration=as_float(entry.get("duration")),
|
||||
error=error,
|
||||
)
|
||||
|
||||
|
||||
@dataclass(frozen=True)
|
||||
class ServiceSpanData:
|
||||
service_name: str
|
||||
call_type: str | None = None
|
||||
error: SpanError | None = None
|
||||
# Caller-supplied attributes to stamp on the service span, passed through
|
||||
# from ``async_service_*_hook(event_metadata=...)``. The mapper owns how
|
||||
# these are namespaced: the canonical vocabulary uses ``litellm.metadata.*``
|
||||
# keys, the semconv-ai / Traceloop vocabulary uses the bare key names.
|
||||
event_metadata: Mapping[str, str] = field(default_factory=dict)
|
||||
|
||||
@classmethod
|
||||
def from_payload(
|
||||
cls,
|
||||
payload: "ServiceLoggerPayload",
|
||||
event_metadata: Mapping[str, object] | None = None,
|
||||
) -> "ServiceSpanData":
|
||||
# ``payload.service`` is a ``ServiceTypes(str, Enum)`` and ``error`` is
|
||||
# ``Optional[str]`` on the Pydantic model — no defensive reads needed.
|
||||
# ``str(value)`` covers every case (``str(None) == "None"``).
|
||||
coerced = {key: str(value) for key, value in (event_metadata or {}).items()}
|
||||
return cls(
|
||||
service_name=payload.service.value,
|
||||
call_type=payload.call_type,
|
||||
error=SpanError(message=payload.error) if payload.error else None,
|
||||
event_metadata=coerced,
|
||||
)
|
||||
|
||||
|
||||
@dataclass(frozen=True)
|
||||
class ProxyRequestSpanData:
|
||||
http_method: str
|
||||
route: str
|
||||
url_path: str | None = None
|
||||
status_code: int | None = None
|
||||
identity: RequestIdentity | None = None
|
||||
|
||||
|
||||
# --- the primary LLM-call model ---------------------------------------------- #
|
||||
|
||||
|
||||
@dataclass(frozen=True)
|
||||
class ToolDefinition:
|
||||
"""A single function/tool declared on a chat-completion request."""
|
||||
|
||||
name: str
|
||||
description: str | None = None
|
||||
parameters_json: str | None = (
|
||||
None # JSON-serialized schema (str so it's an AttrValue)
|
||||
)
|
||||
|
||||
|
||||
@dataclass(frozen=True)
|
||||
class LLMCallSpanData:
|
||||
operation: GenAIOperation
|
||||
provider: str
|
||||
request_model: str
|
||||
response_model: str | None
|
||||
response_id: str | None
|
||||
request_params: LLMRequestParams
|
||||
usage: LLMUsage
|
||||
finish_reasons: tuple[str, ...]
|
||||
error: SpanError | None
|
||||
response_cost: float | None
|
||||
server: ServerInfo | None
|
||||
identity: RequestIdentity
|
||||
is_streaming: bool | None = None
|
||||
tools: tuple[ToolDefinition, ...] = ()
|
||||
# Raw messages and response, needed by vendor mappers (OpenInference,
|
||||
# Langfuse, Weave) that stamp message-level attributes. ``messages_in`` is
|
||||
# the request payload; ``choices_out`` mirrors ``response.choices`` from
|
||||
# the StandardLoggingPayload. Both are tuples of immutable mappings so the
|
||||
# dataclass stays hashable and frozen.
|
||||
messages_in: tuple[Mapping[str, object], ...] = ()
|
||||
choices_out: tuple[Mapping[str, object], ...] = ()
|
||||
system_fingerprint: str | None = None
|
||||
|
||||
@classmethod
|
||||
def from_standard_logging_payload(
|
||||
cls, payload: "StandardLoggingPayload", capture_content: bool = False
|
||||
) -> "LLMCallSpanData":
|
||||
params = cast(Mapping[str, object], payload.get("model_parameters") or {})
|
||||
hidden = cast(Mapping[str, object], payload.get("hidden_params") or {})
|
||||
# Normalize ``response`` to a dict once so every field read below is a
|
||||
# plain ``.get`` — no repeated ``isinstance`` guards.
|
||||
raw_response = payload.get("response")
|
||||
response = cast(
|
||||
Mapping[str, object], raw_response if isinstance(raw_response, dict) else {}
|
||||
)
|
||||
choices_out = _dicts(response.get("choices"))
|
||||
# ``finish_reasons`` is metadata, not content, so derive it from
|
||||
# ``choices_out`` before gating. The raw message/choice bodies are only
|
||||
# retained when content capture is enabled (see ``capture_span_content``);
|
||||
# otherwise the content-bearing mappers receive empty sequences and emit
|
||||
# no prompt/response text.
|
||||
finish_reasons = _finish_reasons(choices_out)
|
||||
return cls(
|
||||
operation=resolve_operation(as_str(payload.get("call_type"))),
|
||||
provider=resolve_provider(as_str(payload.get("custom_llm_provider"))),
|
||||
request_model=as_str(payload.get("model")) or "",
|
||||
response_model=as_str(response.get("model")),
|
||||
response_id=as_str(response.get("id")),
|
||||
request_params=LLMRequestParams.from_model_parameters(params),
|
||||
usage=LLMUsage(
|
||||
input_tokens=as_int(payload.get("prompt_tokens")),
|
||||
output_tokens=as_int(payload.get("completion_tokens")),
|
||||
total_tokens=as_int(payload.get("total_tokens")),
|
||||
),
|
||||
finish_reasons=finish_reasons,
|
||||
error=_parse_error(payload),
|
||||
response_cost=as_float(payload.get("response_cost")),
|
||||
server=ServerInfo.from_api_base(
|
||||
as_str(payload.get("api_base")) or as_str(hidden.get("api_base"))
|
||||
),
|
||||
identity=RequestIdentity.from_payload(payload),
|
||||
is_streaming=as_bool(payload.get("stream")),
|
||||
tools=_extract_tools(params),
|
||||
messages_in=_dicts(payload.get("messages")) if capture_content else (),
|
||||
choices_out=choices_out if capture_content else (),
|
||||
system_fingerprint=as_str(response.get("system_fingerprint")),
|
||||
)
|
||||
|
||||
|
||||
def _json_or_none(value: object) -> str | None:
|
||||
"""JSON-serialize ``value`` (already-string values pass through). ``None`` on failure."""
|
||||
if isinstance(value, str):
|
||||
return value
|
||||
try:
|
||||
return json.dumps(value, default=str)
|
||||
except Exception:
|
||||
return None
|
||||
|
||||
|
||||
def _total_masked_entities(value: object) -> int | None:
|
||||
"""``masked_entity_count`` is a ``{entity_type: count}`` map — sum to a total."""
|
||||
if isinstance(value, Mapping):
|
||||
total = sum(v for v in value.values() if isinstance(v, int))
|
||||
return total or None
|
||||
return as_int(value)
|
||||
|
||||
|
||||
def _dicts(value: object) -> tuple[Mapping[str, object], ...]:
|
||||
"""The dict items of ``value`` (when it's a list), as a tuple. Else empty."""
|
||||
if not isinstance(value, list):
|
||||
return ()
|
||||
return tuple(item for item in value if isinstance(item, dict))
|
||||
|
||||
|
||||
def _finish_reasons(choices: tuple[Mapping[str, object], ...]) -> tuple[str, ...]:
|
||||
"""Non-empty ``finish_reason`` of each response choice."""
|
||||
return tuple(r for c in choices if (r := as_str(c.get("finish_reason"))))
|
||||
|
||||
|
||||
def _parse_error(payload: "StandardLoggingPayload") -> SpanError | None:
|
||||
"""A ``SpanError`` for a failed request, or ``None`` on success."""
|
||||
if payload.get("status") != "failure":
|
||||
return None
|
||||
info = cast(Mapping[str, object], payload.get("error_information") or {})
|
||||
return SpanError(
|
||||
error_type=as_str(info.get("error_class")) or as_str(info.get("error_code")),
|
||||
message=as_str(info.get("error_message")) or as_str(payload.get("error_str")),
|
||||
)
|
||||
|
||||
|
||||
def _tool_from_entry(entry: object) -> ToolDefinition | None:
|
||||
"""One ``tools``/``functions`` entry → ``ToolDefinition``, or ``None`` if unusable."""
|
||||
if not isinstance(entry, dict):
|
||||
return None
|
||||
fn = entry.get("function") if "function" in entry else entry
|
||||
if not isinstance(fn, dict):
|
||||
return None
|
||||
name = as_str(fn.get("name"))
|
||||
if not name:
|
||||
return None
|
||||
params = fn.get("parameters")
|
||||
parameters_json: str | None = None
|
||||
if params is not None:
|
||||
try:
|
||||
parameters_json = json.dumps(params, default=str)
|
||||
except Exception:
|
||||
parameters_json = None
|
||||
return ToolDefinition(
|
||||
name=name,
|
||||
description=as_str(fn.get("description")),
|
||||
parameters_json=parameters_json,
|
||||
)
|
||||
|
||||
|
||||
def _extract_tools(
|
||||
model_parameters: Mapping[str, object],
|
||||
) -> tuple[ToolDefinition, ...]:
|
||||
"""Pull declared tools from request params (OpenAI / Anthropic shape).
|
||||
|
||||
Accepts the chat-completion ``tools=[{"type":"function", "function":
|
||||
{...}}, ...]`` shape, and falls back to the ``functions=[...]`` shape.
|
||||
Returns an empty tuple when neither is present.
|
||||
"""
|
||||
raw_tools = model_parameters.get("tools")
|
||||
if not isinstance(raw_tools, list):
|
||||
raw_tools = model_parameters.get("functions") # ``functions`` shape
|
||||
if not isinstance(raw_tools, list):
|
||||
return ()
|
||||
return tuple(t for entry in raw_tools if (t := _tool_from_entry(entry)) is not None)
|
||||
78
litellm/integrations/otel/presets/__init__.py
Normal file
78
litellm/integrations/otel/presets/__init__.py
Normal file
|
|
@ -0,0 +1,78 @@
|
|||
"""Integration presets — each one returns an :class:`OpenTelemetryV2Config`.
|
||||
|
||||
A preset is a callable that reads an integration's env vars and returns an
|
||||
``OpenTelemetryV2Config`` describing the exporter destination, the mapper
|
||||
vocabularies to apply, and any resource attributes. ``PRESET_BY_CALLBACK``
|
||||
maps a callback name (``"arize"``, ``"langfuse_otel"``, ...) to its preset so
|
||||
the factory in ``litellm_logging`` can resolve a name and build a single
|
||||
``OpenTelemetryV2`` instance from the result.
|
||||
"""
|
||||
|
||||
from typing import Callable
|
||||
|
||||
from litellm.integrations.otel.presets.agentops import agentops_preset
|
||||
from litellm.integrations.otel.presets.arize import arize_dynamic_headers, arize_preset
|
||||
from litellm.integrations.otel.presets.base import Preset
|
||||
from litellm.integrations.otel.presets.langfuse import (
|
||||
langfuse_dynamic_headers,
|
||||
langfuse_preset,
|
||||
)
|
||||
from litellm.integrations.otel.presets.langtrace import langtrace_preset
|
||||
from litellm.integrations.otel.presets.levo import levo_preset
|
||||
from litellm.integrations.otel.presets.phoenix import phoenix_preset
|
||||
from litellm.integrations.otel.presets.weave import weave_dynamic_headers, weave_preset
|
||||
from litellm.types.utils import StandardCallbackDynamicParams
|
||||
|
||||
#: Callback name → preset. The ``Preset`` annotation makes mypy verify every
|
||||
#: registered value matches the preset interface.
|
||||
PRESET_BY_CALLBACK: dict[str, Preset] = {
|
||||
"agentops": agentops_preset,
|
||||
"arize": arize_preset,
|
||||
"arize_phoenix": phoenix_preset,
|
||||
"langfuse_otel": langfuse_preset,
|
||||
"langtrace": langtrace_preset,
|
||||
"levo": levo_preset,
|
||||
"weave_otel": weave_preset,
|
||||
}
|
||||
|
||||
#: Callback name → per-request OTLP header builder (team/key multi-tenant
|
||||
#: routing). Only integrations that support dynamic credentials appear here —
|
||||
#: Arize-Phoenix/Langtrace/Levo/AgentOps don't, so they use the logger's
|
||||
#: default tracer.
|
||||
DYNAMIC_HEADERS_BY_CALLBACK: dict[
|
||||
str, Callable[[StandardCallbackDynamicParams], dict[str, str]]
|
||||
] = {
|
||||
"arize": arize_dynamic_headers,
|
||||
"langfuse_otel": langfuse_dynamic_headers,
|
||||
"weave_otel": weave_dynamic_headers,
|
||||
}
|
||||
|
||||
|
||||
def dynamic_otlp_headers(
|
||||
callback_name: str | None,
|
||||
dynamic_params: StandardCallbackDynamicParams | None,
|
||||
) -> dict[str, str] | None:
|
||||
"""Per-request OTLP headers for ``callback_name``, or ``None`` if N/A.
|
||||
|
||||
``None`` means "no per-request routing" — the caller uses its default tracer.
|
||||
"""
|
||||
builder = DYNAMIC_HEADERS_BY_CALLBACK.get(callback_name or "")
|
||||
if builder is None or not dynamic_params:
|
||||
return None
|
||||
headers = builder(dynamic_params)
|
||||
return headers or None
|
||||
|
||||
|
||||
__all__ = [
|
||||
"PRESET_BY_CALLBACK",
|
||||
"DYNAMIC_HEADERS_BY_CALLBACK",
|
||||
"Preset",
|
||||
"dynamic_otlp_headers",
|
||||
"agentops_preset",
|
||||
"arize_preset",
|
||||
"langfuse_preset",
|
||||
"langtrace_preset",
|
||||
"levo_preset",
|
||||
"phoenix_preset",
|
||||
"weave_preset",
|
||||
]
|
||||
139
litellm/integrations/otel/presets/agentops.py
Normal file
139
litellm/integrations/otel/presets/agentops.py
Normal file
|
|
@ -0,0 +1,139 @@
|
|||
"""AgentOps preset — OTLP/HTTP to AgentOps' endpoint with a lazily-fetched JWT.
|
||||
|
||||
AgentOps authenticates with a short-lived JWT minted from the API key. Fetching
|
||||
it is blocking network I/O, so it must never run on the event loop: callback
|
||||
construction (where presets are built) can run inside the proxy's async startup
|
||||
or, in the SDK, on the first request. Instead of fetching at config-build time,
|
||||
this preset registers a custom exporter (``kind="agentops"``) that mints the JWT
|
||||
**on its first export** — which the ``BatchSpanProcessor`` runs in its own
|
||||
worker thread, off any event loop — and caches it for the process lifetime.
|
||||
"""
|
||||
|
||||
from typing import Any
|
||||
|
||||
import httpx
|
||||
from pydantic import Field
|
||||
from pydantic_settings import BaseSettings, SettingsConfigDict
|
||||
|
||||
from litellm._logging import verbose_logger
|
||||
from litellm.integrations.otel.config import ExporterSpec, OpenTelemetryV2Config
|
||||
from litellm.integrations.otel.providers import register_exporter_factory
|
||||
|
||||
_AGENTOPS_ENDPOINT = "https://otlp.agentops.cloud/v1/traces"
|
||||
_AGENTOPS_AUTH_ENDPOINT = "https://api.agentops.ai/v3/auth/token"
|
||||
_AGENTOPS_EXPORTER_KIND = "agentops"
|
||||
|
||||
|
||||
class _AgentOpsSettings(BaseSettings):
|
||||
model_config = SettingsConfigDict(case_sensitive=False, extra="ignore")
|
||||
|
||||
api_key: str | None = Field(default=None, validation_alias="AGENTOPS_API_KEY")
|
||||
service_name: str = Field(
|
||||
default="agentops", validation_alias="AGENTOPS_SERVICE_NAME"
|
||||
)
|
||||
environment: str | None = Field(
|
||||
default=None, validation_alias="AGENTOPS_ENVIRONMENT"
|
||||
)
|
||||
|
||||
|
||||
def agentops_preset(
|
||||
*,
|
||||
config_overrides: OpenTelemetryV2Config | None = None,
|
||||
) -> OpenTelemetryV2Config:
|
||||
"""Build the AgentOps config without any network I/O.
|
||||
|
||||
The ``agentops`` exporter mints (and caches) the JWT lazily on its first
|
||||
export, so this stays non-blocking. ``project.id`` is therefore not a
|
||||
resource attribute — it is encoded in the JWT, which AgentOps uses to route
|
||||
the trace to the right project.
|
||||
"""
|
||||
settings = _AgentOpsSettings()
|
||||
base = config_overrides or OpenTelemetryV2Config()
|
||||
return base.model_copy(
|
||||
update={
|
||||
"exporters": [
|
||||
*base.exporters,
|
||||
ExporterSpec(
|
||||
kind=_AGENTOPS_EXPORTER_KIND,
|
||||
endpoint=_AGENTOPS_ENDPOINT,
|
||||
options=(
|
||||
{"api_key": settings.api_key} if settings.api_key else None
|
||||
),
|
||||
),
|
||||
],
|
||||
"resource_attributes": {
|
||||
**base.resource_attributes,
|
||||
"service.name": settings.service_name,
|
||||
"telemetry.sdk.name": "agentops",
|
||||
**(
|
||||
{"deployment.environment": settings.environment}
|
||||
if settings.environment
|
||||
else {}
|
||||
),
|
||||
},
|
||||
}
|
||||
)
|
||||
|
||||
|
||||
def _build_agentops_exporter(spec: ExporterSpec) -> Any:
|
||||
"""Factory for the ``agentops`` exporter kind: a lazy-auth OTLP/HTTP exporter."""
|
||||
from opentelemetry.exporter.otlp.proto.http.trace_exporter import (
|
||||
OTLPSpanExporter,
|
||||
)
|
||||
|
||||
class _LazyAuthAgentOpsExporter(OTLPSpanExporter):
|
||||
"""OTLP/HTTP exporter that mints the AgentOps JWT on its first export.
|
||||
|
||||
``export`` runs in the ``BatchSpanProcessor`` worker thread, so the
|
||||
blocking token fetch never touches an event loop. The result is cached
|
||||
after the first attempt (success or failure) so it runs at most once.
|
||||
"""
|
||||
|
||||
def __init__(self, *, endpoint: str | None, api_key: str | None) -> None:
|
||||
super().__init__(endpoint=endpoint)
|
||||
self._agentops_api_key = api_key
|
||||
self._auth_resolved = False
|
||||
|
||||
def _ensure_authenticated(self) -> None:
|
||||
if self._auth_resolved:
|
||||
return
|
||||
self._auth_resolved = True
|
||||
if not self._agentops_api_key:
|
||||
return
|
||||
try:
|
||||
token = _fetch_agentops_jwt(self._agentops_api_key).get("token")
|
||||
if token:
|
||||
# ``_session`` is the requests.Session the base exporter
|
||||
# POSTs through; updating its Authorization header is how the
|
||||
# minted JWT reaches every subsequent export.
|
||||
self._session.headers["Authorization"] = f"Bearer {token}"
|
||||
except Exception as e:
|
||||
verbose_logger.debug("AgentOps JWT fetch failed: %s", e)
|
||||
|
||||
def export(self, spans: Any) -> Any:
|
||||
self._ensure_authenticated()
|
||||
return super().export(spans)
|
||||
|
||||
options = spec.options or {}
|
||||
return _LazyAuthAgentOpsExporter(
|
||||
endpoint=spec.endpoint, api_key=options.get("api_key")
|
||||
)
|
||||
|
||||
|
||||
def _fetch_agentops_jwt(api_key: str) -> dict[str, Any]:
|
||||
# Own a short-lived client rather than ``_get_httpx_client()``: that returns
|
||||
# a process-wide cached ``HTTPHandler`` whose connection pool is shared by
|
||||
# every caller, so closing it here would break concurrent/subsequent
|
||||
# requests. This one-shot auth call gets its own client to close.
|
||||
with httpx.Client(timeout=10) as client:
|
||||
response = client.post(
|
||||
url=_AGENTOPS_AUTH_ENDPOINT,
|
||||
headers={"Content-Type": "application/json", "Connection": "keep-alive"},
|
||||
json={"api_key": api_key},
|
||||
)
|
||||
if response.status_code != 200:
|
||||
raise RuntimeError(f"Failed to fetch AgentOps token: {response.text}")
|
||||
return response.json()
|
||||
|
||||
|
||||
register_exporter_factory(_AGENTOPS_EXPORTER_KIND, _build_agentops_exporter)
|
||||
75
litellm/integrations/otel/presets/arize.py
Normal file
75
litellm/integrations/otel/presets/arize.py
Normal file
|
|
@ -0,0 +1,75 @@
|
|||
"""Arize preset — OTLP exporter to Arize + OpenInference vocabulary."""
|
||||
|
||||
from pydantic import Field
|
||||
from pydantic_settings import BaseSettings, SettingsConfigDict
|
||||
|
||||
from litellm.integrations.arize.arize import ArizeLogger as _V1ArizeLogger
|
||||
from litellm.integrations.otel.config import ExporterSpec, OpenTelemetryV2Config
|
||||
from litellm.integrations.otel.presets.utils import ensure_mappers
|
||||
from litellm.types.utils import StandardCallbackDynamicParams
|
||||
|
||||
|
||||
class _ArizeSettings(BaseSettings):
|
||||
model_config = SettingsConfigDict(case_sensitive=False, extra="ignore")
|
||||
|
||||
# Standard OTLP headers env var, used as the fallback when no Arize
|
||||
# credentials are configured.
|
||||
otlp_traces_headers: str | None = Field(
|
||||
default=None, validation_alias="OTEL_EXPORTER_OTLP_TRACES_HEADERS"
|
||||
)
|
||||
|
||||
|
||||
def arize_preset(
|
||||
*,
|
||||
config_overrides: OpenTelemetryV2Config | None = None,
|
||||
) -> OpenTelemetryV2Config:
|
||||
arize_cfg = _V1ArizeLogger.get_arize_config()
|
||||
headers = _arize_headers(arize_cfg)
|
||||
base = config_overrides or OpenTelemetryV2Config()
|
||||
return base.model_copy(
|
||||
update={
|
||||
"exporters": [
|
||||
*base.exporters,
|
||||
ExporterSpec(
|
||||
kind=arize_cfg.protocol or "otlp_grpc",
|
||||
endpoint=arize_cfg.endpoint or "https://otlp.arize.com/v1",
|
||||
headers=headers,
|
||||
),
|
||||
],
|
||||
"mapper_names": ensure_mappers(base.mapper_names, "openinference"),
|
||||
"resource_attributes": {
|
||||
**base.resource_attributes,
|
||||
**(
|
||||
{"model_id": arize_cfg.project_name}
|
||||
if arize_cfg.project_name
|
||||
else {}
|
||||
),
|
||||
},
|
||||
}
|
||||
)
|
||||
|
||||
|
||||
def _arize_headers(arize_cfg) -> str | None:
|
||||
pieces = []
|
||||
if arize_cfg.space_id or arize_cfg.space_key:
|
||||
pieces.append(f"space_id={arize_cfg.space_id or arize_cfg.space_key}")
|
||||
if arize_cfg.api_key:
|
||||
pieces.append(f"api_key={arize_cfg.api_key}")
|
||||
if not pieces:
|
||||
# Fall back to the standard OTLP headers env var when no Arize
|
||||
# credentials are configured.
|
||||
return _ArizeSettings().otlp_traces_headers
|
||||
return ",".join(pieces)
|
||||
|
||||
|
||||
def arize_dynamic_headers(params: StandardCallbackDynamicParams) -> dict[str, str]:
|
||||
"""Per-request Arize OTLP headers from team/key dynamic params."""
|
||||
headers: dict[str, str] = {}
|
||||
# ``arize_space_key`` is the suggested param and wins over ``arize_space_id``.
|
||||
space = params.get("arize_space_key") or params.get("arize_space_id")
|
||||
if space:
|
||||
headers["arize-space-id"] = space
|
||||
api_key = params.get("arize_api_key")
|
||||
if api_key:
|
||||
headers["api_key"] = api_key
|
||||
return headers
|
||||
25
litellm/integrations/otel/presets/base.py
Normal file
25
litellm/integrations/otel/presets/base.py
Normal file
|
|
@ -0,0 +1,25 @@
|
|||
"""Preset interface.
|
||||
|
||||
A preset is a callable that reads its integration's env vars and produces an
|
||||
:class:`OpenTelemetryV2Config` (exporter list + mapper-name list + resource
|
||||
attributes). This ``Protocol`` pins that contract so ``PRESET_BY_CALLBACK`` and
|
||||
the factory in ``litellm_logging`` are type-checked structurally against it,
|
||||
matching the ``AttributeMapper`` protocol the mappers use.
|
||||
"""
|
||||
|
||||
from typing import Protocol, runtime_checkable
|
||||
|
||||
from litellm.integrations.otel.config import OpenTelemetryV2Config
|
||||
|
||||
|
||||
@runtime_checkable
|
||||
class Preset(Protocol):
|
||||
"""Reads an integration's env config and returns an ``OpenTelemetryV2Config``.
|
||||
|
||||
``config_overrides`` lets one preset layer onto another's config (or onto
|
||||
test-supplied defaults); the factory calls presets with no arguments.
|
||||
"""
|
||||
|
||||
def __call__(
|
||||
self, *, config_overrides: OpenTelemetryV2Config | None = None
|
||||
) -> OpenTelemetryV2Config: ...
|
||||
42
litellm/integrations/otel/presets/langfuse.py
Normal file
42
litellm/integrations/otel/presets/langfuse.py
Normal file
|
|
@ -0,0 +1,42 @@
|
|||
"""Langfuse-OTEL preset."""
|
||||
|
||||
from litellm.integrations.langfuse.langfuse_otel import (
|
||||
LangfuseOtelLogger as _V1Langfuse,
|
||||
)
|
||||
from litellm.integrations.otel.config import ExporterSpec, OpenTelemetryV2Config
|
||||
from litellm.integrations.otel.presets.utils import ensure_mappers
|
||||
from litellm.types.utils import StandardCallbackDynamicParams
|
||||
|
||||
|
||||
def langfuse_preset(
|
||||
*,
|
||||
config_overrides: OpenTelemetryV2Config | None = None,
|
||||
) -> OpenTelemetryV2Config:
|
||||
cfg = _V1Langfuse.get_langfuse_otel_config()
|
||||
base = config_overrides or OpenTelemetryV2Config()
|
||||
return base.model_copy(
|
||||
update={
|
||||
"exporters": [
|
||||
*base.exporters,
|
||||
ExporterSpec(
|
||||
kind=cfg.exporter if hasattr(cfg, "exporter") else "otlp_http",
|
||||
endpoint=cfg.endpoint,
|
||||
headers=cfg.headers,
|
||||
),
|
||||
],
|
||||
"mapper_names": ensure_mappers(base.mapper_names, "langfuse"),
|
||||
}
|
||||
)
|
||||
|
||||
|
||||
def langfuse_dynamic_headers(params: StandardCallbackDynamicParams) -> dict[str, str]:
|
||||
"""Per-request Langfuse OTLP headers from team/key dynamic params."""
|
||||
public_key = params.get("langfuse_public_key")
|
||||
secret_key = params.get("langfuse_secret_key")
|
||||
if public_key and secret_key:
|
||||
return {
|
||||
"Authorization": _V1Langfuse._get_langfuse_authorization_header(
|
||||
public_key=public_key, secret_key=secret_key
|
||||
)
|
||||
}
|
||||
return {}
|
||||
22
litellm/integrations/otel/presets/langtrace.py
Normal file
22
litellm/integrations/otel/presets/langtrace.py
Normal file
|
|
@ -0,0 +1,22 @@
|
|||
"""Langtrace preset — Langtrace consumes generic OTLP + a vendor mapper."""
|
||||
|
||||
from litellm.integrations.otel.config import OpenTelemetryV2Config
|
||||
from litellm.integrations.otel.presets.utils import ensure_mappers
|
||||
|
||||
|
||||
def langtrace_preset(
|
||||
*,
|
||||
config_overrides: OpenTelemetryV2Config | None = None,
|
||||
) -> OpenTelemetryV2Config:
|
||||
"""Compose the Langtrace mapper on top of the customer's OTLP destination.
|
||||
|
||||
Unlike Arize / Phoenix / Langfuse, Langtrace doesn't ship its own endpoint
|
||||
— users point their existing OTLP collector at Langtrace and just
|
||||
need the vendor attribute schema applied to outgoing spans.
|
||||
"""
|
||||
base = config_overrides or OpenTelemetryV2Config()
|
||||
return base.model_copy(
|
||||
update={
|
||||
"mapper_names": ensure_mappers(base.mapper_names, "langtrace"),
|
||||
}
|
||||
)
|
||||
24
litellm/integrations/otel/presets/levo.py
Normal file
24
litellm/integrations/otel/presets/levo.py
Normal file
|
|
@ -0,0 +1,24 @@
|
|||
"""Levo preset — OTLP/HTTP to a Levo collector with org+workspace headers."""
|
||||
|
||||
from litellm.integrations.levo.levo import LevoLogger as _V1Levo
|
||||
from litellm.integrations.otel.config import ExporterSpec, OpenTelemetryV2Config
|
||||
|
||||
|
||||
def levo_preset(
|
||||
*,
|
||||
config_overrides: OpenTelemetryV2Config | None = None,
|
||||
) -> OpenTelemetryV2Config:
|
||||
cfg = _V1Levo.get_levo_config()
|
||||
base = config_overrides or OpenTelemetryV2Config()
|
||||
return base.model_copy(
|
||||
update={
|
||||
"exporters": [
|
||||
*base.exporters,
|
||||
ExporterSpec(
|
||||
kind="otlp_http",
|
||||
endpoint=cfg.endpoint,
|
||||
headers=cfg.otlp_auth_headers,
|
||||
),
|
||||
],
|
||||
}
|
||||
)
|
||||
48
litellm/integrations/otel/presets/phoenix.py
Normal file
48
litellm/integrations/otel/presets/phoenix.py
Normal file
|
|
@ -0,0 +1,48 @@
|
|||
"""Arize-Phoenix preset."""
|
||||
|
||||
from pydantic import AliasChoices, Field
|
||||
from pydantic_settings import BaseSettings, SettingsConfigDict
|
||||
|
||||
from litellm.integrations.arize.arize_phoenix import (
|
||||
ArizePhoenixLogger as _V1Phoenix,
|
||||
)
|
||||
from litellm.integrations.otel.config import ExporterSpec, OpenTelemetryV2Config
|
||||
from litellm.integrations.otel.presets.utils import ensure_mappers
|
||||
|
||||
|
||||
class _PhoenixSettings(BaseSettings):
|
||||
model_config = SettingsConfigDict(case_sensitive=False, extra="ignore")
|
||||
|
||||
project_name: str = Field(
|
||||
default="default",
|
||||
validation_alias=AliasChoices(
|
||||
"PHOENIX_PROJECT_NAME", "PHOENIX_COLLECTOR_PROJECT_NAME"
|
||||
),
|
||||
)
|
||||
|
||||
|
||||
def phoenix_preset(
|
||||
*,
|
||||
config_overrides: OpenTelemetryV2Config | None = None,
|
||||
) -> OpenTelemetryV2Config:
|
||||
cfg = _V1Phoenix.get_arize_phoenix_config()
|
||||
headers = cfg.otlp_auth_headers if hasattr(cfg, "otlp_auth_headers") else None
|
||||
project_name = _PhoenixSettings().project_name
|
||||
base = config_overrides or OpenTelemetryV2Config()
|
||||
return base.model_copy(
|
||||
update={
|
||||
"exporters": [
|
||||
*base.exporters,
|
||||
ExporterSpec(
|
||||
kind=cfg.protocol if hasattr(cfg, "protocol") else "otlp_http",
|
||||
endpoint=cfg.endpoint,
|
||||
headers=headers,
|
||||
),
|
||||
],
|
||||
"mapper_names": ensure_mappers(base.mapper_names, "openinference"),
|
||||
"resource_attributes": {
|
||||
**base.resource_attributes,
|
||||
"openinference.project.name": project_name,
|
||||
},
|
||||
}
|
||||
)
|
||||
16
litellm/integrations/otel/presets/utils.py
Normal file
16
litellm/integrations/otel/presets/utils.py
Normal file
|
|
@ -0,0 +1,16 @@
|
|||
"""Shared helpers for the integration presets."""
|
||||
|
||||
from typing import Iterable
|
||||
|
||||
|
||||
def ensure_mappers(mapper_names: Iterable[str], *names: str) -> list[str]:
|
||||
"""Return ``mapper_names`` with each of ``names`` appended if not already present.
|
||||
|
||||
Order is preserved and duplicates are skipped, so composing several presets
|
||||
(or re-applying one) never double-adds a vocabulary.
|
||||
"""
|
||||
result = list(mapper_names)
|
||||
for name in names:
|
||||
if name not in result:
|
||||
result.append(name)
|
||||
return result
|
||||
43
litellm/integrations/otel/presets/weave.py
Normal file
43
litellm/integrations/otel/presets/weave.py
Normal file
|
|
@ -0,0 +1,43 @@
|
|||
"""Weave (W&B) preset."""
|
||||
|
||||
from litellm.integrations.otel.config import ExporterSpec, OpenTelemetryV2Config
|
||||
from litellm.integrations.otel.presets.utils import ensure_mappers
|
||||
from litellm.integrations.weave.weave_otel import (
|
||||
_get_weave_authorization_header,
|
||||
get_weave_otel_config,
|
||||
)
|
||||
from litellm.types.utils import StandardCallbackDynamicParams
|
||||
|
||||
|
||||
def weave_preset(
|
||||
*,
|
||||
config_overrides: OpenTelemetryV2Config | None = None,
|
||||
) -> OpenTelemetryV2Config:
|
||||
weave_cfg = get_weave_otel_config()
|
||||
base = config_overrides or OpenTelemetryV2Config()
|
||||
return base.model_copy(
|
||||
update={
|
||||
"exporters": [
|
||||
*base.exporters,
|
||||
ExporterSpec(
|
||||
kind=weave_cfg.protocol or "otlp_http",
|
||||
endpoint=weave_cfg.endpoint,
|
||||
headers=weave_cfg.otlp_auth_headers,
|
||||
),
|
||||
],
|
||||
# Weave consumes OpenInference + a small Weave-specific overlay.
|
||||
"mapper_names": ensure_mappers(base.mapper_names, "openinference", "weave"),
|
||||
}
|
||||
)
|
||||
|
||||
|
||||
def weave_dynamic_headers(params: StandardCallbackDynamicParams) -> dict[str, str]:
|
||||
"""Per-request Weave OTLP headers from team/key dynamic params."""
|
||||
headers: dict[str, str] = {}
|
||||
api_key = params.get("wandb_api_key")
|
||||
if api_key:
|
||||
headers["Authorization"] = _get_weave_authorization_header(api_key=api_key)
|
||||
project_id = params.get("weave_project_id")
|
||||
if project_id:
|
||||
headers["project_id"] = project_id
|
||||
return headers
|
||||
220
litellm/integrations/otel/providers.py
Normal file
220
litellm/integrations/otel/providers.py
Normal file
|
|
@ -0,0 +1,220 @@
|
|||
"""Provider / exporter factory + the Baggage span processor."""
|
||||
|
||||
from typing import Callable, Iterable
|
||||
|
||||
from opentelemetry import baggage
|
||||
from opentelemetry.context import Context
|
||||
from opentelemetry.sdk.resources import Resource
|
||||
from opentelemetry.sdk.trace import ReadableSpan, SpanProcessor, TracerProvider
|
||||
from opentelemetry.sdk.trace.export import (
|
||||
BatchSpanProcessor,
|
||||
ConsoleSpanExporter,
|
||||
SimpleSpanProcessor,
|
||||
SpanExporter,
|
||||
)
|
||||
from opentelemetry.sdk.trace.export.in_memory_span_exporter import (
|
||||
InMemorySpanExporter,
|
||||
)
|
||||
from opentelemetry.trace import Span, SpanKind, Tracer
|
||||
|
||||
from litellm.integrations.otel.config import ExporterSpec, OpenTelemetryV2Config
|
||||
from litellm.integrations.otel.semconv import LiteLLM
|
||||
from litellm.integrations.otel.spans import LiteLLMSpanKind
|
||||
|
||||
# Re-exported so ``providers.parse_headers`` remains a stable entry point.
|
||||
from litellm.integrations.otel.utils import parse_headers as parse_headers
|
||||
|
||||
_SPAN_KIND_BY_ROLE_KIND: dict[LiteLLMSpanKind, SpanKind] = {
|
||||
LiteLLMSpanKind.SERVER: SpanKind.SERVER,
|
||||
LiteLLMSpanKind.CLIENT: SpanKind.CLIENT,
|
||||
LiteLLMSpanKind.INTERNAL: SpanKind.INTERNAL,
|
||||
LiteLLMSpanKind.PRODUCER: SpanKind.PRODUCER,
|
||||
LiteLLMSpanKind.CONSUMER: SpanKind.CONSUMER,
|
||||
}
|
||||
|
||||
|
||||
def to_otel_span_kind(kind: LiteLLMSpanKind) -> SpanKind:
|
||||
return _SPAN_KIND_BY_ROLE_KIND[kind]
|
||||
|
||||
|
||||
# Custom exporter factories keyed by ``ExporterSpec.kind``. A preset registers
|
||||
# one here when its destination needs construction logic the built-in kinds
|
||||
# can't express — e.g. an exporter that fetches an auth token lazily on its
|
||||
# first export (off the event loop) instead of blocking at config-build time.
|
||||
# Keeping the registry here lets this module stay vendor-agnostic: the factory
|
||||
# lives with the integration that needs it.
|
||||
_EXPORTER_FACTORIES: dict[str, Callable[[ExporterSpec], SpanExporter]] = {}
|
||||
|
||||
|
||||
def register_exporter_factory(
|
||||
kind: str, factory: Callable[[ExporterSpec], SpanExporter]
|
||||
) -> None:
|
||||
"""Register a custom exporter ``factory`` for the exporter ``kind``."""
|
||||
_EXPORTER_FACTORIES[kind.lower()] = factory
|
||||
|
||||
|
||||
class LiteLLMBaggageSpanProcessor(SpanProcessor):
|
||||
"""Stamps an allowlisted set of Baggage entries onto every span at start."""
|
||||
|
||||
def __init__(
|
||||
self,
|
||||
allowed_keys: Iterable[str],
|
||||
allowed_prefixes: tuple[str, ...] = (LiteLLM.METADATA_PREFIX,),
|
||||
) -> None:
|
||||
self._allowed_keys = frozenset(allowed_keys)
|
||||
self._allowed_prefixes = tuple(allowed_prefixes)
|
||||
|
||||
def _is_allowed(self, key: str) -> bool:
|
||||
return key in self._allowed_keys or any(
|
||||
key.startswith(prefix) for prefix in self._allowed_prefixes
|
||||
)
|
||||
|
||||
def on_start(self, span: Span, parent_context: Context | None = None) -> None:
|
||||
for key, value in baggage.get_all(parent_context).items():
|
||||
if self._is_allowed(key) and isinstance(value, (str, bool, int, float)):
|
||||
span.set_attribute(key, value)
|
||||
|
||||
def on_end(self, span: ReadableSpan) -> None: # noqa: D401 - no-op
|
||||
return None
|
||||
|
||||
def shutdown(self) -> None:
|
||||
return None
|
||||
|
||||
def force_flush(self, timeout_millis: int = 30000) -> bool:
|
||||
return True
|
||||
|
||||
|
||||
def _otlp_traces_endpoint(endpoint: str | None) -> str | None:
|
||||
"""Point an OTLP/HTTP base endpoint at the ``/v1/traces`` signal path.
|
||||
|
||||
``OTEL_EXPORTER_OTLP_ENDPOINT`` is a base URL (e.g. ``http://host:4318``).
|
||||
The OTLP/HTTP exporter only appends the ``/v1/traces`` path when it reads
|
||||
that env var itself; when an endpoint is passed explicitly it is used
|
||||
verbatim, so a base URL would POST to the root and the collector returns
|
||||
404. Append the signal path here (leaving an already-correct path intact).
|
||||
"""
|
||||
if not endpoint:
|
||||
return endpoint
|
||||
endpoint = endpoint.rstrip("/")
|
||||
# Splunk Observability uses ``/v2/trace/otlp``; never rewrite it.
|
||||
if endpoint.endswith("/v1/traces") or "/v2/trace/otlp" in endpoint:
|
||||
return endpoint
|
||||
for other_signal in ("/v1/logs", "/v1/metrics"):
|
||||
if endpoint.endswith(other_signal):
|
||||
return endpoint[: -len(other_signal)] + "/v1/traces"
|
||||
return endpoint + "/v1/traces"
|
||||
|
||||
|
||||
def _exporter_from_spec(spec: ExporterSpec) -> SpanExporter:
|
||||
kind = (spec.kind or "console").lower()
|
||||
factory = _EXPORTER_FACTORIES.get(kind)
|
||||
if factory is not None:
|
||||
return factory(spec)
|
||||
if kind in ("in_memory", "inmemory", "memory"):
|
||||
return InMemorySpanExporter()
|
||||
if kind in ("otlp_http", "http", "http/protobuf", "http/json"):
|
||||
from opentelemetry.exporter.otlp.proto.http.trace_exporter import (
|
||||
OTLPSpanExporter as HTTPExporter,
|
||||
)
|
||||
|
||||
return HTTPExporter(
|
||||
endpoint=_otlp_traces_endpoint(spec.endpoint),
|
||||
headers=parse_headers(spec.headers),
|
||||
)
|
||||
if kind in ("otlp_grpc", "grpc"):
|
||||
from opentelemetry.exporter.otlp.proto.grpc.trace_exporter import (
|
||||
OTLPSpanExporter as GRPCExporter,
|
||||
)
|
||||
|
||||
return GRPCExporter(endpoint=spec.endpoint, headers=parse_headers(spec.headers))
|
||||
return ConsoleSpanExporter()
|
||||
|
||||
|
||||
def _processor_for(exporter: SpanExporter, use_simple: bool | None) -> SpanProcessor:
|
||||
"""Pick a Simple or Batch span processor for ``exporter``.
|
||||
|
||||
When ``use_simple`` is unset, default to Simple for console and in-memory
|
||||
exporters (spans export synchronously, which tests rely on) and Batch for
|
||||
everything else (the right export semantics for production).
|
||||
"""
|
||||
if use_simple is None:
|
||||
use_simple = isinstance(exporter, (ConsoleSpanExporter, InMemorySpanExporter))
|
||||
return SimpleSpanProcessor(exporter) if use_simple else BatchSpanProcessor(exporter)
|
||||
|
||||
|
||||
def build_span_exporter(config: OpenTelemetryV2Config) -> SpanExporter:
|
||||
"""Build a single exporter from the top-level config fields.
|
||||
|
||||
Convenience for the common single-exporter case (and for tests): reads the
|
||||
``exporter`` / ``endpoint`` / ``headers`` fields. To configure multiple
|
||||
exporters, populate ``config.exporters`` directly.
|
||||
"""
|
||||
return _exporter_from_spec(
|
||||
ExporterSpec(
|
||||
kind=config.exporter, endpoint=config.endpoint, headers=config.headers
|
||||
)
|
||||
)
|
||||
|
||||
|
||||
def build_resource(config: OpenTelemetryV2Config) -> Resource:
|
||||
attributes: dict[str, str] = {"service.name": config.service_name}
|
||||
if config.deployment_environment:
|
||||
attributes["deployment.environment"] = config.deployment_environment
|
||||
attributes.update(config.resource_attributes)
|
||||
return Resource.create(attributes)
|
||||
|
||||
|
||||
def build_tracer_provider(
|
||||
config: OpenTelemetryV2Config,
|
||||
exporter: SpanExporter | None = None,
|
||||
baggage_processor: SpanProcessor | None = None,
|
||||
use_simple_processor: bool | None = None,
|
||||
) -> TracerProvider:
|
||||
"""Build the shared :class:`TracerProvider`.
|
||||
|
||||
Attach the Baggage processor first (so identity attributes land on each
|
||||
span before any export decision), then add one ``SpanProcessor`` per
|
||||
``config.exporters`` entry — this is what fans spans out to multiple
|
||||
backends. ``exporter`` and ``use_simple_processor`` are explicit overrides:
|
||||
pass a single exporter to attach exactly that one (used by tests).
|
||||
"""
|
||||
provider = TracerProvider(resource=build_resource(config))
|
||||
if baggage_processor is None:
|
||||
baggage_processor = LiteLLMBaggageSpanProcessor(
|
||||
allowed_keys=config.baggage_promoted_keys
|
||||
)
|
||||
provider.add_span_processor(baggage_processor)
|
||||
|
||||
if exporter is not None:
|
||||
provider.add_span_processor(_processor_for(exporter, use_simple_processor))
|
||||
return provider
|
||||
|
||||
# ``config._normalize`` guarantees at least one spec (it folds the top-level
|
||||
# ``exporter``/``endpoint``/``headers`` fields in when ``exporters`` is empty).
|
||||
for spec in config.exporters:
|
||||
exp = _exporter_from_spec(spec)
|
||||
provider.add_span_processor(
|
||||
_processor_for(
|
||||
exp,
|
||||
(
|
||||
spec.use_simple_processor
|
||||
if spec.use_simple_processor is not None
|
||||
else use_simple_processor
|
||||
),
|
||||
)
|
||||
)
|
||||
return provider
|
||||
|
||||
|
||||
def get_tracer(provider: TracerProvider, name: str = "litellm") -> Tracer:
|
||||
return provider.get_tracer(name)
|
||||
|
||||
|
||||
def in_memory_provider(
|
||||
config: OpenTelemetryV2Config | None = None,
|
||||
) -> tuple[TracerProvider, InMemorySpanExporter]:
|
||||
"""Convenience for tests: a provider exporting to an in-memory buffer."""
|
||||
cfg = config or OpenTelemetryV2Config(exporter="in_memory")
|
||||
exporter = InMemorySpanExporter()
|
||||
provider = build_tracer_provider(cfg, exporter=exporter)
|
||||
return provider, exporter
|
||||
98
litellm/integrations/otel/routing.py
Normal file
98
litellm/integrations/otel/routing.py
Normal file
|
|
@ -0,0 +1,98 @@
|
|||
"""Per-request multi-tenant tracer routing.
|
||||
|
||||
When a request carries team/key vendor credentials in
|
||||
``standard_callback_dynamic_params``, its spans must export through a
|
||||
``TracerProvider`` whose OTLP headers carry those credentials.
|
||||
``TenantTracerCache`` builds and caches one provider per distinct credential
|
||||
set, and otherwise hands back the logger's default tracer. This lets a single
|
||||
logger fan requests out to many tenants without needing a logger per tenant.
|
||||
"""
|
||||
|
||||
from collections import OrderedDict
|
||||
from typing import Any, Mapping
|
||||
|
||||
from opentelemetry.sdk.trace import TracerProvider
|
||||
from opentelemetry.trace import Tracer
|
||||
|
||||
from litellm._logging import verbose_logger
|
||||
from litellm.integrations.otel.config import OpenTelemetryV2Config
|
||||
from litellm.integrations.otel.presets import dynamic_otlp_headers
|
||||
from litellm.integrations.otel.providers import build_tracer_provider, get_tracer
|
||||
|
||||
# Exporter kinds that ignore headers — never rewritten with dynamic credentials.
|
||||
_NON_OTLP_KINDS = ("console", "in_memory", "inmemory", "memory")
|
||||
|
||||
# Cap on distinct credential-scoped providers held at once. ``dynamic_params``
|
||||
# can be populated from request metadata, so an unbounded cache lets a caller
|
||||
# spawn one ``TracerProvider`` (plus its ``BatchSpanProcessor`` background
|
||||
# thread) per unique credential set and exhaust the proxy. The LRU bound keeps
|
||||
# the working set of active tenants resident while flushing and shutting down
|
||||
# evicted providers so their threads are reclaimed.
|
||||
_MAX_CACHED_PROVIDERS = 256
|
||||
|
||||
|
||||
def _shutdown_provider(provider: TracerProvider) -> None:
|
||||
"""Flush + stop an evicted provider's processors (reclaims their threads).
|
||||
|
||||
``TracerProvider.shutdown`` force-flushes each ``SpanProcessor`` before
|
||||
stopping it, so any spans already handed to a ``BatchSpanProcessor`` are
|
||||
exported rather than dropped. Best-effort: a shutdown failure must not break
|
||||
the request that triggered the eviction.
|
||||
"""
|
||||
try:
|
||||
provider.shutdown()
|
||||
except Exception as e: # pragma: no cover - defensive
|
||||
verbose_logger.debug("OTel V2: error shutting down evicted provider: %s", e)
|
||||
|
||||
|
||||
class TenantTracerCache:
|
||||
"""Credential-scoped ``TracerProvider`` cache keyed by the dynamic headers."""
|
||||
|
||||
def __init__(
|
||||
self,
|
||||
config: OpenTelemetryV2Config,
|
||||
callback_name: str | None,
|
||||
tracer_name: str,
|
||||
) -> None:
|
||||
self._config = config
|
||||
self._callback_name = callback_name
|
||||
self._tracer_name = tracer_name
|
||||
self._providers: "OrderedDict[tuple[tuple[str, str], ...], TracerProvider]" = (
|
||||
OrderedDict()
|
||||
)
|
||||
|
||||
def tracer_for(self, default: Tracer, dynamic_params: Any) -> Tracer:
|
||||
"""Return the tracer for this request.
|
||||
|
||||
Use ``default`` unless the request's dynamic credentials require a
|
||||
credential-scoped tracer, in which case build (or reuse) one. The cache
|
||||
is a bounded LRU: the least-recently-used provider is flushed and shut
|
||||
down on overflow so its exporter threads don't accumulate.
|
||||
"""
|
||||
headers = dynamic_otlp_headers(self._callback_name, dynamic_params)
|
||||
if not headers:
|
||||
return default
|
||||
cache_key = tuple(sorted(headers.items()))
|
||||
provider = self._providers.get(cache_key)
|
||||
if provider is not None:
|
||||
self._providers.move_to_end(cache_key)
|
||||
else:
|
||||
provider = build_tracer_provider(self._config_with_headers(headers))
|
||||
self._providers[cache_key] = provider
|
||||
if len(self._providers) > _MAX_CACHED_PROVIDERS:
|
||||
_, evicted = self._providers.popitem(last=False)
|
||||
_shutdown_provider(evicted)
|
||||
return get_tracer(provider, self._tracer_name)
|
||||
|
||||
def _config_with_headers(self, headers: Mapping[str, str]) -> OpenTelemetryV2Config:
|
||||
"""Clone the config, replacing OTLP exporter headers with ``headers``."""
|
||||
header_str = ",".join(f"{key}={value}" for key, value in headers.items())
|
||||
exporters = [
|
||||
(
|
||||
spec
|
||||
if spec.kind.lower() in _NON_OTLP_KINDS
|
||||
else spec.model_copy(update={"headers": header_str})
|
||||
)
|
||||
for spec in self._config.exporters
|
||||
]
|
||||
return self._config.model_copy(update={"exporters": exporters})
|
||||
182
litellm/integrations/otel/semconv.py
Normal file
182
litellm/integrations/otel/semconv.py
Normal file
|
|
@ -0,0 +1,182 @@
|
|||
"""
|
||||
Keys follow the OpenTelemetry GenAI semantic conventions (experimental). Anything
|
||||
without a semconv equivalent lives under the ``litellm.*`` vendor namespace.
|
||||
"""
|
||||
|
||||
from enum import Enum
|
||||
from typing import Final
|
||||
|
||||
|
||||
class GenAIOperation(str, Enum):
|
||||
"""Values for ``gen_ai.operation.name``."""
|
||||
|
||||
CHAT = "chat"
|
||||
TEXT_COMPLETION = "text_completion"
|
||||
EMBEDDINGS = "embeddings"
|
||||
GENERATE_CONTENT = "generate_content"
|
||||
CREATE_AGENT = "create_agent" # reserved for future agent spans
|
||||
INVOKE_AGENT = "invoke_agent" # reserved for future agent spans
|
||||
EXECUTE_TOOL = "execute_tool" # reserved for future tool spans
|
||||
|
||||
|
||||
class GenAIProvider(str, Enum):
|
||||
"""Common values for the ``gen_ai.provider.name`` attribute."""
|
||||
|
||||
OPENAI = "openai"
|
||||
ANTHROPIC = "anthropic"
|
||||
AWS_BEDROCK = "aws.bedrock"
|
||||
AZURE_AI_OPENAI = "azure.ai.openai"
|
||||
AZURE_AI_INFERENCE = "azure.ai.inference"
|
||||
GCP_GEMINI = "gcp.gemini"
|
||||
GCP_VERTEX_AI = "gcp.vertex_ai"
|
||||
COHERE = "cohere"
|
||||
MISTRAL_AI = "mistral_ai"
|
||||
DEEPSEEK = "deepseek"
|
||||
GROQ = "groq"
|
||||
PERPLEXITY = "perplexity"
|
||||
X_AI = "x_ai"
|
||||
IBM_WATSONX_AI = "ibm.watsonx.ai"
|
||||
|
||||
|
||||
class GenAI:
|
||||
"""Canonical OTel GenAI span-attribute keys."""
|
||||
|
||||
# request
|
||||
OPERATION_NAME: Final = "gen_ai.operation.name"
|
||||
PROVIDER_NAME: Final = "gen_ai.provider.name"
|
||||
REQUEST_MODEL: Final = "gen_ai.request.model"
|
||||
REQUEST_TEMPERATURE: Final = "gen_ai.request.temperature"
|
||||
REQUEST_TOP_P: Final = "gen_ai.request.top_p"
|
||||
REQUEST_TOP_K: Final = "gen_ai.request.top_k"
|
||||
REQUEST_MAX_TOKENS: Final = "gen_ai.request.max_tokens"
|
||||
REQUEST_FREQUENCY_PENALTY: Final = "gen_ai.request.frequency_penalty"
|
||||
REQUEST_PRESENCE_PENALTY: Final = "gen_ai.request.presence_penalty"
|
||||
REQUEST_STOP_SEQUENCES: Final = "gen_ai.request.stop_sequences"
|
||||
REQUEST_SEED: Final = "gen_ai.request.seed"
|
||||
REQUEST_CHOICE_COUNT: Final = "gen_ai.request.choice.count"
|
||||
REQUEST_ENCODING_FORMATS: Final = "gen_ai.request.encoding_formats"
|
||||
# response
|
||||
RESPONSE_ID: Final = "gen_ai.response.id"
|
||||
RESPONSE_MODEL: Final = "gen_ai.response.model"
|
||||
RESPONSE_FINISH_REASONS: Final = "gen_ai.response.finish_reasons"
|
||||
# usage
|
||||
USAGE_INPUT_TOKENS: Final = "gen_ai.usage.input_tokens"
|
||||
USAGE_OUTPUT_TOKENS: Final = "gen_ai.usage.output_tokens"
|
||||
# content (opt-in, gated by capture mode)
|
||||
INPUT_MESSAGES: Final = "gen_ai.input.messages"
|
||||
OUTPUT_MESSAGES: Final = "gen_ai.output.messages"
|
||||
SYSTEM_INSTRUCTIONS: Final = "gen_ai.system_instructions"
|
||||
OUTPUT_TYPE: Final = "gen_ai.output.type"
|
||||
CONVERSATION_ID: Final = "gen_ai.conversation.id"
|
||||
# agent / tool (reserved)
|
||||
AGENT_ID: Final = "gen_ai.agent.id"
|
||||
AGENT_NAME: Final = "gen_ai.agent.name"
|
||||
TOOL_NAME: Final = "gen_ai.tool.name"
|
||||
TOOL_CALL_ID: Final = "gen_ai.tool.call.id"
|
||||
|
||||
|
||||
class Error:
|
||||
TYPE: Final = "error.type"
|
||||
|
||||
|
||||
class Server:
|
||||
ADDRESS: Final = "server.address"
|
||||
PORT: Final = "server.port"
|
||||
|
||||
|
||||
class HTTP:
|
||||
"""HTTP server-span keys. Belong on the SERVER span only (never promoted)."""
|
||||
|
||||
REQUEST_METHOD: Final = "http.request.method"
|
||||
ROUTE: Final = "http.route"
|
||||
RESPONSE_STATUS_CODE: Final = "http.response.status_code"
|
||||
URL_PATH: Final = "url.path"
|
||||
|
||||
|
||||
class LiteLLM:
|
||||
"""Vendor-extension keys (no semconv equivalent). Always ``litellm.*``."""
|
||||
|
||||
CALL_ID: Final = "litellm.call_id"
|
||||
COST_PREFIX: Final = "litellm.cost."
|
||||
METADATA_PREFIX: Final = "litellm.metadata."
|
||||
TEAM_ID: Final = "litellm.team.id"
|
||||
TEAM_ALIAS: Final = "litellm.team.alias"
|
||||
KEY_HASH: Final = "litellm.api_key.hash"
|
||||
END_USER: Final = "litellm.end_user.id"
|
||||
REQUEST_STREAMING: Final = "litellm.request.streaming"
|
||||
GUARDRAIL_NAME: Final = "litellm.guardrail.name"
|
||||
GUARDRAIL_MODE: Final = "litellm.guardrail.mode"
|
||||
GUARDRAIL_STATUS: Final = "litellm.guardrail.status"
|
||||
GUARDRAIL_PROVIDER: Final = "litellm.guardrail.provider"
|
||||
GUARDRAIL_ACTION: Final = "litellm.guardrail.action"
|
||||
GUARDRAIL_RESPONSE: Final = "litellm.guardrail.response"
|
||||
GUARDRAIL_VIOLATION_CATEGORIES: Final = "litellm.guardrail.violation_categories"
|
||||
GUARDRAIL_CONFIDENCE_SCORE: Final = "litellm.guardrail.confidence_score"
|
||||
GUARDRAIL_RISK_SCORE: Final = "litellm.guardrail.risk_score"
|
||||
GUARDRAIL_MASKED_ENTITY_COUNT: Final = "litellm.guardrail.masked_entity_count"
|
||||
GUARDRAIL_DURATION: Final = "litellm.guardrail.duration"
|
||||
SERVICE_NAME: Final = "litellm.service.name"
|
||||
SERVICE_CALL_TYPE: Final = "litellm.service.call_type"
|
||||
PREPROCESSING_MS: Final = "litellm.preprocessing.duration_ms"
|
||||
|
||||
|
||||
class Metric:
|
||||
"""GenAI metric instrument names."""
|
||||
|
||||
TOKEN_USAGE: Final = "gen_ai.client.token.usage"
|
||||
OPERATION_DURATION: Final = "gen_ai.client.operation.duration"
|
||||
|
||||
|
||||
# litellm ``custom_llm_provider`` -> ``gen_ai.provider.name`` value.
|
||||
_PROVIDER_BY_LITELLM: dict[str, GenAIProvider] = {
|
||||
"openai": GenAIProvider.OPENAI,
|
||||
"text-completion-openai": GenAIProvider.OPENAI,
|
||||
"azure": GenAIProvider.AZURE_AI_OPENAI,
|
||||
"azure_ai": GenAIProvider.AZURE_AI_INFERENCE,
|
||||
"anthropic": GenAIProvider.ANTHROPIC,
|
||||
"bedrock": GenAIProvider.AWS_BEDROCK,
|
||||
"bedrock_converse": GenAIProvider.AWS_BEDROCK,
|
||||
"vertex_ai": GenAIProvider.GCP_VERTEX_AI,
|
||||
"vertex_ai_beta": GenAIProvider.GCP_VERTEX_AI,
|
||||
"gemini": GenAIProvider.GCP_GEMINI,
|
||||
"cohere": GenAIProvider.COHERE,
|
||||
"cohere_chat": GenAIProvider.COHERE,
|
||||
"mistral": GenAIProvider.MISTRAL_AI,
|
||||
"deepseek": GenAIProvider.DEEPSEEK,
|
||||
"groq": GenAIProvider.GROQ,
|
||||
"perplexity": GenAIProvider.PERPLEXITY,
|
||||
"xai": GenAIProvider.X_AI,
|
||||
"watsonx": GenAIProvider.IBM_WATSONX_AI,
|
||||
}
|
||||
|
||||
# litellm ``call_type`` -> ``gen_ai.operation.name``.
|
||||
_OPERATION_BY_CALL_TYPE: dict[str, GenAIOperation] = {
|
||||
"completion": GenAIOperation.CHAT,
|
||||
"acompletion": GenAIOperation.CHAT,
|
||||
"completion_with_retries": GenAIOperation.CHAT,
|
||||
"text_completion": GenAIOperation.TEXT_COMPLETION,
|
||||
"atext_completion": GenAIOperation.TEXT_COMPLETION,
|
||||
"embedding": GenAIOperation.EMBEDDINGS,
|
||||
"aembedding": GenAIOperation.EMBEDDINGS,
|
||||
"responses": GenAIOperation.CHAT,
|
||||
"aresponses": GenAIOperation.CHAT,
|
||||
}
|
||||
|
||||
|
||||
def resolve_provider(custom_llm_provider: str | None) -> str:
|
||||
"""Map a litellm provider string to a ``gen_ai.provider.name`` value.
|
||||
|
||||
Unknown providers pass through verbatim — the convention explicitly allows
|
||||
provider-specific values, so an unmapped name is still valid.
|
||||
"""
|
||||
if not custom_llm_provider:
|
||||
return ""
|
||||
mapped = _PROVIDER_BY_LITELLM.get(custom_llm_provider.lower())
|
||||
return mapped.value if mapped is not None else custom_llm_provider
|
||||
|
||||
|
||||
def resolve_operation(call_type: str | None) -> GenAIOperation:
|
||||
"""Map a litellm ``call_type`` to a ``gen_ai.operation.name`` value."""
|
||||
if not call_type:
|
||||
return GenAIOperation.CHAT
|
||||
return _OPERATION_BY_CALL_TYPE.get(call_type.lower(), GenAIOperation.CHAT)
|
||||
116
litellm/integrations/otel/spans.py
Normal file
116
litellm/integrations/otel/spans.py
Normal file
|
|
@ -0,0 +1,116 @@
|
|||
"""
|
||||
This module declares every span the instrumentation can emit and the hierarchy.
|
||||
|
||||
Span-name patterns live here as typed builder functions.
|
||||
|
||||
Canonical hierarchy::
|
||||
|
||||
PROXY_REQUEST (SERVER, root) # owned by the FastAPI instrumentor
|
||||
├── LLM_CALL (CLIENT)
|
||||
├── GUARDRAIL (INTERNAL) # request-lifecycle hook, sibling of LLM_CALL
|
||||
└── SERVICE (INTERNAL)
|
||||
|
||||
Guardrails parent to PROXY_REQUEST, not LLM_CALL: pre/during/post-call guardrail
|
||||
hooks are orchestrated by the request lifecycle (a pre-call guardrail runs
|
||||
before the LLM call even starts), so a guardrail is a sibling of the LLM call,
|
||||
not a child of it. The emitter parents every span to the ambient OTel context
|
||||
(the active server span), which matches this.
|
||||
|
||||
Management/admin endpoints are ordinary FastAPI routes — their SERVER spans are
|
||||
owned by the instrumentor too, so they don't appear as a role here.
|
||||
"""
|
||||
|
||||
from dataclasses import dataclass
|
||||
from enum import Enum
|
||||
from typing import TYPE_CHECKING
|
||||
|
||||
if TYPE_CHECKING:
|
||||
from litellm.integrations.otel.payloads import (
|
||||
GuardrailSpanData,
|
||||
LLMCallSpanData,
|
||||
ProxyRequestSpanData,
|
||||
ServiceSpanData,
|
||||
)
|
||||
|
||||
|
||||
class SpanRole(str, Enum):
|
||||
PROXY_REQUEST = "proxy_request"
|
||||
LLM_CALL = "llm_call"
|
||||
GUARDRAIL = "guardrail"
|
||||
SERVICE = "service"
|
||||
|
||||
|
||||
class LiteLLMSpanKind(str, Enum):
|
||||
SERVER = "server"
|
||||
CLIENT = "client"
|
||||
INTERNAL = "internal"
|
||||
PRODUCER = "producer"
|
||||
CONSUMER = "consumer"
|
||||
|
||||
|
||||
@dataclass(frozen=True)
|
||||
class SpanSpec:
|
||||
role: SpanRole
|
||||
kind: LiteLLMSpanKind
|
||||
parent: SpanRole | None
|
||||
|
||||
|
||||
SPAN_REGISTRY: dict[SpanRole, SpanSpec] = {
|
||||
SpanRole.PROXY_REQUEST: SpanSpec(
|
||||
SpanRole.PROXY_REQUEST, LiteLLMSpanKind.SERVER, parent=None
|
||||
),
|
||||
SpanRole.LLM_CALL: SpanSpec(
|
||||
SpanRole.LLM_CALL, LiteLLMSpanKind.CLIENT, parent=SpanRole.PROXY_REQUEST
|
||||
),
|
||||
SpanRole.GUARDRAIL: SpanSpec(
|
||||
SpanRole.GUARDRAIL, LiteLLMSpanKind.INTERNAL, parent=SpanRole.PROXY_REQUEST
|
||||
),
|
||||
SpanRole.SERVICE: SpanSpec(
|
||||
SpanRole.SERVICE, LiteLLMSpanKind.INTERNAL, parent=SpanRole.PROXY_REQUEST
|
||||
),
|
||||
}
|
||||
|
||||
|
||||
# --- span name builders (the naming convention, per role) ------------------- #
|
||||
|
||||
|
||||
def llm_call_span_name(data: "LLMCallSpanData") -> str:
|
||||
"""``"{operation} {model}"`` e.g. ``"chat gpt-4o"`` (GenAI semconv)."""
|
||||
model = data.request_model or ""
|
||||
return f"{data.operation.value} {model}".strip()
|
||||
|
||||
|
||||
def proxy_request_span_name(data: "ProxyRequestSpanData") -> str:
|
||||
"""``"{method} {route}"`` (HTTP semconv)."""
|
||||
return f"{data.http_method} {data.route}".strip()
|
||||
|
||||
|
||||
def guardrail_span_name(data: "GuardrailSpanData") -> str:
|
||||
return f"execute_guardrail {data.guardrail_name}".strip()
|
||||
|
||||
|
||||
def service_span_name(data: "ServiceSpanData") -> str:
|
||||
return data.service_name
|
||||
|
||||
|
||||
def root_roles() -> list[SpanRole]:
|
||||
"""Roles that start a new trace (no in-process parent)."""
|
||||
return [role for role, spec in SPAN_REGISTRY.items() if spec.parent is None]
|
||||
|
||||
|
||||
def child_roles(parent: SpanRole) -> list[SpanRole]:
|
||||
return [role for role, spec in SPAN_REGISTRY.items() if spec.parent == parent]
|
||||
|
||||
|
||||
def validate_registry(
|
||||
registry: dict[SpanRole, SpanSpec] | None = None,
|
||||
) -> None:
|
||||
reg = registry if registry is not None else SPAN_REGISTRY
|
||||
for role, spec in reg.items():
|
||||
if spec.role is not role:
|
||||
raise ValueError(f"SPAN_REGISTRY[{role}] has mismatched role {spec.role}")
|
||||
if spec.parent is not None and spec.parent not in reg:
|
||||
raise ValueError(f"span role {role} declares unknown parent {spec.parent}")
|
||||
missing = [role for role in SpanRole if role not in reg]
|
||||
if missing:
|
||||
raise ValueError(f"SPAN_REGISTRY is missing roles: {missing}")
|
||||
103
litellm/integrations/otel/utils.py
Normal file
103
litellm/integrations/otel/utils.py
Normal file
|
|
@ -0,0 +1,103 @@
|
|||
"""Shared, OpenTelemetry-free helpers for the otel integration.
|
||||
|
||||
Generic value coercion (for reading heterogeneous logging-payload dicts), time
|
||||
conversion, and header parsing — pulled out of the individual modules so they
|
||||
live in one place. Deliberately free of any ``opentelemetry`` import so the
|
||||
OTel-free sources of truth (payloads, semconv, spans, config) can use it too.
|
||||
"""
|
||||
|
||||
from datetime import datetime
|
||||
|
||||
|
||||
def as_str(value: object) -> str | None:
|
||||
if value is None:
|
||||
return None
|
||||
if isinstance(value, str):
|
||||
return value
|
||||
return str(value)
|
||||
|
||||
|
||||
def as_int(value: object) -> int | None:
|
||||
if isinstance(value, bool):
|
||||
return int(value)
|
||||
if isinstance(value, int):
|
||||
return value
|
||||
if isinstance(value, float):
|
||||
return int(value)
|
||||
if isinstance(value, str):
|
||||
try:
|
||||
return int(value)
|
||||
except ValueError:
|
||||
return None
|
||||
return None
|
||||
|
||||
|
||||
def as_float(value: object) -> float | None:
|
||||
if isinstance(value, bool):
|
||||
return float(value)
|
||||
if isinstance(value, (int, float)):
|
||||
return float(value)
|
||||
if isinstance(value, str):
|
||||
try:
|
||||
return float(value)
|
||||
except ValueError:
|
||||
return None
|
||||
return None
|
||||
|
||||
|
||||
def as_bool(value: object) -> bool | None:
|
||||
if value is None:
|
||||
return None
|
||||
if isinstance(value, bool):
|
||||
return value
|
||||
return bool(value)
|
||||
|
||||
|
||||
def as_str_tuple(value: object) -> tuple[str, ...] | None:
|
||||
if value is None:
|
||||
return None
|
||||
if isinstance(value, str):
|
||||
return (value,)
|
||||
if isinstance(value, (list, tuple)):
|
||||
return tuple(str(v) for v in value)
|
||||
return None
|
||||
|
||||
|
||||
def to_ns(value: datetime | float | int | None) -> int | None:
|
||||
"""Coerce a datetime / epoch value to integer nanoseconds."""
|
||||
if value is None:
|
||||
return None
|
||||
if isinstance(value, datetime):
|
||||
return int(value.timestamp() * 1e9)
|
||||
if isinstance(value, (int, float)) and not isinstance(value, bool):
|
||||
return int(float(value) * 1e9)
|
||||
return None
|
||||
|
||||
|
||||
def to_seconds(value: datetime | float | int | str | None) -> float | None:
|
||||
"""Coerce a datetime / epoch / formatted-string value to epoch seconds."""
|
||||
if value is None:
|
||||
return None
|
||||
if isinstance(value, datetime):
|
||||
return value.timestamp()
|
||||
if isinstance(value, (int, float)) and not isinstance(value, bool):
|
||||
return float(value)
|
||||
if isinstance(value, str):
|
||||
for fmt in ("%Y-%m-%d %H:%M:%S.%f", "%Y-%m-%d %H:%M:%S"):
|
||||
try:
|
||||
return datetime.strptime(value, fmt).timestamp()
|
||||
except ValueError:
|
||||
continue
|
||||
return None
|
||||
|
||||
|
||||
def parse_headers(raw: str | None) -> dict[str, str]:
|
||||
"""Parse an OTLP ``"k=v,k=v"`` header string into a dict."""
|
||||
headers: dict[str, str] = {}
|
||||
if not raw:
|
||||
return headers
|
||||
for pair in raw.split(","):
|
||||
if "=" in pair:
|
||||
key, _, value = pair.partition("=")
|
||||
headers[key.strip()] = value.strip()
|
||||
return headers
|
||||
|
|
@ -3718,6 +3718,9 @@ def _init_custom_logger_compatible_class( # noqa: PLR0915
|
|||
try:
|
||||
custom_logger_init_args = custom_logger_init_args or {}
|
||||
if logging_integration == "agentops": # Add AgentOps initialization
|
||||
_v2 = _maybe_construct_otel_v2("agentops", _in_memory_loggers)
|
||||
if _v2 is not None:
|
||||
return _v2 # type: ignore
|
||||
for callback in _in_memory_loggers:
|
||||
if isinstance(callback, AgentOps):
|
||||
return callback # type: ignore
|
||||
|
|
@ -3870,6 +3873,9 @@ def _init_custom_logger_compatible_class( # noqa: PLR0915
|
|||
_in_memory_loggers.append(_opik_logger)
|
||||
return _opik_logger # type: ignore
|
||||
elif logging_integration == "arize":
|
||||
_v2 = _maybe_construct_otel_v2("arize", _in_memory_loggers)
|
||||
if _v2 is not None:
|
||||
return _v2 # type: ignore
|
||||
from litellm.integrations.opentelemetry import (
|
||||
OpenTelemetry,
|
||||
OpenTelemetryConfig,
|
||||
|
|
@ -3899,6 +3905,9 @@ def _init_custom_logger_compatible_class( # noqa: PLR0915
|
|||
_in_memory_loggers.append(_arize_otel_logger)
|
||||
return _arize_otel_logger # type: ignore
|
||||
elif logging_integration == "arize_phoenix":
|
||||
_v2 = _maybe_construct_otel_v2("arize_phoenix", _in_memory_loggers)
|
||||
if _v2 is not None:
|
||||
return _v2 # type: ignore
|
||||
from litellm.integrations.opentelemetry import (
|
||||
OpenTelemetry,
|
||||
OpenTelemetryConfig,
|
||||
|
|
@ -3929,6 +3938,9 @@ def _init_custom_logger_compatible_class( # noqa: PLR0915
|
|||
_in_memory_loggers.append(_arize_phoenix_otel_logger)
|
||||
return _arize_phoenix_otel_logger # type: ignore
|
||||
elif logging_integration == "levo":
|
||||
_v2 = _maybe_construct_otel_v2("levo", _in_memory_loggers)
|
||||
if _v2 is not None:
|
||||
return _v2 # type: ignore
|
||||
from litellm.integrations.levo.levo import LevoLogger
|
||||
from litellm.integrations.opentelemetry import (
|
||||
OpenTelemetry,
|
||||
|
|
@ -3954,6 +3966,28 @@ def _init_custom_logger_compatible_class( # noqa: PLR0915
|
|||
_in_memory_loggers.append(_levo_otel_logger)
|
||||
return _levo_otel_logger # type: ignore
|
||||
elif logging_integration == "otel":
|
||||
# Gate the new typed V2 adapter behind LITELLM_OTEL_V2. When off,
|
||||
# the legacy 3,227-line god-class is used unchanged. The two are
|
||||
# never registered simultaneously — the dedup loop below treats
|
||||
# any module under ``litellm.integrations.otel`` or
|
||||
# ``litellm.integrations.opentelemetry`` as "the OTel callback".
|
||||
from litellm.integrations.otel.config import is_otel_v2_enabled
|
||||
|
||||
if is_otel_v2_enabled():
|
||||
from litellm.integrations.otel.logger import OpenTelemetryV2
|
||||
|
||||
for callback in _in_memory_loggers:
|
||||
if type(callback) is OpenTelemetryV2:
|
||||
return callback # type: ignore
|
||||
otel_logger_v2 = OpenTelemetryV2(
|
||||
**_get_custom_logger_settings_from_proxy_server(
|
||||
callback_name=logging_integration
|
||||
)
|
||||
)
|
||||
_in_memory_loggers.append(otel_logger_v2)
|
||||
_maybe_auto_initialize_arize_phoenix(_in_memory_loggers)
|
||||
return otel_logger_v2 # type: ignore
|
||||
|
||||
from litellm.integrations.opentelemetry import OpenTelemetry
|
||||
|
||||
for callback in _in_memory_loggers:
|
||||
|
|
@ -4092,6 +4126,9 @@ def _init_custom_logger_compatible_class( # noqa: PLR0915
|
|||
elif logging_integration == "langtrace":
|
||||
if "LANGTRACE_API_KEY" not in os.environ:
|
||||
raise ValueError("LANGTRACE_API_KEY not found in environment variables")
|
||||
_v2 = _maybe_construct_otel_v2("langtrace", _in_memory_loggers)
|
||||
if _v2 is not None:
|
||||
return _v2 # type: ignore
|
||||
|
||||
from litellm.integrations.opentelemetry import (
|
||||
OpenTelemetry,
|
||||
|
|
@ -4132,6 +4169,9 @@ def _init_custom_logger_compatible_class( # noqa: PLR0915
|
|||
_in_memory_loggers.append(langfuse_logger)
|
||||
return langfuse_logger # type: ignore
|
||||
elif logging_integration == "langfuse_otel":
|
||||
_v2 = _maybe_construct_otel_v2("langfuse_otel", _in_memory_loggers)
|
||||
if _v2 is not None:
|
||||
return _v2 # type: ignore
|
||||
from litellm.integrations.langfuse.langfuse_otel import LangfuseOtelLogger
|
||||
|
||||
for callback in _in_memory_loggers:
|
||||
|
|
@ -4148,6 +4188,9 @@ def _init_custom_logger_compatible_class( # noqa: PLR0915
|
|||
_in_memory_loggers.append(_otel_logger)
|
||||
return _otel_logger # type: ignore
|
||||
elif logging_integration == "weave_otel":
|
||||
_v2 = _maybe_construct_otel_v2("weave_otel", _in_memory_loggers)
|
||||
if _v2 is not None:
|
||||
return _v2 # type: ignore
|
||||
from litellm.integrations.opentelemetry import OpenTelemetryConfig
|
||||
from litellm.integrations.weave.weave_otel import (
|
||||
WeaveOtelLogger,
|
||||
|
|
@ -4296,6 +4339,42 @@ def _init_custom_logger_compatible_class( # noqa: PLR0915
|
|||
return None
|
||||
|
||||
|
||||
def _maybe_construct_otel_v2(
|
||||
callback_name: str, _in_memory_loggers: list
|
||||
) -> Optional[Any]:
|
||||
"""If ``LITELLM_OTEL_V2`` is on, build (or reuse) a single ``OpenTelemetryV2``
|
||||
instance configured via the preset for ``callback_name``.
|
||||
|
||||
Returns ``None`` when V2 is off OR when there's no preset registered for
|
||||
``callback_name`` — callers should then fall through to the legacy path.
|
||||
"""
|
||||
from litellm.integrations.otel.config import is_otel_v2_enabled
|
||||
|
||||
if not is_otel_v2_enabled():
|
||||
return None
|
||||
from litellm.integrations.otel.logger import OpenTelemetryV2
|
||||
from litellm.integrations.otel.presets import PRESET_BY_CALLBACK
|
||||
|
||||
preset_fn = PRESET_BY_CALLBACK.get(callback_name)
|
||||
if preset_fn is None:
|
||||
return None
|
||||
for callback in _in_memory_loggers:
|
||||
if (
|
||||
isinstance(callback, OpenTelemetryV2)
|
||||
and getattr(callback, "callback_name", None) == callback_name
|
||||
):
|
||||
return callback
|
||||
try:
|
||||
config = preset_fn()
|
||||
except Exception:
|
||||
# If env vars are missing or the preset raises, defer to the legacy path
|
||||
# so customers get the same error story they had before V2 landed.
|
||||
return None
|
||||
v2_logger = OpenTelemetryV2(config=config, callback_name=callback_name)
|
||||
_in_memory_loggers.append(v2_logger)
|
||||
return v2_logger
|
||||
|
||||
|
||||
def _maybe_auto_initialize_arize_phoenix(_in_memory_loggers: list) -> None:
|
||||
"""
|
||||
Auto-initialize ArizePhoenixLogger when Phoenix env vars are detected.
|
||||
|
|
|
|||
|
|
@ -825,6 +825,37 @@ async def proxy_startup_event(app: FastAPI): # noqa: PLR0915
|
|||
if isinstance(worker_config, dict):
|
||||
await initialize(**worker_config)
|
||||
|
||||
## V2 OTEL: now that config (and therefore the callbacks) is loaded, publish
|
||||
## the chosen V2 logger's TracerProvider as the OTel global. The FastAPI
|
||||
## instrumentation mounted at app-creation binds to the global provider, so
|
||||
## this is what makes server spans and gen-ai spans share one provider and
|
||||
## land in the same trace. Prefer an already-registered preset logger
|
||||
## (arize, langfuse, …) so server spans export to that backend too; otherwise
|
||||
## build a generic one from OTEL_* envs. ``set_tracer_provider`` only takes
|
||||
## effect once, so the first configured logger wins.
|
||||
try:
|
||||
from litellm.integrations.otel.config import is_otel_v2_enabled
|
||||
|
||||
if is_otel_v2_enabled():
|
||||
from opentelemetry import trace as _otel_trace
|
||||
|
||||
from litellm.integrations.otel.logger import OpenTelemetryV2
|
||||
|
||||
_otel_v2_logger = (
|
||||
next(
|
||||
(
|
||||
cb
|
||||
for cb in litellm.service_callback
|
||||
if isinstance(cb, OpenTelemetryV2)
|
||||
),
|
||||
None,
|
||||
)
|
||||
or OpenTelemetryV2()
|
||||
)
|
||||
_otel_trace.set_tracer_provider(_otel_v2_logger._tracer_provider)
|
||||
except Exception as e:
|
||||
verbose_proxy_logger.debug("Skipping OTel V2 provider setup: %s", e)
|
||||
|
||||
# check if DATABASE_URL in environment - load from there
|
||||
if prisma_client is None:
|
||||
_db_url: Optional[str] = get_secret("DATABASE_URL", None) # type: ignore
|
||||
|
|
@ -1061,6 +1092,56 @@ def ensure_unique_openapi_operation_ids(
|
|||
return openapi_schema
|
||||
|
||||
|
||||
# Passthrough routes are catch-alls (e.g. "/openai/{endpoint:path}"), so the
|
||||
# default OTel server-span name "{method} {route}" collapses every upstream
|
||||
# endpoint into "POST /openai/{endpoint:path}". Rename those spans to the real
|
||||
# request path so each endpoint is distinguishable. Non-catch-all routes keep
|
||||
# their low-cardinality template name.
|
||||
_OTEL_V2_PASSTHROUGH_PREFIXES = frozenset(
|
||||
{
|
||||
"openai",
|
||||
"openai_passthrough",
|
||||
"anthropic",
|
||||
"azure",
|
||||
"azure_ai",
|
||||
"bedrock",
|
||||
"cohere",
|
||||
"cursor",
|
||||
"gemini",
|
||||
"mistral",
|
||||
"vllm",
|
||||
"vertex_ai",
|
||||
"vertex-ai",
|
||||
"assemblyai",
|
||||
"eu.assemblyai",
|
||||
"milvus",
|
||||
}
|
||||
)
|
||||
|
||||
|
||||
def _otel_v2_passthrough_span_name_hook(span: Any, scope: dict) -> None:
|
||||
"""FastAPI ``server_request_hook``: give passthrough server spans a useful name.
|
||||
|
||||
The instrumentation matches the route at span creation, so both the span name
|
||||
and ``http.route`` are set to the catch-all template (``/openai/{endpoint:path}``)
|
||||
before this hook runs. Rewrite both to the real request path so each upstream
|
||||
endpoint is distinguishable. (The ASGI ``http receive``/``http send`` sub-spans
|
||||
can't be renamed from here — their name is captured at creation — so they are
|
||||
dropped via ``exclude_spans`` at instrumentation time.)
|
||||
"""
|
||||
try:
|
||||
if span is None or not span.is_recording():
|
||||
return
|
||||
path = scope.get("path") or ""
|
||||
method = scope.get("method") or ""
|
||||
first_segment = path.lstrip("/").split("/", 1)[0]
|
||||
if first_segment in _OTEL_V2_PASSTHROUGH_PREFIXES:
|
||||
span.update_name(f"{method} {path}".strip())
|
||||
span.set_attribute("http.route", path)
|
||||
except Exception:
|
||||
pass
|
||||
|
||||
|
||||
app = FastAPI(
|
||||
docs_url=_get_docs_url(),
|
||||
redoc_url=_get_redoc_url(),
|
||||
|
|
@ -1074,6 +1155,41 @@ app = FastAPI(
|
|||
strict_content_type=False,
|
||||
)
|
||||
|
||||
## V2 OTEL: instrument the FastAPI app for server spans (gated by
|
||||
## LITELLM_OTEL_V2; lazy imports keep the package optional). This MUST run at
|
||||
## app-creation time — once the lifespan runs, the middleware stack is frozen
|
||||
## and ``instrument_app`` raises "Cannot add middleware after an application has
|
||||
## started". No TracerProvider is passed, so the instrumentation binds to the
|
||||
## OTel global ``ProxyTracerProvider``; ``proxy_startup_event`` sets the real
|
||||
## provider as the global after config load, and the proxy delegates to it. That
|
||||
## way server spans and gen-ai spans share one provider and the same trace.
|
||||
try:
|
||||
from litellm.integrations.otel.config import is_otel_v2_enabled
|
||||
|
||||
if is_otel_v2_enabled():
|
||||
from opentelemetry.instrumentation.fastapi import FastAPIInstrumentor
|
||||
|
||||
# Drop health-check spans by default — load balancers poll
|
||||
# /health/readiness and /health/liveness constantly, which floods traces
|
||||
# with noise. Honor the standard OTel env var so operators can override
|
||||
# (e.g. set it to "" to trace everything, or add their own paths).
|
||||
_otel_excluded_urls = (
|
||||
os.environ.get("OTEL_PYTHON_FASTAPI_EXCLUDED_URLS")
|
||||
if "OTEL_PYTHON_FASTAPI_EXCLUDED_URLS" in os.environ
|
||||
else "/health"
|
||||
)
|
||||
FastAPIInstrumentor.instrument_app(
|
||||
app,
|
||||
excluded_urls=_otel_excluded_urls,
|
||||
server_request_hook=_otel_v2_passthrough_span_name_hook,
|
||||
# Drop the ASGI "http receive"/"http send" lifecycle sub-spans: they
|
||||
# are low-value noise and (for passthrough) carry the catch-all route
|
||||
# template in their name, which can't be rewritten from a hook.
|
||||
exclude_spans=["receive", "send"],
|
||||
)
|
||||
except Exception as e:
|
||||
verbose_proxy_logger.debug("Skipping OTel V2 FastAPI instrumentation: %s", e)
|
||||
|
||||
vertex_live_passthrough_vertex_base = VertexBase()
|
||||
|
||||
|
||||
|
|
@ -1238,6 +1354,17 @@ def _close_dangling_otel_server_span(request: Request, status_code: int) -> None
|
|||
return
|
||||
if open_telemetry_logger is None:
|
||||
return
|
||||
# Under OTel V2 the FastAPI instrumentor owns the server span (parent_otel_span
|
||||
# is that same span), and it records the error + ends it itself. Ending it here
|
||||
# would end it early — losing the http.* attributes the instrumentor stamps on
|
||||
# completion — and double-end it. Leave it to the instrumentor.
|
||||
try:
|
||||
from litellm.integrations.otel.config import is_otel_v2_enabled
|
||||
|
||||
if is_otel_v2_enabled():
|
||||
return
|
||||
except Exception:
|
||||
pass
|
||||
try:
|
||||
from opentelemetry.trace import Status, StatusCode
|
||||
|
||||
|
|
|
|||
|
|
@ -119,6 +119,7 @@ proxy-runtime = [
|
|||
"opentelemetry-api==1.28.0",
|
||||
"opentelemetry-sdk==1.28.0",
|
||||
"opentelemetry-exporter-otlp==1.28.0",
|
||||
"opentelemetry-instrumentation-fastapi==0.49b0",
|
||||
"ddtrace>=2.19.0,<3.0",
|
||||
"sentry-sdk>=2.21.0,<3.0",
|
||||
"mangum>=0.17.0,<1.0",
|
||||
|
|
@ -160,6 +161,7 @@ dev = [
|
|||
"opentelemetry-api==1.28.0",
|
||||
"opentelemetry-sdk==1.28.0",
|
||||
"opentelemetry-exporter-otlp==1.28.0",
|
||||
"opentelemetry-instrumentation-fastapi==0.49b0",
|
||||
"langfuse==2.59.7",
|
||||
"fastapi-offline==1.7.6",
|
||||
"fakeredis==2.34.1",
|
||||
|
|
@ -178,6 +180,7 @@ proxy-dev = [
|
|||
"opentelemetry-api==1.28.0",
|
||||
"opentelemetry-sdk==1.28.0",
|
||||
"opentelemetry-exporter-otlp==1.28.0",
|
||||
"opentelemetry-instrumentation-fastapi==0.49b0",
|
||||
"azure-identity==1.25.2",
|
||||
"a2a-sdk==0.3.24",
|
||||
]
|
||||
|
|
|
|||
134
tests/test_litellm/integrations/otel/test_otel_v2_baggage.py
Normal file
134
tests/test_litellm/integrations/otel/test_otel_v2_baggage.py
Normal file
|
|
@ -0,0 +1,134 @@
|
|||
"""Tests for Baggage-based promotion of request-scoped attributes onto every span,
|
||||
and the two antipattern boundaries: http.* is never promoted, and the full
|
||||
metadata blob is never promoted (only the bounded allowlist)."""
|
||||
|
||||
import pytest
|
||||
|
||||
pytest.importorskip("opentelemetry")
|
||||
|
||||
from litellm.integrations.otel import ( # noqa: E402
|
||||
GenAI,
|
||||
HTTP,
|
||||
LiteLLM,
|
||||
OpenTelemetryV2Config,
|
||||
promoted_baggage,
|
||||
)
|
||||
from litellm.integrations.otel import context as ctx_mod # noqa: E402
|
||||
from litellm.integrations.otel import providers # noqa: E402
|
||||
from litellm.integrations.otel.emitter import SpanEmitter # noqa: E402
|
||||
from litellm.integrations.otel.payloads import ( # noqa: E402
|
||||
GuardrailSpanData,
|
||||
LLMCallSpanData,
|
||||
ServiceSpanData,
|
||||
)
|
||||
from litellm.integrations.otel.baggage import BAGGAGE_PROMOTED_KEYS # noqa: E402
|
||||
from litellm.integrations.otel.spans import SpanRole # noqa: E402
|
||||
|
||||
|
||||
def _payload():
|
||||
return {
|
||||
"call_type": "acompletion",
|
||||
"custom_llm_provider": "openai",
|
||||
"model": "gpt-4o",
|
||||
"prompt_tokens": 1,
|
||||
"completion_tokens": 1,
|
||||
"total_tokens": 2,
|
||||
"metadata": {
|
||||
"team_id": "t1",
|
||||
"team_alias": "team one",
|
||||
"user_api_key_hash": "hsh",
|
||||
"user_api_key_org_id": "org1",
|
||||
"private_note": "do-not-promote",
|
||||
},
|
||||
"status": "success",
|
||||
"litellm_call_id": "call_1",
|
||||
"hidden_params": {},
|
||||
}
|
||||
|
||||
|
||||
def _engine_and_exporter(config=None):
|
||||
cfg = config or OpenTelemetryV2Config(exporter="in_memory")
|
||||
provider, exporter = providers.in_memory_provider(cfg)
|
||||
tracer = providers.get_tracer(provider, "litellm-baggage-test")
|
||||
return SpanEmitter(tracer, cfg), exporter
|
||||
|
||||
|
||||
def test_identity_promoted_onto_every_span():
|
||||
engine, exporter = _engine_and_exporter()
|
||||
data = LLMCallSpanData.from_standard_logging_payload(_payload())
|
||||
bag = promoted_baggage(data.identity, data.request_model, BAGGAGE_PROMOTED_KEYS)
|
||||
ctx = ctx_mod.set_request_baggage(bag)
|
||||
|
||||
root = engine.start_span(SpanRole.PROXY_REQUEST, "POST /chat/completions", ctx)
|
||||
root_ctx = ctx_mod.context_from_span(root, ctx)
|
||||
engine.emit(SpanRole.LLM_CALL, data, parent_context=root_ctx)
|
||||
engine.emit(
|
||||
SpanRole.GUARDRAIL, GuardrailSpanData("presidio", status="success"), root_ctx
|
||||
)
|
||||
engine.emit(SpanRole.SERVICE, ServiceSpanData("redis", call_type="set"), root_ctx)
|
||||
root.end()
|
||||
|
||||
spans = exporter.get_finished_spans()
|
||||
assert len(spans) == 4
|
||||
for span in spans:
|
||||
assert span.attributes.get(LiteLLM.TEAM_ID) == "t1"
|
||||
assert span.attributes.get(LiteLLM.TEAM_ALIAS) == "team one"
|
||||
assert span.attributes.get(GenAI.REQUEST_MODEL) == "gpt-4o"
|
||||
|
||||
|
||||
def test_allowlisted_metadata_subkey_promoted_blob_excluded():
|
||||
engine, exporter = _engine_and_exporter()
|
||||
data = LLMCallSpanData.from_standard_logging_payload(_payload())
|
||||
bag = promoted_baggage(data.identity, data.request_model, BAGGAGE_PROMOTED_KEYS)
|
||||
ctx = ctx_mod.set_request_baggage(bag)
|
||||
engine.emit(SpanRole.SERVICE, ServiceSpanData("redis", call_type="set"), ctx)
|
||||
(span,) = exporter.get_finished_spans()
|
||||
# allowlisted metadata sub-key is promoted
|
||||
assert (
|
||||
span.attributes.get(f"{LiteLLM.METADATA_PREFIX}user_api_key_org_id") == "org1"
|
||||
)
|
||||
# non-allowlisted metadata is NOT promoted (no full-blob dumping)
|
||||
assert all("private_note" not in k for k in span.attributes)
|
||||
|
||||
|
||||
def test_http_attributes_never_promoted():
|
||||
"""Even if http.* is present in baggage, the processor must not stamp it on
|
||||
child spans (it belongs on the SERVER span only)."""
|
||||
engine, exporter = _engine_and_exporter()
|
||||
ctx = ctx_mod.set_request_baggage(
|
||||
{
|
||||
LiteLLM.TEAM_ID: "t1",
|
||||
HTTP.ROUTE: "/chat/completions",
|
||||
HTTP.REQUEST_METHOD: "POST",
|
||||
}
|
||||
)
|
||||
engine.emit(SpanRole.SERVICE, ServiceSpanData("redis", call_type="set"), ctx)
|
||||
(span,) = exporter.get_finished_spans()
|
||||
assert span.attributes.get(LiteLLM.TEAM_ID) == "t1"
|
||||
assert HTTP.ROUTE not in span.attributes
|
||||
assert HTTP.REQUEST_METHOD not in span.attributes
|
||||
|
||||
|
||||
def test_arbitrary_upstream_baggage_not_promoted():
|
||||
engine, exporter = _engine_and_exporter()
|
||||
ctx = ctx_mod.set_request_baggage(
|
||||
{LiteLLM.TEAM_ID: "t1", "some.upstream.key": "leak"}
|
||||
)
|
||||
engine.emit(SpanRole.SERVICE, ServiceSpanData("redis", call_type="set"), ctx)
|
||||
(span,) = exporter.get_finished_spans()
|
||||
assert span.attributes.get(LiteLLM.TEAM_ID) == "t1"
|
||||
assert "some.upstream.key" not in span.attributes
|
||||
|
||||
|
||||
def test_baggage_processor_allowlist_can_be_widened():
|
||||
cfg = OpenTelemetryV2Config(
|
||||
exporter="in_memory",
|
||||
baggage_promoted_keys=[LiteLLM.TEAM_ID, "custom.key"],
|
||||
)
|
||||
engine, exporter = _engine_and_exporter(cfg)
|
||||
ctx = ctx_mod.set_request_baggage({"custom.key": "v", LiteLLM.TEAM_ALIAS: "ta"})
|
||||
engine.emit(SpanRole.SERVICE, ServiceSpanData("redis"), ctx)
|
||||
(span,) = exporter.get_finished_spans()
|
||||
assert span.attributes.get("custom.key") == "v"
|
||||
# team_alias not in this config's allowlist -> not promoted
|
||||
assert LiteLLM.TEAM_ALIAS not in span.attributes
|
||||
395
tests/test_litellm/integrations/otel/test_otel_v2_components.py
Normal file
395
tests/test_litellm/integrations/otel/test_otel_v2_components.py
Normal file
|
|
@ -0,0 +1,395 @@
|
|||
"""Coverage for the engine-layer components: providers/exporters, context +
|
||||
baggage helpers, metrics, the typed coercion helpers, mapper branches, span-name
|
||||
builders, and the registry validator's failure paths. Needs the OTel SDK."""
|
||||
|
||||
import pytest
|
||||
|
||||
pytest.importorskip("opentelemetry")
|
||||
|
||||
from opentelemetry.sdk.metrics import MeterProvider # noqa: E402
|
||||
from opentelemetry.sdk.metrics.export import InMemoryMetricReader # noqa: E402
|
||||
from opentelemetry.sdk.trace.export import ( # noqa: E402
|
||||
BatchSpanProcessor,
|
||||
ConsoleSpanExporter,
|
||||
SimpleSpanProcessor,
|
||||
)
|
||||
from opentelemetry.sdk.trace.export.in_memory_span_exporter import ( # noqa: E402
|
||||
InMemorySpanExporter,
|
||||
)
|
||||
from opentelemetry.trace import SpanKind # noqa: E402
|
||||
|
||||
from litellm.integrations.otel import context as ctx_mod # noqa: E402
|
||||
from litellm.integrations.otel import providers # noqa: E402
|
||||
from litellm.integrations.otel.config import OpenTelemetryV2Config # noqa: E402
|
||||
from litellm.integrations.otel.mappers.genai import GenAIMapper # noqa: E402
|
||||
from litellm.integrations.otel.mappers.legacy import LegacyMapper # noqa: E402
|
||||
from litellm.integrations.otel.metrics import create_genai_metrics # noqa: E402
|
||||
from litellm.integrations.otel.payloads import ( # noqa: E402
|
||||
GuardrailSpanData,
|
||||
LLMCallSpanData,
|
||||
LLMRequestParams,
|
||||
LLMUsage,
|
||||
ProxyRequestSpanData,
|
||||
RequestIdentity,
|
||||
ServerInfo,
|
||||
ServiceSpanData,
|
||||
SpanError,
|
||||
)
|
||||
from litellm.integrations.otel.semconv import GenAI, GenAIOperation
|
||||
from litellm.integrations.otel.spans import ( # noqa: E402
|
||||
SPAN_REGISTRY,
|
||||
LiteLLMSpanKind,
|
||||
SpanRole,
|
||||
SpanSpec,
|
||||
guardrail_span_name,
|
||||
proxy_request_span_name,
|
||||
service_span_name,
|
||||
validate_registry,
|
||||
)
|
||||
from litellm.integrations.otel.utils import ( # noqa: E402
|
||||
as_bool,
|
||||
as_float,
|
||||
as_int,
|
||||
as_str,
|
||||
as_str_tuple,
|
||||
)
|
||||
|
||||
# --- typed coercion helpers ------------------------------------------------- #
|
||||
|
||||
|
||||
def test_as_str():
|
||||
assert as_str(None) is None
|
||||
assert as_str("x") == "x"
|
||||
assert as_str(5) == "5"
|
||||
|
||||
|
||||
def test_as_int():
|
||||
assert as_int(True) == 1
|
||||
assert as_int(3) == 3
|
||||
assert as_int(3.9) == 3
|
||||
assert as_int("7") == 7
|
||||
assert as_int("nope") is None
|
||||
assert as_int(None) is None
|
||||
|
||||
|
||||
def test_as_float():
|
||||
assert as_float(True) == 1.0
|
||||
assert as_float(2) == 2.0
|
||||
assert as_float("1.5") == 1.5
|
||||
assert as_float("nope") is None
|
||||
assert as_float(None) is None
|
||||
|
||||
|
||||
def test_as_bool():
|
||||
assert as_bool(None) is None
|
||||
assert as_bool(True) is True
|
||||
assert as_bool(1) is True
|
||||
assert as_bool(0) is False
|
||||
|
||||
|
||||
def test_as_str_tuple():
|
||||
assert as_str_tuple(None) is None
|
||||
assert as_str_tuple("a") == ("a",)
|
||||
assert as_str_tuple(["a", 2]) == ("a", "2")
|
||||
assert as_str_tuple(123) is None
|
||||
|
||||
|
||||
def test_request_params_max_completion_tokens_fallback():
|
||||
params = LLMRequestParams.from_model_parameters({"max_completion_tokens": 99})
|
||||
assert params.max_tokens == 99
|
||||
|
||||
|
||||
def test_server_info_from_api_base():
|
||||
assert ServerInfo.from_api_base(None) is None
|
||||
assert ServerInfo.from_api_base("api.host.com:8080") == ServerInfo(
|
||||
"api.host.com", 8080
|
||||
)
|
||||
assert ServerInfo.from_api_base("https://h.com/v1") == ServerInfo("h.com", None)
|
||||
# scheme present but empty netloc -> no hostname
|
||||
assert ServerInfo.from_api_base("http:///v1") is None
|
||||
|
||||
|
||||
def test_service_span_data_from_payload():
|
||||
class _Service:
|
||||
value = "redis"
|
||||
|
||||
class _Payload:
|
||||
service = _Service()
|
||||
call_type = "async_set_cache"
|
||||
error = None
|
||||
|
||||
data = ServiceSpanData.from_payload(_Payload())
|
||||
assert data.service_name == "redis"
|
||||
assert data.call_type == "async_set_cache"
|
||||
assert data.error is None
|
||||
|
||||
class _FailPayload:
|
||||
service = _Service()
|
||||
call_type = "async_set_cache"
|
||||
error = "boom"
|
||||
|
||||
failed = ServiceSpanData.from_payload(_FailPayload())
|
||||
assert failed.error is not None
|
||||
assert failed.error.message == "boom"
|
||||
|
||||
|
||||
# --- span name builders ----------------------------------------------------- #
|
||||
|
||||
|
||||
def test_name_builders():
|
||||
assert (
|
||||
proxy_request_span_name(ProxyRequestSpanData("POST", "/chat/completions"))
|
||||
== "POST /chat/completions"
|
||||
)
|
||||
assert service_span_name(ServiceSpanData("redis")) == "redis"
|
||||
assert (
|
||||
guardrail_span_name(GuardrailSpanData("presidio"))
|
||||
== "execute_guardrail presidio"
|
||||
)
|
||||
|
||||
|
||||
# --- registry validator failure paths --------------------------------------- #
|
||||
|
||||
|
||||
def test_validate_registry_detects_role_mismatch():
|
||||
bad = {SpanRole.LLM_CALL: SpanSpec(SpanRole.SERVICE, LiteLLMSpanKind.CLIENT, None)}
|
||||
with pytest.raises(ValueError, match="mismatched role"):
|
||||
validate_registry(bad)
|
||||
|
||||
|
||||
def test_validate_registry_detects_unknown_parent():
|
||||
bad = {
|
||||
SpanRole.LLM_CALL: SpanSpec(
|
||||
SpanRole.LLM_CALL, LiteLLMSpanKind.CLIENT, parent=SpanRole.PROXY_REQUEST
|
||||
)
|
||||
}
|
||||
with pytest.raises(ValueError, match="unknown parent"):
|
||||
validate_registry(bad)
|
||||
|
||||
|
||||
def test_validate_registry_detects_missing_roles():
|
||||
partial = {
|
||||
SpanRole.PROXY_REQUEST: SPAN_REGISTRY[SpanRole.PROXY_REQUEST],
|
||||
}
|
||||
with pytest.raises(ValueError, match="missing roles"):
|
||||
validate_registry(partial)
|
||||
|
||||
|
||||
# --- mappers (full branch coverage) ----------------------------------------- #
|
||||
|
||||
|
||||
def _full_llm_call():
|
||||
return LLMCallSpanData(
|
||||
operation=GenAIOperation.CHAT,
|
||||
provider="openai",
|
||||
request_model="gpt-4o",
|
||||
response_model="gpt-4o-2024",
|
||||
response_id="resp_1",
|
||||
request_params=LLMRequestParams(
|
||||
temperature=0.7,
|
||||
top_p=0.9,
|
||||
top_k=40,
|
||||
max_tokens=256,
|
||||
frequency_penalty=0.1,
|
||||
presence_penalty=0.2,
|
||||
stop_sequences=("STOP",),
|
||||
seed=42,
|
||||
),
|
||||
usage=LLMUsage(input_tokens=10, output_tokens=5, total_tokens=15),
|
||||
finish_reasons=("stop",),
|
||||
error=None,
|
||||
response_cost=0.002,
|
||||
server=ServerInfo("api.openai.com", 443),
|
||||
identity=RequestIdentity(call_id="c1"),
|
||||
is_streaming=True,
|
||||
)
|
||||
|
||||
|
||||
def test_genai_mapper_all_request_params():
|
||||
attrs = GenAIMapper().map(_full_llm_call())
|
||||
assert attrs[GenAI.REQUEST_TOP_P] == 0.9
|
||||
assert attrs[GenAI.REQUEST_TOP_K] == 40
|
||||
assert attrs[GenAI.REQUEST_MAX_TOKENS] == 256
|
||||
assert attrs[GenAI.REQUEST_FREQUENCY_PENALTY] == 0.1
|
||||
assert attrs[GenAI.REQUEST_PRESENCE_PENALTY] == 0.2
|
||||
assert attrs[GenAI.REQUEST_STOP_SEQUENCES] == ["STOP"]
|
||||
assert attrs[GenAI.REQUEST_SEED] == 42
|
||||
assert attrs["server.port"] == 443
|
||||
|
||||
|
||||
def test_genai_mapper_guardrail_and_service():
|
||||
from litellm.integrations.otel.semconv import LiteLLM
|
||||
|
||||
g = GenAIMapper().map(GuardrailSpanData("presidio", mode="pre"))
|
||||
assert g[LiteLLM.GUARDRAIL_NAME] == "presidio"
|
||||
assert g[LiteLLM.GUARDRAIL_MODE] == "pre"
|
||||
|
||||
s = GenAIMapper().map(ServiceSpanData("redis", call_type="set"))
|
||||
assert s[LiteLLM.SERVICE_NAME] == "redis"
|
||||
assert s[LiteLLM.SERVICE_CALL_TYPE] == "set"
|
||||
|
||||
|
||||
def test_legacy_mapper_all_request_params():
|
||||
attrs = LegacyMapper().map(_full_llm_call())
|
||||
assert attrs["llm.top_k"] == 40
|
||||
assert attrs["llm.frequency_penalty"] == 0.1
|
||||
assert attrs["llm.presence_penalty"] == 0.2
|
||||
assert attrs["llm.chat.stop_sequences"] == ["STOP"]
|
||||
assert attrs["gen_ai.usage.total_tokens"] == 15
|
||||
|
||||
|
||||
def test_legacy_mapper_covers_service_with_v1_bare_keys():
|
||||
"""Service spans dual-emit V1's bare ``service``/``call_type``/``error`` keys."""
|
||||
attrs = LegacyMapper().map(
|
||||
ServiceSpanData("redis", call_type="set", event_metadata={"k": "v"}),
|
||||
)
|
||||
assert attrs["service"] == "redis"
|
||||
assert attrs["call_type"] == "set"
|
||||
assert attrs["k"] == "v" # event_metadata is stamped bare (V1 behavior)
|
||||
|
||||
|
||||
def test_legacy_mapper_skips_guardrail_role():
|
||||
"""Guardrail spans never had a V1 vocabulary; legacy mapper returns ``{}``."""
|
||||
assert LegacyMapper().map(GuardrailSpanData("presidio")) == {}
|
||||
|
||||
|
||||
# --- metrics ---------------------------------------------------------------- #
|
||||
|
||||
|
||||
def test_create_genai_metrics_records():
|
||||
reader = InMemoryMetricReader()
|
||||
meter = MeterProvider(metric_readers=[reader]).get_meter("test")
|
||||
metrics = create_genai_metrics(meter)
|
||||
metrics.token_usage.record(10, {"x": "y"})
|
||||
metrics.operation_duration.record(0.5, {"x": "y"})
|
||||
data = reader.get_metrics_data()
|
||||
assert data is not None
|
||||
|
||||
|
||||
# --- context + baggage helpers ---------------------------------------------- #
|
||||
|
||||
|
||||
def test_extract_traceparent():
|
||||
valid = {"traceparent": "00-0af7651916cd43dd8448eb211c80319c-b7ad6b7169203331-01"}
|
||||
assert ctx_mod.extract_traceparent(valid) is not None
|
||||
assert ctx_mod.extract_traceparent({"x": "y"}) is None
|
||||
|
||||
|
||||
def test_set_request_baggage_empty_returns_context():
|
||||
assert ctx_mod.set_request_baggage({}) is not None
|
||||
|
||||
|
||||
def test_get_baggage_attributes_roundtrip():
|
||||
ctx = ctx_mod.set_request_baggage({"litellm.team.id": "t1"})
|
||||
assert ctx_mod.get_baggage_attributes(ctx)["litellm.team.id"] == "t1"
|
||||
|
||||
|
||||
# --- providers -------------------------------------------------------------- #
|
||||
|
||||
|
||||
def test_to_otel_span_kind_covers_all():
|
||||
assert providers.to_otel_span_kind(LiteLLMSpanKind.SERVER) is SpanKind.SERVER
|
||||
assert providers.to_otel_span_kind(LiteLLMSpanKind.CLIENT) is SpanKind.CLIENT
|
||||
assert providers.to_otel_span_kind(LiteLLMSpanKind.INTERNAL) is SpanKind.INTERNAL
|
||||
assert providers.to_otel_span_kind(LiteLLMSpanKind.PRODUCER) is SpanKind.PRODUCER
|
||||
assert providers.to_otel_span_kind(LiteLLMSpanKind.CONSUMER) is SpanKind.CONSUMER
|
||||
|
||||
|
||||
def test_parse_headers():
|
||||
assert providers.parse_headers(None) == {}
|
||||
assert providers.parse_headers("a=1,b=2") == {"a": "1", "b": "2"}
|
||||
assert providers.parse_headers("no-equals") == {}
|
||||
|
||||
|
||||
def test_otlp_traces_endpoint_normalization():
|
||||
norm = providers._otlp_traces_endpoint
|
||||
# A base endpoint gets the signal path appended (the common OTLP env shape).
|
||||
assert norm("http://collector:4318") == "http://collector:4318/v1/traces"
|
||||
assert norm("http://collector:4318/") == "http://collector:4318/v1/traces"
|
||||
# An already-correct path is left intact.
|
||||
assert norm("http://collector:4318/v1/traces") == "http://collector:4318/v1/traces"
|
||||
# Another signal's path is rewritten to traces.
|
||||
assert norm("http://collector:4318/v1/logs") == "http://collector:4318/v1/traces"
|
||||
# Splunk's path is preserved; None passes through.
|
||||
assert (
|
||||
norm("https://x.splunk.com/v2/trace/otlp")
|
||||
== "https://x.splunk.com/v2/trace/otlp"
|
||||
)
|
||||
assert norm(None) is None
|
||||
|
||||
|
||||
def test_build_span_exporter_variants():
|
||||
assert isinstance(
|
||||
providers.build_span_exporter(OpenTelemetryV2Config(exporter="console")),
|
||||
ConsoleSpanExporter,
|
||||
)
|
||||
assert isinstance(
|
||||
providers.build_span_exporter(OpenTelemetryV2Config(exporter="in_memory")),
|
||||
InMemorySpanExporter,
|
||||
)
|
||||
assert isinstance(
|
||||
providers.build_span_exporter(OpenTelemetryV2Config(exporter="unknown")),
|
||||
ConsoleSpanExporter,
|
||||
)
|
||||
http_exporter = providers.build_span_exporter(
|
||||
OpenTelemetryV2Config(exporter="otlp_http", endpoint="http://h:4318")
|
||||
)
|
||||
assert "OTLPSpanExporter" in type(http_exporter).__name__
|
||||
grpc_exporter = providers.build_span_exporter(
|
||||
OpenTelemetryV2Config(exporter="otlp_grpc", endpoint="http://h:4317")
|
||||
)
|
||||
assert "OTLPSpanExporter" in type(grpc_exporter).__name__
|
||||
|
||||
|
||||
def test_build_resource_includes_deployment_environment():
|
||||
resource = providers.build_resource(
|
||||
OpenTelemetryV2Config(service_name="svc", deployment_environment="prod")
|
||||
)
|
||||
assert resource.attributes["service.name"] == "svc"
|
||||
assert resource.attributes["deployment.environment"] == "prod"
|
||||
|
||||
|
||||
def test_build_tracer_provider_processor_selection():
|
||||
cfg = OpenTelemetryV2Config(exporter="in_memory")
|
||||
simple = providers.build_tracer_provider(cfg, exporter=InMemorySpanExporter())
|
||||
batch = providers.build_tracer_provider(
|
||||
cfg, exporter=ConsoleSpanExporter(), use_simple_processor=False
|
||||
)
|
||||
# both build without error; assert the requested processor type was used
|
||||
simple_procs = simple._active_span_processor._span_processors
|
||||
batch_procs = batch._active_span_processor._span_processors
|
||||
assert any(isinstance(p, SimpleSpanProcessor) for p in simple_procs)
|
||||
assert any(isinstance(p, BatchSpanProcessor) for p in batch_procs)
|
||||
|
||||
|
||||
def test_baggage_processor_lifecycle_noops():
|
||||
proc = providers.LiteLLMBaggageSpanProcessor(allowed_keys=["litellm.team.id"])
|
||||
# no-op lifecycle hooks must not raise
|
||||
assert proc.on_end(None) is None # type: ignore[arg-type]
|
||||
assert proc.shutdown() is None
|
||||
assert proc.force_flush() is True
|
||||
|
||||
|
||||
def test_emitter_without_call_id_is_not_deduped():
|
||||
from litellm.integrations.otel.emitter import SpanEmitter
|
||||
|
||||
cfg = OpenTelemetryV2Config(exporter="in_memory")
|
||||
provider, exporter = providers.in_memory_provider(cfg)
|
||||
engine = SpanEmitter(providers.get_tracer(provider, "t"), cfg)
|
||||
data = LLMCallSpanData(
|
||||
operation=GenAIOperation.CHAT,
|
||||
provider="openai",
|
||||
request_model="gpt-4o",
|
||||
response_model=None,
|
||||
response_id=None,
|
||||
request_params=LLMRequestParams(),
|
||||
usage=LLMUsage(),
|
||||
finish_reasons=(),
|
||||
error=SpanError(error_type="X", message=None),
|
||||
response_cost=None,
|
||||
server=None,
|
||||
identity=RequestIdentity(call_id=None),
|
||||
)
|
||||
engine.emit(SpanRole.LLM_CALL, data)
|
||||
engine.emit(SpanRole.LLM_CALL, data) # no call_id -> not deduped
|
||||
assert len(exporter.get_finished_spans()) == 2
|
||||
131
tests/test_litellm/integrations/otel/test_otel_v2_dynamic.py
Normal file
131
tests/test_litellm/integrations/otel/test_otel_v2_dynamic.py
Normal file
|
|
@ -0,0 +1,131 @@
|
|||
"""Per-request multi-tenant credential routing (V1 parity)."""
|
||||
|
||||
import os
|
||||
import sys
|
||||
|
||||
sys.path.insert(0, os.path.abspath("../../../.."))
|
||||
|
||||
from opentelemetry.trace import NoOpTracer
|
||||
|
||||
from litellm.integrations.otel.config import ExporterSpec, OpenTelemetryV2Config
|
||||
from litellm.integrations.otel.presets import dynamic_otlp_headers
|
||||
from litellm.integrations.otel.routing import TenantTracerCache
|
||||
|
||||
|
||||
def _cache(callback_name, exporters=None):
|
||||
cfg = OpenTelemetryV2Config(exporters=exporters or [ExporterSpec(kind="in_memory")])
|
||||
return TenantTracerCache(cfg, callback_name, "litellm")
|
||||
|
||||
|
||||
# --- header builders mirror the V1 construct_dynamic_otel_headers overrides --- #
|
||||
|
||||
|
||||
def test_arize_dynamic_headers():
|
||||
headers = dynamic_otlp_headers(
|
||||
"arize", {"arize_space_id": "S", "arize_api_key": "K"}
|
||||
)
|
||||
assert headers == {"arize-space-id": "S", "api_key": "K"}
|
||||
|
||||
|
||||
def test_arize_space_key_overrides_space_id():
|
||||
headers = dynamic_otlp_headers(
|
||||
"arize", {"arize_space_id": "S", "arize_space_key": "SK"}
|
||||
)
|
||||
assert headers == {"arize-space-id": "SK"}
|
||||
|
||||
|
||||
def test_langfuse_dynamic_headers_need_both_keys():
|
||||
assert dynamic_otlp_headers("langfuse_otel", {"langfuse_public_key": "pk"}) is None
|
||||
headers = dynamic_otlp_headers(
|
||||
"langfuse_otel", {"langfuse_public_key": "pk", "langfuse_secret_key": "sk"}
|
||||
)
|
||||
assert headers is not None and "Authorization" in headers
|
||||
|
||||
|
||||
def test_weave_dynamic_headers():
|
||||
headers = dynamic_otlp_headers(
|
||||
"weave_otel", {"wandb_api_key": "w", "weave_project_id": "p"}
|
||||
)
|
||||
assert headers is not None
|
||||
assert "Authorization" in headers and headers["project_id"] == "p"
|
||||
|
||||
|
||||
def test_non_participating_callbacks_have_no_routing():
|
||||
# Phoenix subclasses the base in V1 (no override) → no dynamic routing.
|
||||
assert dynamic_otlp_headers("arize_phoenix", {"arize_api_key": "K"}) is None
|
||||
assert dynamic_otlp_headers("langtrace", {"arize_api_key": "K"}) is None
|
||||
assert dynamic_otlp_headers(None, {"arize_api_key": "K"}) is None
|
||||
|
||||
|
||||
def test_no_dynamic_params_is_no_routing():
|
||||
assert dynamic_otlp_headers("arize", None) is None
|
||||
assert dynamic_otlp_headers("arize", {}) is None
|
||||
|
||||
|
||||
# --- TenantTracerCache routes + caches a TracerProvider per credential set --- #
|
||||
|
||||
|
||||
def test_provider_cached_per_credential_set():
|
||||
cache = _cache("arize")
|
||||
default = NoOpTracer()
|
||||
creds_a = {"arize_space_id": "S", "arize_api_key": "K"}
|
||||
creds_b = {"arize_space_id": "S2", "arize_api_key": "K2"}
|
||||
|
||||
cache.tracer_for(default, creds_a)
|
||||
cache.tracer_for(default, creds_a) # same set → reuse, no new provider
|
||||
assert len(cache._providers) == 1
|
||||
cache.tracer_for(default, creds_b) # new set → new provider
|
||||
assert len(cache._providers) == 2
|
||||
|
||||
|
||||
def test_provider_cache_is_bounded_and_evicts_lru(monkeypatch):
|
||||
# The cache key derives from request-supplied dynamic credentials, so it
|
||||
# must be bounded — an unbounded cache lets a caller spawn one provider (and
|
||||
# its background exporter thread) per unique credential set. On overflow the
|
||||
# least-recently-used provider is evicted and shut down.
|
||||
from litellm.integrations.otel import routing as routing_mod
|
||||
|
||||
monkeypatch.setattr(routing_mod, "_MAX_CACHED_PROVIDERS", 2)
|
||||
shut_down = []
|
||||
monkeypatch.setattr(
|
||||
routing_mod, "_shutdown_provider", lambda p: shut_down.append(p)
|
||||
)
|
||||
|
||||
cache = _cache("arize")
|
||||
default = NoOpTracer()
|
||||
|
||||
def creds(space):
|
||||
return {"arize_space_id": space, "arize_api_key": "K"}
|
||||
|
||||
cache.tracer_for(default, creds("1"))
|
||||
cache.tracer_for(default, creds("2"))
|
||||
cache.tracer_for(default, creds("1")) # touch "1" → "2" is now LRU
|
||||
cache.tracer_for(default, creds("3")) # overflow → evict "2"
|
||||
|
||||
assert len(cache._providers) == 2
|
||||
assert len(shut_down) == 1 # exactly the evicted provider was shut down
|
||||
|
||||
|
||||
def test_no_dynamic_params_uses_default_tracer():
|
||||
cache = _cache("arize")
|
||||
default = NoOpTracer()
|
||||
assert cache.tracer_for(default, {}) is default
|
||||
assert cache._providers == {}
|
||||
|
||||
|
||||
def test_non_participating_callback_uses_default_tracer():
|
||||
cache = _cache("arize_phoenix")
|
||||
default = NoOpTracer()
|
||||
assert cache.tracer_for(default, {"arize_api_key": "K"}) is default
|
||||
assert cache._providers == {}
|
||||
|
||||
|
||||
def test_dynamic_headers_applied_to_otlp_exporter_only():
|
||||
cache = _cache(
|
||||
"arize",
|
||||
exporters=[ExporterSpec(kind="otlp_http"), ExporterSpec(kind="in_memory")],
|
||||
)
|
||||
new_cfg = cache._config_with_headers({"arize-space-id": "S", "api_key": "K"})
|
||||
otlp, in_mem = new_cfg.exporters
|
||||
assert otlp.headers == "arize-space-id=S,api_key=K"
|
||||
assert in_mem.headers is None # console/in_memory left untouched
|
||||
220
tests/test_litellm/integrations/otel/test_otel_v2_emitter.py
Normal file
220
tests/test_litellm/integrations/otel/test_otel_v2_emitter.py
Normal file
|
|
@ -0,0 +1,220 @@
|
|||
"""Golden tests for the OTel v2 engine: span shape, kinds, semconv attributes,
|
||||
legacy dual-emit, hierarchy, error status, and idempotency. Needs the OTel SDK."""
|
||||
|
||||
import pytest
|
||||
|
||||
pytest.importorskip("opentelemetry")
|
||||
|
||||
from opentelemetry.trace import SpanKind # noqa: E402
|
||||
from opentelemetry.trace.status import StatusCode # noqa: E402
|
||||
|
||||
from litellm.integrations.otel import ( # noqa: E402
|
||||
GenAI,
|
||||
LiteLLM,
|
||||
OpenTelemetryV2Config,
|
||||
)
|
||||
from litellm.integrations.otel import context as ctx_mod # noqa: E402
|
||||
from litellm.integrations.otel import providers # noqa: E402
|
||||
from litellm.integrations.otel.emitter import SpanEmitter # noqa: E402
|
||||
from litellm.integrations.otel.payloads import ( # noqa: E402
|
||||
GuardrailSpanData,
|
||||
LLMCallSpanData,
|
||||
ServiceSpanData,
|
||||
)
|
||||
from litellm.integrations.otel.spans import SPAN_REGISTRY, SpanRole # noqa: E402
|
||||
|
||||
|
||||
def _payload(**overrides):
|
||||
payload = {
|
||||
"call_type": "acompletion",
|
||||
"custom_llm_provider": "openai",
|
||||
"model": "gpt-4o",
|
||||
"prompt_tokens": 10,
|
||||
"completion_tokens": 5,
|
||||
"total_tokens": 15,
|
||||
"stream": False,
|
||||
"model_parameters": {"temperature": 0.7, "max_tokens": 256, "top_k": 40},
|
||||
"response": {
|
||||
"id": "resp_1",
|
||||
"model": "gpt-4o-2024",
|
||||
"choices": [{"finish_reason": "stop"}],
|
||||
},
|
||||
"metadata": {"team_id": "t1", "team_alias": "team one"},
|
||||
"api_base": "https://api.openai.com:443/v1",
|
||||
"status": "success",
|
||||
"litellm_call_id": "call_1",
|
||||
"response_cost": 0.002,
|
||||
"hidden_params": {},
|
||||
}
|
||||
payload.update(overrides)
|
||||
return payload
|
||||
|
||||
|
||||
def _engine(legacy_compat=True):
|
||||
cfg = OpenTelemetryV2Config(exporter="in_memory", legacy_compat=legacy_compat)
|
||||
provider, exporter = providers.in_memory_provider(cfg)
|
||||
tracer = providers.get_tracer(provider, "litellm-test")
|
||||
return SpanEmitter(tracer, cfg), exporter
|
||||
|
||||
|
||||
def test_llm_call_span_golden():
|
||||
engine, exporter = _engine()
|
||||
data = LLMCallSpanData.from_standard_logging_payload(_payload())
|
||||
engine.emit(SpanRole.LLM_CALL, data)
|
||||
(span,) = exporter.get_finished_spans()
|
||||
assert span.name == "chat gpt-4o"
|
||||
assert span.kind is SpanKind.CLIENT
|
||||
a = span.attributes
|
||||
assert a[GenAI.OPERATION_NAME] == "chat"
|
||||
assert a[GenAI.PROVIDER_NAME] == "openai"
|
||||
assert a[GenAI.REQUEST_MODEL] == "gpt-4o"
|
||||
assert a[GenAI.RESPONSE_MODEL] == "gpt-4o-2024"
|
||||
assert a[GenAI.RESPONSE_ID] == "resp_1"
|
||||
assert a[GenAI.USAGE_INPUT_TOKENS] == 10
|
||||
assert a[GenAI.USAGE_OUTPUT_TOKENS] == 5
|
||||
assert a[GenAI.RESPONSE_FINISH_REASONS] == ("stop",)
|
||||
assert a[GenAI.REQUEST_TEMPERATURE] == 0.7
|
||||
assert a["server.address"] == "api.openai.com"
|
||||
assert a[LiteLLM.CALL_ID] == "call_1"
|
||||
assert a["litellm.cost.total"] == 0.002
|
||||
assert span.status.status_code is StatusCode.OK
|
||||
|
||||
|
||||
def test_legacy_dual_emit_on():
|
||||
engine, exporter = _engine(legacy_compat=True)
|
||||
engine.emit(
|
||||
SpanRole.LLM_CALL, LLMCallSpanData.from_standard_logging_payload(_payload())
|
||||
)
|
||||
(span,) = exporter.get_finished_spans()
|
||||
# canonical AND legacy keys are both present
|
||||
assert span.attributes[GenAI.USAGE_OUTPUT_TOKENS] == 5
|
||||
assert span.attributes["gen_ai.usage.completion_tokens"] == 5
|
||||
assert span.attributes["gen_ai.system"] == "openai"
|
||||
|
||||
|
||||
def test_legacy_dual_emit_off():
|
||||
engine, exporter = _engine(legacy_compat=False)
|
||||
engine.emit(
|
||||
SpanRole.LLM_CALL, LLMCallSpanData.from_standard_logging_payload(_payload())
|
||||
)
|
||||
(span,) = exporter.get_finished_spans()
|
||||
# canonical present, legacy absent
|
||||
assert span.attributes[GenAI.USAGE_OUTPUT_TOKENS] == 5
|
||||
assert "gen_ai.usage.completion_tokens" not in span.attributes
|
||||
assert "gen_ai.system" not in span.attributes
|
||||
|
||||
|
||||
def test_error_span_sets_status_and_error_type():
|
||||
engine, exporter = _engine()
|
||||
payload = _payload(
|
||||
status="failure",
|
||||
error_information={"error_class": "RateLimitError", "error_message": "429"},
|
||||
)
|
||||
engine.emit(
|
||||
SpanRole.LLM_CALL, LLMCallSpanData.from_standard_logging_payload(payload)
|
||||
)
|
||||
(span,) = exporter.get_finished_spans()
|
||||
assert span.status.status_code is StatusCode.ERROR
|
||||
assert span.attributes["error.type"] == "RateLimitError"
|
||||
|
||||
|
||||
def test_hierarchy_and_kinds_match_registry():
|
||||
engine, exporter = _engine()
|
||||
data = LLMCallSpanData.from_standard_logging_payload(_payload())
|
||||
root = engine.start_span(SpanRole.PROXY_REQUEST, "POST /chat/completions")
|
||||
root_ctx = ctx_mod.context_from_span(root)
|
||||
engine.emit(SpanRole.LLM_CALL, data, parent_context=root_ctx)
|
||||
engine.emit(
|
||||
SpanRole.GUARDRAIL, GuardrailSpanData("presidio", status="success"), root_ctx
|
||||
)
|
||||
engine.emit(SpanRole.SERVICE, ServiceSpanData("redis", call_type="set"), root_ctx)
|
||||
root.end()
|
||||
|
||||
by_name = {s.name: s for s in exporter.get_finished_spans()}
|
||||
root_id = root.get_span_context().span_id
|
||||
assert by_name["chat gpt-4o"].parent.span_id == root_id
|
||||
assert by_name["execute_guardrail presidio"].parent.span_id == root_id
|
||||
assert by_name["redis"].parent.span_id == root_id
|
||||
# kinds come straight from the registry
|
||||
assert by_name["chat gpt-4o"].kind is SpanKind.CLIENT
|
||||
assert by_name["execute_guardrail presidio"].kind is SpanKind.INTERNAL
|
||||
assert by_name["redis"].kind is SpanKind.INTERNAL
|
||||
assert by_name["POST /chat/completions"].kind is SpanKind.SERVER
|
||||
|
||||
|
||||
def test_idempotent_dual_fire():
|
||||
engine, exporter = _engine()
|
||||
data = LLMCallSpanData.from_standard_logging_payload(_payload())
|
||||
first = engine.emit(SpanRole.LLM_CALL, data)
|
||||
second = engine.emit(SpanRole.LLM_CALL, data) # same call_id -> deduped
|
||||
assert first is not None
|
||||
assert second is None
|
||||
assert len(exporter.get_finished_spans()) == 1
|
||||
|
||||
|
||||
def test_dedup_cache_is_bounded(monkeypatch):
|
||||
"""The dedup cache only needs to coalesce one request's sync+async fire, so
|
||||
it is a bounded LRU — every unique call_id must not accumulate forever on a
|
||||
long-running proxy."""
|
||||
from litellm.integrations.otel import emitter as emitter_mod
|
||||
|
||||
monkeypatch.setattr(emitter_mod, "_DEDUP_CACHE_MAX", 3)
|
||||
engine, _ = _engine()
|
||||
for i in range(10):
|
||||
engine.emit(
|
||||
SpanRole.LLM_CALL,
|
||||
LLMCallSpanData.from_standard_logging_payload(
|
||||
_payload(litellm_call_id=f"call_{i}")
|
||||
),
|
||||
)
|
||||
assert len(engine._emitted) <= 3
|
||||
|
||||
|
||||
def test_service_error_span():
|
||||
from litellm.integrations.otel.payloads import SpanError
|
||||
|
||||
engine, exporter = _engine()
|
||||
engine.emit(
|
||||
SpanRole.SERVICE,
|
||||
ServiceSpanData(
|
||||
"postgres", call_type="query", error=SpanError("DBError", "boom")
|
||||
),
|
||||
)
|
||||
(span,) = exporter.get_finished_spans()
|
||||
assert span.status.status_code is StatusCode.ERROR
|
||||
assert span.attributes["error.type"] == "DBError"
|
||||
assert span.attributes[LiteLLM.SERVICE_NAME] == "postgres"
|
||||
|
||||
|
||||
def test_guardrail_block_span_is_error_and_carries_verdict():
|
||||
engine, exporter = _engine()
|
||||
data = GuardrailSpanData.from_logging_entry(
|
||||
{
|
||||
"guardrail_name": "openai-moderation",
|
||||
"guardrail_mode": "pre_call",
|
||||
"guardrail_status": "guardrail_intervened",
|
||||
"guardrail_provider": "openai",
|
||||
"guardrail_response": {"violated_categories": ["violence"]},
|
||||
"masked_entity_count": {"EMAIL": 2},
|
||||
}
|
||||
)
|
||||
engine.emit(SpanRole.GUARDRAIL, data)
|
||||
(span,) = exporter.get_finished_spans()
|
||||
assert span.status.status_code is StatusCode.ERROR # intervention → ERROR
|
||||
a = span.attributes
|
||||
assert a[LiteLLM.GUARDRAIL_STATUS] == "guardrail_intervened"
|
||||
assert a[LiteLLM.GUARDRAIL_PROVIDER] == "openai"
|
||||
assert "violence" in a[LiteLLM.GUARDRAIL_RESPONSE] # the verdict rides the span
|
||||
assert a[LiteLLM.GUARDRAIL_MASKED_ENTITY_COUNT] == 2
|
||||
|
||||
|
||||
def test_guardrail_success_span_is_ok():
|
||||
engine, exporter = _engine()
|
||||
engine.emit(
|
||||
SpanRole.GUARDRAIL,
|
||||
GuardrailSpanData.from_logging_entry(
|
||||
{"guardrail_name": "g", "guardrail_status": "success"}
|
||||
),
|
||||
)
|
||||
(span,) = exporter.get_finished_spans()
|
||||
assert span.status.status_code is StatusCode.OK
|
||||
567
tests/test_litellm/integrations/otel/test_otel_v2_logger.py
Normal file
567
tests/test_litellm/integrations/otel/test_otel_v2_logger.py
Normal file
|
|
@ -0,0 +1,567 @@
|
|||
"""Phase 2 / Phase 3 tests for the V2 ``OpenTelemetryV2`` CustomLogger adapter.
|
||||
|
||||
Exercises the callback surface the existing call sites use: LLM-call sync/async
|
||||
success + failure, service hooks, proxy SERVER span lifecycle (start + setters),
|
||||
parent-context resolution (explicit span, traceparent header), and Baggage
|
||||
promotion onto child spans.
|
||||
"""
|
||||
|
||||
import asyncio
|
||||
from datetime import datetime, timezone
|
||||
|
||||
import pytest
|
||||
|
||||
pytest.importorskip("opentelemetry")
|
||||
|
||||
from opentelemetry import trace # noqa: E402
|
||||
from opentelemetry.sdk.trace.export.in_memory_span_exporter import ( # noqa: E402
|
||||
InMemorySpanExporter,
|
||||
)
|
||||
from opentelemetry.trace import SpanKind # noqa: E402
|
||||
from opentelemetry.trace.status import StatusCode # noqa: E402
|
||||
|
||||
from litellm.integrations.otel import ( # noqa: E402
|
||||
GenAI,
|
||||
HTTP,
|
||||
LiteLLM,
|
||||
OpenTelemetryV2Config,
|
||||
)
|
||||
from litellm.integrations.otel import providers # noqa: E402
|
||||
from litellm.integrations.otel.logger import ( # noqa: E402
|
||||
LITELLM_PROXY_REQUEST_SPAN_NAME,
|
||||
OpenTelemetryV2,
|
||||
)
|
||||
from litellm.integrations.otel.spans import SpanRole # noqa: E402
|
||||
from litellm.integrations.otel.utils import to_ns, to_seconds # noqa: E402
|
||||
|
||||
# --------------------------------------------------------------------------- #
|
||||
# Fixtures
|
||||
# --------------------------------------------------------------------------- #
|
||||
|
||||
|
||||
def _payload(**overrides):
|
||||
payload = {
|
||||
"call_type": "acompletion",
|
||||
"custom_llm_provider": "openai",
|
||||
"model": "gpt-4o",
|
||||
"prompt_tokens": 10,
|
||||
"completion_tokens": 5,
|
||||
"total_tokens": 15,
|
||||
"stream": False,
|
||||
"model_parameters": {"temperature": 0.7, "max_tokens": 256},
|
||||
"response": {
|
||||
"id": "resp_1",
|
||||
"model": "gpt-4o-2024",
|
||||
"choices": [{"finish_reason": "stop"}],
|
||||
},
|
||||
"metadata": {
|
||||
"team_id": "t1",
|
||||
"team_alias": "team one",
|
||||
"user_api_key_hash": "hsh",
|
||||
},
|
||||
"api_base": "https://api.openai.com:443/v1",
|
||||
"status": "success",
|
||||
"litellm_call_id": "call_1",
|
||||
"response_cost": 0.002,
|
||||
"hidden_params": {},
|
||||
}
|
||||
payload.update(overrides)
|
||||
return payload
|
||||
|
||||
|
||||
def _kwargs(payload=None):
|
||||
return {
|
||||
"standard_logging_object": payload if payload is not None else _payload(),
|
||||
"litellm_params": {"metadata": {}},
|
||||
}
|
||||
|
||||
|
||||
def _logger(legacy_compat=True):
|
||||
cfg = OpenTelemetryV2Config(exporter="in_memory", legacy_compat=legacy_compat)
|
||||
exporter = InMemorySpanExporter()
|
||||
tracer_provider = providers.build_tracer_provider(cfg, exporter=exporter)
|
||||
return OpenTelemetryV2(config=cfg, tracer_provider=tracer_provider), exporter
|
||||
|
||||
|
||||
# --------------------------------------------------------------------------- #
|
||||
# Time helpers
|
||||
# --------------------------------------------------------------------------- #
|
||||
|
||||
|
||||
def test_to_ns_handles_datetime_and_float():
|
||||
dt = datetime(2026, 5, 26, 12, 0, 0, tzinfo=timezone.utc)
|
||||
assert to_ns(dt) == int(dt.timestamp() * 1e9)
|
||||
assert to_ns(1.5) == 1_500_000_000
|
||||
assert to_ns(None) is None
|
||||
assert to_ns(True) is None # bool is rejected — not a real epoch value
|
||||
|
||||
|
||||
def test_to_seconds_parses_string_formats():
|
||||
assert to_seconds("2026-05-26 12:00:00.123") is not None
|
||||
assert to_seconds("2026-05-26 12:00:00") is not None
|
||||
assert to_seconds("nonsense") is None
|
||||
assert to_seconds(None) is None
|
||||
assert to_seconds(1.5) == 1.5
|
||||
|
||||
|
||||
# --------------------------------------------------------------------------- #
|
||||
# LLM-call callbacks
|
||||
# --------------------------------------------------------------------------- #
|
||||
|
||||
|
||||
def test_async_log_success_event_emits_llm_call_span():
|
||||
logger, exporter = _logger()
|
||||
asyncio.run(logger.async_log_success_event(_kwargs(), None, None, None))
|
||||
(span,) = exporter.get_finished_spans()
|
||||
assert span.name == "chat gpt-4o"
|
||||
assert span.kind is SpanKind.CLIENT
|
||||
assert span.attributes[GenAI.OPERATION_NAME] == "chat"
|
||||
assert span.attributes[GenAI.REQUEST_MODEL] == "gpt-4o"
|
||||
assert span.attributes[LiteLLM.CALL_ID] == "call_1"
|
||||
assert span.status.status_code is StatusCode.OK
|
||||
|
||||
|
||||
def test_async_log_failure_event_marks_error_status():
|
||||
logger, exporter = _logger()
|
||||
payload = _payload(
|
||||
status="failure",
|
||||
error_information={"error_class": "RateLimitError", "error_message": "429"},
|
||||
)
|
||||
asyncio.run(
|
||||
logger.async_log_failure_event(_kwargs(payload=payload), None, None, None)
|
||||
)
|
||||
(span,) = exporter.get_finished_spans()
|
||||
assert span.status.status_code is StatusCode.ERROR
|
||||
assert span.attributes["error.type"] == "RateLimitError"
|
||||
|
||||
|
||||
def test_sync_log_event_is_noop():
|
||||
"""V2 emits async-only; the sync callback runs out-of-context, so it no-ops."""
|
||||
logger, exporter = _logger()
|
||||
logger.log_success_event(_kwargs(), None, None, None)
|
||||
logger.log_failure_event(_kwargs(), None, None, None)
|
||||
assert exporter.get_finished_spans() == ()
|
||||
|
||||
|
||||
def test_missing_standard_logging_object_is_noop():
|
||||
logger, exporter = _logger()
|
||||
asyncio.run(
|
||||
logger.async_log_success_event({"litellm_params": {}}, None, None, None)
|
||||
)
|
||||
assert exporter.get_finished_spans() == ()
|
||||
|
||||
|
||||
def test_pre_call_guardrail_block_suppresses_phantom_llm_span():
|
||||
"""A pre-call guardrail block means the LLM was never called. litellm still
|
||||
emits a failure log, but a CLIENT 'chat …' span would be misleading — so it
|
||||
is suppressed (the guardrail span represents the outcome)."""
|
||||
logger, exporter = _logger()
|
||||
payload = _payload(
|
||||
status="failure",
|
||||
guardrail_information=[
|
||||
{"guardrail_mode": "pre_call", "guardrail_status": "guardrail_intervened"}
|
||||
],
|
||||
)
|
||||
asyncio.run(
|
||||
logger.async_log_failure_event(_kwargs(payload=payload), None, None, None)
|
||||
)
|
||||
assert exporter.get_finished_spans() == () # no phantom LLM span
|
||||
|
||||
|
||||
def test_llm_span_still_emitted_when_guardrail_only_masked():
|
||||
"""A pre-call guardrail that masks (not blocks) lets the call proceed, so the
|
||||
request succeeds and the LLM span must still be emitted."""
|
||||
logger, exporter = _logger()
|
||||
payload = _payload(
|
||||
status="success",
|
||||
guardrail_information=[
|
||||
{"guardrail_mode": "pre_call", "guardrail_status": "guardrail_intervened"}
|
||||
],
|
||||
)
|
||||
asyncio.run(
|
||||
logger.async_log_success_event(_kwargs(payload=payload), None, None, None)
|
||||
)
|
||||
assert len(exporter.get_finished_spans()) == 1 # real LLM call span present
|
||||
|
||||
|
||||
def test_idempotent_on_repeat_call_id():
|
||||
"""Same StandardLoggingPayload (same id) emits once even if the async hook fires twice."""
|
||||
logger, exporter = _logger()
|
||||
kwargs = _kwargs()
|
||||
asyncio.run(logger.async_log_success_event(kwargs, None, None, None))
|
||||
asyncio.run(logger.async_log_success_event(kwargs, None, None, None))
|
||||
assert len(exporter.get_finished_spans()) == 1
|
||||
|
||||
|
||||
# --------------------------------------------------------------------------- #
|
||||
# Parent resolution — ambient context (no metadata threading)
|
||||
# --------------------------------------------------------------------------- #
|
||||
|
||||
|
||||
def test_llm_span_parents_to_ambient_server_span():
|
||||
"""With the FastAPI instrumentor, the server span is the active context; the
|
||||
LLM span nests under it via ambient context (no ``litellm_parent_otel_span``).
|
||||
"""
|
||||
logger, exporter = _logger()
|
||||
server = logger._emitter.start_span(
|
||||
SpanRole.PROXY_REQUEST, LITELLM_PROXY_REQUEST_SPAN_NAME
|
||||
)
|
||||
with trace.use_span(server, end_on_exit=False):
|
||||
asyncio.run(logger.async_log_success_event(_kwargs(), None, None, None))
|
||||
server.end()
|
||||
by_name = {s.name: s for s in exporter.get_finished_spans()}
|
||||
llm_span = by_name["chat gpt-4o"]
|
||||
assert llm_span.parent is not None
|
||||
assert llm_span.parent.span_id == server.get_span_context().span_id
|
||||
|
||||
|
||||
def test_llm_span_is_root_without_ambient_server_span():
|
||||
logger, exporter = _logger()
|
||||
asyncio.run(logger.async_log_success_event(_kwargs(), None, None, None))
|
||||
(span,) = exporter.get_finished_spans()
|
||||
assert span.parent is None # standalone (no proxy server span) → root
|
||||
|
||||
|
||||
# Inbound ``traceparent`` propagation is now the FastAPI instrumentor's job
|
||||
# (see proxy_server's startup mount + ``test_otel_v2_mount``), not the logger's.
|
||||
|
||||
|
||||
# --------------------------------------------------------------------------- #
|
||||
# Baggage promotion (LLM call writes identity into baggage so child spans
|
||||
# inherit team/key/model attrs).
|
||||
# --------------------------------------------------------------------------- #
|
||||
|
||||
|
||||
def test_baggage_identity_promoted_onto_llm_call():
|
||||
logger, exporter = _logger()
|
||||
asyncio.run(logger.async_log_success_event(_kwargs(), None, None, None))
|
||||
(span,) = exporter.get_finished_spans()
|
||||
assert span.attributes[LiteLLM.TEAM_ID] == "t1"
|
||||
assert span.attributes[LiteLLM.TEAM_ALIAS] == "team one"
|
||||
assert span.attributes[GenAI.REQUEST_MODEL] == "gpt-4o"
|
||||
|
||||
|
||||
class _Auth:
|
||||
"""Stub matching the ``UserAPIKeyAuth`` fields the logger reads."""
|
||||
|
||||
team_id = "t1"
|
||||
team_alias = "team one"
|
||||
api_key = "hash1"
|
||||
user_id = "u1"
|
||||
org_id = None
|
||||
key_alias = "k1"
|
||||
end_user_id = None
|
||||
|
||||
|
||||
def test_pre_call_hook_seeds_baggage_onto_server_and_child_spans():
|
||||
"""The pre-call hook seeds identity Baggage in the request context so the
|
||||
server span (stamped directly) AND later child spans (service here, via the
|
||||
Baggage processor) carry identity — not just the LLM-call span."""
|
||||
logger, exporter = _logger()
|
||||
server = logger._emitter.start_span(
|
||||
SpanRole.PROXY_REQUEST, LITELLM_PROXY_REQUEST_SPAN_NAME
|
||||
)
|
||||
|
||||
async def _flow():
|
||||
# pre-call seeds baggage + stamps the active server span
|
||||
await logger.async_pre_call_hook(
|
||||
_Auth(), None, {"model": "gpt-4o"}, "completion"
|
||||
)
|
||||
# a later service call (same task) must inherit the identity
|
||||
await logger.async_service_success_hook(
|
||||
payload=_ServicePayload("redis", "set"), parent_otel_span=server
|
||||
)
|
||||
|
||||
with trace.use_span(server, end_on_exit=False):
|
||||
asyncio.run(_flow())
|
||||
server.end()
|
||||
|
||||
spans = {s.name: s for s in exporter.get_finished_spans()}
|
||||
redis = spans["redis"]
|
||||
assert redis.attributes[LiteLLM.TEAM_ID] == "t1"
|
||||
assert redis.attributes[LiteLLM.KEY_HASH] == "hash1"
|
||||
assert redis.attributes[f"{LiteLLM.METADATA_PREFIX}user_api_key_user_id"] == "u1"
|
||||
srv = spans[LITELLM_PROXY_REQUEST_SPAN_NAME]
|
||||
assert (
|
||||
srv.attributes[LiteLLM.TEAM_ID] == "t1"
|
||||
) # stamped directly on the server span
|
||||
assert srv.attributes[f"{LiteLLM.METADATA_PREFIX}user_api_key_user_id"] == "u1"
|
||||
|
||||
|
||||
# --------------------------------------------------------------------------- #
|
||||
# Service hooks (Phase 3)
|
||||
# --------------------------------------------------------------------------- #
|
||||
|
||||
|
||||
class _Service:
|
||||
"""Stub matching ``ServiceTypes(str, Enum)``."""
|
||||
|
||||
def __init__(self, value):
|
||||
self.value = value
|
||||
|
||||
|
||||
class _ServicePayload:
|
||||
def __init__(self, service="redis", call_type="set", error=None):
|
||||
self.service = _Service(service)
|
||||
self.call_type = call_type
|
||||
self.error = error
|
||||
|
||||
|
||||
def _service_parent(logger):
|
||||
"""Helper: a live PROXY_REQUEST span to parent service spans under."""
|
||||
return logger._emitter.start_span(
|
||||
SpanRole.PROXY_REQUEST, LITELLM_PROXY_REQUEST_SPAN_NAME
|
||||
)
|
||||
|
||||
|
||||
def test_async_service_success_hook_emits_service_span():
|
||||
logger, exporter = _logger()
|
||||
parent = _service_parent(logger)
|
||||
try:
|
||||
asyncio.run(
|
||||
logger.async_service_success_hook(
|
||||
payload=_ServicePayload("redis", "set"),
|
||||
parent_otel_span=parent,
|
||||
event_metadata={"key1": "val1"},
|
||||
)
|
||||
)
|
||||
finally:
|
||||
parent.end()
|
||||
by_name = {s.name: s for s in exporter.get_finished_spans()}
|
||||
span = by_name["redis"]
|
||||
assert span.kind is SpanKind.INTERNAL
|
||||
assert span.attributes[LiteLLM.SERVICE_NAME] == "redis"
|
||||
assert span.attributes[LiteLLM.SERVICE_CALL_TYPE] == "set"
|
||||
# Canonical (V2) namespaced metadata key
|
||||
assert span.attributes[f"{LiteLLM.METADATA_PREFIX}key1"] == "val1"
|
||||
# V1 bare key (legacy dual-emit)
|
||||
assert span.attributes["key1"] == "val1"
|
||||
assert span.attributes["service"] == "redis" # V1 bare key
|
||||
assert span.attributes["call_type"] == "set" # V1 bare key
|
||||
assert span.status.status_code is StatusCode.OK
|
||||
|
||||
|
||||
def test_async_service_failure_hook_marks_error_status():
|
||||
logger, exporter = _logger()
|
||||
parent = _service_parent(logger)
|
||||
try:
|
||||
asyncio.run(
|
||||
logger.async_service_failure_hook(
|
||||
payload=_ServicePayload("postgres", "query"),
|
||||
error="boom",
|
||||
parent_otel_span=parent,
|
||||
)
|
||||
)
|
||||
finally:
|
||||
parent.end()
|
||||
by_name = {s.name: s for s in exporter.get_finished_spans()}
|
||||
span = by_name["postgres"]
|
||||
assert span.status.status_code is StatusCode.ERROR
|
||||
# Without an explicit error_type from the payload, V2 stamps the fallback.
|
||||
assert span.attributes["error.type"] == "error"
|
||||
assert span.attributes[LiteLLM.SERVICE_NAME] == "postgres"
|
||||
|
||||
|
||||
def test_async_service_failure_hook_preserves_payload_error_over_override():
|
||||
"""When the payload itself carries an error, that takes precedence over the override."""
|
||||
logger, exporter = _logger()
|
||||
parent = _service_parent(logger)
|
||||
try:
|
||||
asyncio.run(
|
||||
logger.async_service_failure_hook(
|
||||
payload=_ServicePayload("postgres", "query", error="db-down"),
|
||||
error="override-only-used-when-payload-clean",
|
||||
parent_otel_span=parent,
|
||||
)
|
||||
)
|
||||
finally:
|
||||
parent.end()
|
||||
by_name = {s.name: s for s in exporter.get_finished_spans()}
|
||||
span = by_name["postgres"]
|
||||
assert span.status.status_code is StatusCode.ERROR
|
||||
assert "db-down" in (span.status.description or "")
|
||||
|
||||
|
||||
def test_service_hook_without_parent_is_noop():
|
||||
"""Mirrors V1: no parent OTel span → no service span (no free-standing roots)."""
|
||||
logger, exporter = _logger()
|
||||
asyncio.run(
|
||||
logger.async_service_success_hook(
|
||||
payload=_ServicePayload(), parent_otel_span=None
|
||||
)
|
||||
)
|
||||
assert exporter.get_finished_spans() == ()
|
||||
|
||||
|
||||
def test_service_span_inherits_parent_when_provided():
|
||||
logger, exporter = _logger()
|
||||
parent = logger._emitter.start_span(
|
||||
SpanRole.PROXY_REQUEST, LITELLM_PROXY_REQUEST_SPAN_NAME
|
||||
)
|
||||
try:
|
||||
asyncio.run(
|
||||
logger.async_service_success_hook(
|
||||
payload=_ServicePayload(), parent_otel_span=parent
|
||||
)
|
||||
)
|
||||
finally:
|
||||
parent.end()
|
||||
by_name = {s.name: s for s in exporter.get_finished_spans()}
|
||||
assert (
|
||||
by_name["redis"].parent.span_id
|
||||
== by_name[LITELLM_PROXY_REQUEST_SPAN_NAME].get_span_context().span_id
|
||||
)
|
||||
|
||||
|
||||
# --------------------------------------------------------------------------- #
|
||||
# Proxy SERVER span lifecycle
|
||||
# --------------------------------------------------------------------------- #
|
||||
|
||||
|
||||
def test_create_proxy_request_started_span_returns_ambient_span():
|
||||
"""V2 doesn't create a server span (the instrumentor does), but it returns
|
||||
the active server span so the proxy can thread it as the service-span parent
|
||||
— service logging only fires the OTel hook when that parent is non-None."""
|
||||
logger, exporter = _logger()
|
||||
# No ambient recordable span → None (and creates nothing).
|
||||
assert (
|
||||
logger.create_litellm_proxy_request_started_span(
|
||||
start_time=datetime.now(timezone.utc), headers={"traceparent": "x"}
|
||||
)
|
||||
is None
|
||||
)
|
||||
assert exporter.get_finished_spans() == ()
|
||||
# With an active server span, return it (do NOT create a new one).
|
||||
server = logger._emitter.start_span(
|
||||
SpanRole.PROXY_REQUEST, LITELLM_PROXY_REQUEST_SPAN_NAME
|
||||
)
|
||||
with trace.use_span(server, end_on_exit=False):
|
||||
got = logger.create_litellm_proxy_request_started_span(
|
||||
start_time=datetime.now(timezone.utc), headers=None
|
||||
)
|
||||
server.end()
|
||||
assert got is server
|
||||
|
||||
|
||||
def test_proxy_span_setters_are_noops():
|
||||
"""The FastAPI instrumentor owns the server span; the setters write nothing
|
||||
(and must tolerate any span / None without raising).
|
||||
"""
|
||||
logger, exporter = _logger()
|
||||
span = logger._emitter.start_span(
|
||||
SpanRole.PROXY_REQUEST, LITELLM_PROXY_REQUEST_SPAN_NAME
|
||||
)
|
||||
OpenTelemetryV2.set_proxy_request_route_attributes(
|
||||
span, url_path="/chat/completions", http_route="/chat/completions"
|
||||
)
|
||||
OpenTelemetryV2.set_response_status_code_attribute(span, 200)
|
||||
OpenTelemetryV2.set_preprocessing_duration_attribute(
|
||||
span,
|
||||
{"first_api_call_start_time": 1.0, "metadata": {"litellm_received_at": 0.5}},
|
||||
)
|
||||
span.end()
|
||||
(finished,) = exporter.get_finished_spans()
|
||||
assert HTTP.URL_PATH not in finished.attributes
|
||||
assert HTTP.ROUTE not in finished.attributes
|
||||
assert HTTP.RESPONSE_STATUS_CODE not in finished.attributes
|
||||
assert LiteLLM.PREPROCESSING_MS not in finished.attributes
|
||||
# None span is tolerated too.
|
||||
OpenTelemetryV2.set_proxy_request_route_attributes(None, http_route="/x")
|
||||
OpenTelemetryV2.set_response_status_code_attribute(None, 200)
|
||||
OpenTelemetryV2.set_preprocessing_duration_attribute(None, {})
|
||||
|
||||
|
||||
# --------------------------------------------------------------------------- #
|
||||
# Constructor / proxy global guard
|
||||
# --------------------------------------------------------------------------- #
|
||||
|
||||
|
||||
def test_constructor_accepts_v1_compatible_kwargs():
|
||||
"""Mirrors V1's positional shape — config / callback_name / providers / **kwargs."""
|
||||
cfg = OpenTelemetryV2Config(exporter="in_memory")
|
||||
tp = providers.build_tracer_provider(cfg)
|
||||
logger = OpenTelemetryV2(
|
||||
config=cfg,
|
||||
callback_name="otel",
|
||||
tracer_provider=tp,
|
||||
logger_provider=None,
|
||||
meter_provider=None,
|
||||
turn_off_message_logging=True,
|
||||
)
|
||||
assert logger.callback_name == "otel"
|
||||
assert logger.turn_off_message_logging is True
|
||||
assert logger.tracer is not None
|
||||
|
||||
|
||||
def test_default_config_reads_env(monkeypatch):
|
||||
"""No explicit config → reads env (exporter=console by default)."""
|
||||
monkeypatch.delenv("OTEL_EXPORTER", raising=False)
|
||||
monkeypatch.delenv("OTEL_EXPORTER_OTLP_PROTOCOL", raising=False)
|
||||
logger = OpenTelemetryV2(
|
||||
tracer_provider=providers.build_tracer_provider(
|
||||
OpenTelemetryV2Config(exporter="in_memory")
|
||||
)
|
||||
)
|
||||
assert logger.config.exporter == "console"
|
||||
|
||||
|
||||
def test_proxy_global_first_registered_wins(monkeypatch):
|
||||
"""``_init_otel_logger_on_litellm_proxy`` claims the global only when empty."""
|
||||
proxy_server = pytest.importorskip("litellm.proxy.proxy_server")
|
||||
monkeypatch.setattr(proxy_server, "open_telemetry_logger", None, raising=False)
|
||||
cfg = OpenTelemetryV2Config(exporter="in_memory")
|
||||
tp = providers.build_tracer_provider(cfg)
|
||||
|
||||
first = OpenTelemetryV2(config=cfg, tracer_provider=tp)
|
||||
assert proxy_server.open_telemetry_logger is first
|
||||
|
||||
second = OpenTelemetryV2(config=cfg, tracer_provider=tp)
|
||||
# Global still points at the first registration.
|
||||
assert proxy_server.open_telemetry_logger is first
|
||||
assert second is not first
|
||||
|
||||
|
||||
def test_registers_into_litellm_service_callback(monkeypatch):
|
||||
"""The logger must mutate ``litellm.service_callback`` in place. An empty
|
||||
list is falsy, so a ``getattr(..) or []`` would append to a throwaway local
|
||||
and service spans (Redis, …) would silently never fire on this logger.
|
||||
"""
|
||||
import litellm
|
||||
|
||||
pytest.importorskip("litellm.proxy.proxy_server")
|
||||
monkeypatch.setattr(litellm, "service_callback", [], raising=False)
|
||||
cfg = OpenTelemetryV2Config(exporter="in_memory")
|
||||
tp = providers.build_tracer_provider(cfg)
|
||||
|
||||
first = OpenTelemetryV2(config=cfg, tracer_provider=tp)
|
||||
assert first in litellm.service_callback
|
||||
|
||||
# A second OTel logger sees one is already registered and does not duplicate.
|
||||
OpenTelemetryV2(config=cfg, tracer_provider=tp)
|
||||
otel_registrations = [
|
||||
cb
|
||||
for cb in litellm.service_callback
|
||||
if cb.__class__.__module__.startswith("litellm.integrations.otel")
|
||||
]
|
||||
assert len(otel_registrations) == 1
|
||||
|
||||
|
||||
# --------------------------------------------------------------------------- #
|
||||
# Management endpoint hooks — no-ops: management endpoints are ordinary FastAPI
|
||||
# routes, so the mounted instrumentor spans them. The hooks must not emit.
|
||||
# --------------------------------------------------------------------------- #
|
||||
|
||||
|
||||
def test_management_hooks_are_noops():
|
||||
logger, exporter = _logger()
|
||||
|
||||
class _Payload:
|
||||
route = "/key/generate"
|
||||
request_data = {"models": "gpt-4o"}
|
||||
response = {"key": "sk-123"}
|
||||
exception = ValueError("nope")
|
||||
start_time = end_time = None
|
||||
|
||||
asyncio.run(logger.async_management_endpoint_success_hook(_Payload()))
|
||||
asyncio.run(logger.async_management_endpoint_failure_hook(_Payload()))
|
||||
assert exporter.get_finished_spans() == ()
|
||||
73
tests/test_litellm/integrations/otel/test_otel_v2_mount.py
Normal file
73
tests/test_litellm/integrations/otel/test_otel_v2_mount.py
Normal file
|
|
@ -0,0 +1,73 @@
|
|||
"""V2 entrypoint: the FastAPI instrumentation pattern proxy_server mounts at
|
||||
startup (gated by LITELLM_OTEL_V2). The mount logic itself lives inline in
|
||||
``proxy_server.proxy_startup_event``; this exercises the same pattern in
|
||||
isolation so the server-span + shared-provider behavior stays covered.
|
||||
"""
|
||||
|
||||
import os
|
||||
import sys
|
||||
|
||||
import pytest
|
||||
|
||||
sys.path.insert(0, os.path.abspath("../../../.."))
|
||||
|
||||
pytest.importorskip("opentelemetry")
|
||||
pytest.importorskip("opentelemetry.instrumentation.fastapi")
|
||||
fastapi = pytest.importorskip("fastapi")
|
||||
|
||||
from fastapi.testclient import TestClient # noqa: E402
|
||||
from opentelemetry.instrumentation.fastapi import FastAPIInstrumentor # noqa: E402
|
||||
from opentelemetry.sdk.trace.export import SimpleSpanProcessor # noqa: E402
|
||||
from opentelemetry.sdk.trace.export.in_memory_span_exporter import ( # noqa: E402
|
||||
InMemorySpanExporter,
|
||||
)
|
||||
from opentelemetry.trace import SpanKind # noqa: E402
|
||||
|
||||
from litellm.integrations.otel.config import ( # noqa: E402
|
||||
OpenTelemetryV2Config,
|
||||
is_otel_v2_enabled,
|
||||
)
|
||||
from litellm.integrations.otel.logger import OpenTelemetryV2 # noqa: E402
|
||||
|
||||
|
||||
def _instrumented_app():
|
||||
"""Mirror proxy_server's startup mount: a logger builds the shared provider,
|
||||
and the FastAPI instrumentor is attached to it."""
|
||||
app = fastapi.FastAPI()
|
||||
|
||||
@app.get("/ping")
|
||||
def ping():
|
||||
return {"ok": True}
|
||||
|
||||
logger = OpenTelemetryV2(config=OpenTelemetryV2Config(exporter="in_memory"))
|
||||
FastAPIInstrumentor.instrument_app(app, tracer_provider=logger._tracer_provider)
|
||||
return app, logger
|
||||
|
||||
|
||||
def test_gate_toggles_with_env(monkeypatch):
|
||||
"""The startup mount is guarded by this flag."""
|
||||
monkeypatch.delenv("LITELLM_OTEL_V2", raising=False)
|
||||
assert is_otel_v2_enabled() is False
|
||||
monkeypatch.setenv("LITELLM_OTEL_V2", "1")
|
||||
assert is_otel_v2_enabled() is True
|
||||
|
||||
|
||||
def test_instrumented_app_emits_server_span():
|
||||
app, logger = _instrumented_app()
|
||||
exporter = InMemorySpanExporter()
|
||||
logger._tracer_provider.add_span_processor(SimpleSpanProcessor(exporter))
|
||||
|
||||
TestClient(app).get("/ping")
|
||||
|
||||
server_spans = [
|
||||
s for s in exporter.get_finished_spans() if s.kind is SpanKind.SERVER
|
||||
]
|
||||
assert server_spans, "FastAPI instrumentor should emit a SERVER span per request"
|
||||
attrs = server_spans[0].attributes or {}
|
||||
assert any("route" in k or "method" in k for k in attrs)
|
||||
|
||||
|
||||
def test_logger_and_instrumentor_share_provider():
|
||||
"""Gen-ai spans (logger) and server spans (instrumentor) write to one provider."""
|
||||
_, logger = _instrumented_app()
|
||||
assert logger._emitter._tracer is logger.tracer
|
||||
|
|
@ -0,0 +1,89 @@
|
|||
"""Multi-backend fan-out: one TracerProvider, *N* SpanProcessors.
|
||||
|
||||
V1 needed a separate ``TracerProvider`` per integration to avoid stepping on
|
||||
the global. V2 attaches a ``SpanProcessor`` per exporter to the *same*
|
||||
provider, so the same trace ID lights up every backend — no duplicate spans,
|
||||
no per-integration provider caches.
|
||||
"""
|
||||
|
||||
import pytest
|
||||
|
||||
pytest.importorskip("opentelemetry")
|
||||
|
||||
from opentelemetry.sdk.trace.export.in_memory_span_exporter import (
|
||||
InMemorySpanExporter,
|
||||
)
|
||||
|
||||
from litellm.integrations.otel.config import ExporterSpec, OpenTelemetryV2Config
|
||||
from litellm.integrations.otel.providers import build_tracer_provider
|
||||
|
||||
|
||||
def test_two_exporters_receive_the_same_span():
|
||||
"""A single ``span.end()`` lands in BOTH exporters with the same span ID."""
|
||||
exporter_a = InMemorySpanExporter()
|
||||
exporter_b = InMemorySpanExporter()
|
||||
cfg = OpenTelemetryV2Config(
|
||||
exporters=[
|
||||
ExporterSpec(kind="in_memory"),
|
||||
ExporterSpec(kind="in_memory"),
|
||||
]
|
||||
)
|
||||
# Override the auto-built exporters with our test ones by swapping
|
||||
# processors after construction (the test's purpose is to exercise the
|
||||
# multi-processor wiring, not to negotiate the in-memory pipe).
|
||||
from opentelemetry.sdk.trace.export import SimpleSpanProcessor
|
||||
|
||||
provider = build_tracer_provider(cfg)
|
||||
# Clear out any auto-built export processors and attach our pair.
|
||||
while provider._active_span_processor._span_processors:
|
||||
provider._active_span_processor._span_processors = (
|
||||
provider._active_span_processor._span_processors[:-1]
|
||||
)
|
||||
provider.add_span_processor(SimpleSpanProcessor(exporter_a))
|
||||
provider.add_span_processor(SimpleSpanProcessor(exporter_b))
|
||||
|
||||
tracer = provider.get_tracer("test")
|
||||
span = tracer.start_span("multi-backend")
|
||||
span.set_attribute("test.marker", "yes")
|
||||
span.end()
|
||||
|
||||
spans_a = exporter_a.get_finished_spans()
|
||||
spans_b = exporter_b.get_finished_spans()
|
||||
assert len(spans_a) == 1
|
||||
assert len(spans_b) == 1
|
||||
assert spans_a[0].context.span_id == spans_b[0].context.span_id
|
||||
|
||||
|
||||
def test_resource_attributes_apply_to_all_exporters():
|
||||
"""``resource_attributes`` flow through the shared TracerProvider."""
|
||||
cfg = OpenTelemetryV2Config(
|
||||
exporters=[ExporterSpec(kind="in_memory")],
|
||||
resource_attributes={"openinference.project.name": "phoenix-test"},
|
||||
)
|
||||
provider = build_tracer_provider(cfg)
|
||||
assert provider.resource.attributes["openinference.project.name"] == "phoenix-test"
|
||||
|
||||
|
||||
def test_config_normalizer_inserts_genai_first():
|
||||
"""The validator pins ``genai`` at the head + appends ``legacy`` on legacy_compat."""
|
||||
cfg = OpenTelemetryV2Config(mapper_names=["openinference", "langfuse"])
|
||||
assert cfg.mapper_names[0] == "genai"
|
||||
assert "openinference" in cfg.mapper_names
|
||||
assert "langfuse" in cfg.mapper_names
|
||||
assert cfg.mapper_names[-1] == "legacy" # legacy_compat=True by default
|
||||
|
||||
|
||||
def test_config_normalizer_no_legacy_when_compat_off():
|
||||
cfg = OpenTelemetryV2Config(legacy_compat=False, mapper_names=["openinference"])
|
||||
assert "legacy" not in cfg.mapper_names
|
||||
assert cfg.mapper_names[0] == "genai"
|
||||
|
||||
|
||||
def test_config_folds_legacy_exporter_triple_into_exporters_list():
|
||||
"""When ``exporters`` is empty, the validator folds the legacy single triple."""
|
||||
cfg = OpenTelemetryV2Config(
|
||||
exporter="otlp_http", endpoint="https://api.example.com", headers="k=v"
|
||||
)
|
||||
assert len(cfg.exporters) == 1
|
||||
assert cfg.exporters[0].kind == "otlp_http"
|
||||
assert cfg.exporters[0].endpoint == "https://api.example.com"
|
||||
122
tests/test_litellm/integrations/otel/test_otel_v2_presets.py
Normal file
122
tests/test_litellm/integrations/otel/test_otel_v2_presets.py
Normal file
|
|
@ -0,0 +1,122 @@
|
|||
"""Preset tests. Focused on the AgentOps JWT fetch, which must never block the
|
||||
event loop: the preset does no network I/O, and a custom exporter mints the JWT
|
||||
lazily on its first export (in the BatchSpanProcessor worker thread)."""
|
||||
|
||||
import httpx
|
||||
import pytest
|
||||
|
||||
from litellm.integrations.otel import providers
|
||||
from litellm.integrations.otel.config import ExporterSpec
|
||||
from litellm.integrations.otel.presets import agentops as agentops_mod
|
||||
from litellm.integrations.otel.presets.agentops import (
|
||||
_AGENTOPS_ENDPOINT,
|
||||
_AGENTOPS_EXPORTER_KIND,
|
||||
_build_agentops_exporter,
|
||||
_fetch_agentops_jwt,
|
||||
agentops_preset,
|
||||
)
|
||||
|
||||
|
||||
def test_agentops_preset_does_no_network_io(monkeypatch):
|
||||
# The preset must not fetch the JWT at build time — that would block the
|
||||
# event loop during callback construction. It only describes the exporter.
|
||||
def _boom(*_a, **_k):
|
||||
raise AssertionError("agentops_preset must not fetch the JWT eagerly")
|
||||
|
||||
monkeypatch.setattr(agentops_mod, "_fetch_agentops_jwt", _boom)
|
||||
monkeypatch.setenv("AGENTOPS_API_KEY", "ak-123")
|
||||
cfg = agentops_preset()
|
||||
agentops_exporters = [e for e in cfg.exporters if e.kind == _AGENTOPS_EXPORTER_KIND]
|
||||
assert len(agentops_exporters) == 1
|
||||
spec = agentops_exporters[0]
|
||||
assert spec.endpoint == _AGENTOPS_ENDPOINT
|
||||
assert spec.options == {"api_key": "ak-123"} # carried to the lazy exporter
|
||||
|
||||
|
||||
def test_agentops_preset_without_key_omits_options(monkeypatch):
|
||||
monkeypatch.delenv("AGENTOPS_API_KEY", raising=False)
|
||||
cfg = agentops_preset()
|
||||
spec = next(e for e in cfg.exporters if e.kind == _AGENTOPS_EXPORTER_KIND)
|
||||
assert spec.options is None
|
||||
|
||||
|
||||
def test_agentops_exporter_factory_is_registered():
|
||||
assert _AGENTOPS_EXPORTER_KIND in providers._EXPORTER_FACTORIES
|
||||
|
||||
|
||||
def test_agentops_exporter_mints_jwt_lazily(monkeypatch):
|
||||
pytest.importorskip("opentelemetry.exporter.otlp.proto.http.trace_exporter")
|
||||
monkeypatch.setattr(
|
||||
agentops_mod, "_fetch_agentops_jwt", lambda _k: {"token": "jwt-xyz"}
|
||||
)
|
||||
spec = ExporterSpec(
|
||||
kind=_AGENTOPS_EXPORTER_KIND,
|
||||
endpoint=_AGENTOPS_ENDPOINT,
|
||||
options={"api_key": "ak"},
|
||||
)
|
||||
exporter = _build_agentops_exporter(spec)
|
||||
|
||||
# No auth header until the first export triggers the (off-loop) fetch.
|
||||
assert "Authorization" not in exporter._session.headers
|
||||
exporter._ensure_authenticated()
|
||||
assert exporter._session.headers["Authorization"] == "Bearer jwt-xyz"
|
||||
|
||||
# Cached: a second resolution does not re-fetch.
|
||||
calls = []
|
||||
monkeypatch.setattr(
|
||||
agentops_mod,
|
||||
"_fetch_agentops_jwt",
|
||||
lambda k: calls.append(k) or {"token": "again"},
|
||||
)
|
||||
exporter._ensure_authenticated()
|
||||
assert calls == []
|
||||
|
||||
|
||||
def test_agentops_exporter_tolerates_fetch_failure(monkeypatch):
|
||||
pytest.importorskip("opentelemetry.exporter.otlp.proto.http.trace_exporter")
|
||||
|
||||
def _raise(_k):
|
||||
raise RuntimeError("auth down")
|
||||
|
||||
monkeypatch.setattr(agentops_mod, "_fetch_agentops_jwt", _raise)
|
||||
exporter = _build_agentops_exporter(
|
||||
ExporterSpec(
|
||||
kind=_AGENTOPS_EXPORTER_KIND,
|
||||
endpoint=_AGENTOPS_ENDPOINT,
|
||||
options={"api_key": "ak"},
|
||||
)
|
||||
)
|
||||
exporter._ensure_authenticated() # must not raise
|
||||
assert "Authorization" not in exporter._session.headers
|
||||
|
||||
|
||||
def test_fetch_jwt_uses_owned_client_not_shared_pool(monkeypatch):
|
||||
"""The fetch owns a short-lived client and closes it, rather than closing
|
||||
the process-wide cached ``_get_httpx_client`` pool shared by other callers."""
|
||||
closed = {"n": 0}
|
||||
|
||||
class _FakeResponse:
|
||||
status_code = 200
|
||||
|
||||
def json(self):
|
||||
return {"token": "jwt-123"}
|
||||
|
||||
class _FakeClient:
|
||||
def __init__(self, *_a, **_k):
|
||||
pass
|
||||
|
||||
def __enter__(self):
|
||||
return self
|
||||
|
||||
def __exit__(self, *_a):
|
||||
closed["n"] += 1
|
||||
|
||||
def post(self, *_a, **_k):
|
||||
return _FakeResponse()
|
||||
|
||||
monkeypatch.setattr(httpx, "Client", _FakeClient)
|
||||
assert not hasattr(agentops_mod, "_get_httpx_client")
|
||||
|
||||
result = _fetch_agentops_jwt("api-key")
|
||||
assert result == {"token": "jwt-123"}
|
||||
assert closed["n"] == 1 # the owned client was closed
|
||||
|
|
@ -0,0 +1,400 @@
|
|||
"""Tests for the OTel v2 sources of truth: span registry, semconv keys, config,
|
||||
and the typed StandardLoggingPayload adapter. These need no OTel SDK."""
|
||||
|
||||
import pytest
|
||||
|
||||
from litellm.integrations.otel import (
|
||||
BAGGAGE_PROMOTED_KEYS,
|
||||
Error,
|
||||
GenAI,
|
||||
GenAIOperation,
|
||||
HTTP,
|
||||
LiteLLM,
|
||||
OpenTelemetryV2Config,
|
||||
Server,
|
||||
is_otel_v2_enabled,
|
||||
promoted_baggage,
|
||||
resolve_operation,
|
||||
resolve_provider,
|
||||
)
|
||||
from litellm.integrations.otel import spans as spans_mod
|
||||
from litellm.integrations.otel.payloads import LLMCallSpanData, RequestIdentity
|
||||
from litellm.integrations.otel.spans import (
|
||||
SPAN_REGISTRY,
|
||||
LiteLLMSpanKind,
|
||||
SpanRole,
|
||||
child_roles,
|
||||
root_roles,
|
||||
validate_registry,
|
||||
)
|
||||
|
||||
|
||||
def _sample_payload(**overrides):
|
||||
payload = {
|
||||
"call_type": "acompletion",
|
||||
"custom_llm_provider": "openai",
|
||||
"model": "gpt-4o",
|
||||
"prompt_tokens": 10,
|
||||
"completion_tokens": 5,
|
||||
"total_tokens": 15,
|
||||
"stream": False,
|
||||
"model_parameters": {
|
||||
"temperature": 0.7,
|
||||
"max_tokens": 256,
|
||||
"top_p": 0.9,
|
||||
"top_k": 40,
|
||||
"frequency_penalty": 0.1,
|
||||
"presence_penalty": 0.2,
|
||||
"stop": ["STOP"],
|
||||
"seed": 42,
|
||||
},
|
||||
"response": {
|
||||
"id": "resp_1",
|
||||
"model": "gpt-4o-2024",
|
||||
"choices": [{"finish_reason": "stop"}],
|
||||
},
|
||||
"metadata": {
|
||||
"team_id": "t1",
|
||||
"team_alias": "team one",
|
||||
"user_api_key_hash": "hsh",
|
||||
"user_api_key_org_id": "org1",
|
||||
},
|
||||
"api_base": "https://api.openai.com:443/v1",
|
||||
"status": "success",
|
||||
"litellm_call_id": "call_1",
|
||||
"end_user": "u1",
|
||||
"response_cost": 0.002,
|
||||
"hidden_params": {},
|
||||
}
|
||||
payload.update(overrides)
|
||||
return payload
|
||||
|
||||
|
||||
# --- span registry (source of truth #2) ------------------------------------- #
|
||||
|
||||
|
||||
def test_registry_validates_and_is_complete():
|
||||
validate_registry() # raises on inconsistency
|
||||
assert set(SPAN_REGISTRY) == set(SpanRole)
|
||||
|
||||
|
||||
def test_registry_parent_integrity_no_orphans():
|
||||
for role, spec in SPAN_REGISTRY.items():
|
||||
assert spec.role is role
|
||||
if spec.parent is not None:
|
||||
assert spec.parent in SPAN_REGISTRY
|
||||
|
||||
|
||||
def test_registry_hierarchy_shape():
|
||||
assert set(root_roles()) == {SpanRole.PROXY_REQUEST}
|
||||
# Guardrails parent to the request span, not the LLM call: a pre-call
|
||||
# guardrail runs before the LLM call exists, so it's a sibling of it.
|
||||
assert set(child_roles(SpanRole.PROXY_REQUEST)) == {
|
||||
SpanRole.LLM_CALL,
|
||||
SpanRole.GUARDRAIL,
|
||||
SpanRole.SERVICE,
|
||||
}
|
||||
assert SPAN_REGISTRY[SpanRole.LLM_CALL].kind is LiteLLMSpanKind.CLIENT
|
||||
assert SPAN_REGISTRY[SpanRole.PROXY_REQUEST].kind is LiteLLMSpanKind.SERVER
|
||||
assert SPAN_REGISTRY[SpanRole.GUARDRAIL].parent is SpanRole.PROXY_REQUEST
|
||||
|
||||
|
||||
def test_llm_call_span_name():
|
||||
data = LLMCallSpanData.from_standard_logging_payload(_sample_payload())
|
||||
assert spans_mod.llm_call_span_name(data) == "chat gpt-4o"
|
||||
|
||||
|
||||
# --- semconv (source of truth #1) ------------------------------------------- #
|
||||
|
||||
|
||||
def _all_constants(cls):
|
||||
return {
|
||||
getattr(cls, name)
|
||||
for name in vars(cls)
|
||||
if not name.startswith("__") and isinstance(getattr(cls, name), str)
|
||||
}
|
||||
|
||||
|
||||
def test_attribute_keys_are_unique_across_namespaces():
|
||||
# prefixes are allowed to be substrings; exact keys must not collide.
|
||||
exact = set()
|
||||
for cls in (GenAI, Error, Server, HTTP):
|
||||
for key in _all_constants(cls):
|
||||
assert key not in exact, f"duplicate attribute key {key}"
|
||||
exact.add(key)
|
||||
|
||||
|
||||
def test_provider_resolution():
|
||||
assert resolve_provider("openai") == "openai"
|
||||
assert resolve_provider("bedrock") == "aws.bedrock"
|
||||
assert resolve_provider("vertex_ai") == "gcp.vertex_ai"
|
||||
# unknown providers pass through verbatim (semconv allows provider-specific)
|
||||
assert resolve_provider("my_custom_llm") == "my_custom_llm"
|
||||
assert resolve_provider(None) == ""
|
||||
|
||||
|
||||
def test_operation_resolution():
|
||||
assert resolve_operation("acompletion") is GenAIOperation.CHAT
|
||||
assert resolve_operation("aembedding") is GenAIOperation.EMBEDDINGS
|
||||
assert resolve_operation("atext_completion") is GenAIOperation.TEXT_COMPLETION
|
||||
assert resolve_operation(None) is GenAIOperation.CHAT
|
||||
|
||||
|
||||
# --- typed adapter (source of truth #3) ------------------------------------- #
|
||||
|
||||
|
||||
def test_llm_call_adapter_extracts_all_fields():
|
||||
data = LLMCallSpanData.from_standard_logging_payload(_sample_payload())
|
||||
assert data.operation is GenAIOperation.CHAT
|
||||
assert data.provider == "openai"
|
||||
assert data.request_model == "gpt-4o"
|
||||
assert data.response_model == "gpt-4o-2024"
|
||||
assert data.response_id == "resp_1"
|
||||
assert data.finish_reasons == ("stop",)
|
||||
assert (data.usage.input_tokens, data.usage.output_tokens) == (10, 5)
|
||||
assert data.request_params.temperature == 0.7
|
||||
assert data.request_params.top_k == 40
|
||||
assert data.request_params.stop_sequences == ("STOP",)
|
||||
assert data.request_params.seed == 42
|
||||
assert data.server is not None
|
||||
assert data.server.address == "api.openai.com"
|
||||
assert data.server.port == 443
|
||||
assert data.response_cost == 0.002
|
||||
assert data.error is None
|
||||
assert data.identity.team_id == "t1"
|
||||
assert data.identity.key_hash == "hsh"
|
||||
|
||||
|
||||
def test_llm_call_adapter_failure_path():
|
||||
payload = _sample_payload(
|
||||
status="failure",
|
||||
error_information={
|
||||
"error_class": "RateLimitError",
|
||||
"error_message": "429 slow down",
|
||||
},
|
||||
)
|
||||
data = LLMCallSpanData.from_standard_logging_payload(payload)
|
||||
assert data.error is not None
|
||||
assert data.error.error_type == "RateLimitError"
|
||||
assert data.error.message == "429 slow down"
|
||||
|
||||
|
||||
def test_adapter_is_resilient_to_minimal_payload():
|
||||
data = LLMCallSpanData.from_standard_logging_payload({})
|
||||
assert data.request_model == ""
|
||||
assert data.operation is GenAIOperation.CHAT
|
||||
assert data.server is None
|
||||
assert data.usage.input_tokens is None
|
||||
|
||||
|
||||
def test_content_capture_gated_off_by_default():
|
||||
# ``capture_content`` defaults off: prompt/response bodies must not reach the
|
||||
# span data (and so no vendor mapper can export them) unless explicitly
|
||||
# opted in. Non-content metadata (finish reasons) is still derived.
|
||||
payload = _sample_payload(
|
||||
messages=[{"role": "user", "content": "secret prompt"}],
|
||||
)
|
||||
payload["response"]["choices"] = [
|
||||
{"finish_reason": "stop", "message": {"role": "assistant", "content": "secret"}}
|
||||
]
|
||||
data = LLMCallSpanData.from_standard_logging_payload(payload)
|
||||
assert data.messages_in == ()
|
||||
assert data.choices_out == ()
|
||||
assert data.finish_reasons == ("stop",)
|
||||
|
||||
|
||||
def test_request_identity_prefers_canonical_team_keys():
|
||||
from litellm.integrations.otel.payloads import RequestIdentity
|
||||
|
||||
payload = _sample_payload(
|
||||
metadata={
|
||||
"user_api_key_team_id": "team-canonical",
|
||||
"user_api_key_team_alias": "alias-canonical",
|
||||
"user_api_key_hash": "hsh",
|
||||
"team_id": "legacy-ignored", # legacy alias loses to the canonical key
|
||||
}
|
||||
)
|
||||
ident = RequestIdentity.from_payload(payload)
|
||||
assert ident.team_id == "team-canonical"
|
||||
assert ident.team_alias == "alias-canonical"
|
||||
assert ident.key_hash == "hsh"
|
||||
|
||||
|
||||
def test_request_identity_falls_back_to_legacy_team_keys():
|
||||
from litellm.integrations.otel.payloads import RequestIdentity
|
||||
|
||||
payload = _sample_payload(
|
||||
metadata={"team_id": "legacy-team", "team_alias": "legacy"}
|
||||
)
|
||||
ident = RequestIdentity.from_payload(payload)
|
||||
assert ident.team_id == "legacy-team"
|
||||
assert ident.team_alias == "legacy"
|
||||
|
||||
|
||||
def test_guardrail_span_data_block_carries_verdict_and_error():
|
||||
from litellm.integrations.otel.payloads import GuardrailSpanData
|
||||
|
||||
entry = {
|
||||
"guardrail_name": "openai-moderation",
|
||||
"guardrail_mode": "pre_call",
|
||||
"guardrail_status": "guardrail_intervened",
|
||||
"guardrail_provider": "openai",
|
||||
"guardrail_action": "BLOCKED",
|
||||
"guardrail_response": {"violated_categories": ["violence"]},
|
||||
"violation_categories": ["violence"],
|
||||
"masked_entity_count": {"EMAIL": 2, "PHONE": 1},
|
||||
"duration": 0.05,
|
||||
}
|
||||
d = GuardrailSpanData.from_logging_entry(entry)
|
||||
assert d.guardrail_name == "openai-moderation"
|
||||
assert d.status == "guardrail_intervened"
|
||||
assert d.provider == "openai"
|
||||
assert d.action == "BLOCKED"
|
||||
assert '"violence"' in (d.response_json or "")
|
||||
assert d.violation_categories == ("violence",)
|
||||
assert d.masked_entity_count == 3 # summed across entity types
|
||||
assert d.duration == 0.05
|
||||
assert d.error is not None # intervention → span marked ERROR
|
||||
|
||||
|
||||
def test_guardrail_span_data_success_has_no_error():
|
||||
from litellm.integrations.otel.payloads import GuardrailSpanData
|
||||
|
||||
d = GuardrailSpanData.from_logging_entry(
|
||||
{
|
||||
"guardrail_name": "g",
|
||||
"guardrail_mode": "pre_call",
|
||||
"guardrail_status": "success",
|
||||
}
|
||||
)
|
||||
assert d.error is None
|
||||
assert d.status == "success"
|
||||
|
||||
|
||||
def test_request_identity_from_user_api_key_auth():
|
||||
from litellm.integrations.otel.payloads import RequestIdentity
|
||||
|
||||
class _Auth:
|
||||
team_id = "t9"
|
||||
team_alias = "team nine"
|
||||
api_key = "hashed-key"
|
||||
user_id = "u9"
|
||||
org_id = "o9"
|
||||
key_alias = "my-key"
|
||||
end_user_id = "eu9"
|
||||
|
||||
ident = RequestIdentity.from_user_api_key_auth(_Auth())
|
||||
assert (ident.team_id, ident.team_alias, ident.key_hash) == (
|
||||
"t9",
|
||||
"team nine",
|
||||
"hashed-key",
|
||||
)
|
||||
assert ident.end_user == "eu9"
|
||||
assert ident.metadata["user_api_key_user_id"] == "u9"
|
||||
assert ident.metadata["user_api_key_org_id"] == "o9"
|
||||
assert ident.metadata["user_api_key_alias"] == "my-key"
|
||||
assert ident.metadata["user_api_key_end_user_id"] == "eu9"
|
||||
|
||||
|
||||
def test_content_capture_opt_in_retains_bodies():
|
||||
payload = _sample_payload(
|
||||
messages=[{"role": "user", "content": "secret prompt"}],
|
||||
)
|
||||
payload["response"]["choices"] = [
|
||||
{"finish_reason": "stop", "message": {"role": "assistant", "content": "hi"}}
|
||||
]
|
||||
data = LLMCallSpanData.from_standard_logging_payload(payload, capture_content=True)
|
||||
assert data.messages_in and data.messages_in[0]["content"] == "secret prompt"
|
||||
assert data.choices_out and data.choices_out[0]["message"]["content"] == "hi"
|
||||
|
||||
|
||||
# --- config ----------------------------------------------------------------- #
|
||||
|
||||
|
||||
def test_capture_span_content_resolves_modes():
|
||||
from litellm.integrations.otel.config import (
|
||||
CaptureMessageContent,
|
||||
OpenTelemetryV2Config,
|
||||
)
|
||||
|
||||
# default (no_content) → off
|
||||
assert OpenTelemetryV2Config().capture_span_content is False
|
||||
assert (
|
||||
OpenTelemetryV2Config(
|
||||
capture_message_content=CaptureMessageContent.SPAN_ONLY
|
||||
).capture_span_content
|
||||
is True
|
||||
)
|
||||
assert (
|
||||
OpenTelemetryV2Config(
|
||||
capture_message_content=CaptureMessageContent.SPAN_AND_EVENT
|
||||
).capture_span_content
|
||||
is True
|
||||
)
|
||||
# event-only does not authorize span-attribute content
|
||||
assert (
|
||||
OpenTelemetryV2Config(
|
||||
capture_message_content=CaptureMessageContent.EVENT_ONLY
|
||||
).capture_span_content
|
||||
is False
|
||||
)
|
||||
|
||||
|
||||
def test_v2_flag_is_off_by_default(monkeypatch):
|
||||
monkeypatch.delenv("LITELLM_OTEL_V2", raising=False)
|
||||
assert is_otel_v2_enabled() is False
|
||||
monkeypatch.setenv("LITELLM_OTEL_V2", "true")
|
||||
assert is_otel_v2_enabled() is True
|
||||
|
||||
|
||||
def test_config_from_env(monkeypatch):
|
||||
for var in (
|
||||
"OTEL_EXPORTER",
|
||||
"OTEL_EXPORTER_OTLP_PROTOCOL",
|
||||
"OTEL_ENDPOINT",
|
||||
"OTEL_EXPORTER_OTLP_ENDPOINT",
|
||||
"OTEL_HEADERS",
|
||||
"OTEL_EXPORTER_OTLP_HEADERS",
|
||||
"OTEL_SERVICE_NAME",
|
||||
"LITELLM_OTEL_LEGACY_COMPAT",
|
||||
):
|
||||
monkeypatch.delenv(var, raising=False)
|
||||
|
||||
monkeypatch.setenv("OTEL_EXPORTER_OTLP_ENDPOINT", "https://collector:4318")
|
||||
monkeypatch.setenv("OTEL_SERVICE_NAME", "my-svc")
|
||||
cfg = OpenTelemetryV2Config.from_env()
|
||||
# endpoint with no explicit exporter implies OTLP/HTTP
|
||||
assert cfg.exporter == "otlp_http"
|
||||
assert cfg.endpoint == "https://collector:4318"
|
||||
assert cfg.service_name == "my-svc"
|
||||
assert cfg.legacy_compat is True # dual-emit default during deprecation window
|
||||
|
||||
|
||||
def test_config_legacy_compat_env_toggle(monkeypatch):
|
||||
monkeypatch.setenv("LITELLM_OTEL_LEGACY_COMPAT", "false")
|
||||
assert OpenTelemetryV2Config.from_env().legacy_compat is False
|
||||
|
||||
|
||||
# --- baggage allowlist (the antipattern boundary) --------------------------- #
|
||||
|
||||
|
||||
def test_promoted_baggage_is_bounded_allowlist():
|
||||
identity = RequestIdentity(
|
||||
call_id="c1",
|
||||
team_id="t1",
|
||||
team_alias="team one",
|
||||
key_hash="hsh",
|
||||
end_user="u1",
|
||||
metadata={"user_api_key_org_id": "org1", "secret_blob": "should-not-promote"},
|
||||
)
|
||||
promoted = promoted_baggage(identity, "gpt-4o", BAGGAGE_PROMOTED_KEYS)
|
||||
assert promoted[LiteLLM.TEAM_ID] == "t1"
|
||||
assert promoted[LiteLLM.TEAM_ALIAS] == "team one"
|
||||
assert promoted[GenAI.REQUEST_MODEL] == "gpt-4o"
|
||||
# allowlisted metadata sub-key is promoted under the litellm.metadata.* prefix
|
||||
assert promoted[f"{LiteLLM.METADATA_PREFIX}user_api_key_org_id"] == "org1"
|
||||
# full metadata blob is NOT promoted
|
||||
assert all("secret_blob" not in key for key in promoted)
|
||||
# http.* is never a promoted key
|
||||
assert HTTP.ROUTE not in promoted
|
||||
assert HTTP.REQUEST_METHOD not in promoted
|
||||
|
|
@ -0,0 +1,196 @@
|
|||
"""Tests for the vendor mappers (OpenInference, Langfuse, Weave, Langtrace).
|
||||
|
||||
Composition over inheritance: each vendor's vocabulary is a mapper. Layering
|
||||
mappers on the same span carries multiple naming schemes for different
|
||||
backends, so one trace lights up every configured destination.
|
||||
"""
|
||||
|
||||
import json
|
||||
|
||||
import pytest
|
||||
|
||||
from litellm.integrations.otel import GenAIOperation
|
||||
from litellm.integrations.otel.mappers import (
|
||||
GenAIMapper,
|
||||
LangfuseMapper,
|
||||
LangtraceMapper,
|
||||
OpenInferenceMapper,
|
||||
WeaveMapper,
|
||||
resolve_mappers,
|
||||
)
|
||||
from litellm.integrations.otel.payloads import (
|
||||
LLMCallSpanData,
|
||||
LLMRequestParams,
|
||||
LLMUsage,
|
||||
RequestIdentity,
|
||||
ServerInfo,
|
||||
ToolDefinition,
|
||||
)
|
||||
|
||||
|
||||
def _llm_call(**overrides):
|
||||
base = dict(
|
||||
operation=GenAIOperation.CHAT,
|
||||
provider="openai",
|
||||
request_model="gpt-4o",
|
||||
response_model="gpt-4o-2024",
|
||||
response_id="resp_1",
|
||||
request_params=LLMRequestParams(temperature=0.5, top_p=0.9, max_tokens=128),
|
||||
usage=LLMUsage(input_tokens=12, output_tokens=8, total_tokens=20),
|
||||
finish_reasons=("stop",),
|
||||
error=None,
|
||||
response_cost=0.001,
|
||||
server=ServerInfo("api.openai.com", 443),
|
||||
identity=RequestIdentity(call_id="c1", team_id="t1", team_alias="team one"),
|
||||
is_streaming=False,
|
||||
tools=(
|
||||
ToolDefinition(
|
||||
name="lookup_weather",
|
||||
description="Get weather",
|
||||
parameters_json='{"type":"object"}',
|
||||
),
|
||||
),
|
||||
messages_in=(
|
||||
{"role": "system", "content": "Be concise."},
|
||||
{"role": "user", "content": "What's the weather?"},
|
||||
),
|
||||
choices_out=(
|
||||
{
|
||||
"finish_reason": "stop",
|
||||
"message": {"role": "assistant", "content": "Sunny."},
|
||||
},
|
||||
),
|
||||
system_fingerprint="fp_abc",
|
||||
)
|
||||
base.update(overrides)
|
||||
return LLMCallSpanData(**base)
|
||||
|
||||
|
||||
# --------------------------------------------------------------------------- #
|
||||
# OpenInference (Arize + Phoenix shared vocabulary)
|
||||
# --------------------------------------------------------------------------- #
|
||||
|
||||
|
||||
def test_openinference_mapper_input_output_messages():
|
||||
attrs = OpenInferenceMapper().map(_llm_call())
|
||||
assert attrs["openinference.span.kind"] == "LLM"
|
||||
assert attrs["llm.model_name"] == "gpt-4o"
|
||||
assert attrs["llm.provider"] == "openai"
|
||||
assert attrs["llm.input_messages.0.message.role"] == "system"
|
||||
assert attrs["llm.input_messages.0.message.content"] == "Be concise."
|
||||
assert attrs["llm.input_messages.1.message.role"] == "user"
|
||||
assert attrs["llm.output_messages.0.message.role"] == "assistant"
|
||||
assert attrs["llm.output_messages.0.message.content"] == "Sunny."
|
||||
assert attrs["llm.token_count.prompt"] == 12
|
||||
assert attrs["llm.token_count.completion"] == 8
|
||||
assert attrs["llm.token_count.total"] == 20
|
||||
# tool definitions ride the OpenInference schema
|
||||
assert attrs["llm.tools.0.tool.name"] == "lookup_weather"
|
||||
# invocation_parameters is JSON-serialized
|
||||
params = json.loads(attrs["llm.invocation_parameters"])
|
||||
assert params["temperature"] == 0.5
|
||||
assert params["max_tokens"] == 128
|
||||
|
||||
|
||||
def test_openinference_mapper_skips_non_llm_roles():
|
||||
from litellm.integrations.otel.payloads import GuardrailSpanData
|
||||
|
||||
assert OpenInferenceMapper().map(GuardrailSpanData("presidio")) == {}
|
||||
|
||||
|
||||
def test_openinference_multimodal_content_text_only():
|
||||
data = _llm_call(
|
||||
messages_in=(
|
||||
{
|
||||
"role": "user",
|
||||
"content": [
|
||||
{"type": "text", "text": "hi "},
|
||||
{"type": "image_url", "image_url": {"url": "x"}},
|
||||
{"type": "text", "text": "there"},
|
||||
],
|
||||
},
|
||||
)
|
||||
)
|
||||
attrs = OpenInferenceMapper().map(data)
|
||||
assert attrs["llm.input_messages.0.message.content"] == "hi there"
|
||||
|
||||
|
||||
# --------------------------------------------------------------------------- #
|
||||
# Langfuse
|
||||
# --------------------------------------------------------------------------- #
|
||||
|
||||
|
||||
def test_langfuse_mapper_observation_attrs():
|
||||
attrs = LangfuseMapper().map(_llm_call())
|
||||
assert attrs["langfuse.observation.type"] == "generation"
|
||||
assert attrs["langfuse.observation.model.name"] == "gpt-4o"
|
||||
assert attrs["langfuse.observation.metadata.provider"] == "openai"
|
||||
usage = json.loads(attrs["langfuse.observation.usage_details"])
|
||||
assert usage["input"] == 12 and usage["output"] == 8
|
||||
params = json.loads(attrs["langfuse.observation.model.parameters"])
|
||||
assert params["temperature"] == 0.5
|
||||
cost = json.loads(attrs["langfuse.observation.cost_details"])
|
||||
assert cost["total"] == 0.001
|
||||
assert attrs["langfuse.trace.metadata.team_id"] == "t1"
|
||||
|
||||
|
||||
def test_langfuse_mapper_skips_when_no_messages():
|
||||
data = _llm_call(messages_in=(), choices_out=())
|
||||
attrs = LangfuseMapper().map(data)
|
||||
assert "langfuse.observation.input" not in attrs
|
||||
assert "langfuse.observation.output" not in attrs
|
||||
|
||||
|
||||
# --------------------------------------------------------------------------- #
|
||||
# Weave
|
||||
# --------------------------------------------------------------------------- #
|
||||
|
||||
|
||||
def test_weave_mapper_display_and_output():
|
||||
attrs = WeaveMapper().map(_llm_call())
|
||||
assert attrs["weave.display_name"] == "chat gpt-4o"
|
||||
assert attrs["weave.call_id"] == "c1"
|
||||
decoded = json.loads(attrs["weave.output"])
|
||||
assert decoded[0]["message"]["content"] == "Sunny."
|
||||
|
||||
|
||||
# --------------------------------------------------------------------------- #
|
||||
# Langtrace
|
||||
# --------------------------------------------------------------------------- #
|
||||
|
||||
|
||||
def test_langtrace_mapper_attrs():
|
||||
attrs = LangtraceMapper().map(_llm_call())
|
||||
assert attrs["gen_ai.operation.name"] == "chat"
|
||||
assert attrs["langtrace.service.name"] == "openai"
|
||||
assert attrs["llm.model"] == "gpt-4o"
|
||||
assert attrs["gen_ai.response.model"] == "gpt-4o-2024"
|
||||
assert attrs["gen_ai.system_fingerprint"] == "fp_abc"
|
||||
assert attrs["llm.temperature"] == 0.5
|
||||
assert attrs["llm.token.counts.total"] == 20
|
||||
|
||||
|
||||
# --------------------------------------------------------------------------- #
|
||||
# Composition (the V2 punchline)
|
||||
# --------------------------------------------------------------------------- #
|
||||
|
||||
|
||||
def test_resolve_mappers_composition_layers_vocabularies():
|
||||
"""One span, three vocabularies — Arize + Langfuse + canonical together."""
|
||||
chain = resolve_mappers(["genai", "openinference", "langfuse"])
|
||||
data = _llm_call()
|
||||
union: dict = {}
|
||||
for mapper in chain:
|
||||
union.update(mapper.map(data))
|
||||
# Canonical
|
||||
assert union["gen_ai.operation.name"] == "chat"
|
||||
# OpenInference
|
||||
assert union["llm.model_name"] == "gpt-4o"
|
||||
assert union["openinference.span.kind"] == "LLM"
|
||||
# Langfuse
|
||||
assert union["langfuse.observation.type"] == "generation"
|
||||
|
||||
|
||||
def test_resolve_mappers_rejects_unknown_name():
|
||||
with pytest.raises(ValueError, match="unknown mapper name 'nope'"):
|
||||
resolve_mappers(["genai", "nope"])
|
||||
52
uv.lock
generated
52
uv.lock
generated
|
|
@ -9,7 +9,7 @@ resolution-markers = [
|
|||
]
|
||||
|
||||
[options]
|
||||
exclude-newer = "2026-05-25T20:42:18.420988002Z"
|
||||
exclude-newer = "0001-01-01T00:00:00Z" # This has no effect and is included for backwards compatibility when using relative exclude-newer values.
|
||||
exclude-newer-span = "P3D"
|
||||
|
||||
[manifest]
|
||||
|
|
@ -294,6 +294,18 @@ wheels = [
|
|||
{ url = "https://files.pythonhosted.org/packages/ee/82/82745642d3c46e7cea25e1885b014b033f4693346ce46b7f47483cf5d448/argon2_cffi_bindings-25.1.0-pp310-pypy310_pp73-win_amd64.whl", hash = "sha256:da0c79c23a63723aa5d782250fbf51b768abca630285262fb5144ba5ae01e520", size = 29187, upload-time = "2025-07-30T10:02:03.674Z" },
|
||||
]
|
||||
|
||||
[[package]]
|
||||
name = "asgiref"
|
||||
version = "3.11.1"
|
||||
source = { registry = "https://pypi.org/simple" }
|
||||
dependencies = [
|
||||
{ name = "typing-extensions", marker = "python_full_version < '3.11'" },
|
||||
]
|
||||
sdist = { url = "https://files.pythonhosted.org/packages/63/40/f03da1264ae8f7cfdbf9146542e5e7e8100a4c66ab48e791df9a03d3f6c0/asgiref-3.11.1.tar.gz", hash = "sha256:5f184dc43b7e763efe848065441eac62229c9f7b0475f41f80e207a114eda4ce", size = 38550, upload-time = "2026-02-03T13:30:14.33Z" }
|
||||
wheels = [
|
||||
{ url = "https://files.pythonhosted.org/packages/5c/0a/a72d10ed65068e115044937873362e6e32fab1b7dce0046aeb224682c989/asgiref-3.11.1-py3-none-any.whl", hash = "sha256:e8667a091e69529631969fd45dc268fa79b99c92c5fcdda727757e52146ec133", size = 24345, upload-time = "2026-02-03T13:30:13.039Z" },
|
||||
]
|
||||
|
||||
[[package]]
|
||||
name = "assemblyai"
|
||||
version = "0.52.4"
|
||||
|
|
@ -3349,6 +3361,7 @@ proxy-runtime = [
|
|||
{ name = "mangum" },
|
||||
{ name = "opentelemetry-api" },
|
||||
{ name = "opentelemetry-exporter-otlp" },
|
||||
{ name = "opentelemetry-instrumentation-fastapi" },
|
||||
{ name = "opentelemetry-sdk" },
|
||||
{ name = "prometheus-client" },
|
||||
{ name = "pypdf" },
|
||||
|
|
@ -3411,6 +3424,7 @@ dev = [
|
|||
{ name = "openapi-core" },
|
||||
{ name = "opentelemetry-api" },
|
||||
{ name = "opentelemetry-exporter-otlp" },
|
||||
{ name = "opentelemetry-instrumentation-fastapi" },
|
||||
{ name = "opentelemetry-sdk" },
|
||||
{ name = "parameterized" },
|
||||
{ name = "psycopg" },
|
||||
|
|
@ -3444,6 +3458,7 @@ proxy-dev = [
|
|||
{ name = "hypercorn" },
|
||||
{ name = "opentelemetry-api" },
|
||||
{ name = "opentelemetry-exporter-otlp" },
|
||||
{ name = "opentelemetry-instrumentation-fastapi" },
|
||||
{ name = "opentelemetry-sdk" },
|
||||
{ name = "prisma" },
|
||||
{ name = "prometheus-client" },
|
||||
|
|
@ -3499,6 +3514,7 @@ requires-dist = [
|
|||
{ name = "openai", specifier = ">=2.20.0,<3.0.0" },
|
||||
{ name = "opentelemetry-api", marker = "extra == 'proxy-runtime'", specifier = "==1.28.0" },
|
||||
{ name = "opentelemetry-exporter-otlp", marker = "extra == 'proxy-runtime'", specifier = "==1.28.0" },
|
||||
{ name = "opentelemetry-instrumentation-fastapi", marker = "extra == 'proxy-runtime'", specifier = "==0.49b0" },
|
||||
{ name = "opentelemetry-sdk", marker = "extra == 'proxy-runtime'", specifier = "==1.28.0" },
|
||||
{ name = "orjson", marker = "extra == 'proxy'", specifier = ">=3.11.6,<4.0" },
|
||||
{ name = "polars", marker = "extra == 'proxy'", specifier = ">=1.38.1,<2.0" },
|
||||
|
|
@ -3573,6 +3589,7 @@ dev = [
|
|||
{ name = "openapi-core", marker = "python_full_version < '3.14'", specifier = "==0.22.0" },
|
||||
{ name = "opentelemetry-api", specifier = "==1.28.0" },
|
||||
{ name = "opentelemetry-exporter-otlp", specifier = "==1.28.0" },
|
||||
{ name = "opentelemetry-instrumentation-fastapi", specifier = "==0.49b0" },
|
||||
{ name = "opentelemetry-sdk", specifier = "==1.28.0" },
|
||||
{ name = "parameterized", specifier = "==0.9.0" },
|
||||
{ name = "psycopg", specifier = "==3.3.3" },
|
||||
|
|
@ -3606,6 +3623,7 @@ proxy-dev = [
|
|||
{ name = "hypercorn", specifier = "==0.17.3" },
|
||||
{ name = "opentelemetry-api", specifier = "==1.28.0" },
|
||||
{ name = "opentelemetry-exporter-otlp", specifier = "==1.28.0" },
|
||||
{ name = "opentelemetry-instrumentation-fastapi", specifier = "==0.49b0" },
|
||||
{ name = "opentelemetry-sdk", specifier = "==1.28.0" },
|
||||
{ name = "prisma", specifier = "==0.11.0" },
|
||||
{ name = "prometheus-client", specifier = "==0.20.0" },
|
||||
|
|
@ -4561,6 +4579,22 @@ wheels = [
|
|||
{ url = "https://files.pythonhosted.org/packages/ba/46/ba2dc8d18b04acae3d34facd8fe1e5e0cdc9fe64292d45eca9d1d4a8a298/opentelemetry_instrumentation_anthropic-0.33.12-py3-none-any.whl", hash = "sha256:b31618d12a429045db14ed982a142a25df0f0f1dbf03d756e8d597f25b9a053d", size = 11024, upload-time = "2024-11-13T20:27:14.622Z" },
|
||||
]
|
||||
|
||||
[[package]]
|
||||
name = "opentelemetry-instrumentation-asgi"
|
||||
version = "0.49b0"
|
||||
source = { registry = "https://pypi.org/simple" }
|
||||
dependencies = [
|
||||
{ name = "asgiref" },
|
||||
{ name = "opentelemetry-api" },
|
||||
{ name = "opentelemetry-instrumentation" },
|
||||
{ name = "opentelemetry-semantic-conventions" },
|
||||
{ name = "opentelemetry-util-http" },
|
||||
]
|
||||
sdist = { url = "https://files.pythonhosted.org/packages/e8/55/693c3d0938ba5fead5c3aa4ac7022a992b4ff99a8e9979800d0feb843ff4/opentelemetry_instrumentation_asgi-0.49b0.tar.gz", hash = "sha256:959fd9b1345c92f20c6ef1d42f92ef6a76b3c3083fbc4104d59da6859b15b083", size = 24117, upload-time = "2024-11-05T19:21:46.769Z" }
|
||||
wheels = [
|
||||
{ url = "https://files.pythonhosted.org/packages/2c/0b/7900c782a1dfaa584588d724bc3bbdf8405a32497537dd96b3fcbf8461b9/opentelemetry_instrumentation_asgi-0.49b0-py3-none-any.whl", hash = "sha256:722a90856457c81956c88f35a6db606cc7db3231046b708aae2ddde065723dbe", size = 16326, upload-time = "2024-11-05T19:20:46.176Z" },
|
||||
]
|
||||
|
||||
[[package]]
|
||||
name = "opentelemetry-instrumentation-bedrock"
|
||||
version = "0.33.12"
|
||||
|
|
@ -4607,6 +4641,22 @@ wheels = [
|
|||
{ url = "https://files.pythonhosted.org/packages/bf/08/dce2b7926ace0204ce7946563348e1ff755873e387833484791e4ed391c8/opentelemetry_instrumentation_cohere-0.33.12-py3-none-any.whl", hash = "sha256:3bee3f7f7105259c85145be8c3b68612421860c95ad170f4d03144a3b8c07418", size = 5589, upload-time = "2024-11-13T20:27:21.317Z" },
|
||||
]
|
||||
|
||||
[[package]]
|
||||
name = "opentelemetry-instrumentation-fastapi"
|
||||
version = "0.49b0"
|
||||
source = { registry = "https://pypi.org/simple" }
|
||||
dependencies = [
|
||||
{ name = "opentelemetry-api" },
|
||||
{ name = "opentelemetry-instrumentation" },
|
||||
{ name = "opentelemetry-instrumentation-asgi" },
|
||||
{ name = "opentelemetry-semantic-conventions" },
|
||||
{ name = "opentelemetry-util-http" },
|
||||
]
|
||||
sdist = { url = "https://files.pythonhosted.org/packages/fe/bf/8e6d2a4807360f2203192017eb4845f5628dbeaf0597adf3d141cc5c24e1/opentelemetry_instrumentation_fastapi-0.49b0.tar.gz", hash = "sha256:6d14935c41fd3e49328188b6a59dd4c37bd17a66b01c15b0c64afa9714a1f905", size = 19230, upload-time = "2024-11-05T19:21:59.361Z" }
|
||||
wheels = [
|
||||
{ url = "https://files.pythonhosted.org/packages/b1/f4/0895b9410c10abf987c90dee1b7688a8f2214a284fe15e575648f6a1473a/opentelemetry_instrumentation_fastapi-0.49b0-py3-none-any.whl", hash = "sha256:646e1b18523cbe6860ae9711eb2c7b9c85466c3c7697cd6b8fb5180d85d3fe6e", size = 12101, upload-time = "2024-11-05T19:21:01.805Z" },
|
||||
]
|
||||
|
||||
[[package]]
|
||||
name = "opentelemetry-instrumentation-google-generativeai"
|
||||
version = "0.33.12"
|
||||
|
|
|
|||
Loading…
Add table
Reference in a new issue