diff --git a/litellm/integrations/otel/README.md b/litellm/integrations/otel/README.md new file mode 100644 index 00000000000..41c09344b33 --- /dev/null +++ b/litellm/integrations/otel/README.md @@ -0,0 +1,155 @@ +# OpenTelemetry instrumentation + +This package produces OpenTelemetry traces for LiteLLM. It is enabled by the +`LITELLM_OTEL_V2` environment variable (`is_otel_v2_enabled()` in +[`config.py`](./config.py)); when unset, nothing in this package runs. + +## What gets traced + +A traced proxy request produces one trace with two kinds of spans: + +``` +SERVER span "POST /v1/chat/completions" ← FastAPI instrumentation +├── CLIENT span "chat gpt-4o" ← LLM call ┐ +├── INTERNAL span "execute_guardrail …" ← guardrail │ this package +└── INTERNAL span "redis" … ← service call ┘ +``` + +The gen-ai spans are siblings under the server span. In particular the guardrail +span is a sibling of the LLM call, not a child of it: pre/during/post-call +guardrail hooks are part of the request lifecycle (a pre-call guardrail runs +before the LLM call even starts), so they parent to the server span via the +ambient OpenTelemetry context, alongside the LLM call. + +- **Server spans** (one per HTTP route) are created by the + `opentelemetry-instrumentation-fastapi` package. It stamps `http.*` attributes + and extracts inbound `traceparent` headers. This package does **not** create + or modify server spans — request routes never touch spans. +- **Gen-AI spans** (LLM calls, guardrails, internal service calls) are created + by this package from LiteLLM's logging callbacks and parent to the active + server span via ambient OpenTelemetry context. + +Both kinds share a single `TracerProvider`, so they belong to the same trace +and export through the same configured exporters. FastAPI middleware can only be +added before the app starts serving, so the app is instrumented at +import time **without** a provider — it binds to the OTel global +`ProxyTracerProvider`. Once config (and the callbacks) is loaded, the proxy +publishes the chosen logger's `TracerProvider` as the global via +`trace.set_tracer_provider(...)`, and the server spans delegate to it. When a +preset callback (`arize`, `langfuse_otel`, …) is configured, its provider +becomes the global, so server spans export to that backend too. + +## How a request flows + +1. **App creation** (`proxy_server` import): when the gate is on, + `FastAPIInstrumentor.instrument_app(app)` is called with no provider (the + middleware stack is frozen once the app serves, so this can't wait for + startup). It binds to the OTel global `ProxyTracerProvider`. Health-check + routes (`/health*`) are excluded by default so load-balancer polling doesn't + flood traces; set `OTEL_PYTHON_FASTAPI_EXCLUDED_URLS` to override (e.g. `""` + to trace everything, or your own comma-separated path list). +2. **Startup** (`proxy_server.proxy_startup_event`): after the config (and + callbacks) is loaded, the already-registered preset `OpenTelemetryV2` logger + is reused — or a generic one reading `OTEL_*` envs is built when no preset is + configured — and its `TracerProvider` is published as the OTel global with + `trace.set_tracer_provider(...)`. The proxy tracer then delegates to it, so + server spans and gen-ai spans share one provider and the same trace. +3. **Request**: the FastAPI instrumentation starts the server span and makes it + the active context for the request task. +4. **LLM call logging**: LiteLLM's async logging worker copies the request's + context when it enqueues the success/failure callback, so + `OpenTelemetryV2.async_log_success_event` runs with the server span as the + ambient parent. It builds an `LLMCallSpanData` from the request's + `standard_logging_object` and hands it to the engine, which creates the LLM + span as a child of the server span. Emission is **async-only**; the + synchronous callback runs in a worker thread without the request context and + is a no-op. +5. **Guardrails / services**: the post-call and service hooks emit guardrail and + service spans the same way — typed data → engine → span. +6. **Export**: each span ends and is handed to the provider's span processors, + which export to the configured backends (OTLP, console, in-memory, …). + +## Components + +### Sources of truth (no OpenTelemetry import) + +These define the shape of a span without depending on the OTel SDK, so they can +be imported anywhere: + +- [`semconv.py`](./semconv.py) — attribute-key constants (`gen_ai.*`, `http.*`, + `litellm.*`), the GenAI operation/provider enums, and the functions that map + LiteLLM provider/call-type strings onto convention values. +- [`spans.py`](./spans.py) — the span registry: every span role, its OTel span + kind, its place in the hierarchy, and its name builder. +- [`payloads.py`](./payloads.py) — frozen dataclasses (`LLMCallSpanData`, + `GuardrailSpanData`, `ServiceSpanData`, …) built from heterogeneous logging + payloads via `from_*` classmethods. +- [`config.py`](./config.py) — `OpenTelemetryV2Config`, a pydantic-settings + model that reads `OTEL_*` / `LITELLM_OTEL_*` env vars, plus the feature gate. + `capture_span_content` gates whether prompt/response bodies may be written as + span attributes; it defaults **off** (`no_content`). + +### Engine + +- [`emitter.py`](./emitter.py) — `SpanEmitter.emit(role, data)`: dedupe → start + the span → run the mapper chain to stamp attributes → set status → end. It + owns no attribute keys. The dedupe set (which coalesces the sync+async firing + of one request) is a bounded LRU so it can't grow without limit. +- [`mappers/`](./mappers) — each mapper turns typed span data into a flat + `{attribute key: value}` dict. They compose: listing several mapper names in + the config layers multiple attribute vocabularies onto the same span. + - `genai` — the canonical OpenTelemetry GenAI vocabulary, always present. + - `legacy` — an additional vocabulary using the older semconv-ai / Traceloop + attribute key names, for backends that read those. + - `openinference`, `langfuse`, `weave`, `langtrace` — vendor vocabularies. + - `resolve_mappers(names)` turns config names into mapper instances. + +### Plumbing + +- [`providers.py`](./providers.py) — builds the `TracerProvider`, its exporters + (from `ExporterSpec`s), and the span processor that copies allowlisted Baggage + entries onto every span. `register_exporter_factory(kind, factory)` lets a + preset contribute a custom exporter `kind` (e.g. one that fetches an auth + token lazily) without coupling this module to any vendor. +- [`context.py`](./context.py) — trace-context and Baggage read/write helpers. +- [`baggage.py`](./baggage.py) — the single definition of which request-identity + values are promoted into Baggage (so child spans inherit them) and under which + attribute keys. +- [`routing.py`](./routing.py) — `TenantTracerCache`: when a request carries + team/key-scoped vendor credentials, route its spans through a credential-keyed + `TracerProvider` so one logger serves many tenants. The cache is a bounded LRU + that flushes + shuts down evicted providers, since the key derives from + request-supplied credentials and must not grow (or leak threads) without limit. +- [`utils.py`](./utils.py) — value coercion, JSON serialization, and + extractor-table application, shared across the package. +- [`metrics.py`](./metrics.py) — GenAI client metric instruments. + +### Adapter + +- [`logger.py`](./logger.py) — `OpenTelemetryV2`, a `CustomLogger` that + translates LiteLLM's logging callbacks into typed span data and hands them to + the engine. + +### Presets + +- [`presets/`](./presets) — each preset reads one integration's env vars and + returns an `OpenTelemetryV2Config` (exporter destination + mapper vocabularies + + resource attributes). `PRESET_BY_CALLBACK` maps a callback name (`"arize"`, + `"langfuse_otel"`, …) to its preset. Integrations that support team/key-scoped + credentials also provide a per-request OTLP header builder + (`DYNAMIC_HEADERS_BY_CALLBACK`). Presets do **no** network I/O at build time: + AgentOps, for example, mints its JWT lazily inside a custom exporter on the + first export (in the `BatchSpanProcessor` worker thread), never on the event + loop. + +## Extending + +- **A new attribute vocabulary for a backend**: add a mapper in `mappers/` + (a class with a `map(data) -> AttributeMap` method, typically built from + `key -> extractor` tables) and register it in `mappers/__init__._MAPPER_BY_NAME`. +- **A new integration**: add a preset in `presets/` that returns an + `OpenTelemetryV2Config`, and register it in `presets/__init__.PRESET_BY_CALLBACK`. + If it supports dynamic credentials, add a header builder to + `DYNAMIC_HEADERS_BY_CALLBACK`. +- **A new span kind**: add a role to `spans.py` (registry entry + name builder), + a payload dataclass in `payloads.py`, and a branch in the relevant mapper(s). diff --git a/litellm/integrations/otel/__init__.py b/litellm/integrations/otel/__init__.py new file mode 100644 index 00000000000..2134b2edba1 --- /dev/null +++ b/litellm/integrations/otel/__init__.py @@ -0,0 +1,92 @@ +"""Typed, semconv-aligned OpenTelemetry instrumentation for LiteLLM. + +The three sources of truth — attribute keys (:mod:`semconv`), the span and +hierarchy registry (:mod:`spans`), and the typed span-data inputs +(:mod:`payloads`) — plus :mod:`config` are exported here and are free of any +``opentelemetry`` import. The engine layer (``emitter``, ``providers``, +``context``, ``metrics``) and the ``CustomLogger`` adapter (``logger``) are +reached via their submodule paths so that importing this package never +requires the OTel SDK. + +The ``LITELLM_OTEL_V2`` env var gates whether the factory in +``litellm_core_utils.litellm_logging`` constructs the ``OpenTelemetryV2`` +class (from :mod:`logger`). +""" + +from litellm.integrations.otel.config import ( + OTEL_V2_ENV, + OpenTelemetryV2Config, + is_otel_v2_enabled, +) +from litellm.integrations.otel.baggage import ( + BAGGAGE_PROMOTED_KEYS, + DEFAULT_BAGGAGE_METADATA_KEYS, + promoted_baggage, +) +from litellm.integrations.otel.payloads import ( + GuardrailSpanData, + LLMCallSpanData, + LLMRequestParams, + LLMUsage, + ProxyRequestSpanData, + RequestIdentity, + ServerInfo, + ServiceSpanData, + SpanError, +) +from litellm.integrations.otel.semconv import ( + Error, + GenAI, + GenAIOperation, + GenAIProvider, + HTTP, + LiteLLM, + Metric, + Server, + resolve_operation, + resolve_provider, +) +from litellm.integrations.otel.spans import ( + SPAN_REGISTRY, + LiteLLMSpanKind, + SpanRole, + SpanSpec, + validate_registry, +) + +__all__ = [ + # config + "OTEL_V2_ENV", + "OpenTelemetryV2Config", + "is_otel_v2_enabled", + # semconv + "BAGGAGE_PROMOTED_KEYS", + "DEFAULT_BAGGAGE_METADATA_KEYS", + "Error", + "GenAI", + "GenAIOperation", + "GenAIProvider", + "HTTP", + "LiteLLM", + "Metric", + "Server", + "resolve_operation", + "resolve_provider", + # spans + "SPAN_REGISTRY", + "LiteLLMSpanKind", + "SpanRole", + "SpanSpec", + "validate_registry", + # payloads + "GuardrailSpanData", + "LLMCallSpanData", + "LLMRequestParams", + "LLMUsage", + "ProxyRequestSpanData", + "RequestIdentity", + "ServerInfo", + "ServiceSpanData", + "SpanError", + "promoted_baggage", +] diff --git a/litellm/integrations/otel/baggage.py b/litellm/integrations/otel/baggage.py new file mode 100644 index 00000000000..12a3d6d1176 --- /dev/null +++ b/litellm/integrations/otel/baggage.py @@ -0,0 +1,72 @@ +"""Baggage promotion: request-identity values carried across child spans. + +A bounded set of identity values is written into OpenTelemetry Baggage on the +LLM-call span so that child spans (guardrail, service) inherit them. +``providers.LiteLLMBaggageSpanProcessor`` reads Baggage at span start and stamps +the allowlisted keys onto every span. + +This module is the single place baggage is defined: ``_PROMOTABLE`` maps each +promotable attribute key to how its value is read, and the two ``*_KEYS`` +defaults select what is promoted unless the config overrides them. +""" + +from collections.abc import Callable +from typing import Final + +from litellm.integrations.otel.payloads import RequestIdentity +from litellm.integrations.otel.semconv import GenAI, LiteLLM + +# Attribute key -> value extractor over (identity, request_model). The single +# definition of what may be promoted and under which key. +_PROMOTABLE: Final[dict[str, Callable[[RequestIdentity, str | None], str | None]]] = { + LiteLLM.TEAM_ID: lambda identity, model: identity.team_id, + LiteLLM.TEAM_ALIAS: lambda identity, model: identity.team_alias, + LiteLLM.KEY_HASH: lambda identity, model: identity.key_hash, + LiteLLM.END_USER: lambda identity, model: identity.end_user, + GenAI.REQUEST_MODEL: lambda identity, model: model, +} + +# Keys promoted by default (a subset of ``_PROMOTABLE``). ``END_USER`` is +# promotable but off by default — it identifies an individual user, so stamping +# it onto every span is opt-in via ``config.baggage_promoted_keys``. +BAGGAGE_PROMOTED_KEYS: Final[tuple[str, ...]] = ( + LiteLLM.TEAM_ID, + LiteLLM.TEAM_ALIAS, + LiteLLM.KEY_HASH, + GenAI.REQUEST_MODEL, +) + +# Metadata sub-keys eligible for promotion under the ``litellm.metadata.*`` +# namespace. The full metadata blob is never promoted; only this allowlist is. +DEFAULT_BAGGAGE_METADATA_KEYS: Final[tuple[str, ...]] = ( + "user_api_key_org_id", + "user_api_key_user_id", + "user_api_key_alias", + "user_api_key_end_user_id", + "requester_ip_address", +) + + +def promoted_baggage( + identity: RequestIdentity, + request_model: str | None, + promoted_keys: tuple[str, ...], + metadata_keys: tuple[str, ...] = DEFAULT_BAGGAGE_METADATA_KEYS, +) -> dict[str, str]: + """Identity values to write into Baggage, filtered to ``promoted_keys``. + + ``promoted_keys`` selects from ``_PROMOTABLE``; ``metadata_keys`` selects + sub-keys of ``identity.metadata`` to promote under ``litellm.metadata.*``. + Empty values are dropped. + """ + out: dict[str, str] = {} + for key, extract in _PROMOTABLE.items(): + if key in promoted_keys: + value = extract(identity, request_model) + if value: + out[key] = value + for meta_key in metadata_keys: + value = identity.metadata.get(meta_key) + if value: + out[f"{LiteLLM.METADATA_PREFIX}{meta_key}"] = value + return out diff --git a/litellm/integrations/otel/config.py b/litellm/integrations/otel/config.py new file mode 100644 index 00000000000..e8749e31343 --- /dev/null +++ b/litellm/integrations/otel/config.py @@ -0,0 +1,194 @@ +"""Typed configuration for the OpenTelemetry instrumentation.""" + +from pydantic import AliasChoices, BaseModel, Field, model_validator +from pydantic_settings import BaseSettings, SettingsConfigDict + +from litellm.integrations.otel.baggage import ( + BAGGAGE_PROMOTED_KEYS, + DEFAULT_BAGGAGE_METADATA_KEYS, +) + +#: Master feature-flag env var. The logger is inert until this is truthy. +OTEL_V2_ENV = "LITELLM_OTEL_V2" + + +class CaptureMessageContent(str): + NO_CONTENT = "no_content" + SPAN_ONLY = "span_only" + EVENT_ONLY = "event_only" + SPAN_AND_EVENT = "span_and_event" + + +class _OTelV2Flag(BaseSettings): + model_config = SettingsConfigDict(extra="ignore") + + enabled: bool = Field(default=False, validation_alias=AliasChoices(OTEL_V2_ENV)) + + +def is_otel_v2_enabled() -> bool: + return _OTelV2Flag().enabled + + +class ExporterSpec(BaseModel): + """One span-export destination. + + The shared ``TracerProvider`` attaches one ``SpanProcessor`` per spec, so + listing several specs sends every span to all of them at once (e.g. Arize + + Phoenix + your own Honeycomb). + """ + + model_config = {"extra": "forbid"} + + kind: str = Field( + default="console", + description="console | in_memory | otlp_http | otlp_grpc | ", + ) + endpoint: str | None = None + headers: str | None = None + options: dict[str, str] | None = Field( + default=None, + description=( + "Factory-specific configuration for a custom exporter ``kind`` " + "registered via ``providers.register_exporter_factory`` (e.g. an " + "API key a lazy-auth exporter fetches a token with). Ignored by the " + "built-in console/in_memory/otlp exporters." + ), + ) + use_simple_processor: bool | None = Field( + default=None, + description=( + "Force SimpleSpanProcessor regardless of exporter kind. Default: " + "auto (Simple for console/in_memory, Batch otherwise)." + ), + ) + + +class OpenTelemetryV2Config(BaseSettings): + model_config = SettingsConfigDict(populate_by_name=True, extra="ignore") + + # ----- single-destination shorthand, read from standard OTEL_* envs ----- # + exporter: str = Field( + default="console", + validation_alias=AliasChoices("OTEL_EXPORTER", "OTEL_EXPORTER_OTLP_PROTOCOL"), + description=( + "Exporter kind for the single-destination shorthand. The model " + "validator folds this (with ``endpoint`` / ``headers``) into a " + "one-entry ``exporters`` list when ``exporters`` is empty; set " + "``exporters`` directly for multiple destinations." + ), + ) + endpoint: str | None = Field( + default=None, + validation_alias=AliasChoices("OTEL_ENDPOINT", "OTEL_EXPORTER_OTLP_ENDPOINT"), + ) + headers: str | None = Field( + default=None, + validation_alias=AliasChoices("OTEL_HEADERS", "OTEL_EXPORTER_OTLP_HEADERS"), + ) + service_name: str = Field( + default="litellm", validation_alias=AliasChoices("OTEL_SERVICE_NAME") + ) + deployment_environment: str | None = Field( + default=None, validation_alias=AliasChoices("OTEL_ENVIRONMENT_NAME") + ) + + enable_metrics: bool = Field( + default=False, + validation_alias=AliasChoices("LITELLM_OTEL_INTEGRATION_ENABLE_METRICS"), + ) + enable_events: bool = Field( + default=False, + validation_alias=AliasChoices("LITELLM_OTEL_INTEGRATION_ENABLE_EVENTS"), + ) + capture_message_content: str = Field( + default=CaptureMessageContent.NO_CONTENT, + validation_alias=AliasChoices( + "OTEL_INSTRUMENTATION_GENAI_CAPTURE_MESSAGE_CONTENT" + ), + ) + legacy_compat: bool = Field( + default=True, validation_alias=AliasChoices("LITELLM_OTEL_LEGACY_COMPAT") + ) + + # ----- explicit multi-destination / vocabulary configuration ------------ # + + exporters: list[ExporterSpec] = Field( + default_factory=list, + description=( + "One destination per spec. The shared TracerProvider attaches a " + "SpanProcessor per entry. When empty, the model validator folds " + "the ``exporter`` / ``endpoint`` / ``headers`` shorthand into a " + "single spec so there is always at least one destination." + ), + ) + + mapper_names: list[str] = Field( + default_factory=lambda: ["genai"], + description=( + "Ordered attribute vocabularies to emit. ``genai`` is the " + "canonical OTel GenAI vocabulary and is always placed first. " + "Vendor names: ``openinference`` (Arize + Phoenix), ``langfuse``, " + "``weave``, ``langtrace``." + ), + ) + + resource_attributes: dict[str, str] = Field( + default_factory=dict, + description=( + "Extra Resource attributes beyond ``service.name`` and " + "``deployment.environment`` (e.g. integration-specific markers)." + ), + ) + + baggage_promoted_keys: list[str] = Field( + default_factory=lambda: list(BAGGAGE_PROMOTED_KEYS) + ) + baggage_metadata_keys: list[str] = Field( + default_factory=lambda: list(DEFAULT_BAGGAGE_METADATA_KEYS) + ) + + @model_validator(mode="after") + def _normalize(self) -> "OpenTelemetryV2Config": + # An endpoint with the default exporter kind implies OTLP/HTTP. + if self.endpoint and self.exporter == "console": + self.exporter = "otlp_http" + # When no explicit destinations are given, fold the single-destination + # shorthand into one spec so the provider always has a destination. + if not self.exporters: + self.exporters = [ + ExporterSpec( + kind=self.exporter, + endpoint=self.endpoint, + headers=self.headers, + ) + ] + # Ensure ``genai`` is always present and first. + names = list(self.mapper_names) + if "genai" in names: + names = ["genai"] + [n for n in names if n != "genai"] + else: + names = ["genai"] + names + # When enabled, also emit attribute keys under their semconv-ai / + # Traceloop names via the ``legacy`` mapper. Append it at the tail so + # the canonical ``genai`` keys win on any conflict. + if self.legacy_compat and "legacy" not in names: + names.append("legacy") + self.mapper_names = names + return self + + @property + def capture_span_content(self) -> bool: + """Whether prompt/response content may be stamped as span attributes. + + Defaults off (``no_content``): an operator must opt in before message + bodies leave the process, so a user request can never force its prompt + or completion into the configured backend while capture is disabled. + """ + return self.capture_message_content in ( + CaptureMessageContent.SPAN_ONLY, + CaptureMessageContent.SPAN_AND_EVENT, + ) + + @classmethod + def from_env(cls) -> "OpenTelemetryV2Config": + return cls() diff --git a/litellm/integrations/otel/context.py b/litellm/integrations/otel/context.py new file mode 100644 index 00000000000..d7e01e91ea7 --- /dev/null +++ b/litellm/integrations/otel/context.py @@ -0,0 +1,51 @@ +"""Trace-context + Baggage helpers.""" + +from typing import Mapping + +from opentelemetry import baggage +from opentelemetry.context import Context, get_current +from opentelemetry.trace import Span, set_span_in_context +from opentelemetry.trace.propagation.tracecontext import ( + TraceContextTextMapPropagator, +) + +_PROPAGATOR = TraceContextTextMapPropagator() + + +def set_request_baggage( + values: Mapping[str, str], context: Context | None = None +) -> Context: + """Return a context with ``values`` written into Baggage.""" + ctx = context + for key, value in values.items(): + ctx = baggage.set_baggage(key, value, context=ctx) + return ctx if ctx is not None else (context or get_current()) + + +def get_baggage_attributes(context: Context | None = None) -> dict[str, str]: + """All Baggage entries on ``context`` as strings.""" + return {key: str(value) for key, value in baggage.get_all(context).items()} + + +def context_from_span(span: Span, context: Context | None = None) -> Context: + """A context with ``span`` as the active span (for explicit parenting).""" + return set_span_in_context(span, context=context) + + +def is_recordable_span(obj: object) -> bool: + """True if ``obj`` is a live span with a valid context (safe to parent under).""" + if not isinstance(obj, Span): + return False + try: + ctx = obj.get_span_context() + except Exception: + return False + return ctx is not None and ctx.is_valid + + +def extract_traceparent(headers: Mapping[str, str]) -> Context | None: + """Extract a remote parent context from incoming HTTP headers, if present.""" + if not any(key.lower() == "traceparent" for key in headers): + return None + carrier = {str(key).lower(): value for key, value in headers.items()} + return _PROPAGATOR.extract(carrier) diff --git a/litellm/integrations/otel/emitter.py b/litellm/integrations/otel/emitter.py new file mode 100644 index 00000000000..ca035851245 --- /dev/null +++ b/litellm/integrations/otel/emitter.py @@ -0,0 +1,150 @@ +"""The span engine: dedup, start, run the mapper chain, set status, end.""" + +from collections import OrderedDict +from typing import Callable, Sequence + +from opentelemetry.context import Context +from opentelemetry.trace import Span, Tracer +from opentelemetry.trace.status import Status, StatusCode + +from litellm.integrations.otel.config import OpenTelemetryV2Config +from litellm.integrations.otel.mappers import resolve_mappers +from litellm.integrations.otel.mappers.base import AttributeMapper, SpanData +from litellm.integrations.otel.payloads import ( + GuardrailSpanData, + LLMCallSpanData, + ServiceSpanData, +) +from litellm.integrations.otel.providers import to_otel_span_kind +from litellm.integrations.otel.semconv import Error +from litellm.integrations.otel.spans import ( + SPAN_REGISTRY, + SpanRole, + guardrail_span_name, + llm_call_span_name, + service_span_name, +) + +# Roles emit() knows how to name and emit. PROXY_REQUEST and the management +# routes are SERVER spans owned by the mounted FastAPI instrumentor, so they +# have no builder here. +_NAME_BUILDERS: dict[SpanRole, Callable[..., str]] = { + SpanRole.LLM_CALL: llm_call_span_name, + SpanRole.GUARDRAIL: guardrail_span_name, + SpanRole.SERVICE: service_span_name, +} + +# Cap on the dedup cache. It only needs to coalesce the sync+async firing window +# of a single in-flight request, so a bounded LRU keeps memory flat on a +# long-running proxy while still covering every concurrently-open call. +_DEDUP_CACHE_MAX = 10_000 + + +class SpanEmitter: + def __init__( + self, + tracer: Tracer, + config: OpenTelemetryV2Config, + mappers: Sequence[AttributeMapper] | None = None, + ) -> None: + self._tracer = tracer + self._config = config + # The mapper chain is the sole source of span attributes. When not + # passed in, resolve it from the config so there's one source of truth. + self._mappers: list[AttributeMapper] = ( + list(mappers) + if mappers is not None + else resolve_mappers(config.mapper_names) + ) + # Bounded LRU (ordered by insertion / most-recent touch). Storing keys + # only — the value is unused — so it behaves like a capped set. + self._emitted: "OrderedDict[tuple[str, SpanRole], None]" = OrderedDict() + + # -- low-level helpers --------------------------------------------------- # + + def start_span( + self, + role: SpanRole, + name: str, + parent_context: Context | None = None, + start_time_ns: int | None = None, + *, + tracer: Tracer | None = None, + ) -> Span: + """Start a span for ``role`` without dedup or attribute mapping. + + For callers that own and manage their own span lifecycle. ``tracer`` + overrides the bound tracer for this span only, used for per-request + multi-tenant credential routing. + """ + return (tracer or self._tracer).start_span( + name, + context=parent_context, + kind=to_otel_span_kind(SPAN_REGISTRY[role].kind), + start_time=start_time_ns, + ) + + def _seen(self, dedup_key: str | None, role: SpanRole) -> bool: + """Return True once a ``(dedup_key, role)`` pair has been emitted. + + Guards against emitting the same span twice when a streaming call + fires both a sync and an async logging callback. + """ + if not dedup_key: + return False + marker = (dedup_key, role) + if marker in self._emitted: + self._emitted.move_to_end(marker) + return True + self._emitted[marker] = None + if len(self._emitted) > _DEDUP_CACHE_MAX: + self._emitted.popitem(last=False) # evict least-recently-used + return False + + # -- the engine ---------------------------------------------------------- # + + def emit( + self, + role: SpanRole, + data: SpanData, + parent_context: Context | None = None, + *, + start_time_ns: int | None = None, + end_time_ns: int | None = None, + tracer: Tracer | None = None, + ) -> Span | None: + """Emit one complete span: dedup, start, map attributes, status, end. + + Return the span, or ``None`` if it was deduplicated away. ``tracer`` + overrides the bound tracer for this span, used for per-request routing. + """ + # Only LLM-call spans carry a dedup key; LLM-call and service spans + # carry an ``error`` field. ``isinstance`` narrows the type for mypy and + # keeps the engine free of duck-typed attribute reads. + dedup_key = data.identity.call_id if isinstance(data, LLMCallSpanData) else None + if self._seen(dedup_key, role): + return None + span = self.start_span( + role, + _NAME_BUILDERS[role](data), + parent_context=parent_context, + start_time_ns=start_time_ns, + tracer=tracer, + ) + for mapper in self._mappers: + for key, value in mapper.map(data).items(): + span.set_attribute(key, value) + error = ( + data.error + if isinstance(data, (LLMCallSpanData, ServiceSpanData, GuardrailSpanData)) + else None + ) + if error and (error.error_type or error.message): + span.set_attribute(Error.TYPE, error.error_type or "error") + span.set_status( + Status(StatusCode.ERROR, error.message or error.error_type or "error") + ) + else: + span.set_status(Status(StatusCode.OK)) + span.end(end_time=end_time_ns) + return span diff --git a/litellm/integrations/otel/logger.py b/litellm/integrations/otel/logger.py new file mode 100644 index 00000000000..c55732e6d28 --- /dev/null +++ b/litellm/integrations/otel/logger.py @@ -0,0 +1,425 @@ +"""``CustomLogger`` adapter on the OpenTelemetry span engine. + +Thin adapter: it translates litellm's logging callbacks into typed ``*SpanData`` +and hands them to the engine (:mod:`emitter`), with multi-tenant tracer routing +in :mod:`routing`. It emits the gen-ai spans (LLM call, guardrail, service). + +The proxy server span is NOT owned here. It is created by the FastAPI +instrumentation mounted in ``proxy_server``'s startup event, which stamps the +``http.*`` attributes and handles inbound context propagation. The proxy-span +methods below are therefore no-ops: routes never modify spans. + +Gen-ai spans parent to that server span via the ambient OTel context rather +than an explicitly threaded span. litellm's async logging worker copies the +request's context at enqueue time, so ``async_log_success_event`` runs with the +server span active. Emission is therefore async-only — the sync callback runs +in an out-of-context thread, where there is no parent span, so it is a no-op. +""" + +from datetime import datetime +from typing import Any, Mapping, cast + +from opentelemetry.context import attach, get_current +from opentelemetry.sdk.trace import TracerProvider +from opentelemetry.trace import Span, Tracer, get_current_span + +import litellm +from litellm.integrations.custom_logger import CustomLogger +from litellm.integrations.otel.baggage import promoted_baggage +from litellm.integrations.otel.config import OpenTelemetryV2Config +from litellm.integrations.otel.context import ( + context_from_span, + is_recordable_span, + set_request_baggage, +) +from litellm.integrations.otel.emitter import SpanEmitter +from litellm.integrations.otel.mappers import resolve_mappers +from litellm.integrations.otel.payloads import ( + GuardrailSpanData, + LLMCallSpanData, + RequestIdentity, + ServiceSpanData, + SpanError, +) +from litellm.integrations.otel.providers import build_tracer_provider, get_tracer +from litellm.integrations.otel.routing import TenantTracerCache +from litellm.integrations.otel.spans import SpanRole +from litellm.integrations.otel.utils import to_ns + +LITELLM_TRACER_NAME = "litellm" +LITELLM_PROXY_REQUEST_SPAN_NAME = "Received Proxy Server Request" + +# Any callback whose class belongs to one of these modules is "the OTel +# callback" for proxy-global-registration purposes. +_OTEL_MODULES = ( + "litellm.integrations.otel", + "litellm.integrations.opentelemetry", +) + + +def _pre_call_guardrail_blocked(payload: Mapping[str, Any]) -> bool: + """True when a pre-call guardrail blocked the request (no LLM call happened). + + A blocked pre-call guardrail raises before the upstream call, yet litellm + still emits a failure log — which would otherwise produce a phantom CLIENT + span for a call that never occurred. We detect the case (request failed AND a + ``pre_call`` guardrail intervened) so the caller can skip that span. A + pre-call guardrail that merely *masks* lets the call proceed, so the request + succeeds and this returns False — only genuine blocks fail the request. + """ + if payload.get("status") != "failure": + return False + info = payload.get("guardrail_information") + if not isinstance(info, list): + return False + for entry in info: + if not isinstance(entry, dict): + continue + mode = entry.get("guardrail_mode") + is_pre_call = mode == "pre_call" or ( + isinstance(mode, (list, tuple)) and "pre_call" in mode + ) + if is_pre_call and entry.get("guardrail_status") == "guardrail_intervened": + return True + return False + + +class OpenTelemetryV2(CustomLogger): + """The ``CustomLogger`` for OpenTelemetry. + + The constructor accepts an optional config, callback name, and pre-built + OTel providers; when a provider is omitted it is built from the config. + ``logger_provider`` and ``meter_provider`` are accepted but reserved for + future OTel logs and metrics support. + """ + + def __init__( + self, + config: OpenTelemetryV2Config | None = None, + callback_name: str | None = None, + tracer_provider: TracerProvider | None = None, + logger_provider: Any | None = None, # reserved for OTel logs + meter_provider: Any | None = None, # reserved for metrics + **kwargs: Any, + ) -> None: + super().__init__(**kwargs) + self.config: OpenTelemetryV2Config = config or OpenTelemetryV2Config() + self.callback_name = callback_name + self._tracer_provider: TracerProvider = ( + tracer_provider + if tracer_provider is not None + else build_tracer_provider(self.config) + ) + self.tracer: Tracer = get_tracer(self._tracer_provider, LITELLM_TRACER_NAME) + self._emitter = SpanEmitter( + self.tracer, self.config, mappers=resolve_mappers(self.config.mapper_names) + ) + self._tenant_tracers = TenantTracerCache( + self.config, callback_name, LITELLM_TRACER_NAME + ) + self._init_otel_logger_on_litellm_proxy() + + # ====================================================================== # + # Proxy global registration + # ====================================================================== # + + def _init_otel_logger_on_litellm_proxy(self) -> None: + """Claim ``proxy_server.open_telemetry_logger`` if no one else has.""" + try: + from litellm.proxy import proxy_server + except Exception: + return + try: + # Mutate ``litellm.service_callback`` in place. ``getattr(..) or []`` + # would bind a throwaway local when the list is empty (an empty list + # is falsy), so the append would never reach the global and service + # spans (Redis, Postgres, …) would be silently dropped from traces. + service_callback = litellm.service_callback + already_otel = any( + cb.__class__.__module__.startswith(_OTEL_MODULES) + for cb in service_callback + if hasattr(cb, "__class__") + ) + if not already_otel: + service_callback.append(self) + except Exception: + pass + if getattr(proxy_server, "open_telemetry_logger", None) is None: + setattr(proxy_server, "open_telemetry_logger", self) + + # ====================================================================== # + # LLM-call callbacks + # ====================================================================== # + + # Async-only: the async path runs inside the request's restored OTel context + # (the logging worker copies it at enqueue), so the span parents to the + # instrumentor's server span via ambient context. The sync path runs in an + # out-of-context thread with no parent span, so it is a no-op. + + def log_success_event(self, kwargs, response_obj, start_time, end_time): + return None + + def log_failure_event(self, kwargs, response_obj, start_time, end_time): + return None + + async def async_log_success_event(self, kwargs, response_obj, start_time, end_time): + self._emit_llm_call(kwargs, start_time, end_time) + + async def async_log_failure_event(self, kwargs, response_obj, start_time, end_time): + self._emit_llm_call(kwargs, start_time, end_time) + + def _emit_llm_call( + self, + kwargs: Mapping[str, Any], + start_time: datetime | float | None, + end_time: datetime | float | None, + ) -> Span | None: + payload = kwargs.get("standard_logging_object") + if not payload: + return None + if _pre_call_guardrail_blocked(cast("Mapping[str, Any]", payload)): + # A pre-call guardrail blocked the request, so the upstream LLM was + # never called — litellm still emits a failure log, but a CLIENT + # "chat …" span for a call that didn't happen is misleading. Skip it; + # the guardrail span (ERROR, with the verdict) is the real outcome. + return None + data = LLMCallSpanData.from_standard_logging_payload( + cast("Any", payload), capture_content=self.config.capture_span_content + ) + # Parent is the ambient context (the instrumentor's server span, + # restored by the logging worker); no span is threaded through metadata. + parent_ctx = get_current() + # Write identity into Baggage so child spans (guardrails, services) + # inherit it. + bag = promoted_baggage( + data.identity, + data.request_model, + promoted_keys=tuple(self.config.baggage_promoted_keys), + metadata_keys=tuple(self.config.baggage_metadata_keys), + ) + if bag: + parent_ctx = set_request_baggage(bag, context=parent_ctx) + return self._emitter.emit( + SpanRole.LLM_CALL, + data, + parent_context=parent_ctx, + start_time_ns=to_ns(start_time), + end_time_ns=to_ns(end_time), + tracer=self._tenant_tracers.tracer_for( + self.tracer, kwargs.get("standard_callback_dynamic_params") + ), + ) + + # ====================================================================== # + # Service hooks + # ====================================================================== # + + async def async_service_success_hook( + self, + payload: Any, + parent_otel_span: Span | None = None, + start_time: datetime | float | None = None, + end_time: datetime | float | None = None, + event_metadata: dict | None = None, + ) -> None: + self._emit_service( + payload, + parent_otel_span=parent_otel_span, + start_time=start_time, + end_time=end_time, + event_metadata=event_metadata, + error_override=None, + ) + + async def async_service_failure_hook( + self, + payload: Any, + error: str | None = "", + parent_otel_span: Span | None = None, + start_time: datetime | float | None = None, + end_time: datetime | float | None = None, + event_metadata: dict | None = None, + ) -> None: + self._emit_service( + payload, + parent_otel_span=parent_otel_span, + start_time=start_time, + end_time=end_time, + event_metadata=event_metadata, + error_override=error or "error", + ) + + def _emit_service( + self, + payload: Any, + *, + parent_otel_span: Span | None, + start_time: datetime | float | None, + end_time: datetime | float | None, + event_metadata: dict | None, + error_override: str | None, + ) -> Span | None: + if not is_recordable_span(parent_otel_span): + return None + data = ServiceSpanData.from_payload(payload, event_metadata=event_metadata) + if error_override is not None and data.error is None: + data = ServiceSpanData( + service_name=data.service_name, + call_type=data.call_type, + error=SpanError(message=error_override), + event_metadata=data.event_metadata, + ) + # Parent to the server span, but layer it over the ambient context so the + # identity Baggage seeded in ``async_pre_call_hook`` rides along and the + # service span gets the same identity attributes as the LLM-call span. + parent_context = context_from_span( + cast(Span, parent_otel_span), context=get_current() + ) + return self._emitter.emit( + SpanRole.SERVICE, + data, + parent_context=parent_context, + start_time_ns=to_ns(start_time), + end_time_ns=to_ns(end_time), + ) + + # ====================================================================== # + # async_post_call_* hooks — emit guardrail spans. The server span's status + # / errors are the FastAPI instrumentor's job, so we don't touch it here. + # ====================================================================== # + + async def async_pre_call_hook( + self, + user_api_key_dict: Any, + cache: Any, + data: dict, + call_type: Any, + ) -> dict: + """Seed request identity into Baggage at the start of the request. + + This runs in the request task (the server span is the ambient context), + so attaching the identity Baggage here makes **every** span emitted for + the request — LLM call, guardrail, and service — inherit it via + ``LiteLLMBaggageSpanProcessor``. Without this, only the LLM-call span got + identity (it promoted Baggage locally) and the guardrail/service spans, + which parent to the server span, had none. The async logging worker + copies this context at enqueue time, so the LLM-call span inherits it too. + """ + try: + identity = RequestIdentity.from_user_api_key_auth(user_api_key_dict) + bag = promoted_baggage( + identity, + data.get("model") if isinstance(data, dict) else None, + promoted_keys=tuple(self.config.baggage_promoted_keys), + metadata_keys=tuple(self.config.baggage_metadata_keys), + ) + if bag: + # Attach (no detach): the contextvar is scoped to this request's + # asyncio task and is reclaimed when the task ends. + attach(set_request_baggage(bag, context=get_current())) + # The server span was started by the instrumentor before this + # hook ran, so the Baggage processor (which only fires at span + # start) won't backfill it — stamp identity on it directly. + server_span = get_current_span() + if is_recordable_span(server_span): + for key, value in bag.items(): + server_span.set_attribute(key, value) + except Exception: + pass + return data + + async def async_post_call_success_hook( + self, + data: Mapping[str, Any], + user_api_key_dict: Any, + response: Any, + ) -> Any: + self._emit_guardrail_spans(data) + return response + + async def async_post_call_failure_hook( + self, + request_data: Mapping[str, Any], + original_exception: BaseException | None, + user_api_key_dict: Any, + traceback_str: str | None = None, + ) -> None: + self._emit_guardrail_spans(request_data) + + def _emit_guardrail_spans(self, request_data: Mapping[str, Any]) -> None: + # Post-call hooks run in the request task, so the ambient context is the + # server span; guardrail spans parent to it implicitly via that context. + metadata = request_data.get("metadata") + guardrails: list[Any] = [] + if isinstance(metadata, dict): + info = metadata.get("standard_logging_guardrail_information") + if isinstance(info, list): + guardrails = info + elif isinstance(info, dict): + guardrails = [info] + for entry in guardrails: + if not isinstance(entry, dict): + continue + self._emitter.emit( + SpanRole.GUARDRAIL, GuardrailSpanData.from_logging_entry(entry) + ) + + # ====================================================================== # + # Management endpoint hooks — no-ops. Management endpoints are ordinary + # FastAPI routes, so the mounted instrumentor already spans them. + # ====================================================================== # + + async def async_management_endpoint_success_hook( + self, + logging_payload: Any, + parent_otel_span: Span | None = None, + ) -> None: + return None + + async def async_management_endpoint_failure_hook( + self, + logging_payload: Any, + parent_otel_span: Span | None = None, + ) -> None: + return None + + # ====================================================================== # + # Proxy SERVER-span API — no-ops. The FastAPI instrumentor owns the server + # span (creation, http.* attributes, inbound propagation) and gen-ai spans + # parent to it via ambient context. These methods are the surface the + # proxy and auth call sites invoke; they intentionally do nothing. + # ====================================================================== # + + def create_litellm_proxy_request_started_span( + self, start_time: datetime, headers: Mapping[str, str] | None + ) -> Span | None: + """Return the active server span instead of creating one. + + The FastAPI instrumentor owns the server span, so V2 creates nothing + here. But the proxy threads this return value as ``litellm_parent_otel_span`` + — and service logging (Redis, Postgres, …) only invokes the OTel service + hook when that parent is non-None. Returning the ambient server span lets + service spans nest under it. The proxy must NOT ``.end()`` this span (the + instrumentor does); ``_close_dangling_otel_server_span`` skips it under V2. + """ + span = get_current_span() + return span if is_recordable_span(span) else None + + @staticmethod + def set_proxy_request_route_attributes( + span: Span | None, + *, + url_path: str | None = None, + http_route: str | None = None, + ) -> None: + """No-op: the FastAPI instrumentor stamps ``http.route`` / ``url.path``.""" + + @staticmethod + def set_response_status_code_attribute( + span: Span | None, status_code: int | None + ) -> None: + """No-op: the FastAPI instrumentor stamps ``http.response.status_code``.""" + + @staticmethod + def set_preprocessing_duration_attribute(span: Span | None, container: Any) -> None: + """No-op: the server span belongs to the FastAPI instrumentor.""" diff --git a/litellm/integrations/otel/mappers/__init__.py b/litellm/integrations/otel/mappers/__init__.py new file mode 100644 index 00000000000..012e63f1bee --- /dev/null +++ b/litellm/integrations/otel/mappers/__init__.py @@ -0,0 +1,58 @@ +"""Attribute mappers: pure ``LLMCallSpanData -> {attribute key: value}`` functions. + +Composition over inheritance: vocabularies layer onto the same span. Listing +``["genai", "openinference"]`` in ``config.mapper_names`` makes every span +carry both the canonical ``gen_ai.*`` keys and the OpenInference (Arize + +Phoenix) keys. Add ``"langfuse"`` and it works for all three backends at once. +""" + +from typing import Callable, Iterable + +from litellm.integrations.otel.mappers.base import ( + AttributeMap, + AttributeMapper, + AttrValue, +) +from litellm.integrations.otel.mappers.genai import GenAIMapper +from litellm.integrations.otel.mappers.langfuse import LangfuseMapper +from litellm.integrations.otel.mappers.langtrace import LangtraceMapper +from litellm.integrations.otel.mappers.legacy import LegacyMapper +from litellm.integrations.otel.mappers.openinference import OpenInferenceMapper +from litellm.integrations.otel.mappers.weave import WeaveMapper + +# Registry keyed by ``config.mapper_names`` entries. +_MAPPER_BY_NAME: dict[str, Callable[[], AttributeMapper]] = { + "genai": GenAIMapper, + "legacy": LegacyMapper, + "openinference": OpenInferenceMapper, + "langfuse": LangfuseMapper, + "weave": WeaveMapper, + "langtrace": LangtraceMapper, +} + + +def resolve_mappers(names: Iterable[str]) -> list[AttributeMapper]: + """Resolve mapper names to instances. Unknown names raise ``ValueError``.""" + out: list[AttributeMapper] = [] + for name in names: + factory = _MAPPER_BY_NAME.get(name) + if factory is None: + raise ValueError( + f"unknown mapper name {name!r}; known: " f"{sorted(_MAPPER_BY_NAME)}" + ) + out.append(factory()) + return out + + +__all__ = [ + "AttributeMap", + "AttributeMapper", + "AttrValue", + "GenAIMapper", + "LangfuseMapper", + "LangtraceMapper", + "LegacyMapper", + "OpenInferenceMapper", + "WeaveMapper", + "resolve_mappers", +] diff --git a/litellm/integrations/otel/mappers/base.py b/litellm/integrations/otel/mappers/base.py new file mode 100644 index 00000000000..0209ff3342b --- /dev/null +++ b/litellm/integrations/otel/mappers/base.py @@ -0,0 +1,36 @@ +"""Mapper protocol and attribute value types.""" + +from typing import Sequence + +from typing_extensions import Protocol, runtime_checkable + +from litellm.integrations.otel.payloads import ( + GuardrailSpanData, + LLMCallSpanData, + ServiceSpanData, +) + +AttrScalar = str | bool | int | float +# Mirrors ``opentelemetry.util.types.AttributeValue`` (homogeneous sequences) +# without importing the SDK, so mappers stay OTel-free. +AttrValue = ( + AttrScalar | Sequence[str] | Sequence[bool] | Sequence[int] | Sequence[float] +) +AttributeMap = dict[str, AttrValue] + +# The closed set of span-data types the engine routes through the mapper chain. +# Server spans (PROXY_REQUEST + management routes) belong to the mounted FastAPI +# instrumentor, not the mapper chain. +SpanData = LLMCallSpanData | GuardrailSpanData | ServiceSpanData + + +@runtime_checkable +class AttributeMapper(Protocol): + """Maps a typed span input to a flat dict of OTel span attributes. + + One method per mapper, dispatched internally on the ``data`` type. The + engine calls this uniformly for every span kind — mappers that don't speak + a given type return ``{}``. This is why the engine contains no attribute keys. + """ + + def map(self, data: SpanData) -> AttributeMap: ... diff --git a/litellm/integrations/otel/mappers/genai.py b/litellm/integrations/otel/mappers/genai.py new file mode 100644 index 00000000000..0ceb2e5126c --- /dev/null +++ b/litellm/integrations/otel/mappers/genai.py @@ -0,0 +1,121 @@ +"""Canonical OpenTelemetry GenAI semantic-convention mapper (always active). + +Owns the attribute schema for every span kind the engine emits — LLM call, +guardrail, and service — so the engine itself never references attribute keys. + +Each span kind declares its schema as a flat ``attribute key -> extractor`` +table: one lambda per mapping operation, applied against the typed span data. +""" + +from typing import Callable + +from litellm.integrations.otel.mappers.base import AttributeMap, AttrValue, SpanData +from litellm.integrations.otel.mappers.utils import collect, drop_none +from litellm.integrations.otel.payloads import ( + GuardrailSpanData, + LLMCallSpanData, + ServiceSpanData, + ToolDefinition, +) +from litellm.integrations.otel.semconv import Error, GenAI, LiteLLM, Server + + +class GenAIMapper: + + _LLM_CALL_ATTRS: dict[str, Callable[[LLMCallSpanData], AttrValue | None]] = { + GenAI.OPERATION_NAME: lambda d: d.operation.value, + GenAI.PROVIDER_NAME: lambda d: d.provider or None, + GenAI.REQUEST_MODEL: lambda d: d.request_model or None, + GenAI.REQUEST_TEMPERATURE: lambda d: d.request_params.temperature, + GenAI.REQUEST_TOP_P: lambda d: d.request_params.top_p, + GenAI.REQUEST_TOP_K: lambda d: d.request_params.top_k, + GenAI.REQUEST_MAX_TOKENS: lambda d: d.request_params.max_tokens, + GenAI.REQUEST_FREQUENCY_PENALTY: lambda d: d.request_params.frequency_penalty, + GenAI.REQUEST_PRESENCE_PENALTY: lambda d: d.request_params.presence_penalty, + GenAI.REQUEST_STOP_SEQUENCES: lambda d: ( + list(d.request_params.stop_sequences) + if d.request_params.stop_sequences + else None + ), + GenAI.REQUEST_SEED: lambda d: d.request_params.seed, + GenAI.RESPONSE_MODEL: lambda d: d.response_model, + GenAI.RESPONSE_ID: lambda d: d.response_id, + GenAI.RESPONSE_FINISH_REASONS: lambda d: ( + list(d.finish_reasons) if d.finish_reasons else None + ), + GenAI.USAGE_INPUT_TOKENS: lambda d: d.usage.input_tokens, + GenAI.USAGE_OUTPUT_TOKENS: lambda d: d.usage.output_tokens, + Error.TYPE: lambda d: d.error.error_type if d.error else None, + Server.ADDRESS: lambda d: d.server.address if d.server else None, + Server.PORT: lambda d: d.server.port if d.server else None, + LiteLLM.CALL_ID: lambda d: d.identity.call_id or None, + f"{LiteLLM.COST_PREFIX}total": lambda d: d.response_cost, + LiteLLM.REQUEST_STREAMING: lambda d: d.is_streaming, + } + + _TOOL_ATTRS: dict[str, Callable[[ToolDefinition], AttrValue | None]] = { + "name": lambda t: t.name, + "description": lambda t: t.description or None, + "parameters": lambda t: t.parameters_json or None, + } + + _GUARDRAIL_ATTRS: dict[str, Callable[[GuardrailSpanData], AttrValue | None]] = { + LiteLLM.GUARDRAIL_NAME: lambda d: d.guardrail_name, + LiteLLM.GUARDRAIL_MODE: lambda d: d.mode, + LiteLLM.GUARDRAIL_STATUS: lambda d: d.status, + LiteLLM.GUARDRAIL_PROVIDER: lambda d: d.provider, + LiteLLM.GUARDRAIL_ACTION: lambda d: d.action, + LiteLLM.GUARDRAIL_RESPONSE: lambda d: d.response_json, + LiteLLM.GUARDRAIL_VIOLATION_CATEGORIES: lambda d: ( + list(d.violation_categories) if d.violation_categories else None + ), + LiteLLM.GUARDRAIL_CONFIDENCE_SCORE: lambda d: d.confidence_score, + LiteLLM.GUARDRAIL_RISK_SCORE: lambda d: d.risk_score, + LiteLLM.GUARDRAIL_MASKED_ENTITY_COUNT: lambda d: d.masked_entity_count, + LiteLLM.GUARDRAIL_DURATION: lambda d: d.duration, + } + + _SERVICE_ATTRS: dict[str, Callable[[ServiceSpanData], AttrValue | None]] = { + LiteLLM.SERVICE_NAME: lambda d: d.service_name, + LiteLLM.SERVICE_CALL_TYPE: lambda d: d.call_type, + } + + def map(self, data: SpanData) -> AttributeMap: + match data: + case LLMCallSpanData(): + return self._llm_call(data) + case GuardrailSpanData(): + return self._guardrail(data) + case ServiceSpanData(): + return self._service(data) + case _: + return {} + + @classmethod + def _llm_call(cls, data: LLMCallSpanData) -> AttributeMap: + attrs = collect(cls._LLM_CALL_ATTRS, data) + attrs.update( + drop_none( + { + f"gen_ai.tool.{idx}.{suffix}": extract(tool) + for idx, tool in enumerate(data.tools) + for suffix, extract in cls._TOOL_ATTRS.items() + } + ) + ) + return attrs + + @classmethod + def _guardrail(cls, data: GuardrailSpanData) -> AttributeMap: + return collect(cls._GUARDRAIL_ATTRS, data) + + @classmethod + def _service(cls, data: ServiceSpanData) -> AttributeMap: + attrs = collect(cls._SERVICE_ATTRS, data) + attrs.update( + { + f"{LiteLLM.METADATA_PREFIX}{key}": value + for key, value in data.event_metadata.items() + } + ) + return attrs diff --git a/litellm/integrations/otel/mappers/langfuse.py b/litellm/integrations/otel/mappers/langfuse.py new file mode 100644 index 00000000000..6f8361390fc --- /dev/null +++ b/litellm/integrations/otel/mappers/langfuse.py @@ -0,0 +1,84 @@ +"""Langfuse OTLP attribute mapper. + +Langfuse ingests OTLP spans and reads from its own vendor namespace +(``langfuse.observation.*``, ``langfuse.trace.*``). Compose this mapper after +``GenAIMapper`` to send canonical + Langfuse-flavored spans simultaneously. + +Every attribute is declared as a ``key -> extractor`` table entry (one callable +per mapping operation): ``_LLM_CALL_ATTRS`` for scalars and ``_BLOB_ATTRS`` for +the JSON-serialized payloads. ``_llm_call`` just applies both tables. +""" + +import json +from typing import Callable + +from litellm.integrations.otel.mappers.base import AttributeMap, AttrValue, SpanData +from litellm.integrations.otel.mappers.utils import ( + collect, + json_if, + output_messages, + serialize_messages, +) +from litellm.integrations.otel.payloads import ( + LLMCallSpanData, + LLMRequestParams, + LLMUsage, +) + + +class LangfuseMapper: + + _LLM_CALL_ATTRS: dict[str, Callable[[LLMCallSpanData], AttrValue | None]] = { + "langfuse.observation.type": lambda d: "generation", + "langfuse.observation.model.name": lambda d: d.request_model or None, + "langfuse.observation.metadata.provider": lambda d: d.provider or None, + "langfuse.observation.id": lambda d: d.identity.call_id or None, + "langfuse.trace.metadata.team_id": lambda d: d.identity.team_id or None, + "langfuse.trace.metadata.team_alias": lambda d: d.identity.team_alias or None, + } + + # Sub-tables folded into their respective JSON blobs. + _MODEL_PARAMS: dict[str, Callable[[LLMRequestParams], AttrValue | None]] = { + "temperature": lambda rp: rp.temperature, + "top_p": lambda rp: rp.top_p, + "max_tokens": lambda rp: rp.max_tokens, + "frequency_penalty": lambda rp: rp.frequency_penalty, + "presence_penalty": lambda rp: rp.presence_penalty, + "seed": lambda rp: rp.seed, + } + _USAGE_FIELDS: dict[str, Callable[[LLMUsage], AttrValue | None]] = { + "input": lambda u: u.input_tokens, + "output": lambda u: u.output_tokens, + "total": lambda u: u.total_tokens, + } + + # JSON-payload attributes: each builder returns the serialized blob or None. + _BLOB_ATTRS: dict[str, Callable[[LLMCallSpanData], AttrValue | None]] = { + "langfuse.observation.model.parameters": lambda d: json_if( + collect(LangfuseMapper._MODEL_PARAMS, d.request_params) + ), + "langfuse.observation.input": lambda d: serialize_messages(d.messages_in), + "langfuse.observation.output": lambda d: serialize_messages(output_messages(d)), + "langfuse.observation.usage_details": lambda d: json_if( + collect(LangfuseMapper._USAGE_FIELDS, d.usage) + ), + "langfuse.observation.cost_details": lambda d: ( + json.dumps({"total": d.response_cost}) + if d.response_cost is not None + else None + ), + } + + def map(self, data: SpanData) -> AttributeMap: + match data: + case LLMCallSpanData(): + return self._llm_call(data) + case _: + return {} + + @classmethod + def _llm_call(cls, data: LLMCallSpanData) -> AttributeMap: + return { + **collect(cls._LLM_CALL_ATTRS, data), + **collect(cls._BLOB_ATTRS, data), + } diff --git a/litellm/integrations/otel/mappers/langtrace.py b/litellm/integrations/otel/mappers/langtrace.py new file mode 100644 index 00000000000..7f80457e13b --- /dev/null +++ b/litellm/integrations/otel/mappers/langtrace.py @@ -0,0 +1,64 @@ +"""Langtrace attribute mapper. + +Produces Langtrace's attribute vocabulary so a span can be ingested by a +Langtrace backend. Compose it alongside other mappers like any other +vocabulary. + +Scalar attributes are declared as a flat ``key -> extractor`` table (one lambda +per mapping operation); the prompt/completion blobs are serialized as a tail. +""" + +from typing import Callable + +from litellm.integrations.otel.mappers.base import AttributeMap, AttrValue, SpanData +from litellm.integrations.otel.mappers.utils import ( + collect, + json_or_none, + output_messages, +) +from litellm.integrations.otel.payloads import LLMCallSpanData + + +class LangtraceMapper: + + _LLM_CALL_ATTRS: dict[str, Callable[[LLMCallSpanData], AttrValue | None]] = { + "gen_ai.operation.name": lambda d: "chat", + "langtrace.service.name": lambda d: d.provider or None, + "llm.model": lambda d: d.request_model or None, + "gen_ai.response.model": lambda d: d.response_model or None, + "gen_ai.response_id": lambda d: d.response_id or None, + "gen_ai.system_fingerprint": lambda d: d.system_fingerprint or None, + "llm.temperature": lambda d: d.request_params.temperature, + "llm.top_p": lambda d: d.request_params.top_p, + "llm.top_k": lambda d: d.request_params.top_k, + "llm.max_tokens": lambda d: d.request_params.max_tokens, + "llm.frequency_penalty": lambda d: d.request_params.frequency_penalty, + "llm.presence_penalty": lambda d: d.request_params.presence_penalty, + "llm.stream": lambda d: d.is_streaming, + "llm.token.counts.prompt": lambda d: d.usage.input_tokens, + "llm.token.counts.completion": lambda d: d.usage.output_tokens, + "llm.token.counts.total": lambda d: d.usage.total_tokens, + } + + _BLOB_ATTRS: dict[str, Callable[[LLMCallSpanData], AttrValue | None]] = { + "llm.prompts": lambda d: ( + json_or_none(list(d.messages_in)) if d.messages_in else None + ), + "llm.completions": lambda d: ( + json_or_none(output_messages(d)) if d.choices_out else None + ), + } + + def map(self, data: SpanData) -> AttributeMap: + match data: + case LLMCallSpanData(): + return self._llm_call(data) + case _: + return {} + + @classmethod + def _llm_call(cls, data: LLMCallSpanData) -> AttributeMap: + return { + **collect(cls._LLM_CALL_ATTRS, data), + **collect(cls._BLOB_ATTRS, data), + } diff --git a/litellm/integrations/otel/mappers/legacy.py b/litellm/integrations/otel/mappers/legacy.py new file mode 100644 index 00000000000..949e27e51d5 --- /dev/null +++ b/litellm/integrations/otel/mappers/legacy.py @@ -0,0 +1,97 @@ +"""Mapper for the older semantic-convention attribute vocabulary. + +Emits attributes under the semconv-ai / Traceloop key names (e.g. +``gen_ai.system``, ``gen_ai.usage.prompt_tokens``, ``llm.is_streaming``) plus a +few bare, unprefixed service keys (``service``, ``call_type``, ``error``), for +backends that consume those names. + +Like ``GenAIMapper``, each span kind declares its schema as a flat +``attribute key -> extractor`` table: one lambda per mapping operation. +""" + +from typing import Callable, Final + +from litellm.integrations.otel.mappers.base import AttributeMap, AttrValue, SpanData +from litellm.integrations.otel.mappers.utils import collect, drop_none +from litellm.integrations.otel.payloads import ( + LLMCallSpanData, + ServiceSpanData, + ToolDefinition, +) + +# Attribute keys in the semconv-ai / Traceloop vocabulary. +_LEGACY_SYSTEM: Final = "gen_ai.system" +_LEGACY_PROMPT_TOKENS: Final = "gen_ai.usage.prompt_tokens" +_LEGACY_COMPLETION_TOKENS: Final = "gen_ai.usage.completion_tokens" +_LEGACY_TOTAL_TOKENS: Final = "gen_ai.usage.total_tokens" +_LEGACY_IS_STREAMING: Final = "llm.is_streaming" +_LEGACY_TOP_K: Final = "llm.top_k" +_LEGACY_FREQUENCY_PENALTY: Final = "llm.frequency_penalty" +_LEGACY_PRESENCE_PENALTY: Final = "llm.presence_penalty" +_LEGACY_STOP_SEQUENCES: Final = "llm.chat.stop_sequences" +_LEGACY_SERVICE: Final = "service" +_LEGACY_CALL_TYPE: Final = "call_type" +_LEGACY_ERROR: Final = "error" + + +class LegacyMapper: + """Emits LLM-call and service attributes under the older key names.""" + + _LLM_CALL_ATTRS: dict[str, Callable[[LLMCallSpanData], AttrValue | None]] = { + _LEGACY_SYSTEM: lambda d: d.provider or None, + _LEGACY_PROMPT_TOKENS: lambda d: d.usage.input_tokens, + _LEGACY_COMPLETION_TOKENS: lambda d: d.usage.output_tokens, + _LEGACY_TOTAL_TOKENS: lambda d: d.usage.total_tokens, + _LEGACY_IS_STREAMING: lambda d: d.is_streaming, + _LEGACY_TOP_K: lambda d: d.request_params.top_k, + _LEGACY_FREQUENCY_PENALTY: lambda d: d.request_params.frequency_penalty, + _LEGACY_PRESENCE_PENALTY: lambda d: d.request_params.presence_penalty, + _LEGACY_STOP_SEQUENCES: lambda d: ( + list(d.request_params.stop_sequences) + if d.request_params.stop_sequences + else None + ), + } + + _TOOL_ATTRS: dict[str, Callable[[ToolDefinition], AttrValue | None]] = { + "name": lambda t: t.name, + "description": lambda t: t.description or None, + "parameters": lambda t: t.parameters_json or None, + } + + _SERVICE_ATTRS: dict[str, Callable[[ServiceSpanData], AttrValue | None]] = { + _LEGACY_SERVICE: lambda d: d.service_name, + _LEGACY_CALL_TYPE: lambda d: d.call_type, + _LEGACY_ERROR: lambda d: ( + d.error.message if d.error is not None and d.error.message else None + ), + } + + def map(self, data: SpanData) -> AttributeMap: + match data: + case LLMCallSpanData(): + return self._llm_call(data) + case ServiceSpanData(): + return self._service(data) + case _: + return {} + + @classmethod + def _llm_call(cls, data: LLMCallSpanData) -> AttributeMap: + attrs = collect(cls._LLM_CALL_ATTRS, data) + attrs.update( + drop_none( + { + f"llm.request.functions.{idx}.{suffix}": extract(tool) + for idx, tool in enumerate(data.tools) + for suffix, extract in cls._TOOL_ATTRS.items() + } + ) + ) + return attrs + + @classmethod + def _service(cls, data: ServiceSpanData) -> AttributeMap: + attrs = collect(cls._SERVICE_ATTRS, data) + attrs.update(dict(data.event_metadata)) + return attrs diff --git a/litellm/integrations/otel/mappers/openinference.py b/litellm/integrations/otel/mappers/openinference.py new file mode 100644 index 00000000000..ff3f3dd8e85 --- /dev/null +++ b/litellm/integrations/otel/mappers/openinference.py @@ -0,0 +1,128 @@ +"""OpenInference attribute mapper (Arize + Arize-Phoenix shared vocabulary). + +Spec: https://github.com/Arize-ai/openinference/tree/main/spec — the standard +both Arize and Phoenix consume. Composing this mapper after ``GenAIMapper`` +gives the same span both vocabularies, so a single trace lights up Arize + +Phoenix + any other OpenInference-aware backend simultaneously. +""" + +import json +from typing import Callable, Sequence + +from litellm.integrations.otel.mappers.base import AttributeMap, AttrValue, SpanData +from litellm.integrations.otel.mappers.utils import ( + collect, + drop_none, + json_if, + message_content, + output_messages, +) +from litellm.integrations.otel.payloads import ( + LLMCallSpanData, + LLMRequestParams, + ToolDefinition, +) + + +class OpenInferenceMapper: + """Emits OpenInference attributes for LLM_CALL spans. + + Key families (per the OpenInference spec): + - ``openinference.span.kind`` — discriminator (``"LLM"`` here) + - ``llm.model_name`` / ``llm.provider`` / ``llm.invocation_parameters`` + - ``llm.input_messages.{i}.message.role`` / ``...content`` + - ``llm.output_messages.{i}.message.role`` / ``...content`` + - ``llm.token_count.prompt`` / ``...completion`` / ``...total`` + - ``input.value`` / ``output.value`` — JSON-serialized request / response + """ + + _LLM_CALL_ATTRS: dict[str, Callable[[LLMCallSpanData], AttrValue | None]] = { + "openinference.span.kind": lambda d: "LLM", + "llm.model_name": lambda d: d.request_model or None, + "llm.provider": lambda d: d.provider or None, + "llm.token_count.prompt": lambda d: d.usage.input_tokens, + "llm.token_count.completion": lambda d: d.usage.output_tokens, + "llm.token_count.total": lambda d: d.usage.total_tokens, + } + + # Folded into the ``llm.invocation_parameters`` JSON blob. + _INVOCATION_PARAMS: dict[str, Callable[[LLMRequestParams], AttrValue | None]] = { + "temperature": lambda rp: rp.temperature, + "top_p": lambda rp: rp.top_p, + "top_k": lambda rp: rp.top_k, + "max_tokens": lambda rp: rp.max_tokens, + "frequency_penalty": lambda rp: rp.frequency_penalty, + "presence_penalty": lambda rp: rp.presence_penalty, + "seed": lambda rp: rp.seed, + } + + # Per-tool extractors, keyed by the ``llm.tools.{idx}.*`` suffix. + _TOOL_ATTRS: dict[str, Callable[[ToolDefinition], AttrValue | None]] = { + "tool.name": lambda t: t.name, + "tool.description": lambda t: t.description or None, + "tool.json_schema": lambda t: t.parameters_json or None, + } + + # JSON-payload attributes: each builder returns the serialized blob or None. + _BLOB_ATTRS: dict[str, Callable[[LLMCallSpanData], AttrValue | None]] = { + "llm.invocation_parameters": lambda d: json_if( + collect(OpenInferenceMapper._INVOCATION_PARAMS, d.request_params) + ), + } + + def map(self, data: SpanData) -> AttributeMap: + match data: + case LLMCallSpanData(): + return self._llm_call(data) + case _: + return {} + + @classmethod + def _llm_call(cls, data: LLMCallSpanData) -> AttributeMap: + return { + **collect(cls._LLM_CALL_ATTRS, data), + **collect(cls._BLOB_ATTRS, data), + **cls._messages("llm.input_messages", "input.value", data.messages_in), + **cls._messages( + "llm.output_messages", "output.value", output_messages(data) + ), + **cls._tools(data), + } + + @staticmethod + def _messages( + prefix: str, value_key: str, messages: Sequence[object] + ) -> AttributeMap: + """Per-message ``{prefix}.{idx}.message.*`` keys + the ``value_key`` blob.""" + parsed = [ + (m.get("role") if isinstance(m, dict) else None, message_content(m)) + for m in messages + ] + attrs = drop_none( + { + key: value + for idx, (role, content) in enumerate(parsed) + for key, value in ( + ( + f"{prefix}.{idx}.message.role", + role if isinstance(role, str) else None, + ), + (f"{prefix}.{idx}.message.content", content), + ) + } + ) + if parsed: + attrs[value_key] = json.dumps( + [{"role": role, "content": content} for role, content in parsed] + ) + return attrs + + @classmethod + def _tools(cls, data: LLMCallSpanData) -> AttributeMap: + return drop_none( + { + f"llm.tools.{idx}.{suffix}": extract(tool) + for idx, tool in enumerate(data.tools) + for suffix, extract in cls._TOOL_ATTRS.items() + } + ) diff --git a/litellm/integrations/otel/mappers/utils.py b/litellm/integrations/otel/mappers/utils.py new file mode 100644 index 00000000000..f7986e05790 --- /dev/null +++ b/litellm/integrations/otel/mappers/utils.py @@ -0,0 +1,76 @@ +"""Shared helpers for the attribute mappers. + +Small, mapper-agnostic utilities — JSON serialization, message extraction, and +extractor-table application — pulled out of the individual mapper modules so +they live in one place. +""" + +import json +from typing import Callable, Mapping, Sequence + +from litellm.integrations.otel.mappers.base import AttributeMap, AttrValue +from litellm.integrations.otel.payloads import LLMCallSpanData + + +def drop_none(values: Mapping[str, AttrValue | None]) -> AttributeMap: + """Return ``values`` with ``None``-valued entries removed.""" + return {k: v for k, v in values.items() if v is not None} + + +def collect(table: Mapping[str, Callable], source: object) -> AttributeMap: + """Apply an extractor table to ``source``, dropping ``None`` results.""" + return drop_none({key: extract(source) for key, extract in table.items()}) + + +def json_if(payload: Mapping[str, object]) -> str | None: + """JSON-serialize ``payload`` only when it's non-empty; else ``None``.""" + return json.dumps(payload) if payload else None + + +def json_or_none(value: object) -> str | None: + """JSON-serialize ``value`` (falling back to ``str``); ``None`` on failure.""" + try: + return json.dumps(value, default=str) + except Exception: + return None + + +def stringify_message(message: object) -> str | None: + """JSON-serialize a chat message dict; ``None`` if not a dict or on failure.""" + if not isinstance(message, dict): + return None + try: + return json.dumps(message, default=str) + except Exception: + return None + + +def serialize_messages(messages: Sequence[object]) -> str | None: + """Round-trip a sequence of message dicts through ``stringify_message``.""" + serialized = [ + json.loads(s) for s in (stringify_message(m) for m in messages) if s is not None + ] + return json.dumps(serialized) if serialized else None + + +def message_content(message: object) -> str | None: + """Extract the textual ``content`` from a chat message dict.""" + if not isinstance(message, dict): + return None + content = message.get("content") + if isinstance(content, str): + return content + if isinstance(content, list): + # multimodal: concatenate text parts only + parts = [ + part.get("text", "") + for part in content + if isinstance(part, dict) and part.get("type") == "text" + ] + return "".join(p for p in parts if isinstance(p, str)) or None + return None + + +def output_messages(data: LLMCallSpanData) -> list: + """The ``message`` payload of each response choice.""" + return [c.get("message") for c in data.choices_out if isinstance(c, dict)] diff --git a/litellm/integrations/otel/mappers/weave.py b/litellm/integrations/otel/mappers/weave.py new file mode 100644 index 00000000000..c906a97716a --- /dev/null +++ b/litellm/integrations/otel/mappers/weave.py @@ -0,0 +1,48 @@ +"""Weave (W&B) attribute mapper. + +Weave consumes OpenInference + a small set of Weave-specific keys (display +name, thread id, output value). This mapper layers the latter on top of +OpenInference's vocabulary — compose ``["genai", "openinference", "weave"]`` +to feed a Weave backend. +""" + +from typing import Callable + +from litellm.integrations.otel.mappers.base import AttributeMap, AttrValue, SpanData +from litellm.integrations.otel.mappers.utils import collect, json_or_none +from litellm.integrations.otel.payloads import LLMCallSpanData + + +class WeaveMapper: + """Maps ``LLMCallSpanData`` to Weave's vendor attributes.""" + + _LLM_CALL_ATTRS: dict[str, Callable[[LLMCallSpanData], AttrValue | None]] = { + # ``display_name`` has the form ``"{operation} {model}"``. The span + # name already covers that, but Weave reads this attribute too. + "weave.display_name": lambda d: ( + f"{d.operation.value} {d.request_model}" if d.request_model else None + ), + "weave.call_id": lambda d: d.identity.call_id or None, + } + + # JSON-payload attributes: each builder returns the serialized blob or None. + _BLOB_ATTRS: dict[str, Callable[[LLMCallSpanData], AttrValue | None]] = { + # Weave treats the response choices as the "output" payload. + "weave.output": lambda d: ( + json_or_none(list(d.choices_out)) if d.choices_out else None + ), + } + + def map(self, data: SpanData) -> AttributeMap: + match data: + case LLMCallSpanData(): + return self._llm_call(data) + case _: + return {} + + @classmethod + def _llm_call(cls, data: LLMCallSpanData) -> AttributeMap: + return { + **collect(cls._LLM_CALL_ATTRS, data), + **collect(cls._BLOB_ATTRS, data), + } diff --git a/litellm/integrations/otel/metrics.py b/litellm/integrations/otel/metrics.py new file mode 100644 index 00000000000..791416f18fe --- /dev/null +++ b/litellm/integrations/otel/metrics.py @@ -0,0 +1,28 @@ +"""GenAI client metrics (token usage + operation duration histograms).""" + +from dataclasses import dataclass + +from opentelemetry.metrics import Histogram, Meter + +from litellm.integrations.otel.semconv import Metric + + +@dataclass(frozen=True) +class GenAIMetrics: + token_usage: Histogram + operation_duration: Histogram + + +def create_genai_metrics(meter: Meter) -> GenAIMetrics: + return GenAIMetrics( + token_usage=meter.create_histogram( + name=Metric.TOKEN_USAGE, + unit="{token}", + description="Number of tokens used per GenAI request.", + ), + operation_duration=meter.create_histogram( + name=Metric.OPERATION_DURATION, + unit="s", + description="GenAI operation duration.", + ), + ) diff --git a/litellm/integrations/otel/payloads.py b/litellm/integrations/otel/payloads.py new file mode 100644 index 00000000000..b91509c19ab --- /dev/null +++ b/litellm/integrations/otel/payloads.py @@ -0,0 +1,409 @@ +"""Typed span-data inputs: frozen dataclasses the engine and mappers consume.""" + +from __future__ import annotations + +import json +from dataclasses import dataclass, field +from typing import TYPE_CHECKING, ClassVar, Mapping, cast +from urllib.parse import urlsplit + +from litellm.integrations.otel.semconv import ( + GenAIOperation, + resolve_operation, + resolve_provider, +) +from litellm.integrations.otel.utils import ( + as_bool, + as_float, + as_int, + as_str, + as_str_tuple, +) + +if TYPE_CHECKING: + from litellm.types.services import ServiceLoggerPayload + from litellm.types.utils import StandardLoggingPayload + + +# --- typed sub-structures ---------------------------------------------------- # + + +@dataclass(frozen=True) +class LLMRequestParams: + temperature: float | None = None + top_p: float | None = None + top_k: int | None = None + max_tokens: int | None = None + frequency_penalty: float | None = None + presence_penalty: float | None = None + stop_sequences: tuple[str, ...] | None = None + seed: int | None = None + + @classmethod + def from_model_parameters(cls, params: Mapping[str, object]) -> "LLMRequestParams": + max_tokens = as_int(params.get("max_tokens")) + if max_tokens is None: + max_tokens = as_int(params.get("max_completion_tokens")) + return cls( + temperature=as_float(params.get("temperature")), + top_p=as_float(params.get("top_p")), + top_k=as_int(params.get("top_k")), + max_tokens=max_tokens, + frequency_penalty=as_float(params.get("frequency_penalty")), + presence_penalty=as_float(params.get("presence_penalty")), + stop_sequences=as_str_tuple(params.get("stop")), + seed=as_int(params.get("seed")), + ) + + +@dataclass(frozen=True) +class LLMUsage: + input_tokens: int | None = None + output_tokens: int | None = None + total_tokens: int | None = None + + +@dataclass(frozen=True) +class SpanError: + error_type: str | None = None + message: str | None = None + + +@dataclass(frozen=True) +class ServerInfo: + address: str | None = None + port: int | None = None + + @classmethod + def from_api_base(cls, api_base: str | None) -> ServerInfo | None: + if not api_base: + return None + parsed = urlsplit(api_base if "://" in api_base else f"//{api_base}") + if not parsed.hostname: + return None + return cls(address=parsed.hostname, port=parsed.port) + + +@dataclass(frozen=True) +class RequestIdentity: + call_id: str | None = None + team_id: str | None = None + team_alias: str | None = None + key_hash: str | None = None + end_user: str | None = None + metadata: Mapping[str, str] = field(default_factory=dict) + + @classmethod + def from_payload(cls, payload: "StandardLoggingPayload") -> "RequestIdentity": + raw_meta = cast(Mapping[str, object], payload.get("metadata") or {}) + metadata = { + key: str(value) + for key, value in raw_meta.items() + if isinstance(value, (str, bool, int, float)) + } + return cls( + call_id=as_str(payload.get("litellm_call_id")) or as_str(payload.get("id")), + # StandardLoggingMetadata's canonical key is ``user_api_key_team_id``; + # the bare ``team_id`` is a legacy alias and is often empty, so prefer + # the canonical key and fall back to the alias. + team_id=as_str(raw_meta.get("user_api_key_team_id")) + or as_str(raw_meta.get("team_id")), + team_alias=as_str(raw_meta.get("user_api_key_team_alias")) + or as_str(raw_meta.get("team_alias")), + key_hash=as_str(raw_meta.get("user_api_key_hash")), + end_user=as_str(payload.get("end_user")) + or as_str(raw_meta.get("user_api_key_end_user_id")), + metadata=metadata, + ) + + @classmethod + def from_user_api_key_auth(cls, auth: object) -> "RequestIdentity": + """Identity from a ``UserAPIKeyAuth`` (duck-typed to keep this module + free of a proxy import). + + Used in the pre-call hook to seed Baggage early — before any LLM, + guardrail, or service span is created — so the whole request's spans + inherit identity, not just the LLM-call span. Metadata sub-keys use the + ``user_api_key_*`` names that ``baggage.DEFAULT_BAGGAGE_METADATA_KEYS`` + promotes. + """ + get = lambda name: getattr(auth, name, None) # noqa: E731 + metadata = { + meta_key: str(value) + for meta_key, attr in ( + ("user_api_key_user_id", "user_id"), + ("user_api_key_org_id", "org_id"), + ("user_api_key_alias", "key_alias"), + ("user_api_key_end_user_id", "end_user_id"), + ) + if (value := get(attr)) + } + return cls( + team_id=as_str(get("team_id")), + team_alias=as_str(get("team_alias")), + key_hash=as_str(get("api_key")), + end_user=as_str(get("end_user_id")), + metadata=metadata, + ) + + +@dataclass(frozen=True) +class GuardrailSpanData: + guardrail_name: str + mode: str | None = None + status: str | None = None + masked_entity_count: int | None = None + provider: str | None = None + action: str | None = None + # The guardrail verdict / provider response (e.g. the moderation result), + # JSON-serialized. This is the detail that belongs on the guardrail span. + response_json: str | None = None + violation_categories: tuple[str, ...] = () + confidence_score: float | None = None + risk_score: float | None = None + duration: float | None = None + # Set when the guardrail intervened/blocked or failed, so the emitter marks + # the span ERROR — a blocking guardrail is an error outcome for that span. + error: SpanError | None = None + + # Guardrail statuses that mean the guardrail did not pass the request through. + _ERROR_STATUSES: ClassVar[frozenset[str]] = frozenset( + {"guardrail_intervened", "guardrail_failed_to_respond"} + ) + + @classmethod + def from_logging_entry(cls, entry: Mapping[str, object]) -> "GuardrailSpanData": + """Build from one ``standard_logging_guardrail_information`` entry.""" + name = ( + as_str(entry.get("guardrail_name")) + or as_str(entry.get("name")) + or "guardrail" + ) + status = as_str(entry.get("guardrail_status")) or as_str(entry.get("status")) + response = entry.get("guardrail_response") + error = ( + SpanError(error_type=status, message=as_str(entry.get("guardrail_action"))) + if status in cls._ERROR_STATUSES + else None + ) + return cls( + guardrail_name=name, + mode=as_str(entry.get("guardrail_mode")) or as_str(entry.get("mode")), + status=status, + masked_entity_count=_total_masked_entities( + entry.get("masked_entity_count") + ), + provider=as_str(entry.get("guardrail_provider")), + action=as_str(entry.get("guardrail_action")), + response_json=_json_or_none(response) if response is not None else None, + violation_categories=as_str_tuple(entry.get("violation_categories")) or (), + confidence_score=as_float(entry.get("confidence_score")), + risk_score=as_float(entry.get("risk_score")), + duration=as_float(entry.get("duration")), + error=error, + ) + + +@dataclass(frozen=True) +class ServiceSpanData: + service_name: str + call_type: str | None = None + error: SpanError | None = None + # Caller-supplied attributes to stamp on the service span, passed through + # from ``async_service_*_hook(event_metadata=...)``. The mapper owns how + # these are namespaced: the canonical vocabulary uses ``litellm.metadata.*`` + # keys, the semconv-ai / Traceloop vocabulary uses the bare key names. + event_metadata: Mapping[str, str] = field(default_factory=dict) + + @classmethod + def from_payload( + cls, + payload: "ServiceLoggerPayload", + event_metadata: Mapping[str, object] | None = None, + ) -> "ServiceSpanData": + # ``payload.service`` is a ``ServiceTypes(str, Enum)`` and ``error`` is + # ``Optional[str]`` on the Pydantic model — no defensive reads needed. + # ``str(value)`` covers every case (``str(None) == "None"``). + coerced = {key: str(value) for key, value in (event_metadata or {}).items()} + return cls( + service_name=payload.service.value, + call_type=payload.call_type, + error=SpanError(message=payload.error) if payload.error else None, + event_metadata=coerced, + ) + + +@dataclass(frozen=True) +class ProxyRequestSpanData: + http_method: str + route: str + url_path: str | None = None + status_code: int | None = None + identity: RequestIdentity | None = None + + +# --- the primary LLM-call model ---------------------------------------------- # + + +@dataclass(frozen=True) +class ToolDefinition: + """A single function/tool declared on a chat-completion request.""" + + name: str + description: str | None = None + parameters_json: str | None = ( + None # JSON-serialized schema (str so it's an AttrValue) + ) + + +@dataclass(frozen=True) +class LLMCallSpanData: + operation: GenAIOperation + provider: str + request_model: str + response_model: str | None + response_id: str | None + request_params: LLMRequestParams + usage: LLMUsage + finish_reasons: tuple[str, ...] + error: SpanError | None + response_cost: float | None + server: ServerInfo | None + identity: RequestIdentity + is_streaming: bool | None = None + tools: tuple[ToolDefinition, ...] = () + # Raw messages and response, needed by vendor mappers (OpenInference, + # Langfuse, Weave) that stamp message-level attributes. ``messages_in`` is + # the request payload; ``choices_out`` mirrors ``response.choices`` from + # the StandardLoggingPayload. Both are tuples of immutable mappings so the + # dataclass stays hashable and frozen. + messages_in: tuple[Mapping[str, object], ...] = () + choices_out: tuple[Mapping[str, object], ...] = () + system_fingerprint: str | None = None + + @classmethod + def from_standard_logging_payload( + cls, payload: "StandardLoggingPayload", capture_content: bool = False + ) -> "LLMCallSpanData": + params = cast(Mapping[str, object], payload.get("model_parameters") or {}) + hidden = cast(Mapping[str, object], payload.get("hidden_params") or {}) + # Normalize ``response`` to a dict once so every field read below is a + # plain ``.get`` — no repeated ``isinstance`` guards. + raw_response = payload.get("response") + response = cast( + Mapping[str, object], raw_response if isinstance(raw_response, dict) else {} + ) + choices_out = _dicts(response.get("choices")) + # ``finish_reasons`` is metadata, not content, so derive it from + # ``choices_out`` before gating. The raw message/choice bodies are only + # retained when content capture is enabled (see ``capture_span_content``); + # otherwise the content-bearing mappers receive empty sequences and emit + # no prompt/response text. + finish_reasons = _finish_reasons(choices_out) + return cls( + operation=resolve_operation(as_str(payload.get("call_type"))), + provider=resolve_provider(as_str(payload.get("custom_llm_provider"))), + request_model=as_str(payload.get("model")) or "", + response_model=as_str(response.get("model")), + response_id=as_str(response.get("id")), + request_params=LLMRequestParams.from_model_parameters(params), + usage=LLMUsage( + input_tokens=as_int(payload.get("prompt_tokens")), + output_tokens=as_int(payload.get("completion_tokens")), + total_tokens=as_int(payload.get("total_tokens")), + ), + finish_reasons=finish_reasons, + error=_parse_error(payload), + response_cost=as_float(payload.get("response_cost")), + server=ServerInfo.from_api_base( + as_str(payload.get("api_base")) or as_str(hidden.get("api_base")) + ), + identity=RequestIdentity.from_payload(payload), + is_streaming=as_bool(payload.get("stream")), + tools=_extract_tools(params), + messages_in=_dicts(payload.get("messages")) if capture_content else (), + choices_out=choices_out if capture_content else (), + system_fingerprint=as_str(response.get("system_fingerprint")), + ) + + +def _json_or_none(value: object) -> str | None: + """JSON-serialize ``value`` (already-string values pass through). ``None`` on failure.""" + if isinstance(value, str): + return value + try: + return json.dumps(value, default=str) + except Exception: + return None + + +def _total_masked_entities(value: object) -> int | None: + """``masked_entity_count`` is a ``{entity_type: count}`` map — sum to a total.""" + if isinstance(value, Mapping): + total = sum(v for v in value.values() if isinstance(v, int)) + return total or None + return as_int(value) + + +def _dicts(value: object) -> tuple[Mapping[str, object], ...]: + """The dict items of ``value`` (when it's a list), as a tuple. Else empty.""" + if not isinstance(value, list): + return () + return tuple(item for item in value if isinstance(item, dict)) + + +def _finish_reasons(choices: tuple[Mapping[str, object], ...]) -> tuple[str, ...]: + """Non-empty ``finish_reason`` of each response choice.""" + return tuple(r for c in choices if (r := as_str(c.get("finish_reason")))) + + +def _parse_error(payload: "StandardLoggingPayload") -> SpanError | None: + """A ``SpanError`` for a failed request, or ``None`` on success.""" + if payload.get("status") != "failure": + return None + info = cast(Mapping[str, object], payload.get("error_information") or {}) + return SpanError( + error_type=as_str(info.get("error_class")) or as_str(info.get("error_code")), + message=as_str(info.get("error_message")) or as_str(payload.get("error_str")), + ) + + +def _tool_from_entry(entry: object) -> ToolDefinition | None: + """One ``tools``/``functions`` entry → ``ToolDefinition``, or ``None`` if unusable.""" + if not isinstance(entry, dict): + return None + fn = entry.get("function") if "function" in entry else entry + if not isinstance(fn, dict): + return None + name = as_str(fn.get("name")) + if not name: + return None + params = fn.get("parameters") + parameters_json: str | None = None + if params is not None: + try: + parameters_json = json.dumps(params, default=str) + except Exception: + parameters_json = None + return ToolDefinition( + name=name, + description=as_str(fn.get("description")), + parameters_json=parameters_json, + ) + + +def _extract_tools( + model_parameters: Mapping[str, object], +) -> tuple[ToolDefinition, ...]: + """Pull declared tools from request params (OpenAI / Anthropic shape). + + Accepts the chat-completion ``tools=[{"type":"function", "function": + {...}}, ...]`` shape, and falls back to the ``functions=[...]`` shape. + Returns an empty tuple when neither is present. + """ + raw_tools = model_parameters.get("tools") + if not isinstance(raw_tools, list): + raw_tools = model_parameters.get("functions") # ``functions`` shape + if not isinstance(raw_tools, list): + return () + return tuple(t for entry in raw_tools if (t := _tool_from_entry(entry)) is not None) diff --git a/litellm/integrations/otel/presets/__init__.py b/litellm/integrations/otel/presets/__init__.py new file mode 100644 index 00000000000..c69d257ab52 --- /dev/null +++ b/litellm/integrations/otel/presets/__init__.py @@ -0,0 +1,78 @@ +"""Integration presets — each one returns an :class:`OpenTelemetryV2Config`. + +A preset is a callable that reads an integration's env vars and returns an +``OpenTelemetryV2Config`` describing the exporter destination, the mapper +vocabularies to apply, and any resource attributes. ``PRESET_BY_CALLBACK`` +maps a callback name (``"arize"``, ``"langfuse_otel"``, ...) to its preset so +the factory in ``litellm_logging`` can resolve a name and build a single +``OpenTelemetryV2`` instance from the result. +""" + +from typing import Callable + +from litellm.integrations.otel.presets.agentops import agentops_preset +from litellm.integrations.otel.presets.arize import arize_dynamic_headers, arize_preset +from litellm.integrations.otel.presets.base import Preset +from litellm.integrations.otel.presets.langfuse import ( + langfuse_dynamic_headers, + langfuse_preset, +) +from litellm.integrations.otel.presets.langtrace import langtrace_preset +from litellm.integrations.otel.presets.levo import levo_preset +from litellm.integrations.otel.presets.phoenix import phoenix_preset +from litellm.integrations.otel.presets.weave import weave_dynamic_headers, weave_preset +from litellm.types.utils import StandardCallbackDynamicParams + +#: Callback name → preset. The ``Preset`` annotation makes mypy verify every +#: registered value matches the preset interface. +PRESET_BY_CALLBACK: dict[str, Preset] = { + "agentops": agentops_preset, + "arize": arize_preset, + "arize_phoenix": phoenix_preset, + "langfuse_otel": langfuse_preset, + "langtrace": langtrace_preset, + "levo": levo_preset, + "weave_otel": weave_preset, +} + +#: Callback name → per-request OTLP header builder (team/key multi-tenant +#: routing). Only integrations that support dynamic credentials appear here — +#: Arize-Phoenix/Langtrace/Levo/AgentOps don't, so they use the logger's +#: default tracer. +DYNAMIC_HEADERS_BY_CALLBACK: dict[ + str, Callable[[StandardCallbackDynamicParams], dict[str, str]] +] = { + "arize": arize_dynamic_headers, + "langfuse_otel": langfuse_dynamic_headers, + "weave_otel": weave_dynamic_headers, +} + + +def dynamic_otlp_headers( + callback_name: str | None, + dynamic_params: StandardCallbackDynamicParams | None, +) -> dict[str, str] | None: + """Per-request OTLP headers for ``callback_name``, or ``None`` if N/A. + + ``None`` means "no per-request routing" — the caller uses its default tracer. + """ + builder = DYNAMIC_HEADERS_BY_CALLBACK.get(callback_name or "") + if builder is None or not dynamic_params: + return None + headers = builder(dynamic_params) + return headers or None + + +__all__ = [ + "PRESET_BY_CALLBACK", + "DYNAMIC_HEADERS_BY_CALLBACK", + "Preset", + "dynamic_otlp_headers", + "agentops_preset", + "arize_preset", + "langfuse_preset", + "langtrace_preset", + "levo_preset", + "phoenix_preset", + "weave_preset", +] diff --git a/litellm/integrations/otel/presets/agentops.py b/litellm/integrations/otel/presets/agentops.py new file mode 100644 index 00000000000..05f8fb06354 --- /dev/null +++ b/litellm/integrations/otel/presets/agentops.py @@ -0,0 +1,139 @@ +"""AgentOps preset — OTLP/HTTP to AgentOps' endpoint with a lazily-fetched JWT. + +AgentOps authenticates with a short-lived JWT minted from the API key. Fetching +it is blocking network I/O, so it must never run on the event loop: callback +construction (where presets are built) can run inside the proxy's async startup +or, in the SDK, on the first request. Instead of fetching at config-build time, +this preset registers a custom exporter (``kind="agentops"``) that mints the JWT +**on its first export** — which the ``BatchSpanProcessor`` runs in its own +worker thread, off any event loop — and caches it for the process lifetime. +""" + +from typing import Any + +import httpx +from pydantic import Field +from pydantic_settings import BaseSettings, SettingsConfigDict + +from litellm._logging import verbose_logger +from litellm.integrations.otel.config import ExporterSpec, OpenTelemetryV2Config +from litellm.integrations.otel.providers import register_exporter_factory + +_AGENTOPS_ENDPOINT = "https://otlp.agentops.cloud/v1/traces" +_AGENTOPS_AUTH_ENDPOINT = "https://api.agentops.ai/v3/auth/token" +_AGENTOPS_EXPORTER_KIND = "agentops" + + +class _AgentOpsSettings(BaseSettings): + model_config = SettingsConfigDict(case_sensitive=False, extra="ignore") + + api_key: str | None = Field(default=None, validation_alias="AGENTOPS_API_KEY") + service_name: str = Field( + default="agentops", validation_alias="AGENTOPS_SERVICE_NAME" + ) + environment: str | None = Field( + default=None, validation_alias="AGENTOPS_ENVIRONMENT" + ) + + +def agentops_preset( + *, + config_overrides: OpenTelemetryV2Config | None = None, +) -> OpenTelemetryV2Config: + """Build the AgentOps config without any network I/O. + + The ``agentops`` exporter mints (and caches) the JWT lazily on its first + export, so this stays non-blocking. ``project.id`` is therefore not a + resource attribute — it is encoded in the JWT, which AgentOps uses to route + the trace to the right project. + """ + settings = _AgentOpsSettings() + base = config_overrides or OpenTelemetryV2Config() + return base.model_copy( + update={ + "exporters": [ + *base.exporters, + ExporterSpec( + kind=_AGENTOPS_EXPORTER_KIND, + endpoint=_AGENTOPS_ENDPOINT, + options=( + {"api_key": settings.api_key} if settings.api_key else None + ), + ), + ], + "resource_attributes": { + **base.resource_attributes, + "service.name": settings.service_name, + "telemetry.sdk.name": "agentops", + **( + {"deployment.environment": settings.environment} + if settings.environment + else {} + ), + }, + } + ) + + +def _build_agentops_exporter(spec: ExporterSpec) -> Any: + """Factory for the ``agentops`` exporter kind: a lazy-auth OTLP/HTTP exporter.""" + from opentelemetry.exporter.otlp.proto.http.trace_exporter import ( + OTLPSpanExporter, + ) + + class _LazyAuthAgentOpsExporter(OTLPSpanExporter): + """OTLP/HTTP exporter that mints the AgentOps JWT on its first export. + + ``export`` runs in the ``BatchSpanProcessor`` worker thread, so the + blocking token fetch never touches an event loop. The result is cached + after the first attempt (success or failure) so it runs at most once. + """ + + def __init__(self, *, endpoint: str | None, api_key: str | None) -> None: + super().__init__(endpoint=endpoint) + self._agentops_api_key = api_key + self._auth_resolved = False + + def _ensure_authenticated(self) -> None: + if self._auth_resolved: + return + self._auth_resolved = True + if not self._agentops_api_key: + return + try: + token = _fetch_agentops_jwt(self._agentops_api_key).get("token") + if token: + # ``_session`` is the requests.Session the base exporter + # POSTs through; updating its Authorization header is how the + # minted JWT reaches every subsequent export. + self._session.headers["Authorization"] = f"Bearer {token}" + except Exception as e: + verbose_logger.debug("AgentOps JWT fetch failed: %s", e) + + def export(self, spans: Any) -> Any: + self._ensure_authenticated() + return super().export(spans) + + options = spec.options or {} + return _LazyAuthAgentOpsExporter( + endpoint=spec.endpoint, api_key=options.get("api_key") + ) + + +def _fetch_agentops_jwt(api_key: str) -> dict[str, Any]: + # Own a short-lived client rather than ``_get_httpx_client()``: that returns + # a process-wide cached ``HTTPHandler`` whose connection pool is shared by + # every caller, so closing it here would break concurrent/subsequent + # requests. This one-shot auth call gets its own client to close. + with httpx.Client(timeout=10) as client: + response = client.post( + url=_AGENTOPS_AUTH_ENDPOINT, + headers={"Content-Type": "application/json", "Connection": "keep-alive"}, + json={"api_key": api_key}, + ) + if response.status_code != 200: + raise RuntimeError(f"Failed to fetch AgentOps token: {response.text}") + return response.json() + + +register_exporter_factory(_AGENTOPS_EXPORTER_KIND, _build_agentops_exporter) diff --git a/litellm/integrations/otel/presets/arize.py b/litellm/integrations/otel/presets/arize.py new file mode 100644 index 00000000000..acaab62df8b --- /dev/null +++ b/litellm/integrations/otel/presets/arize.py @@ -0,0 +1,75 @@ +"""Arize preset — OTLP exporter to Arize + OpenInference vocabulary.""" + +from pydantic import Field +from pydantic_settings import BaseSettings, SettingsConfigDict + +from litellm.integrations.arize.arize import ArizeLogger as _V1ArizeLogger +from litellm.integrations.otel.config import ExporterSpec, OpenTelemetryV2Config +from litellm.integrations.otel.presets.utils import ensure_mappers +from litellm.types.utils import StandardCallbackDynamicParams + + +class _ArizeSettings(BaseSettings): + model_config = SettingsConfigDict(case_sensitive=False, extra="ignore") + + # Standard OTLP headers env var, used as the fallback when no Arize + # credentials are configured. + otlp_traces_headers: str | None = Field( + default=None, validation_alias="OTEL_EXPORTER_OTLP_TRACES_HEADERS" + ) + + +def arize_preset( + *, + config_overrides: OpenTelemetryV2Config | None = None, +) -> OpenTelemetryV2Config: + arize_cfg = _V1ArizeLogger.get_arize_config() + headers = _arize_headers(arize_cfg) + base = config_overrides or OpenTelemetryV2Config() + return base.model_copy( + update={ + "exporters": [ + *base.exporters, + ExporterSpec( + kind=arize_cfg.protocol or "otlp_grpc", + endpoint=arize_cfg.endpoint or "https://otlp.arize.com/v1", + headers=headers, + ), + ], + "mapper_names": ensure_mappers(base.mapper_names, "openinference"), + "resource_attributes": { + **base.resource_attributes, + **( + {"model_id": arize_cfg.project_name} + if arize_cfg.project_name + else {} + ), + }, + } + ) + + +def _arize_headers(arize_cfg) -> str | None: + pieces = [] + if arize_cfg.space_id or arize_cfg.space_key: + pieces.append(f"space_id={arize_cfg.space_id or arize_cfg.space_key}") + if arize_cfg.api_key: + pieces.append(f"api_key={arize_cfg.api_key}") + if not pieces: + # Fall back to the standard OTLP headers env var when no Arize + # credentials are configured. + return _ArizeSettings().otlp_traces_headers + return ",".join(pieces) + + +def arize_dynamic_headers(params: StandardCallbackDynamicParams) -> dict[str, str]: + """Per-request Arize OTLP headers from team/key dynamic params.""" + headers: dict[str, str] = {} + # ``arize_space_key`` is the suggested param and wins over ``arize_space_id``. + space = params.get("arize_space_key") or params.get("arize_space_id") + if space: + headers["arize-space-id"] = space + api_key = params.get("arize_api_key") + if api_key: + headers["api_key"] = api_key + return headers diff --git a/litellm/integrations/otel/presets/base.py b/litellm/integrations/otel/presets/base.py new file mode 100644 index 00000000000..6dae2603c84 --- /dev/null +++ b/litellm/integrations/otel/presets/base.py @@ -0,0 +1,25 @@ +"""Preset interface. + +A preset is a callable that reads its integration's env vars and produces an +:class:`OpenTelemetryV2Config` (exporter list + mapper-name list + resource +attributes). This ``Protocol`` pins that contract so ``PRESET_BY_CALLBACK`` and +the factory in ``litellm_logging`` are type-checked structurally against it, +matching the ``AttributeMapper`` protocol the mappers use. +""" + +from typing import Protocol, runtime_checkable + +from litellm.integrations.otel.config import OpenTelemetryV2Config + + +@runtime_checkable +class Preset(Protocol): + """Reads an integration's env config and returns an ``OpenTelemetryV2Config``. + + ``config_overrides`` lets one preset layer onto another's config (or onto + test-supplied defaults); the factory calls presets with no arguments. + """ + + def __call__( + self, *, config_overrides: OpenTelemetryV2Config | None = None + ) -> OpenTelemetryV2Config: ... diff --git a/litellm/integrations/otel/presets/langfuse.py b/litellm/integrations/otel/presets/langfuse.py new file mode 100644 index 00000000000..7c7f2768438 --- /dev/null +++ b/litellm/integrations/otel/presets/langfuse.py @@ -0,0 +1,42 @@ +"""Langfuse-OTEL preset.""" + +from litellm.integrations.langfuse.langfuse_otel import ( + LangfuseOtelLogger as _V1Langfuse, +) +from litellm.integrations.otel.config import ExporterSpec, OpenTelemetryV2Config +from litellm.integrations.otel.presets.utils import ensure_mappers +from litellm.types.utils import StandardCallbackDynamicParams + + +def langfuse_preset( + *, + config_overrides: OpenTelemetryV2Config | None = None, +) -> OpenTelemetryV2Config: + cfg = _V1Langfuse.get_langfuse_otel_config() + base = config_overrides or OpenTelemetryV2Config() + return base.model_copy( + update={ + "exporters": [ + *base.exporters, + ExporterSpec( + kind=cfg.exporter if hasattr(cfg, "exporter") else "otlp_http", + endpoint=cfg.endpoint, + headers=cfg.headers, + ), + ], + "mapper_names": ensure_mappers(base.mapper_names, "langfuse"), + } + ) + + +def langfuse_dynamic_headers(params: StandardCallbackDynamicParams) -> dict[str, str]: + """Per-request Langfuse OTLP headers from team/key dynamic params.""" + public_key = params.get("langfuse_public_key") + secret_key = params.get("langfuse_secret_key") + if public_key and secret_key: + return { + "Authorization": _V1Langfuse._get_langfuse_authorization_header( + public_key=public_key, secret_key=secret_key + ) + } + return {} diff --git a/litellm/integrations/otel/presets/langtrace.py b/litellm/integrations/otel/presets/langtrace.py new file mode 100644 index 00000000000..c8ba5e95863 --- /dev/null +++ b/litellm/integrations/otel/presets/langtrace.py @@ -0,0 +1,22 @@ +"""Langtrace preset — Langtrace consumes generic OTLP + a vendor mapper.""" + +from litellm.integrations.otel.config import OpenTelemetryV2Config +from litellm.integrations.otel.presets.utils import ensure_mappers + + +def langtrace_preset( + *, + config_overrides: OpenTelemetryV2Config | None = None, +) -> OpenTelemetryV2Config: + """Compose the Langtrace mapper on top of the customer's OTLP destination. + + Unlike Arize / Phoenix / Langfuse, Langtrace doesn't ship its own endpoint + — users point their existing OTLP collector at Langtrace and just + need the vendor attribute schema applied to outgoing spans. + """ + base = config_overrides or OpenTelemetryV2Config() + return base.model_copy( + update={ + "mapper_names": ensure_mappers(base.mapper_names, "langtrace"), + } + ) diff --git a/litellm/integrations/otel/presets/levo.py b/litellm/integrations/otel/presets/levo.py new file mode 100644 index 00000000000..b244895a26a --- /dev/null +++ b/litellm/integrations/otel/presets/levo.py @@ -0,0 +1,24 @@ +"""Levo preset — OTLP/HTTP to a Levo collector with org+workspace headers.""" + +from litellm.integrations.levo.levo import LevoLogger as _V1Levo +from litellm.integrations.otel.config import ExporterSpec, OpenTelemetryV2Config + + +def levo_preset( + *, + config_overrides: OpenTelemetryV2Config | None = None, +) -> OpenTelemetryV2Config: + cfg = _V1Levo.get_levo_config() + base = config_overrides or OpenTelemetryV2Config() + return base.model_copy( + update={ + "exporters": [ + *base.exporters, + ExporterSpec( + kind="otlp_http", + endpoint=cfg.endpoint, + headers=cfg.otlp_auth_headers, + ), + ], + } + ) diff --git a/litellm/integrations/otel/presets/phoenix.py b/litellm/integrations/otel/presets/phoenix.py new file mode 100644 index 00000000000..92481b2abea --- /dev/null +++ b/litellm/integrations/otel/presets/phoenix.py @@ -0,0 +1,48 @@ +"""Arize-Phoenix preset.""" + +from pydantic import AliasChoices, Field +from pydantic_settings import BaseSettings, SettingsConfigDict + +from litellm.integrations.arize.arize_phoenix import ( + ArizePhoenixLogger as _V1Phoenix, +) +from litellm.integrations.otel.config import ExporterSpec, OpenTelemetryV2Config +from litellm.integrations.otel.presets.utils import ensure_mappers + + +class _PhoenixSettings(BaseSettings): + model_config = SettingsConfigDict(case_sensitive=False, extra="ignore") + + project_name: str = Field( + default="default", + validation_alias=AliasChoices( + "PHOENIX_PROJECT_NAME", "PHOENIX_COLLECTOR_PROJECT_NAME" + ), + ) + + +def phoenix_preset( + *, + config_overrides: OpenTelemetryV2Config | None = None, +) -> OpenTelemetryV2Config: + cfg = _V1Phoenix.get_arize_phoenix_config() + headers = cfg.otlp_auth_headers if hasattr(cfg, "otlp_auth_headers") else None + project_name = _PhoenixSettings().project_name + base = config_overrides or OpenTelemetryV2Config() + return base.model_copy( + update={ + "exporters": [ + *base.exporters, + ExporterSpec( + kind=cfg.protocol if hasattr(cfg, "protocol") else "otlp_http", + endpoint=cfg.endpoint, + headers=headers, + ), + ], + "mapper_names": ensure_mappers(base.mapper_names, "openinference"), + "resource_attributes": { + **base.resource_attributes, + "openinference.project.name": project_name, + }, + } + ) diff --git a/litellm/integrations/otel/presets/utils.py b/litellm/integrations/otel/presets/utils.py new file mode 100644 index 00000000000..fdf8184441d --- /dev/null +++ b/litellm/integrations/otel/presets/utils.py @@ -0,0 +1,16 @@ +"""Shared helpers for the integration presets.""" + +from typing import Iterable + + +def ensure_mappers(mapper_names: Iterable[str], *names: str) -> list[str]: + """Return ``mapper_names`` with each of ``names`` appended if not already present. + + Order is preserved and duplicates are skipped, so composing several presets + (or re-applying one) never double-adds a vocabulary. + """ + result = list(mapper_names) + for name in names: + if name not in result: + result.append(name) + return result diff --git a/litellm/integrations/otel/presets/weave.py b/litellm/integrations/otel/presets/weave.py new file mode 100644 index 00000000000..c19f8aa0b25 --- /dev/null +++ b/litellm/integrations/otel/presets/weave.py @@ -0,0 +1,43 @@ +"""Weave (W&B) preset.""" + +from litellm.integrations.otel.config import ExporterSpec, OpenTelemetryV2Config +from litellm.integrations.otel.presets.utils import ensure_mappers +from litellm.integrations.weave.weave_otel import ( + _get_weave_authorization_header, + get_weave_otel_config, +) +from litellm.types.utils import StandardCallbackDynamicParams + + +def weave_preset( + *, + config_overrides: OpenTelemetryV2Config | None = None, +) -> OpenTelemetryV2Config: + weave_cfg = get_weave_otel_config() + base = config_overrides or OpenTelemetryV2Config() + return base.model_copy( + update={ + "exporters": [ + *base.exporters, + ExporterSpec( + kind=weave_cfg.protocol or "otlp_http", + endpoint=weave_cfg.endpoint, + headers=weave_cfg.otlp_auth_headers, + ), + ], + # Weave consumes OpenInference + a small Weave-specific overlay. + "mapper_names": ensure_mappers(base.mapper_names, "openinference", "weave"), + } + ) + + +def weave_dynamic_headers(params: StandardCallbackDynamicParams) -> dict[str, str]: + """Per-request Weave OTLP headers from team/key dynamic params.""" + headers: dict[str, str] = {} + api_key = params.get("wandb_api_key") + if api_key: + headers["Authorization"] = _get_weave_authorization_header(api_key=api_key) + project_id = params.get("weave_project_id") + if project_id: + headers["project_id"] = project_id + return headers diff --git a/litellm/integrations/otel/providers.py b/litellm/integrations/otel/providers.py new file mode 100644 index 00000000000..5ad0b1eae03 --- /dev/null +++ b/litellm/integrations/otel/providers.py @@ -0,0 +1,220 @@ +"""Provider / exporter factory + the Baggage span processor.""" + +from typing import Callable, Iterable + +from opentelemetry import baggage +from opentelemetry.context import Context +from opentelemetry.sdk.resources import Resource +from opentelemetry.sdk.trace import ReadableSpan, SpanProcessor, TracerProvider +from opentelemetry.sdk.trace.export import ( + BatchSpanProcessor, + ConsoleSpanExporter, + SimpleSpanProcessor, + SpanExporter, +) +from opentelemetry.sdk.trace.export.in_memory_span_exporter import ( + InMemorySpanExporter, +) +from opentelemetry.trace import Span, SpanKind, Tracer + +from litellm.integrations.otel.config import ExporterSpec, OpenTelemetryV2Config +from litellm.integrations.otel.semconv import LiteLLM +from litellm.integrations.otel.spans import LiteLLMSpanKind + +# Re-exported so ``providers.parse_headers`` remains a stable entry point. +from litellm.integrations.otel.utils import parse_headers as parse_headers + +_SPAN_KIND_BY_ROLE_KIND: dict[LiteLLMSpanKind, SpanKind] = { + LiteLLMSpanKind.SERVER: SpanKind.SERVER, + LiteLLMSpanKind.CLIENT: SpanKind.CLIENT, + LiteLLMSpanKind.INTERNAL: SpanKind.INTERNAL, + LiteLLMSpanKind.PRODUCER: SpanKind.PRODUCER, + LiteLLMSpanKind.CONSUMER: SpanKind.CONSUMER, +} + + +def to_otel_span_kind(kind: LiteLLMSpanKind) -> SpanKind: + return _SPAN_KIND_BY_ROLE_KIND[kind] + + +# Custom exporter factories keyed by ``ExporterSpec.kind``. A preset registers +# one here when its destination needs construction logic the built-in kinds +# can't express — e.g. an exporter that fetches an auth token lazily on its +# first export (off the event loop) instead of blocking at config-build time. +# Keeping the registry here lets this module stay vendor-agnostic: the factory +# lives with the integration that needs it. +_EXPORTER_FACTORIES: dict[str, Callable[[ExporterSpec], SpanExporter]] = {} + + +def register_exporter_factory( + kind: str, factory: Callable[[ExporterSpec], SpanExporter] +) -> None: + """Register a custom exporter ``factory`` for the exporter ``kind``.""" + _EXPORTER_FACTORIES[kind.lower()] = factory + + +class LiteLLMBaggageSpanProcessor(SpanProcessor): + """Stamps an allowlisted set of Baggage entries onto every span at start.""" + + def __init__( + self, + allowed_keys: Iterable[str], + allowed_prefixes: tuple[str, ...] = (LiteLLM.METADATA_PREFIX,), + ) -> None: + self._allowed_keys = frozenset(allowed_keys) + self._allowed_prefixes = tuple(allowed_prefixes) + + def _is_allowed(self, key: str) -> bool: + return key in self._allowed_keys or any( + key.startswith(prefix) for prefix in self._allowed_prefixes + ) + + def on_start(self, span: Span, parent_context: Context | None = None) -> None: + for key, value in baggage.get_all(parent_context).items(): + if self._is_allowed(key) and isinstance(value, (str, bool, int, float)): + span.set_attribute(key, value) + + def on_end(self, span: ReadableSpan) -> None: # noqa: D401 - no-op + return None + + def shutdown(self) -> None: + return None + + def force_flush(self, timeout_millis: int = 30000) -> bool: + return True + + +def _otlp_traces_endpoint(endpoint: str | None) -> str | None: + """Point an OTLP/HTTP base endpoint at the ``/v1/traces`` signal path. + + ``OTEL_EXPORTER_OTLP_ENDPOINT`` is a base URL (e.g. ``http://host:4318``). + The OTLP/HTTP exporter only appends the ``/v1/traces`` path when it reads + that env var itself; when an endpoint is passed explicitly it is used + verbatim, so a base URL would POST to the root and the collector returns + 404. Append the signal path here (leaving an already-correct path intact). + """ + if not endpoint: + return endpoint + endpoint = endpoint.rstrip("/") + # Splunk Observability uses ``/v2/trace/otlp``; never rewrite it. + if endpoint.endswith("/v1/traces") or "/v2/trace/otlp" in endpoint: + return endpoint + for other_signal in ("/v1/logs", "/v1/metrics"): + if endpoint.endswith(other_signal): + return endpoint[: -len(other_signal)] + "/v1/traces" + return endpoint + "/v1/traces" + + +def _exporter_from_spec(spec: ExporterSpec) -> SpanExporter: + kind = (spec.kind or "console").lower() + factory = _EXPORTER_FACTORIES.get(kind) + if factory is not None: + return factory(spec) + if kind in ("in_memory", "inmemory", "memory"): + return InMemorySpanExporter() + if kind in ("otlp_http", "http", "http/protobuf", "http/json"): + from opentelemetry.exporter.otlp.proto.http.trace_exporter import ( + OTLPSpanExporter as HTTPExporter, + ) + + return HTTPExporter( + endpoint=_otlp_traces_endpoint(spec.endpoint), + headers=parse_headers(spec.headers), + ) + if kind in ("otlp_grpc", "grpc"): + from opentelemetry.exporter.otlp.proto.grpc.trace_exporter import ( + OTLPSpanExporter as GRPCExporter, + ) + + return GRPCExporter(endpoint=spec.endpoint, headers=parse_headers(spec.headers)) + return ConsoleSpanExporter() + + +def _processor_for(exporter: SpanExporter, use_simple: bool | None) -> SpanProcessor: + """Pick a Simple or Batch span processor for ``exporter``. + + When ``use_simple`` is unset, default to Simple for console and in-memory + exporters (spans export synchronously, which tests rely on) and Batch for + everything else (the right export semantics for production). + """ + if use_simple is None: + use_simple = isinstance(exporter, (ConsoleSpanExporter, InMemorySpanExporter)) + return SimpleSpanProcessor(exporter) if use_simple else BatchSpanProcessor(exporter) + + +def build_span_exporter(config: OpenTelemetryV2Config) -> SpanExporter: + """Build a single exporter from the top-level config fields. + + Convenience for the common single-exporter case (and for tests): reads the + ``exporter`` / ``endpoint`` / ``headers`` fields. To configure multiple + exporters, populate ``config.exporters`` directly. + """ + return _exporter_from_spec( + ExporterSpec( + kind=config.exporter, endpoint=config.endpoint, headers=config.headers + ) + ) + + +def build_resource(config: OpenTelemetryV2Config) -> Resource: + attributes: dict[str, str] = {"service.name": config.service_name} + if config.deployment_environment: + attributes["deployment.environment"] = config.deployment_environment + attributes.update(config.resource_attributes) + return Resource.create(attributes) + + +def build_tracer_provider( + config: OpenTelemetryV2Config, + exporter: SpanExporter | None = None, + baggage_processor: SpanProcessor | None = None, + use_simple_processor: bool | None = None, +) -> TracerProvider: + """Build the shared :class:`TracerProvider`. + + Attach the Baggage processor first (so identity attributes land on each + span before any export decision), then add one ``SpanProcessor`` per + ``config.exporters`` entry — this is what fans spans out to multiple + backends. ``exporter`` and ``use_simple_processor`` are explicit overrides: + pass a single exporter to attach exactly that one (used by tests). + """ + provider = TracerProvider(resource=build_resource(config)) + if baggage_processor is None: + baggage_processor = LiteLLMBaggageSpanProcessor( + allowed_keys=config.baggage_promoted_keys + ) + provider.add_span_processor(baggage_processor) + + if exporter is not None: + provider.add_span_processor(_processor_for(exporter, use_simple_processor)) + return provider + + # ``config._normalize`` guarantees at least one spec (it folds the top-level + # ``exporter``/``endpoint``/``headers`` fields in when ``exporters`` is empty). + for spec in config.exporters: + exp = _exporter_from_spec(spec) + provider.add_span_processor( + _processor_for( + exp, + ( + spec.use_simple_processor + if spec.use_simple_processor is not None + else use_simple_processor + ), + ) + ) + return provider + + +def get_tracer(provider: TracerProvider, name: str = "litellm") -> Tracer: + return provider.get_tracer(name) + + +def in_memory_provider( + config: OpenTelemetryV2Config | None = None, +) -> tuple[TracerProvider, InMemorySpanExporter]: + """Convenience for tests: a provider exporting to an in-memory buffer.""" + cfg = config or OpenTelemetryV2Config(exporter="in_memory") + exporter = InMemorySpanExporter() + provider = build_tracer_provider(cfg, exporter=exporter) + return provider, exporter diff --git a/litellm/integrations/otel/routing.py b/litellm/integrations/otel/routing.py new file mode 100644 index 00000000000..2c90980db94 --- /dev/null +++ b/litellm/integrations/otel/routing.py @@ -0,0 +1,98 @@ +"""Per-request multi-tenant tracer routing. + +When a request carries team/key vendor credentials in +``standard_callback_dynamic_params``, its spans must export through a +``TracerProvider`` whose OTLP headers carry those credentials. +``TenantTracerCache`` builds and caches one provider per distinct credential +set, and otherwise hands back the logger's default tracer. This lets a single +logger fan requests out to many tenants without needing a logger per tenant. +""" + +from collections import OrderedDict +from typing import Any, Mapping + +from opentelemetry.sdk.trace import TracerProvider +from opentelemetry.trace import Tracer + +from litellm._logging import verbose_logger +from litellm.integrations.otel.config import OpenTelemetryV2Config +from litellm.integrations.otel.presets import dynamic_otlp_headers +from litellm.integrations.otel.providers import build_tracer_provider, get_tracer + +# Exporter kinds that ignore headers — never rewritten with dynamic credentials. +_NON_OTLP_KINDS = ("console", "in_memory", "inmemory", "memory") + +# Cap on distinct credential-scoped providers held at once. ``dynamic_params`` +# can be populated from request metadata, so an unbounded cache lets a caller +# spawn one ``TracerProvider`` (plus its ``BatchSpanProcessor`` background +# thread) per unique credential set and exhaust the proxy. The LRU bound keeps +# the working set of active tenants resident while flushing and shutting down +# evicted providers so their threads are reclaimed. +_MAX_CACHED_PROVIDERS = 256 + + +def _shutdown_provider(provider: TracerProvider) -> None: + """Flush + stop an evicted provider's processors (reclaims their threads). + + ``TracerProvider.shutdown`` force-flushes each ``SpanProcessor`` before + stopping it, so any spans already handed to a ``BatchSpanProcessor`` are + exported rather than dropped. Best-effort: a shutdown failure must not break + the request that triggered the eviction. + """ + try: + provider.shutdown() + except Exception as e: # pragma: no cover - defensive + verbose_logger.debug("OTel V2: error shutting down evicted provider: %s", e) + + +class TenantTracerCache: + """Credential-scoped ``TracerProvider`` cache keyed by the dynamic headers.""" + + def __init__( + self, + config: OpenTelemetryV2Config, + callback_name: str | None, + tracer_name: str, + ) -> None: + self._config = config + self._callback_name = callback_name + self._tracer_name = tracer_name + self._providers: "OrderedDict[tuple[tuple[str, str], ...], TracerProvider]" = ( + OrderedDict() + ) + + def tracer_for(self, default: Tracer, dynamic_params: Any) -> Tracer: + """Return the tracer for this request. + + Use ``default`` unless the request's dynamic credentials require a + credential-scoped tracer, in which case build (or reuse) one. The cache + is a bounded LRU: the least-recently-used provider is flushed and shut + down on overflow so its exporter threads don't accumulate. + """ + headers = dynamic_otlp_headers(self._callback_name, dynamic_params) + if not headers: + return default + cache_key = tuple(sorted(headers.items())) + provider = self._providers.get(cache_key) + if provider is not None: + self._providers.move_to_end(cache_key) + else: + provider = build_tracer_provider(self._config_with_headers(headers)) + self._providers[cache_key] = provider + if len(self._providers) > _MAX_CACHED_PROVIDERS: + _, evicted = self._providers.popitem(last=False) + _shutdown_provider(evicted) + return get_tracer(provider, self._tracer_name) + + def _config_with_headers(self, headers: Mapping[str, str]) -> OpenTelemetryV2Config: + """Clone the config, replacing OTLP exporter headers with ``headers``.""" + header_str = ",".join(f"{key}={value}" for key, value in headers.items()) + exporters = [ + ( + spec + if spec.kind.lower() in _NON_OTLP_KINDS + else spec.model_copy(update={"headers": header_str}) + ) + for spec in self._config.exporters + ] + return self._config.model_copy(update={"exporters": exporters}) diff --git a/litellm/integrations/otel/semconv.py b/litellm/integrations/otel/semconv.py new file mode 100644 index 00000000000..f6e39ce5a0e --- /dev/null +++ b/litellm/integrations/otel/semconv.py @@ -0,0 +1,182 @@ +""" +Keys follow the OpenTelemetry GenAI semantic conventions (experimental). Anything +without a semconv equivalent lives under the ``litellm.*`` vendor namespace. +""" + +from enum import Enum +from typing import Final + + +class GenAIOperation(str, Enum): + """Values for ``gen_ai.operation.name``.""" + + CHAT = "chat" + TEXT_COMPLETION = "text_completion" + EMBEDDINGS = "embeddings" + GENERATE_CONTENT = "generate_content" + CREATE_AGENT = "create_agent" # reserved for future agent spans + INVOKE_AGENT = "invoke_agent" # reserved for future agent spans + EXECUTE_TOOL = "execute_tool" # reserved for future tool spans + + +class GenAIProvider(str, Enum): + """Common values for the ``gen_ai.provider.name`` attribute.""" + + OPENAI = "openai" + ANTHROPIC = "anthropic" + AWS_BEDROCK = "aws.bedrock" + AZURE_AI_OPENAI = "azure.ai.openai" + AZURE_AI_INFERENCE = "azure.ai.inference" + GCP_GEMINI = "gcp.gemini" + GCP_VERTEX_AI = "gcp.vertex_ai" + COHERE = "cohere" + MISTRAL_AI = "mistral_ai" + DEEPSEEK = "deepseek" + GROQ = "groq" + PERPLEXITY = "perplexity" + X_AI = "x_ai" + IBM_WATSONX_AI = "ibm.watsonx.ai" + + +class GenAI: + """Canonical OTel GenAI span-attribute keys.""" + + # request + OPERATION_NAME: Final = "gen_ai.operation.name" + PROVIDER_NAME: Final = "gen_ai.provider.name" + REQUEST_MODEL: Final = "gen_ai.request.model" + REQUEST_TEMPERATURE: Final = "gen_ai.request.temperature" + REQUEST_TOP_P: Final = "gen_ai.request.top_p" + REQUEST_TOP_K: Final = "gen_ai.request.top_k" + REQUEST_MAX_TOKENS: Final = "gen_ai.request.max_tokens" + REQUEST_FREQUENCY_PENALTY: Final = "gen_ai.request.frequency_penalty" + REQUEST_PRESENCE_PENALTY: Final = "gen_ai.request.presence_penalty" + REQUEST_STOP_SEQUENCES: Final = "gen_ai.request.stop_sequences" + REQUEST_SEED: Final = "gen_ai.request.seed" + REQUEST_CHOICE_COUNT: Final = "gen_ai.request.choice.count" + REQUEST_ENCODING_FORMATS: Final = "gen_ai.request.encoding_formats" + # response + RESPONSE_ID: Final = "gen_ai.response.id" + RESPONSE_MODEL: Final = "gen_ai.response.model" + RESPONSE_FINISH_REASONS: Final = "gen_ai.response.finish_reasons" + # usage + USAGE_INPUT_TOKENS: Final = "gen_ai.usage.input_tokens" + USAGE_OUTPUT_TOKENS: Final = "gen_ai.usage.output_tokens" + # content (opt-in, gated by capture mode) + INPUT_MESSAGES: Final = "gen_ai.input.messages" + OUTPUT_MESSAGES: Final = "gen_ai.output.messages" + SYSTEM_INSTRUCTIONS: Final = "gen_ai.system_instructions" + OUTPUT_TYPE: Final = "gen_ai.output.type" + CONVERSATION_ID: Final = "gen_ai.conversation.id" + # agent / tool (reserved) + AGENT_ID: Final = "gen_ai.agent.id" + AGENT_NAME: Final = "gen_ai.agent.name" + TOOL_NAME: Final = "gen_ai.tool.name" + TOOL_CALL_ID: Final = "gen_ai.tool.call.id" + + +class Error: + TYPE: Final = "error.type" + + +class Server: + ADDRESS: Final = "server.address" + PORT: Final = "server.port" + + +class HTTP: + """HTTP server-span keys. Belong on the SERVER span only (never promoted).""" + + REQUEST_METHOD: Final = "http.request.method" + ROUTE: Final = "http.route" + RESPONSE_STATUS_CODE: Final = "http.response.status_code" + URL_PATH: Final = "url.path" + + +class LiteLLM: + """Vendor-extension keys (no semconv equivalent). Always ``litellm.*``.""" + + CALL_ID: Final = "litellm.call_id" + COST_PREFIX: Final = "litellm.cost." + METADATA_PREFIX: Final = "litellm.metadata." + TEAM_ID: Final = "litellm.team.id" + TEAM_ALIAS: Final = "litellm.team.alias" + KEY_HASH: Final = "litellm.api_key.hash" + END_USER: Final = "litellm.end_user.id" + REQUEST_STREAMING: Final = "litellm.request.streaming" + GUARDRAIL_NAME: Final = "litellm.guardrail.name" + GUARDRAIL_MODE: Final = "litellm.guardrail.mode" + GUARDRAIL_STATUS: Final = "litellm.guardrail.status" + GUARDRAIL_PROVIDER: Final = "litellm.guardrail.provider" + GUARDRAIL_ACTION: Final = "litellm.guardrail.action" + GUARDRAIL_RESPONSE: Final = "litellm.guardrail.response" + GUARDRAIL_VIOLATION_CATEGORIES: Final = "litellm.guardrail.violation_categories" + GUARDRAIL_CONFIDENCE_SCORE: Final = "litellm.guardrail.confidence_score" + GUARDRAIL_RISK_SCORE: Final = "litellm.guardrail.risk_score" + GUARDRAIL_MASKED_ENTITY_COUNT: Final = "litellm.guardrail.masked_entity_count" + GUARDRAIL_DURATION: Final = "litellm.guardrail.duration" + SERVICE_NAME: Final = "litellm.service.name" + SERVICE_CALL_TYPE: Final = "litellm.service.call_type" + PREPROCESSING_MS: Final = "litellm.preprocessing.duration_ms" + + +class Metric: + """GenAI metric instrument names.""" + + TOKEN_USAGE: Final = "gen_ai.client.token.usage" + OPERATION_DURATION: Final = "gen_ai.client.operation.duration" + + +# litellm ``custom_llm_provider`` -> ``gen_ai.provider.name`` value. +_PROVIDER_BY_LITELLM: dict[str, GenAIProvider] = { + "openai": GenAIProvider.OPENAI, + "text-completion-openai": GenAIProvider.OPENAI, + "azure": GenAIProvider.AZURE_AI_OPENAI, + "azure_ai": GenAIProvider.AZURE_AI_INFERENCE, + "anthropic": GenAIProvider.ANTHROPIC, + "bedrock": GenAIProvider.AWS_BEDROCK, + "bedrock_converse": GenAIProvider.AWS_BEDROCK, + "vertex_ai": GenAIProvider.GCP_VERTEX_AI, + "vertex_ai_beta": GenAIProvider.GCP_VERTEX_AI, + "gemini": GenAIProvider.GCP_GEMINI, + "cohere": GenAIProvider.COHERE, + "cohere_chat": GenAIProvider.COHERE, + "mistral": GenAIProvider.MISTRAL_AI, + "deepseek": GenAIProvider.DEEPSEEK, + "groq": GenAIProvider.GROQ, + "perplexity": GenAIProvider.PERPLEXITY, + "xai": GenAIProvider.X_AI, + "watsonx": GenAIProvider.IBM_WATSONX_AI, +} + +# litellm ``call_type`` -> ``gen_ai.operation.name``. +_OPERATION_BY_CALL_TYPE: dict[str, GenAIOperation] = { + "completion": GenAIOperation.CHAT, + "acompletion": GenAIOperation.CHAT, + "completion_with_retries": GenAIOperation.CHAT, + "text_completion": GenAIOperation.TEXT_COMPLETION, + "atext_completion": GenAIOperation.TEXT_COMPLETION, + "embedding": GenAIOperation.EMBEDDINGS, + "aembedding": GenAIOperation.EMBEDDINGS, + "responses": GenAIOperation.CHAT, + "aresponses": GenAIOperation.CHAT, +} + + +def resolve_provider(custom_llm_provider: str | None) -> str: + """Map a litellm provider string to a ``gen_ai.provider.name`` value. + + Unknown providers pass through verbatim — the convention explicitly allows + provider-specific values, so an unmapped name is still valid. + """ + if not custom_llm_provider: + return "" + mapped = _PROVIDER_BY_LITELLM.get(custom_llm_provider.lower()) + return mapped.value if mapped is not None else custom_llm_provider + + +def resolve_operation(call_type: str | None) -> GenAIOperation: + """Map a litellm ``call_type`` to a ``gen_ai.operation.name`` value.""" + if not call_type: + return GenAIOperation.CHAT + return _OPERATION_BY_CALL_TYPE.get(call_type.lower(), GenAIOperation.CHAT) diff --git a/litellm/integrations/otel/spans.py b/litellm/integrations/otel/spans.py new file mode 100644 index 00000000000..58205f3e7a5 --- /dev/null +++ b/litellm/integrations/otel/spans.py @@ -0,0 +1,116 @@ +""" +This module declares every span the instrumentation can emit and the hierarchy. + +Span-name patterns live here as typed builder functions. + +Canonical hierarchy:: + + PROXY_REQUEST (SERVER, root) # owned by the FastAPI instrumentor + ├── LLM_CALL (CLIENT) + ├── GUARDRAIL (INTERNAL) # request-lifecycle hook, sibling of LLM_CALL + └── SERVICE (INTERNAL) + +Guardrails parent to PROXY_REQUEST, not LLM_CALL: pre/during/post-call guardrail +hooks are orchestrated by the request lifecycle (a pre-call guardrail runs +before the LLM call even starts), so a guardrail is a sibling of the LLM call, +not a child of it. The emitter parents every span to the ambient OTel context +(the active server span), which matches this. + +Management/admin endpoints are ordinary FastAPI routes — their SERVER spans are +owned by the instrumentor too, so they don't appear as a role here. +""" + +from dataclasses import dataclass +from enum import Enum +from typing import TYPE_CHECKING + +if TYPE_CHECKING: + from litellm.integrations.otel.payloads import ( + GuardrailSpanData, + LLMCallSpanData, + ProxyRequestSpanData, + ServiceSpanData, + ) + + +class SpanRole(str, Enum): + PROXY_REQUEST = "proxy_request" + LLM_CALL = "llm_call" + GUARDRAIL = "guardrail" + SERVICE = "service" + + +class LiteLLMSpanKind(str, Enum): + SERVER = "server" + CLIENT = "client" + INTERNAL = "internal" + PRODUCER = "producer" + CONSUMER = "consumer" + + +@dataclass(frozen=True) +class SpanSpec: + role: SpanRole + kind: LiteLLMSpanKind + parent: SpanRole | None + + +SPAN_REGISTRY: dict[SpanRole, SpanSpec] = { + SpanRole.PROXY_REQUEST: SpanSpec( + SpanRole.PROXY_REQUEST, LiteLLMSpanKind.SERVER, parent=None + ), + SpanRole.LLM_CALL: SpanSpec( + SpanRole.LLM_CALL, LiteLLMSpanKind.CLIENT, parent=SpanRole.PROXY_REQUEST + ), + SpanRole.GUARDRAIL: SpanSpec( + SpanRole.GUARDRAIL, LiteLLMSpanKind.INTERNAL, parent=SpanRole.PROXY_REQUEST + ), + SpanRole.SERVICE: SpanSpec( + SpanRole.SERVICE, LiteLLMSpanKind.INTERNAL, parent=SpanRole.PROXY_REQUEST + ), +} + + +# --- span name builders (the naming convention, per role) ------------------- # + + +def llm_call_span_name(data: "LLMCallSpanData") -> str: + """``"{operation} {model}"`` e.g. ``"chat gpt-4o"`` (GenAI semconv).""" + model = data.request_model or "" + return f"{data.operation.value} {model}".strip() + + +def proxy_request_span_name(data: "ProxyRequestSpanData") -> str: + """``"{method} {route}"`` (HTTP semconv).""" + return f"{data.http_method} {data.route}".strip() + + +def guardrail_span_name(data: "GuardrailSpanData") -> str: + return f"execute_guardrail {data.guardrail_name}".strip() + + +def service_span_name(data: "ServiceSpanData") -> str: + return data.service_name + + +def root_roles() -> list[SpanRole]: + """Roles that start a new trace (no in-process parent).""" + return [role for role, spec in SPAN_REGISTRY.items() if spec.parent is None] + + +def child_roles(parent: SpanRole) -> list[SpanRole]: + return [role for role, spec in SPAN_REGISTRY.items() if spec.parent == parent] + + +def validate_registry( + registry: dict[SpanRole, SpanSpec] | None = None, +) -> None: + reg = registry if registry is not None else SPAN_REGISTRY + for role, spec in reg.items(): + if spec.role is not role: + raise ValueError(f"SPAN_REGISTRY[{role}] has mismatched role {spec.role}") + if spec.parent is not None and spec.parent not in reg: + raise ValueError(f"span role {role} declares unknown parent {spec.parent}") + missing = [role for role in SpanRole if role not in reg] + if missing: + raise ValueError(f"SPAN_REGISTRY is missing roles: {missing}") diff --git a/litellm/integrations/otel/utils.py b/litellm/integrations/otel/utils.py new file mode 100644 index 00000000000..f37afc97879 --- /dev/null +++ b/litellm/integrations/otel/utils.py @@ -0,0 +1,103 @@ +"""Shared, OpenTelemetry-free helpers for the otel integration. + +Generic value coercion (for reading heterogeneous logging-payload dicts), time +conversion, and header parsing — pulled out of the individual modules so they +live in one place. Deliberately free of any ``opentelemetry`` import so the +OTel-free sources of truth (payloads, semconv, spans, config) can use it too. +""" + +from datetime import datetime + + +def as_str(value: object) -> str | None: + if value is None: + return None + if isinstance(value, str): + return value + return str(value) + + +def as_int(value: object) -> int | None: + if isinstance(value, bool): + return int(value) + if isinstance(value, int): + return value + if isinstance(value, float): + return int(value) + if isinstance(value, str): + try: + return int(value) + except ValueError: + return None + return None + + +def as_float(value: object) -> float | None: + if isinstance(value, bool): + return float(value) + if isinstance(value, (int, float)): + return float(value) + if isinstance(value, str): + try: + return float(value) + except ValueError: + return None + return None + + +def as_bool(value: object) -> bool | None: + if value is None: + return None + if isinstance(value, bool): + return value + return bool(value) + + +def as_str_tuple(value: object) -> tuple[str, ...] | None: + if value is None: + return None + if isinstance(value, str): + return (value,) + if isinstance(value, (list, tuple)): + return tuple(str(v) for v in value) + return None + + +def to_ns(value: datetime | float | int | None) -> int | None: + """Coerce a datetime / epoch value to integer nanoseconds.""" + if value is None: + return None + if isinstance(value, datetime): + return int(value.timestamp() * 1e9) + if isinstance(value, (int, float)) and not isinstance(value, bool): + return int(float(value) * 1e9) + return None + + +def to_seconds(value: datetime | float | int | str | None) -> float | None: + """Coerce a datetime / epoch / formatted-string value to epoch seconds.""" + if value is None: + return None + if isinstance(value, datetime): + return value.timestamp() + if isinstance(value, (int, float)) and not isinstance(value, bool): + return float(value) + if isinstance(value, str): + for fmt in ("%Y-%m-%d %H:%M:%S.%f", "%Y-%m-%d %H:%M:%S"): + try: + return datetime.strptime(value, fmt).timestamp() + except ValueError: + continue + return None + + +def parse_headers(raw: str | None) -> dict[str, str]: + """Parse an OTLP ``"k=v,k=v"`` header string into a dict.""" + headers: dict[str, str] = {} + if not raw: + return headers + for pair in raw.split(","): + if "=" in pair: + key, _, value = pair.partition("=") + headers[key.strip()] = value.strip() + return headers diff --git a/litellm/litellm_core_utils/litellm_logging.py b/litellm/litellm_core_utils/litellm_logging.py index 97266096ef9..eb61bf7731b 100644 --- a/litellm/litellm_core_utils/litellm_logging.py +++ b/litellm/litellm_core_utils/litellm_logging.py @@ -3718,6 +3718,9 @@ def _init_custom_logger_compatible_class( # noqa: PLR0915 try: custom_logger_init_args = custom_logger_init_args or {} if logging_integration == "agentops": # Add AgentOps initialization + _v2 = _maybe_construct_otel_v2("agentops", _in_memory_loggers) + if _v2 is not None: + return _v2 # type: ignore for callback in _in_memory_loggers: if isinstance(callback, AgentOps): return callback # type: ignore @@ -3870,6 +3873,9 @@ def _init_custom_logger_compatible_class( # noqa: PLR0915 _in_memory_loggers.append(_opik_logger) return _opik_logger # type: ignore elif logging_integration == "arize": + _v2 = _maybe_construct_otel_v2("arize", _in_memory_loggers) + if _v2 is not None: + return _v2 # type: ignore from litellm.integrations.opentelemetry import ( OpenTelemetry, OpenTelemetryConfig, @@ -3899,6 +3905,9 @@ def _init_custom_logger_compatible_class( # noqa: PLR0915 _in_memory_loggers.append(_arize_otel_logger) return _arize_otel_logger # type: ignore elif logging_integration == "arize_phoenix": + _v2 = _maybe_construct_otel_v2("arize_phoenix", _in_memory_loggers) + if _v2 is not None: + return _v2 # type: ignore from litellm.integrations.opentelemetry import ( OpenTelemetry, OpenTelemetryConfig, @@ -3929,6 +3938,9 @@ def _init_custom_logger_compatible_class( # noqa: PLR0915 _in_memory_loggers.append(_arize_phoenix_otel_logger) return _arize_phoenix_otel_logger # type: ignore elif logging_integration == "levo": + _v2 = _maybe_construct_otel_v2("levo", _in_memory_loggers) + if _v2 is not None: + return _v2 # type: ignore from litellm.integrations.levo.levo import LevoLogger from litellm.integrations.opentelemetry import ( OpenTelemetry, @@ -3954,6 +3966,28 @@ def _init_custom_logger_compatible_class( # noqa: PLR0915 _in_memory_loggers.append(_levo_otel_logger) return _levo_otel_logger # type: ignore elif logging_integration == "otel": + # Gate the new typed V2 adapter behind LITELLM_OTEL_V2. When off, + # the legacy 3,227-line god-class is used unchanged. The two are + # never registered simultaneously — the dedup loop below treats + # any module under ``litellm.integrations.otel`` or + # ``litellm.integrations.opentelemetry`` as "the OTel callback". + from litellm.integrations.otel.config import is_otel_v2_enabled + + if is_otel_v2_enabled(): + from litellm.integrations.otel.logger import OpenTelemetryV2 + + for callback in _in_memory_loggers: + if type(callback) is OpenTelemetryV2: + return callback # type: ignore + otel_logger_v2 = OpenTelemetryV2( + **_get_custom_logger_settings_from_proxy_server( + callback_name=logging_integration + ) + ) + _in_memory_loggers.append(otel_logger_v2) + _maybe_auto_initialize_arize_phoenix(_in_memory_loggers) + return otel_logger_v2 # type: ignore + from litellm.integrations.opentelemetry import OpenTelemetry for callback in _in_memory_loggers: @@ -4092,6 +4126,9 @@ def _init_custom_logger_compatible_class( # noqa: PLR0915 elif logging_integration == "langtrace": if "LANGTRACE_API_KEY" not in os.environ: raise ValueError("LANGTRACE_API_KEY not found in environment variables") + _v2 = _maybe_construct_otel_v2("langtrace", _in_memory_loggers) + if _v2 is not None: + return _v2 # type: ignore from litellm.integrations.opentelemetry import ( OpenTelemetry, @@ -4132,6 +4169,9 @@ def _init_custom_logger_compatible_class( # noqa: PLR0915 _in_memory_loggers.append(langfuse_logger) return langfuse_logger # type: ignore elif logging_integration == "langfuse_otel": + _v2 = _maybe_construct_otel_v2("langfuse_otel", _in_memory_loggers) + if _v2 is not None: + return _v2 # type: ignore from litellm.integrations.langfuse.langfuse_otel import LangfuseOtelLogger for callback in _in_memory_loggers: @@ -4148,6 +4188,9 @@ def _init_custom_logger_compatible_class( # noqa: PLR0915 _in_memory_loggers.append(_otel_logger) return _otel_logger # type: ignore elif logging_integration == "weave_otel": + _v2 = _maybe_construct_otel_v2("weave_otel", _in_memory_loggers) + if _v2 is not None: + return _v2 # type: ignore from litellm.integrations.opentelemetry import OpenTelemetryConfig from litellm.integrations.weave.weave_otel import ( WeaveOtelLogger, @@ -4296,6 +4339,42 @@ def _init_custom_logger_compatible_class( # noqa: PLR0915 return None +def _maybe_construct_otel_v2( + callback_name: str, _in_memory_loggers: list +) -> Optional[Any]: + """If ``LITELLM_OTEL_V2`` is on, build (or reuse) a single ``OpenTelemetryV2`` + instance configured via the preset for ``callback_name``. + + Returns ``None`` when V2 is off OR when there's no preset registered for + ``callback_name`` — callers should then fall through to the legacy path. + """ + from litellm.integrations.otel.config import is_otel_v2_enabled + + if not is_otel_v2_enabled(): + return None + from litellm.integrations.otel.logger import OpenTelemetryV2 + from litellm.integrations.otel.presets import PRESET_BY_CALLBACK + + preset_fn = PRESET_BY_CALLBACK.get(callback_name) + if preset_fn is None: + return None + for callback in _in_memory_loggers: + if ( + isinstance(callback, OpenTelemetryV2) + and getattr(callback, "callback_name", None) == callback_name + ): + return callback + try: + config = preset_fn() + except Exception: + # If env vars are missing or the preset raises, defer to the legacy path + # so customers get the same error story they had before V2 landed. + return None + v2_logger = OpenTelemetryV2(config=config, callback_name=callback_name) + _in_memory_loggers.append(v2_logger) + return v2_logger + + def _maybe_auto_initialize_arize_phoenix(_in_memory_loggers: list) -> None: """ Auto-initialize ArizePhoenixLogger when Phoenix env vars are detected. diff --git a/litellm/proxy/proxy_server.py b/litellm/proxy/proxy_server.py index 8fbe6d97dbc..4243f93eb31 100644 --- a/litellm/proxy/proxy_server.py +++ b/litellm/proxy/proxy_server.py @@ -825,6 +825,37 @@ async def proxy_startup_event(app: FastAPI): # noqa: PLR0915 if isinstance(worker_config, dict): await initialize(**worker_config) + ## V2 OTEL: now that config (and therefore the callbacks) is loaded, publish + ## the chosen V2 logger's TracerProvider as the OTel global. The FastAPI + ## instrumentation mounted at app-creation binds to the global provider, so + ## this is what makes server spans and gen-ai spans share one provider and + ## land in the same trace. Prefer an already-registered preset logger + ## (arize, langfuse, …) so server spans export to that backend too; otherwise + ## build a generic one from OTEL_* envs. ``set_tracer_provider`` only takes + ## effect once, so the first configured logger wins. + try: + from litellm.integrations.otel.config import is_otel_v2_enabled + + if is_otel_v2_enabled(): + from opentelemetry import trace as _otel_trace + + from litellm.integrations.otel.logger import OpenTelemetryV2 + + _otel_v2_logger = ( + next( + ( + cb + for cb in litellm.service_callback + if isinstance(cb, OpenTelemetryV2) + ), + None, + ) + or OpenTelemetryV2() + ) + _otel_trace.set_tracer_provider(_otel_v2_logger._tracer_provider) + except Exception as e: + verbose_proxy_logger.debug("Skipping OTel V2 provider setup: %s", e) + # check if DATABASE_URL in environment - load from there if prisma_client is None: _db_url: Optional[str] = get_secret("DATABASE_URL", None) # type: ignore @@ -1061,6 +1092,56 @@ def ensure_unique_openapi_operation_ids( return openapi_schema +# Passthrough routes are catch-alls (e.g. "/openai/{endpoint:path}"), so the +# default OTel server-span name "{method} {route}" collapses every upstream +# endpoint into "POST /openai/{endpoint:path}". Rename those spans to the real +# request path so each endpoint is distinguishable. Non-catch-all routes keep +# their low-cardinality template name. +_OTEL_V2_PASSTHROUGH_PREFIXES = frozenset( + { + "openai", + "openai_passthrough", + "anthropic", + "azure", + "azure_ai", + "bedrock", + "cohere", + "cursor", + "gemini", + "mistral", + "vllm", + "vertex_ai", + "vertex-ai", + "assemblyai", + "eu.assemblyai", + "milvus", + } +) + + +def _otel_v2_passthrough_span_name_hook(span: Any, scope: dict) -> None: + """FastAPI ``server_request_hook``: give passthrough server spans a useful name. + + The instrumentation matches the route at span creation, so both the span name + and ``http.route`` are set to the catch-all template (``/openai/{endpoint:path}``) + before this hook runs. Rewrite both to the real request path so each upstream + endpoint is distinguishable. (The ASGI ``http receive``/``http send`` sub-spans + can't be renamed from here — their name is captured at creation — so they are + dropped via ``exclude_spans`` at instrumentation time.) + """ + try: + if span is None or not span.is_recording(): + return + path = scope.get("path") or "" + method = scope.get("method") or "" + first_segment = path.lstrip("/").split("/", 1)[0] + if first_segment in _OTEL_V2_PASSTHROUGH_PREFIXES: + span.update_name(f"{method} {path}".strip()) + span.set_attribute("http.route", path) + except Exception: + pass + + app = FastAPI( docs_url=_get_docs_url(), redoc_url=_get_redoc_url(), @@ -1074,6 +1155,41 @@ app = FastAPI( strict_content_type=False, ) +## V2 OTEL: instrument the FastAPI app for server spans (gated by +## LITELLM_OTEL_V2; lazy imports keep the package optional). This MUST run at +## app-creation time — once the lifespan runs, the middleware stack is frozen +## and ``instrument_app`` raises "Cannot add middleware after an application has +## started". No TracerProvider is passed, so the instrumentation binds to the +## OTel global ``ProxyTracerProvider``; ``proxy_startup_event`` sets the real +## provider as the global after config load, and the proxy delegates to it. That +## way server spans and gen-ai spans share one provider and the same trace. +try: + from litellm.integrations.otel.config import is_otel_v2_enabled + + if is_otel_v2_enabled(): + from opentelemetry.instrumentation.fastapi import FastAPIInstrumentor + + # Drop health-check spans by default — load balancers poll + # /health/readiness and /health/liveness constantly, which floods traces + # with noise. Honor the standard OTel env var so operators can override + # (e.g. set it to "" to trace everything, or add their own paths). + _otel_excluded_urls = ( + os.environ.get("OTEL_PYTHON_FASTAPI_EXCLUDED_URLS") + if "OTEL_PYTHON_FASTAPI_EXCLUDED_URLS" in os.environ + else "/health" + ) + FastAPIInstrumentor.instrument_app( + app, + excluded_urls=_otel_excluded_urls, + server_request_hook=_otel_v2_passthrough_span_name_hook, + # Drop the ASGI "http receive"/"http send" lifecycle sub-spans: they + # are low-value noise and (for passthrough) carry the catch-all route + # template in their name, which can't be rewritten from a hook. + exclude_spans=["receive", "send"], + ) +except Exception as e: + verbose_proxy_logger.debug("Skipping OTel V2 FastAPI instrumentation: %s", e) + vertex_live_passthrough_vertex_base = VertexBase() @@ -1238,6 +1354,17 @@ def _close_dangling_otel_server_span(request: Request, status_code: int) -> None return if open_telemetry_logger is None: return + # Under OTel V2 the FastAPI instrumentor owns the server span (parent_otel_span + # is that same span), and it records the error + ends it itself. Ending it here + # would end it early — losing the http.* attributes the instrumentor stamps on + # completion — and double-end it. Leave it to the instrumentor. + try: + from litellm.integrations.otel.config import is_otel_v2_enabled + + if is_otel_v2_enabled(): + return + except Exception: + pass try: from opentelemetry.trace import Status, StatusCode diff --git a/pyproject.toml b/pyproject.toml index 92ac526f570..8de24d2c2ea 100644 --- a/pyproject.toml +++ b/pyproject.toml @@ -119,6 +119,7 @@ proxy-runtime = [ "opentelemetry-api==1.28.0", "opentelemetry-sdk==1.28.0", "opentelemetry-exporter-otlp==1.28.0", + "opentelemetry-instrumentation-fastapi==0.49b0", "ddtrace>=2.19.0,<3.0", "sentry-sdk>=2.21.0,<3.0", "mangum>=0.17.0,<1.0", @@ -160,6 +161,7 @@ dev = [ "opentelemetry-api==1.28.0", "opentelemetry-sdk==1.28.0", "opentelemetry-exporter-otlp==1.28.0", + "opentelemetry-instrumentation-fastapi==0.49b0", "langfuse==2.59.7", "fastapi-offline==1.7.6", "fakeredis==2.34.1", @@ -178,6 +180,7 @@ proxy-dev = [ "opentelemetry-api==1.28.0", "opentelemetry-sdk==1.28.0", "opentelemetry-exporter-otlp==1.28.0", + "opentelemetry-instrumentation-fastapi==0.49b0", "azure-identity==1.25.2", "a2a-sdk==0.3.24", ] diff --git a/tests/test_litellm/integrations/otel/test_otel_v2_baggage.py b/tests/test_litellm/integrations/otel/test_otel_v2_baggage.py new file mode 100644 index 00000000000..74c59cb3b59 --- /dev/null +++ b/tests/test_litellm/integrations/otel/test_otel_v2_baggage.py @@ -0,0 +1,134 @@ +"""Tests for Baggage-based promotion of request-scoped attributes onto every span, +and the two antipattern boundaries: http.* is never promoted, and the full +metadata blob is never promoted (only the bounded allowlist).""" + +import pytest + +pytest.importorskip("opentelemetry") + +from litellm.integrations.otel import ( # noqa: E402 + GenAI, + HTTP, + LiteLLM, + OpenTelemetryV2Config, + promoted_baggage, +) +from litellm.integrations.otel import context as ctx_mod # noqa: E402 +from litellm.integrations.otel import providers # noqa: E402 +from litellm.integrations.otel.emitter import SpanEmitter # noqa: E402 +from litellm.integrations.otel.payloads import ( # noqa: E402 + GuardrailSpanData, + LLMCallSpanData, + ServiceSpanData, +) +from litellm.integrations.otel.baggage import BAGGAGE_PROMOTED_KEYS # noqa: E402 +from litellm.integrations.otel.spans import SpanRole # noqa: E402 + + +def _payload(): + return { + "call_type": "acompletion", + "custom_llm_provider": "openai", + "model": "gpt-4o", + "prompt_tokens": 1, + "completion_tokens": 1, + "total_tokens": 2, + "metadata": { + "team_id": "t1", + "team_alias": "team one", + "user_api_key_hash": "hsh", + "user_api_key_org_id": "org1", + "private_note": "do-not-promote", + }, + "status": "success", + "litellm_call_id": "call_1", + "hidden_params": {}, + } + + +def _engine_and_exporter(config=None): + cfg = config or OpenTelemetryV2Config(exporter="in_memory") + provider, exporter = providers.in_memory_provider(cfg) + tracer = providers.get_tracer(provider, "litellm-baggage-test") + return SpanEmitter(tracer, cfg), exporter + + +def test_identity_promoted_onto_every_span(): + engine, exporter = _engine_and_exporter() + data = LLMCallSpanData.from_standard_logging_payload(_payload()) + bag = promoted_baggage(data.identity, data.request_model, BAGGAGE_PROMOTED_KEYS) + ctx = ctx_mod.set_request_baggage(bag) + + root = engine.start_span(SpanRole.PROXY_REQUEST, "POST /chat/completions", ctx) + root_ctx = ctx_mod.context_from_span(root, ctx) + engine.emit(SpanRole.LLM_CALL, data, parent_context=root_ctx) + engine.emit( + SpanRole.GUARDRAIL, GuardrailSpanData("presidio", status="success"), root_ctx + ) + engine.emit(SpanRole.SERVICE, ServiceSpanData("redis", call_type="set"), root_ctx) + root.end() + + spans = exporter.get_finished_spans() + assert len(spans) == 4 + for span in spans: + assert span.attributes.get(LiteLLM.TEAM_ID) == "t1" + assert span.attributes.get(LiteLLM.TEAM_ALIAS) == "team one" + assert span.attributes.get(GenAI.REQUEST_MODEL) == "gpt-4o" + + +def test_allowlisted_metadata_subkey_promoted_blob_excluded(): + engine, exporter = _engine_and_exporter() + data = LLMCallSpanData.from_standard_logging_payload(_payload()) + bag = promoted_baggage(data.identity, data.request_model, BAGGAGE_PROMOTED_KEYS) + ctx = ctx_mod.set_request_baggage(bag) + engine.emit(SpanRole.SERVICE, ServiceSpanData("redis", call_type="set"), ctx) + (span,) = exporter.get_finished_spans() + # allowlisted metadata sub-key is promoted + assert ( + span.attributes.get(f"{LiteLLM.METADATA_PREFIX}user_api_key_org_id") == "org1" + ) + # non-allowlisted metadata is NOT promoted (no full-blob dumping) + assert all("private_note" not in k for k in span.attributes) + + +def test_http_attributes_never_promoted(): + """Even if http.* is present in baggage, the processor must not stamp it on + child spans (it belongs on the SERVER span only).""" + engine, exporter = _engine_and_exporter() + ctx = ctx_mod.set_request_baggage( + { + LiteLLM.TEAM_ID: "t1", + HTTP.ROUTE: "/chat/completions", + HTTP.REQUEST_METHOD: "POST", + } + ) + engine.emit(SpanRole.SERVICE, ServiceSpanData("redis", call_type="set"), ctx) + (span,) = exporter.get_finished_spans() + assert span.attributes.get(LiteLLM.TEAM_ID) == "t1" + assert HTTP.ROUTE not in span.attributes + assert HTTP.REQUEST_METHOD not in span.attributes + + +def test_arbitrary_upstream_baggage_not_promoted(): + engine, exporter = _engine_and_exporter() + ctx = ctx_mod.set_request_baggage( + {LiteLLM.TEAM_ID: "t1", "some.upstream.key": "leak"} + ) + engine.emit(SpanRole.SERVICE, ServiceSpanData("redis", call_type="set"), ctx) + (span,) = exporter.get_finished_spans() + assert span.attributes.get(LiteLLM.TEAM_ID) == "t1" + assert "some.upstream.key" not in span.attributes + + +def test_baggage_processor_allowlist_can_be_widened(): + cfg = OpenTelemetryV2Config( + exporter="in_memory", + baggage_promoted_keys=[LiteLLM.TEAM_ID, "custom.key"], + ) + engine, exporter = _engine_and_exporter(cfg) + ctx = ctx_mod.set_request_baggage({"custom.key": "v", LiteLLM.TEAM_ALIAS: "ta"}) + engine.emit(SpanRole.SERVICE, ServiceSpanData("redis"), ctx) + (span,) = exporter.get_finished_spans() + assert span.attributes.get("custom.key") == "v" + # team_alias not in this config's allowlist -> not promoted + assert LiteLLM.TEAM_ALIAS not in span.attributes diff --git a/tests/test_litellm/integrations/otel/test_otel_v2_components.py b/tests/test_litellm/integrations/otel/test_otel_v2_components.py new file mode 100644 index 00000000000..7caac405880 --- /dev/null +++ b/tests/test_litellm/integrations/otel/test_otel_v2_components.py @@ -0,0 +1,395 @@ +"""Coverage for the engine-layer components: providers/exporters, context + +baggage helpers, metrics, the typed coercion helpers, mapper branches, span-name +builders, and the registry validator's failure paths. Needs the OTel SDK.""" + +import pytest + +pytest.importorskip("opentelemetry") + +from opentelemetry.sdk.metrics import MeterProvider # noqa: E402 +from opentelemetry.sdk.metrics.export import InMemoryMetricReader # noqa: E402 +from opentelemetry.sdk.trace.export import ( # noqa: E402 + BatchSpanProcessor, + ConsoleSpanExporter, + SimpleSpanProcessor, +) +from opentelemetry.sdk.trace.export.in_memory_span_exporter import ( # noqa: E402 + InMemorySpanExporter, +) +from opentelemetry.trace import SpanKind # noqa: E402 + +from litellm.integrations.otel import context as ctx_mod # noqa: E402 +from litellm.integrations.otel import providers # noqa: E402 +from litellm.integrations.otel.config import OpenTelemetryV2Config # noqa: E402 +from litellm.integrations.otel.mappers.genai import GenAIMapper # noqa: E402 +from litellm.integrations.otel.mappers.legacy import LegacyMapper # noqa: E402 +from litellm.integrations.otel.metrics import create_genai_metrics # noqa: E402 +from litellm.integrations.otel.payloads import ( # noqa: E402 + GuardrailSpanData, + LLMCallSpanData, + LLMRequestParams, + LLMUsage, + ProxyRequestSpanData, + RequestIdentity, + ServerInfo, + ServiceSpanData, + SpanError, +) +from litellm.integrations.otel.semconv import GenAI, GenAIOperation +from litellm.integrations.otel.spans import ( # noqa: E402 + SPAN_REGISTRY, + LiteLLMSpanKind, + SpanRole, + SpanSpec, + guardrail_span_name, + proxy_request_span_name, + service_span_name, + validate_registry, +) +from litellm.integrations.otel.utils import ( # noqa: E402 + as_bool, + as_float, + as_int, + as_str, + as_str_tuple, +) + +# --- typed coercion helpers ------------------------------------------------- # + + +def test_as_str(): + assert as_str(None) is None + assert as_str("x") == "x" + assert as_str(5) == "5" + + +def test_as_int(): + assert as_int(True) == 1 + assert as_int(3) == 3 + assert as_int(3.9) == 3 + assert as_int("7") == 7 + assert as_int("nope") is None + assert as_int(None) is None + + +def test_as_float(): + assert as_float(True) == 1.0 + assert as_float(2) == 2.0 + assert as_float("1.5") == 1.5 + assert as_float("nope") is None + assert as_float(None) is None + + +def test_as_bool(): + assert as_bool(None) is None + assert as_bool(True) is True + assert as_bool(1) is True + assert as_bool(0) is False + + +def test_as_str_tuple(): + assert as_str_tuple(None) is None + assert as_str_tuple("a") == ("a",) + assert as_str_tuple(["a", 2]) == ("a", "2") + assert as_str_tuple(123) is None + + +def test_request_params_max_completion_tokens_fallback(): + params = LLMRequestParams.from_model_parameters({"max_completion_tokens": 99}) + assert params.max_tokens == 99 + + +def test_server_info_from_api_base(): + assert ServerInfo.from_api_base(None) is None + assert ServerInfo.from_api_base("api.host.com:8080") == ServerInfo( + "api.host.com", 8080 + ) + assert ServerInfo.from_api_base("https://h.com/v1") == ServerInfo("h.com", None) + # scheme present but empty netloc -> no hostname + assert ServerInfo.from_api_base("http:///v1") is None + + +def test_service_span_data_from_payload(): + class _Service: + value = "redis" + + class _Payload: + service = _Service() + call_type = "async_set_cache" + error = None + + data = ServiceSpanData.from_payload(_Payload()) + assert data.service_name == "redis" + assert data.call_type == "async_set_cache" + assert data.error is None + + class _FailPayload: + service = _Service() + call_type = "async_set_cache" + error = "boom" + + failed = ServiceSpanData.from_payload(_FailPayload()) + assert failed.error is not None + assert failed.error.message == "boom" + + +# --- span name builders ----------------------------------------------------- # + + +def test_name_builders(): + assert ( + proxy_request_span_name(ProxyRequestSpanData("POST", "/chat/completions")) + == "POST /chat/completions" + ) + assert service_span_name(ServiceSpanData("redis")) == "redis" + assert ( + guardrail_span_name(GuardrailSpanData("presidio")) + == "execute_guardrail presidio" + ) + + +# --- registry validator failure paths --------------------------------------- # + + +def test_validate_registry_detects_role_mismatch(): + bad = {SpanRole.LLM_CALL: SpanSpec(SpanRole.SERVICE, LiteLLMSpanKind.CLIENT, None)} + with pytest.raises(ValueError, match="mismatched role"): + validate_registry(bad) + + +def test_validate_registry_detects_unknown_parent(): + bad = { + SpanRole.LLM_CALL: SpanSpec( + SpanRole.LLM_CALL, LiteLLMSpanKind.CLIENT, parent=SpanRole.PROXY_REQUEST + ) + } + with pytest.raises(ValueError, match="unknown parent"): + validate_registry(bad) + + +def test_validate_registry_detects_missing_roles(): + partial = { + SpanRole.PROXY_REQUEST: SPAN_REGISTRY[SpanRole.PROXY_REQUEST], + } + with pytest.raises(ValueError, match="missing roles"): + validate_registry(partial) + + +# --- mappers (full branch coverage) ----------------------------------------- # + + +def _full_llm_call(): + return LLMCallSpanData( + operation=GenAIOperation.CHAT, + provider="openai", + request_model="gpt-4o", + response_model="gpt-4o-2024", + response_id="resp_1", + request_params=LLMRequestParams( + temperature=0.7, + top_p=0.9, + top_k=40, + max_tokens=256, + frequency_penalty=0.1, + presence_penalty=0.2, + stop_sequences=("STOP",), + seed=42, + ), + usage=LLMUsage(input_tokens=10, output_tokens=5, total_tokens=15), + finish_reasons=("stop",), + error=None, + response_cost=0.002, + server=ServerInfo("api.openai.com", 443), + identity=RequestIdentity(call_id="c1"), + is_streaming=True, + ) + + +def test_genai_mapper_all_request_params(): + attrs = GenAIMapper().map(_full_llm_call()) + assert attrs[GenAI.REQUEST_TOP_P] == 0.9 + assert attrs[GenAI.REQUEST_TOP_K] == 40 + assert attrs[GenAI.REQUEST_MAX_TOKENS] == 256 + assert attrs[GenAI.REQUEST_FREQUENCY_PENALTY] == 0.1 + assert attrs[GenAI.REQUEST_PRESENCE_PENALTY] == 0.2 + assert attrs[GenAI.REQUEST_STOP_SEQUENCES] == ["STOP"] + assert attrs[GenAI.REQUEST_SEED] == 42 + assert attrs["server.port"] == 443 + + +def test_genai_mapper_guardrail_and_service(): + from litellm.integrations.otel.semconv import LiteLLM + + g = GenAIMapper().map(GuardrailSpanData("presidio", mode="pre")) + assert g[LiteLLM.GUARDRAIL_NAME] == "presidio" + assert g[LiteLLM.GUARDRAIL_MODE] == "pre" + + s = GenAIMapper().map(ServiceSpanData("redis", call_type="set")) + assert s[LiteLLM.SERVICE_NAME] == "redis" + assert s[LiteLLM.SERVICE_CALL_TYPE] == "set" + + +def test_legacy_mapper_all_request_params(): + attrs = LegacyMapper().map(_full_llm_call()) + assert attrs["llm.top_k"] == 40 + assert attrs["llm.frequency_penalty"] == 0.1 + assert attrs["llm.presence_penalty"] == 0.2 + assert attrs["llm.chat.stop_sequences"] == ["STOP"] + assert attrs["gen_ai.usage.total_tokens"] == 15 + + +def test_legacy_mapper_covers_service_with_v1_bare_keys(): + """Service spans dual-emit V1's bare ``service``/``call_type``/``error`` keys.""" + attrs = LegacyMapper().map( + ServiceSpanData("redis", call_type="set", event_metadata={"k": "v"}), + ) + assert attrs["service"] == "redis" + assert attrs["call_type"] == "set" + assert attrs["k"] == "v" # event_metadata is stamped bare (V1 behavior) + + +def test_legacy_mapper_skips_guardrail_role(): + """Guardrail spans never had a V1 vocabulary; legacy mapper returns ``{}``.""" + assert LegacyMapper().map(GuardrailSpanData("presidio")) == {} + + +# --- metrics ---------------------------------------------------------------- # + + +def test_create_genai_metrics_records(): + reader = InMemoryMetricReader() + meter = MeterProvider(metric_readers=[reader]).get_meter("test") + metrics = create_genai_metrics(meter) + metrics.token_usage.record(10, {"x": "y"}) + metrics.operation_duration.record(0.5, {"x": "y"}) + data = reader.get_metrics_data() + assert data is not None + + +# --- context + baggage helpers ---------------------------------------------- # + + +def test_extract_traceparent(): + valid = {"traceparent": "00-0af7651916cd43dd8448eb211c80319c-b7ad6b7169203331-01"} + assert ctx_mod.extract_traceparent(valid) is not None + assert ctx_mod.extract_traceparent({"x": "y"}) is None + + +def test_set_request_baggage_empty_returns_context(): + assert ctx_mod.set_request_baggage({}) is not None + + +def test_get_baggage_attributes_roundtrip(): + ctx = ctx_mod.set_request_baggage({"litellm.team.id": "t1"}) + assert ctx_mod.get_baggage_attributes(ctx)["litellm.team.id"] == "t1" + + +# --- providers -------------------------------------------------------------- # + + +def test_to_otel_span_kind_covers_all(): + assert providers.to_otel_span_kind(LiteLLMSpanKind.SERVER) is SpanKind.SERVER + assert providers.to_otel_span_kind(LiteLLMSpanKind.CLIENT) is SpanKind.CLIENT + assert providers.to_otel_span_kind(LiteLLMSpanKind.INTERNAL) is SpanKind.INTERNAL + assert providers.to_otel_span_kind(LiteLLMSpanKind.PRODUCER) is SpanKind.PRODUCER + assert providers.to_otel_span_kind(LiteLLMSpanKind.CONSUMER) is SpanKind.CONSUMER + + +def test_parse_headers(): + assert providers.parse_headers(None) == {} + assert providers.parse_headers("a=1,b=2") == {"a": "1", "b": "2"} + assert providers.parse_headers("no-equals") == {} + + +def test_otlp_traces_endpoint_normalization(): + norm = providers._otlp_traces_endpoint + # A base endpoint gets the signal path appended (the common OTLP env shape). + assert norm("http://collector:4318") == "http://collector:4318/v1/traces" + assert norm("http://collector:4318/") == "http://collector:4318/v1/traces" + # An already-correct path is left intact. + assert norm("http://collector:4318/v1/traces") == "http://collector:4318/v1/traces" + # Another signal's path is rewritten to traces. + assert norm("http://collector:4318/v1/logs") == "http://collector:4318/v1/traces" + # Splunk's path is preserved; None passes through. + assert ( + norm("https://x.splunk.com/v2/trace/otlp") + == "https://x.splunk.com/v2/trace/otlp" + ) + assert norm(None) is None + + +def test_build_span_exporter_variants(): + assert isinstance( + providers.build_span_exporter(OpenTelemetryV2Config(exporter="console")), + ConsoleSpanExporter, + ) + assert isinstance( + providers.build_span_exporter(OpenTelemetryV2Config(exporter="in_memory")), + InMemorySpanExporter, + ) + assert isinstance( + providers.build_span_exporter(OpenTelemetryV2Config(exporter="unknown")), + ConsoleSpanExporter, + ) + http_exporter = providers.build_span_exporter( + OpenTelemetryV2Config(exporter="otlp_http", endpoint="http://h:4318") + ) + assert "OTLPSpanExporter" in type(http_exporter).__name__ + grpc_exporter = providers.build_span_exporter( + OpenTelemetryV2Config(exporter="otlp_grpc", endpoint="http://h:4317") + ) + assert "OTLPSpanExporter" in type(grpc_exporter).__name__ + + +def test_build_resource_includes_deployment_environment(): + resource = providers.build_resource( + OpenTelemetryV2Config(service_name="svc", deployment_environment="prod") + ) + assert resource.attributes["service.name"] == "svc" + assert resource.attributes["deployment.environment"] == "prod" + + +def test_build_tracer_provider_processor_selection(): + cfg = OpenTelemetryV2Config(exporter="in_memory") + simple = providers.build_tracer_provider(cfg, exporter=InMemorySpanExporter()) + batch = providers.build_tracer_provider( + cfg, exporter=ConsoleSpanExporter(), use_simple_processor=False + ) + # both build without error; assert the requested processor type was used + simple_procs = simple._active_span_processor._span_processors + batch_procs = batch._active_span_processor._span_processors + assert any(isinstance(p, SimpleSpanProcessor) for p in simple_procs) + assert any(isinstance(p, BatchSpanProcessor) for p in batch_procs) + + +def test_baggage_processor_lifecycle_noops(): + proc = providers.LiteLLMBaggageSpanProcessor(allowed_keys=["litellm.team.id"]) + # no-op lifecycle hooks must not raise + assert proc.on_end(None) is None # type: ignore[arg-type] + assert proc.shutdown() is None + assert proc.force_flush() is True + + +def test_emitter_without_call_id_is_not_deduped(): + from litellm.integrations.otel.emitter import SpanEmitter + + cfg = OpenTelemetryV2Config(exporter="in_memory") + provider, exporter = providers.in_memory_provider(cfg) + engine = SpanEmitter(providers.get_tracer(provider, "t"), cfg) + data = LLMCallSpanData( + operation=GenAIOperation.CHAT, + provider="openai", + request_model="gpt-4o", + response_model=None, + response_id=None, + request_params=LLMRequestParams(), + usage=LLMUsage(), + finish_reasons=(), + error=SpanError(error_type="X", message=None), + response_cost=None, + server=None, + identity=RequestIdentity(call_id=None), + ) + engine.emit(SpanRole.LLM_CALL, data) + engine.emit(SpanRole.LLM_CALL, data) # no call_id -> not deduped + assert len(exporter.get_finished_spans()) == 2 diff --git a/tests/test_litellm/integrations/otel/test_otel_v2_dynamic.py b/tests/test_litellm/integrations/otel/test_otel_v2_dynamic.py new file mode 100644 index 00000000000..cccddf38b86 --- /dev/null +++ b/tests/test_litellm/integrations/otel/test_otel_v2_dynamic.py @@ -0,0 +1,131 @@ +"""Per-request multi-tenant credential routing (V1 parity).""" + +import os +import sys + +sys.path.insert(0, os.path.abspath("../../../..")) + +from opentelemetry.trace import NoOpTracer + +from litellm.integrations.otel.config import ExporterSpec, OpenTelemetryV2Config +from litellm.integrations.otel.presets import dynamic_otlp_headers +from litellm.integrations.otel.routing import TenantTracerCache + + +def _cache(callback_name, exporters=None): + cfg = OpenTelemetryV2Config(exporters=exporters or [ExporterSpec(kind="in_memory")]) + return TenantTracerCache(cfg, callback_name, "litellm") + + +# --- header builders mirror the V1 construct_dynamic_otel_headers overrides --- # + + +def test_arize_dynamic_headers(): + headers = dynamic_otlp_headers( + "arize", {"arize_space_id": "S", "arize_api_key": "K"} + ) + assert headers == {"arize-space-id": "S", "api_key": "K"} + + +def test_arize_space_key_overrides_space_id(): + headers = dynamic_otlp_headers( + "arize", {"arize_space_id": "S", "arize_space_key": "SK"} + ) + assert headers == {"arize-space-id": "SK"} + + +def test_langfuse_dynamic_headers_need_both_keys(): + assert dynamic_otlp_headers("langfuse_otel", {"langfuse_public_key": "pk"}) is None + headers = dynamic_otlp_headers( + "langfuse_otel", {"langfuse_public_key": "pk", "langfuse_secret_key": "sk"} + ) + assert headers is not None and "Authorization" in headers + + +def test_weave_dynamic_headers(): + headers = dynamic_otlp_headers( + "weave_otel", {"wandb_api_key": "w", "weave_project_id": "p"} + ) + assert headers is not None + assert "Authorization" in headers and headers["project_id"] == "p" + + +def test_non_participating_callbacks_have_no_routing(): + # Phoenix subclasses the base in V1 (no override) → no dynamic routing. + assert dynamic_otlp_headers("arize_phoenix", {"arize_api_key": "K"}) is None + assert dynamic_otlp_headers("langtrace", {"arize_api_key": "K"}) is None + assert dynamic_otlp_headers(None, {"arize_api_key": "K"}) is None + + +def test_no_dynamic_params_is_no_routing(): + assert dynamic_otlp_headers("arize", None) is None + assert dynamic_otlp_headers("arize", {}) is None + + +# --- TenantTracerCache routes + caches a TracerProvider per credential set --- # + + +def test_provider_cached_per_credential_set(): + cache = _cache("arize") + default = NoOpTracer() + creds_a = {"arize_space_id": "S", "arize_api_key": "K"} + creds_b = {"arize_space_id": "S2", "arize_api_key": "K2"} + + cache.tracer_for(default, creds_a) + cache.tracer_for(default, creds_a) # same set → reuse, no new provider + assert len(cache._providers) == 1 + cache.tracer_for(default, creds_b) # new set → new provider + assert len(cache._providers) == 2 + + +def test_provider_cache_is_bounded_and_evicts_lru(monkeypatch): + # The cache key derives from request-supplied dynamic credentials, so it + # must be bounded — an unbounded cache lets a caller spawn one provider (and + # its background exporter thread) per unique credential set. On overflow the + # least-recently-used provider is evicted and shut down. + from litellm.integrations.otel import routing as routing_mod + + monkeypatch.setattr(routing_mod, "_MAX_CACHED_PROVIDERS", 2) + shut_down = [] + monkeypatch.setattr( + routing_mod, "_shutdown_provider", lambda p: shut_down.append(p) + ) + + cache = _cache("arize") + default = NoOpTracer() + + def creds(space): + return {"arize_space_id": space, "arize_api_key": "K"} + + cache.tracer_for(default, creds("1")) + cache.tracer_for(default, creds("2")) + cache.tracer_for(default, creds("1")) # touch "1" → "2" is now LRU + cache.tracer_for(default, creds("3")) # overflow → evict "2" + + assert len(cache._providers) == 2 + assert len(shut_down) == 1 # exactly the evicted provider was shut down + + +def test_no_dynamic_params_uses_default_tracer(): + cache = _cache("arize") + default = NoOpTracer() + assert cache.tracer_for(default, {}) is default + assert cache._providers == {} + + +def test_non_participating_callback_uses_default_tracer(): + cache = _cache("arize_phoenix") + default = NoOpTracer() + assert cache.tracer_for(default, {"arize_api_key": "K"}) is default + assert cache._providers == {} + + +def test_dynamic_headers_applied_to_otlp_exporter_only(): + cache = _cache( + "arize", + exporters=[ExporterSpec(kind="otlp_http"), ExporterSpec(kind="in_memory")], + ) + new_cfg = cache._config_with_headers({"arize-space-id": "S", "api_key": "K"}) + otlp, in_mem = new_cfg.exporters + assert otlp.headers == "arize-space-id=S,api_key=K" + assert in_mem.headers is None # console/in_memory left untouched diff --git a/tests/test_litellm/integrations/otel/test_otel_v2_emitter.py b/tests/test_litellm/integrations/otel/test_otel_v2_emitter.py new file mode 100644 index 00000000000..cc4c96d9ad9 --- /dev/null +++ b/tests/test_litellm/integrations/otel/test_otel_v2_emitter.py @@ -0,0 +1,220 @@ +"""Golden tests for the OTel v2 engine: span shape, kinds, semconv attributes, +legacy dual-emit, hierarchy, error status, and idempotency. Needs the OTel SDK.""" + +import pytest + +pytest.importorskip("opentelemetry") + +from opentelemetry.trace import SpanKind # noqa: E402 +from opentelemetry.trace.status import StatusCode # noqa: E402 + +from litellm.integrations.otel import ( # noqa: E402 + GenAI, + LiteLLM, + OpenTelemetryV2Config, +) +from litellm.integrations.otel import context as ctx_mod # noqa: E402 +from litellm.integrations.otel import providers # noqa: E402 +from litellm.integrations.otel.emitter import SpanEmitter # noqa: E402 +from litellm.integrations.otel.payloads import ( # noqa: E402 + GuardrailSpanData, + LLMCallSpanData, + ServiceSpanData, +) +from litellm.integrations.otel.spans import SPAN_REGISTRY, SpanRole # noqa: E402 + + +def _payload(**overrides): + payload = { + "call_type": "acompletion", + "custom_llm_provider": "openai", + "model": "gpt-4o", + "prompt_tokens": 10, + "completion_tokens": 5, + "total_tokens": 15, + "stream": False, + "model_parameters": {"temperature": 0.7, "max_tokens": 256, "top_k": 40}, + "response": { + "id": "resp_1", + "model": "gpt-4o-2024", + "choices": [{"finish_reason": "stop"}], + }, + "metadata": {"team_id": "t1", "team_alias": "team one"}, + "api_base": "https://api.openai.com:443/v1", + "status": "success", + "litellm_call_id": "call_1", + "response_cost": 0.002, + "hidden_params": {}, + } + payload.update(overrides) + return payload + + +def _engine(legacy_compat=True): + cfg = OpenTelemetryV2Config(exporter="in_memory", legacy_compat=legacy_compat) + provider, exporter = providers.in_memory_provider(cfg) + tracer = providers.get_tracer(provider, "litellm-test") + return SpanEmitter(tracer, cfg), exporter + + +def test_llm_call_span_golden(): + engine, exporter = _engine() + data = LLMCallSpanData.from_standard_logging_payload(_payload()) + engine.emit(SpanRole.LLM_CALL, data) + (span,) = exporter.get_finished_spans() + assert span.name == "chat gpt-4o" + assert span.kind is SpanKind.CLIENT + a = span.attributes + assert a[GenAI.OPERATION_NAME] == "chat" + assert a[GenAI.PROVIDER_NAME] == "openai" + assert a[GenAI.REQUEST_MODEL] == "gpt-4o" + assert a[GenAI.RESPONSE_MODEL] == "gpt-4o-2024" + assert a[GenAI.RESPONSE_ID] == "resp_1" + assert a[GenAI.USAGE_INPUT_TOKENS] == 10 + assert a[GenAI.USAGE_OUTPUT_TOKENS] == 5 + assert a[GenAI.RESPONSE_FINISH_REASONS] == ("stop",) + assert a[GenAI.REQUEST_TEMPERATURE] == 0.7 + assert a["server.address"] == "api.openai.com" + assert a[LiteLLM.CALL_ID] == "call_1" + assert a["litellm.cost.total"] == 0.002 + assert span.status.status_code is StatusCode.OK + + +def test_legacy_dual_emit_on(): + engine, exporter = _engine(legacy_compat=True) + engine.emit( + SpanRole.LLM_CALL, LLMCallSpanData.from_standard_logging_payload(_payload()) + ) + (span,) = exporter.get_finished_spans() + # canonical AND legacy keys are both present + assert span.attributes[GenAI.USAGE_OUTPUT_TOKENS] == 5 + assert span.attributes["gen_ai.usage.completion_tokens"] == 5 + assert span.attributes["gen_ai.system"] == "openai" + + +def test_legacy_dual_emit_off(): + engine, exporter = _engine(legacy_compat=False) + engine.emit( + SpanRole.LLM_CALL, LLMCallSpanData.from_standard_logging_payload(_payload()) + ) + (span,) = exporter.get_finished_spans() + # canonical present, legacy absent + assert span.attributes[GenAI.USAGE_OUTPUT_TOKENS] == 5 + assert "gen_ai.usage.completion_tokens" not in span.attributes + assert "gen_ai.system" not in span.attributes + + +def test_error_span_sets_status_and_error_type(): + engine, exporter = _engine() + payload = _payload( + status="failure", + error_information={"error_class": "RateLimitError", "error_message": "429"}, + ) + engine.emit( + SpanRole.LLM_CALL, LLMCallSpanData.from_standard_logging_payload(payload) + ) + (span,) = exporter.get_finished_spans() + assert span.status.status_code is StatusCode.ERROR + assert span.attributes["error.type"] == "RateLimitError" + + +def test_hierarchy_and_kinds_match_registry(): + engine, exporter = _engine() + data = LLMCallSpanData.from_standard_logging_payload(_payload()) + root = engine.start_span(SpanRole.PROXY_REQUEST, "POST /chat/completions") + root_ctx = ctx_mod.context_from_span(root) + engine.emit(SpanRole.LLM_CALL, data, parent_context=root_ctx) + engine.emit( + SpanRole.GUARDRAIL, GuardrailSpanData("presidio", status="success"), root_ctx + ) + engine.emit(SpanRole.SERVICE, ServiceSpanData("redis", call_type="set"), root_ctx) + root.end() + + by_name = {s.name: s for s in exporter.get_finished_spans()} + root_id = root.get_span_context().span_id + assert by_name["chat gpt-4o"].parent.span_id == root_id + assert by_name["execute_guardrail presidio"].parent.span_id == root_id + assert by_name["redis"].parent.span_id == root_id + # kinds come straight from the registry + assert by_name["chat gpt-4o"].kind is SpanKind.CLIENT + assert by_name["execute_guardrail presidio"].kind is SpanKind.INTERNAL + assert by_name["redis"].kind is SpanKind.INTERNAL + assert by_name["POST /chat/completions"].kind is SpanKind.SERVER + + +def test_idempotent_dual_fire(): + engine, exporter = _engine() + data = LLMCallSpanData.from_standard_logging_payload(_payload()) + first = engine.emit(SpanRole.LLM_CALL, data) + second = engine.emit(SpanRole.LLM_CALL, data) # same call_id -> deduped + assert first is not None + assert second is None + assert len(exporter.get_finished_spans()) == 1 + + +def test_dedup_cache_is_bounded(monkeypatch): + """The dedup cache only needs to coalesce one request's sync+async fire, so + it is a bounded LRU — every unique call_id must not accumulate forever on a + long-running proxy.""" + from litellm.integrations.otel import emitter as emitter_mod + + monkeypatch.setattr(emitter_mod, "_DEDUP_CACHE_MAX", 3) + engine, _ = _engine() + for i in range(10): + engine.emit( + SpanRole.LLM_CALL, + LLMCallSpanData.from_standard_logging_payload( + _payload(litellm_call_id=f"call_{i}") + ), + ) + assert len(engine._emitted) <= 3 + + +def test_service_error_span(): + from litellm.integrations.otel.payloads import SpanError + + engine, exporter = _engine() + engine.emit( + SpanRole.SERVICE, + ServiceSpanData( + "postgres", call_type="query", error=SpanError("DBError", "boom") + ), + ) + (span,) = exporter.get_finished_spans() + assert span.status.status_code is StatusCode.ERROR + assert span.attributes["error.type"] == "DBError" + assert span.attributes[LiteLLM.SERVICE_NAME] == "postgres" + + +def test_guardrail_block_span_is_error_and_carries_verdict(): + engine, exporter = _engine() + data = GuardrailSpanData.from_logging_entry( + { + "guardrail_name": "openai-moderation", + "guardrail_mode": "pre_call", + "guardrail_status": "guardrail_intervened", + "guardrail_provider": "openai", + "guardrail_response": {"violated_categories": ["violence"]}, + "masked_entity_count": {"EMAIL": 2}, + } + ) + engine.emit(SpanRole.GUARDRAIL, data) + (span,) = exporter.get_finished_spans() + assert span.status.status_code is StatusCode.ERROR # intervention → ERROR + a = span.attributes + assert a[LiteLLM.GUARDRAIL_STATUS] == "guardrail_intervened" + assert a[LiteLLM.GUARDRAIL_PROVIDER] == "openai" + assert "violence" in a[LiteLLM.GUARDRAIL_RESPONSE] # the verdict rides the span + assert a[LiteLLM.GUARDRAIL_MASKED_ENTITY_COUNT] == 2 + + +def test_guardrail_success_span_is_ok(): + engine, exporter = _engine() + engine.emit( + SpanRole.GUARDRAIL, + GuardrailSpanData.from_logging_entry( + {"guardrail_name": "g", "guardrail_status": "success"} + ), + ) + (span,) = exporter.get_finished_spans() + assert span.status.status_code is StatusCode.OK diff --git a/tests/test_litellm/integrations/otel/test_otel_v2_logger.py b/tests/test_litellm/integrations/otel/test_otel_v2_logger.py new file mode 100644 index 00000000000..249081de500 --- /dev/null +++ b/tests/test_litellm/integrations/otel/test_otel_v2_logger.py @@ -0,0 +1,567 @@ +"""Phase 2 / Phase 3 tests for the V2 ``OpenTelemetryV2`` CustomLogger adapter. + +Exercises the callback surface the existing call sites use: LLM-call sync/async +success + failure, service hooks, proxy SERVER span lifecycle (start + setters), +parent-context resolution (explicit span, traceparent header), and Baggage +promotion onto child spans. +""" + +import asyncio +from datetime import datetime, timezone + +import pytest + +pytest.importorskip("opentelemetry") + +from opentelemetry import trace # noqa: E402 +from opentelemetry.sdk.trace.export.in_memory_span_exporter import ( # noqa: E402 + InMemorySpanExporter, +) +from opentelemetry.trace import SpanKind # noqa: E402 +from opentelemetry.trace.status import StatusCode # noqa: E402 + +from litellm.integrations.otel import ( # noqa: E402 + GenAI, + HTTP, + LiteLLM, + OpenTelemetryV2Config, +) +from litellm.integrations.otel import providers # noqa: E402 +from litellm.integrations.otel.logger import ( # noqa: E402 + LITELLM_PROXY_REQUEST_SPAN_NAME, + OpenTelemetryV2, +) +from litellm.integrations.otel.spans import SpanRole # noqa: E402 +from litellm.integrations.otel.utils import to_ns, to_seconds # noqa: E402 + +# --------------------------------------------------------------------------- # +# Fixtures +# --------------------------------------------------------------------------- # + + +def _payload(**overrides): + payload = { + "call_type": "acompletion", + "custom_llm_provider": "openai", + "model": "gpt-4o", + "prompt_tokens": 10, + "completion_tokens": 5, + "total_tokens": 15, + "stream": False, + "model_parameters": {"temperature": 0.7, "max_tokens": 256}, + "response": { + "id": "resp_1", + "model": "gpt-4o-2024", + "choices": [{"finish_reason": "stop"}], + }, + "metadata": { + "team_id": "t1", + "team_alias": "team one", + "user_api_key_hash": "hsh", + }, + "api_base": "https://api.openai.com:443/v1", + "status": "success", + "litellm_call_id": "call_1", + "response_cost": 0.002, + "hidden_params": {}, + } + payload.update(overrides) + return payload + + +def _kwargs(payload=None): + return { + "standard_logging_object": payload if payload is not None else _payload(), + "litellm_params": {"metadata": {}}, + } + + +def _logger(legacy_compat=True): + cfg = OpenTelemetryV2Config(exporter="in_memory", legacy_compat=legacy_compat) + exporter = InMemorySpanExporter() + tracer_provider = providers.build_tracer_provider(cfg, exporter=exporter) + return OpenTelemetryV2(config=cfg, tracer_provider=tracer_provider), exporter + + +# --------------------------------------------------------------------------- # +# Time helpers +# --------------------------------------------------------------------------- # + + +def test_to_ns_handles_datetime_and_float(): + dt = datetime(2026, 5, 26, 12, 0, 0, tzinfo=timezone.utc) + assert to_ns(dt) == int(dt.timestamp() * 1e9) + assert to_ns(1.5) == 1_500_000_000 + assert to_ns(None) is None + assert to_ns(True) is None # bool is rejected — not a real epoch value + + +def test_to_seconds_parses_string_formats(): + assert to_seconds("2026-05-26 12:00:00.123") is not None + assert to_seconds("2026-05-26 12:00:00") is not None + assert to_seconds("nonsense") is None + assert to_seconds(None) is None + assert to_seconds(1.5) == 1.5 + + +# --------------------------------------------------------------------------- # +# LLM-call callbacks +# --------------------------------------------------------------------------- # + + +def test_async_log_success_event_emits_llm_call_span(): + logger, exporter = _logger() + asyncio.run(logger.async_log_success_event(_kwargs(), None, None, None)) + (span,) = exporter.get_finished_spans() + assert span.name == "chat gpt-4o" + assert span.kind is SpanKind.CLIENT + assert span.attributes[GenAI.OPERATION_NAME] == "chat" + assert span.attributes[GenAI.REQUEST_MODEL] == "gpt-4o" + assert span.attributes[LiteLLM.CALL_ID] == "call_1" + assert span.status.status_code is StatusCode.OK + + +def test_async_log_failure_event_marks_error_status(): + logger, exporter = _logger() + payload = _payload( + status="failure", + error_information={"error_class": "RateLimitError", "error_message": "429"}, + ) + asyncio.run( + logger.async_log_failure_event(_kwargs(payload=payload), None, None, None) + ) + (span,) = exporter.get_finished_spans() + assert span.status.status_code is StatusCode.ERROR + assert span.attributes["error.type"] == "RateLimitError" + + +def test_sync_log_event_is_noop(): + """V2 emits async-only; the sync callback runs out-of-context, so it no-ops.""" + logger, exporter = _logger() + logger.log_success_event(_kwargs(), None, None, None) + logger.log_failure_event(_kwargs(), None, None, None) + assert exporter.get_finished_spans() == () + + +def test_missing_standard_logging_object_is_noop(): + logger, exporter = _logger() + asyncio.run( + logger.async_log_success_event({"litellm_params": {}}, None, None, None) + ) + assert exporter.get_finished_spans() == () + + +def test_pre_call_guardrail_block_suppresses_phantom_llm_span(): + """A pre-call guardrail block means the LLM was never called. litellm still + emits a failure log, but a CLIENT 'chat …' span would be misleading — so it + is suppressed (the guardrail span represents the outcome).""" + logger, exporter = _logger() + payload = _payload( + status="failure", + guardrail_information=[ + {"guardrail_mode": "pre_call", "guardrail_status": "guardrail_intervened"} + ], + ) + asyncio.run( + logger.async_log_failure_event(_kwargs(payload=payload), None, None, None) + ) + assert exporter.get_finished_spans() == () # no phantom LLM span + + +def test_llm_span_still_emitted_when_guardrail_only_masked(): + """A pre-call guardrail that masks (not blocks) lets the call proceed, so the + request succeeds and the LLM span must still be emitted.""" + logger, exporter = _logger() + payload = _payload( + status="success", + guardrail_information=[ + {"guardrail_mode": "pre_call", "guardrail_status": "guardrail_intervened"} + ], + ) + asyncio.run( + logger.async_log_success_event(_kwargs(payload=payload), None, None, None) + ) + assert len(exporter.get_finished_spans()) == 1 # real LLM call span present + + +def test_idempotent_on_repeat_call_id(): + """Same StandardLoggingPayload (same id) emits once even if the async hook fires twice.""" + logger, exporter = _logger() + kwargs = _kwargs() + asyncio.run(logger.async_log_success_event(kwargs, None, None, None)) + asyncio.run(logger.async_log_success_event(kwargs, None, None, None)) + assert len(exporter.get_finished_spans()) == 1 + + +# --------------------------------------------------------------------------- # +# Parent resolution — ambient context (no metadata threading) +# --------------------------------------------------------------------------- # + + +def test_llm_span_parents_to_ambient_server_span(): + """With the FastAPI instrumentor, the server span is the active context; the + LLM span nests under it via ambient context (no ``litellm_parent_otel_span``). + """ + logger, exporter = _logger() + server = logger._emitter.start_span( + SpanRole.PROXY_REQUEST, LITELLM_PROXY_REQUEST_SPAN_NAME + ) + with trace.use_span(server, end_on_exit=False): + asyncio.run(logger.async_log_success_event(_kwargs(), None, None, None)) + server.end() + by_name = {s.name: s for s in exporter.get_finished_spans()} + llm_span = by_name["chat gpt-4o"] + assert llm_span.parent is not None + assert llm_span.parent.span_id == server.get_span_context().span_id + + +def test_llm_span_is_root_without_ambient_server_span(): + logger, exporter = _logger() + asyncio.run(logger.async_log_success_event(_kwargs(), None, None, None)) + (span,) = exporter.get_finished_spans() + assert span.parent is None # standalone (no proxy server span) → root + + +# Inbound ``traceparent`` propagation is now the FastAPI instrumentor's job +# (see proxy_server's startup mount + ``test_otel_v2_mount``), not the logger's. + + +# --------------------------------------------------------------------------- # +# Baggage promotion (LLM call writes identity into baggage so child spans +# inherit team/key/model attrs). +# --------------------------------------------------------------------------- # + + +def test_baggage_identity_promoted_onto_llm_call(): + logger, exporter = _logger() + asyncio.run(logger.async_log_success_event(_kwargs(), None, None, None)) + (span,) = exporter.get_finished_spans() + assert span.attributes[LiteLLM.TEAM_ID] == "t1" + assert span.attributes[LiteLLM.TEAM_ALIAS] == "team one" + assert span.attributes[GenAI.REQUEST_MODEL] == "gpt-4o" + + +class _Auth: + """Stub matching the ``UserAPIKeyAuth`` fields the logger reads.""" + + team_id = "t1" + team_alias = "team one" + api_key = "hash1" + user_id = "u1" + org_id = None + key_alias = "k1" + end_user_id = None + + +def test_pre_call_hook_seeds_baggage_onto_server_and_child_spans(): + """The pre-call hook seeds identity Baggage in the request context so the + server span (stamped directly) AND later child spans (service here, via the + Baggage processor) carry identity — not just the LLM-call span.""" + logger, exporter = _logger() + server = logger._emitter.start_span( + SpanRole.PROXY_REQUEST, LITELLM_PROXY_REQUEST_SPAN_NAME + ) + + async def _flow(): + # pre-call seeds baggage + stamps the active server span + await logger.async_pre_call_hook( + _Auth(), None, {"model": "gpt-4o"}, "completion" + ) + # a later service call (same task) must inherit the identity + await logger.async_service_success_hook( + payload=_ServicePayload("redis", "set"), parent_otel_span=server + ) + + with trace.use_span(server, end_on_exit=False): + asyncio.run(_flow()) + server.end() + + spans = {s.name: s for s in exporter.get_finished_spans()} + redis = spans["redis"] + assert redis.attributes[LiteLLM.TEAM_ID] == "t1" + assert redis.attributes[LiteLLM.KEY_HASH] == "hash1" + assert redis.attributes[f"{LiteLLM.METADATA_PREFIX}user_api_key_user_id"] == "u1" + srv = spans[LITELLM_PROXY_REQUEST_SPAN_NAME] + assert ( + srv.attributes[LiteLLM.TEAM_ID] == "t1" + ) # stamped directly on the server span + assert srv.attributes[f"{LiteLLM.METADATA_PREFIX}user_api_key_user_id"] == "u1" + + +# --------------------------------------------------------------------------- # +# Service hooks (Phase 3) +# --------------------------------------------------------------------------- # + + +class _Service: + """Stub matching ``ServiceTypes(str, Enum)``.""" + + def __init__(self, value): + self.value = value + + +class _ServicePayload: + def __init__(self, service="redis", call_type="set", error=None): + self.service = _Service(service) + self.call_type = call_type + self.error = error + + +def _service_parent(logger): + """Helper: a live PROXY_REQUEST span to parent service spans under.""" + return logger._emitter.start_span( + SpanRole.PROXY_REQUEST, LITELLM_PROXY_REQUEST_SPAN_NAME + ) + + +def test_async_service_success_hook_emits_service_span(): + logger, exporter = _logger() + parent = _service_parent(logger) + try: + asyncio.run( + logger.async_service_success_hook( + payload=_ServicePayload("redis", "set"), + parent_otel_span=parent, + event_metadata={"key1": "val1"}, + ) + ) + finally: + parent.end() + by_name = {s.name: s for s in exporter.get_finished_spans()} + span = by_name["redis"] + assert span.kind is SpanKind.INTERNAL + assert span.attributes[LiteLLM.SERVICE_NAME] == "redis" + assert span.attributes[LiteLLM.SERVICE_CALL_TYPE] == "set" + # Canonical (V2) namespaced metadata key + assert span.attributes[f"{LiteLLM.METADATA_PREFIX}key1"] == "val1" + # V1 bare key (legacy dual-emit) + assert span.attributes["key1"] == "val1" + assert span.attributes["service"] == "redis" # V1 bare key + assert span.attributes["call_type"] == "set" # V1 bare key + assert span.status.status_code is StatusCode.OK + + +def test_async_service_failure_hook_marks_error_status(): + logger, exporter = _logger() + parent = _service_parent(logger) + try: + asyncio.run( + logger.async_service_failure_hook( + payload=_ServicePayload("postgres", "query"), + error="boom", + parent_otel_span=parent, + ) + ) + finally: + parent.end() + by_name = {s.name: s for s in exporter.get_finished_spans()} + span = by_name["postgres"] + assert span.status.status_code is StatusCode.ERROR + # Without an explicit error_type from the payload, V2 stamps the fallback. + assert span.attributes["error.type"] == "error" + assert span.attributes[LiteLLM.SERVICE_NAME] == "postgres" + + +def test_async_service_failure_hook_preserves_payload_error_over_override(): + """When the payload itself carries an error, that takes precedence over the override.""" + logger, exporter = _logger() + parent = _service_parent(logger) + try: + asyncio.run( + logger.async_service_failure_hook( + payload=_ServicePayload("postgres", "query", error="db-down"), + error="override-only-used-when-payload-clean", + parent_otel_span=parent, + ) + ) + finally: + parent.end() + by_name = {s.name: s for s in exporter.get_finished_spans()} + span = by_name["postgres"] + assert span.status.status_code is StatusCode.ERROR + assert "db-down" in (span.status.description or "") + + +def test_service_hook_without_parent_is_noop(): + """Mirrors V1: no parent OTel span → no service span (no free-standing roots).""" + logger, exporter = _logger() + asyncio.run( + logger.async_service_success_hook( + payload=_ServicePayload(), parent_otel_span=None + ) + ) + assert exporter.get_finished_spans() == () + + +def test_service_span_inherits_parent_when_provided(): + logger, exporter = _logger() + parent = logger._emitter.start_span( + SpanRole.PROXY_REQUEST, LITELLM_PROXY_REQUEST_SPAN_NAME + ) + try: + asyncio.run( + logger.async_service_success_hook( + payload=_ServicePayload(), parent_otel_span=parent + ) + ) + finally: + parent.end() + by_name = {s.name: s for s in exporter.get_finished_spans()} + assert ( + by_name["redis"].parent.span_id + == by_name[LITELLM_PROXY_REQUEST_SPAN_NAME].get_span_context().span_id + ) + + +# --------------------------------------------------------------------------- # +# Proxy SERVER span lifecycle +# --------------------------------------------------------------------------- # + + +def test_create_proxy_request_started_span_returns_ambient_span(): + """V2 doesn't create a server span (the instrumentor does), but it returns + the active server span so the proxy can thread it as the service-span parent + — service logging only fires the OTel hook when that parent is non-None.""" + logger, exporter = _logger() + # No ambient recordable span → None (and creates nothing). + assert ( + logger.create_litellm_proxy_request_started_span( + start_time=datetime.now(timezone.utc), headers={"traceparent": "x"} + ) + is None + ) + assert exporter.get_finished_spans() == () + # With an active server span, return it (do NOT create a new one). + server = logger._emitter.start_span( + SpanRole.PROXY_REQUEST, LITELLM_PROXY_REQUEST_SPAN_NAME + ) + with trace.use_span(server, end_on_exit=False): + got = logger.create_litellm_proxy_request_started_span( + start_time=datetime.now(timezone.utc), headers=None + ) + server.end() + assert got is server + + +def test_proxy_span_setters_are_noops(): + """The FastAPI instrumentor owns the server span; the setters write nothing + (and must tolerate any span / None without raising). + """ + logger, exporter = _logger() + span = logger._emitter.start_span( + SpanRole.PROXY_REQUEST, LITELLM_PROXY_REQUEST_SPAN_NAME + ) + OpenTelemetryV2.set_proxy_request_route_attributes( + span, url_path="/chat/completions", http_route="/chat/completions" + ) + OpenTelemetryV2.set_response_status_code_attribute(span, 200) + OpenTelemetryV2.set_preprocessing_duration_attribute( + span, + {"first_api_call_start_time": 1.0, "metadata": {"litellm_received_at": 0.5}}, + ) + span.end() + (finished,) = exporter.get_finished_spans() + assert HTTP.URL_PATH not in finished.attributes + assert HTTP.ROUTE not in finished.attributes + assert HTTP.RESPONSE_STATUS_CODE not in finished.attributes + assert LiteLLM.PREPROCESSING_MS not in finished.attributes + # None span is tolerated too. + OpenTelemetryV2.set_proxy_request_route_attributes(None, http_route="/x") + OpenTelemetryV2.set_response_status_code_attribute(None, 200) + OpenTelemetryV2.set_preprocessing_duration_attribute(None, {}) + + +# --------------------------------------------------------------------------- # +# Constructor / proxy global guard +# --------------------------------------------------------------------------- # + + +def test_constructor_accepts_v1_compatible_kwargs(): + """Mirrors V1's positional shape — config / callback_name / providers / **kwargs.""" + cfg = OpenTelemetryV2Config(exporter="in_memory") + tp = providers.build_tracer_provider(cfg) + logger = OpenTelemetryV2( + config=cfg, + callback_name="otel", + tracer_provider=tp, + logger_provider=None, + meter_provider=None, + turn_off_message_logging=True, + ) + assert logger.callback_name == "otel" + assert logger.turn_off_message_logging is True + assert logger.tracer is not None + + +def test_default_config_reads_env(monkeypatch): + """No explicit config → reads env (exporter=console by default).""" + monkeypatch.delenv("OTEL_EXPORTER", raising=False) + monkeypatch.delenv("OTEL_EXPORTER_OTLP_PROTOCOL", raising=False) + logger = OpenTelemetryV2( + tracer_provider=providers.build_tracer_provider( + OpenTelemetryV2Config(exporter="in_memory") + ) + ) + assert logger.config.exporter == "console" + + +def test_proxy_global_first_registered_wins(monkeypatch): + """``_init_otel_logger_on_litellm_proxy`` claims the global only when empty.""" + proxy_server = pytest.importorskip("litellm.proxy.proxy_server") + monkeypatch.setattr(proxy_server, "open_telemetry_logger", None, raising=False) + cfg = OpenTelemetryV2Config(exporter="in_memory") + tp = providers.build_tracer_provider(cfg) + + first = OpenTelemetryV2(config=cfg, tracer_provider=tp) + assert proxy_server.open_telemetry_logger is first + + second = OpenTelemetryV2(config=cfg, tracer_provider=tp) + # Global still points at the first registration. + assert proxy_server.open_telemetry_logger is first + assert second is not first + + +def test_registers_into_litellm_service_callback(monkeypatch): + """The logger must mutate ``litellm.service_callback`` in place. An empty + list is falsy, so a ``getattr(..) or []`` would append to a throwaway local + and service spans (Redis, …) would silently never fire on this logger. + """ + import litellm + + pytest.importorskip("litellm.proxy.proxy_server") + monkeypatch.setattr(litellm, "service_callback", [], raising=False) + cfg = OpenTelemetryV2Config(exporter="in_memory") + tp = providers.build_tracer_provider(cfg) + + first = OpenTelemetryV2(config=cfg, tracer_provider=tp) + assert first in litellm.service_callback + + # A second OTel logger sees one is already registered and does not duplicate. + OpenTelemetryV2(config=cfg, tracer_provider=tp) + otel_registrations = [ + cb + for cb in litellm.service_callback + if cb.__class__.__module__.startswith("litellm.integrations.otel") + ] + assert len(otel_registrations) == 1 + + +# --------------------------------------------------------------------------- # +# Management endpoint hooks — no-ops: management endpoints are ordinary FastAPI +# routes, so the mounted instrumentor spans them. The hooks must not emit. +# --------------------------------------------------------------------------- # + + +def test_management_hooks_are_noops(): + logger, exporter = _logger() + + class _Payload: + route = "/key/generate" + request_data = {"models": "gpt-4o"} + response = {"key": "sk-123"} + exception = ValueError("nope") + start_time = end_time = None + + asyncio.run(logger.async_management_endpoint_success_hook(_Payload())) + asyncio.run(logger.async_management_endpoint_failure_hook(_Payload())) + assert exporter.get_finished_spans() == () diff --git a/tests/test_litellm/integrations/otel/test_otel_v2_mount.py b/tests/test_litellm/integrations/otel/test_otel_v2_mount.py new file mode 100644 index 00000000000..824f1e773e9 --- /dev/null +++ b/tests/test_litellm/integrations/otel/test_otel_v2_mount.py @@ -0,0 +1,73 @@ +"""V2 entrypoint: the FastAPI instrumentation pattern proxy_server mounts at +startup (gated by LITELLM_OTEL_V2). The mount logic itself lives inline in +``proxy_server.proxy_startup_event``; this exercises the same pattern in +isolation so the server-span + shared-provider behavior stays covered. +""" + +import os +import sys + +import pytest + +sys.path.insert(0, os.path.abspath("../../../..")) + +pytest.importorskip("opentelemetry") +pytest.importorskip("opentelemetry.instrumentation.fastapi") +fastapi = pytest.importorskip("fastapi") + +from fastapi.testclient import TestClient # noqa: E402 +from opentelemetry.instrumentation.fastapi import FastAPIInstrumentor # noqa: E402 +from opentelemetry.sdk.trace.export import SimpleSpanProcessor # noqa: E402 +from opentelemetry.sdk.trace.export.in_memory_span_exporter import ( # noqa: E402 + InMemorySpanExporter, +) +from opentelemetry.trace import SpanKind # noqa: E402 + +from litellm.integrations.otel.config import ( # noqa: E402 + OpenTelemetryV2Config, + is_otel_v2_enabled, +) +from litellm.integrations.otel.logger import OpenTelemetryV2 # noqa: E402 + + +def _instrumented_app(): + """Mirror proxy_server's startup mount: a logger builds the shared provider, + and the FastAPI instrumentor is attached to it.""" + app = fastapi.FastAPI() + + @app.get("/ping") + def ping(): + return {"ok": True} + + logger = OpenTelemetryV2(config=OpenTelemetryV2Config(exporter="in_memory")) + FastAPIInstrumentor.instrument_app(app, tracer_provider=logger._tracer_provider) + return app, logger + + +def test_gate_toggles_with_env(monkeypatch): + """The startup mount is guarded by this flag.""" + monkeypatch.delenv("LITELLM_OTEL_V2", raising=False) + assert is_otel_v2_enabled() is False + monkeypatch.setenv("LITELLM_OTEL_V2", "1") + assert is_otel_v2_enabled() is True + + +def test_instrumented_app_emits_server_span(): + app, logger = _instrumented_app() + exporter = InMemorySpanExporter() + logger._tracer_provider.add_span_processor(SimpleSpanProcessor(exporter)) + + TestClient(app).get("/ping") + + server_spans = [ + s for s in exporter.get_finished_spans() if s.kind is SpanKind.SERVER + ] + assert server_spans, "FastAPI instrumentor should emit a SERVER span per request" + attrs = server_spans[0].attributes or {} + assert any("route" in k or "method" in k for k in attrs) + + +def test_logger_and_instrumentor_share_provider(): + """Gen-ai spans (logger) and server spans (instrumentor) write to one provider.""" + _, logger = _instrumented_app() + assert logger._emitter._tracer is logger.tracer diff --git a/tests/test_litellm/integrations/otel/test_otel_v2_multibackend.py b/tests/test_litellm/integrations/otel/test_otel_v2_multibackend.py new file mode 100644 index 00000000000..dee0aeeec64 --- /dev/null +++ b/tests/test_litellm/integrations/otel/test_otel_v2_multibackend.py @@ -0,0 +1,89 @@ +"""Multi-backend fan-out: one TracerProvider, *N* SpanProcessors. + +V1 needed a separate ``TracerProvider`` per integration to avoid stepping on +the global. V2 attaches a ``SpanProcessor`` per exporter to the *same* +provider, so the same trace ID lights up every backend — no duplicate spans, +no per-integration provider caches. +""" + +import pytest + +pytest.importorskip("opentelemetry") + +from opentelemetry.sdk.trace.export.in_memory_span_exporter import ( + InMemorySpanExporter, +) + +from litellm.integrations.otel.config import ExporterSpec, OpenTelemetryV2Config +from litellm.integrations.otel.providers import build_tracer_provider + + +def test_two_exporters_receive_the_same_span(): + """A single ``span.end()`` lands in BOTH exporters with the same span ID.""" + exporter_a = InMemorySpanExporter() + exporter_b = InMemorySpanExporter() + cfg = OpenTelemetryV2Config( + exporters=[ + ExporterSpec(kind="in_memory"), + ExporterSpec(kind="in_memory"), + ] + ) + # Override the auto-built exporters with our test ones by swapping + # processors after construction (the test's purpose is to exercise the + # multi-processor wiring, not to negotiate the in-memory pipe). + from opentelemetry.sdk.trace.export import SimpleSpanProcessor + + provider = build_tracer_provider(cfg) + # Clear out any auto-built export processors and attach our pair. + while provider._active_span_processor._span_processors: + provider._active_span_processor._span_processors = ( + provider._active_span_processor._span_processors[:-1] + ) + provider.add_span_processor(SimpleSpanProcessor(exporter_a)) + provider.add_span_processor(SimpleSpanProcessor(exporter_b)) + + tracer = provider.get_tracer("test") + span = tracer.start_span("multi-backend") + span.set_attribute("test.marker", "yes") + span.end() + + spans_a = exporter_a.get_finished_spans() + spans_b = exporter_b.get_finished_spans() + assert len(spans_a) == 1 + assert len(spans_b) == 1 + assert spans_a[0].context.span_id == spans_b[0].context.span_id + + +def test_resource_attributes_apply_to_all_exporters(): + """``resource_attributes`` flow through the shared TracerProvider.""" + cfg = OpenTelemetryV2Config( + exporters=[ExporterSpec(kind="in_memory")], + resource_attributes={"openinference.project.name": "phoenix-test"}, + ) + provider = build_tracer_provider(cfg) + assert provider.resource.attributes["openinference.project.name"] == "phoenix-test" + + +def test_config_normalizer_inserts_genai_first(): + """The validator pins ``genai`` at the head + appends ``legacy`` on legacy_compat.""" + cfg = OpenTelemetryV2Config(mapper_names=["openinference", "langfuse"]) + assert cfg.mapper_names[0] == "genai" + assert "openinference" in cfg.mapper_names + assert "langfuse" in cfg.mapper_names + assert cfg.mapper_names[-1] == "legacy" # legacy_compat=True by default + + +def test_config_normalizer_no_legacy_when_compat_off(): + cfg = OpenTelemetryV2Config(legacy_compat=False, mapper_names=["openinference"]) + assert "legacy" not in cfg.mapper_names + assert cfg.mapper_names[0] == "genai" + + +def test_config_folds_legacy_exporter_triple_into_exporters_list(): + """When ``exporters`` is empty, the validator folds the legacy single triple.""" + cfg = OpenTelemetryV2Config( + exporter="otlp_http", endpoint="https://api.example.com", headers="k=v" + ) + assert len(cfg.exporters) == 1 + assert cfg.exporters[0].kind == "otlp_http" + assert cfg.exporters[0].endpoint == "https://api.example.com" diff --git a/tests/test_litellm/integrations/otel/test_otel_v2_presets.py b/tests/test_litellm/integrations/otel/test_otel_v2_presets.py new file mode 100644 index 00000000000..c47dcebfaac --- /dev/null +++ b/tests/test_litellm/integrations/otel/test_otel_v2_presets.py @@ -0,0 +1,122 @@ +"""Preset tests. Focused on the AgentOps JWT fetch, which must never block the +event loop: the preset does no network I/O, and a custom exporter mints the JWT +lazily on its first export (in the BatchSpanProcessor worker thread).""" + +import httpx +import pytest + +from litellm.integrations.otel import providers +from litellm.integrations.otel.config import ExporterSpec +from litellm.integrations.otel.presets import agentops as agentops_mod +from litellm.integrations.otel.presets.agentops import ( + _AGENTOPS_ENDPOINT, + _AGENTOPS_EXPORTER_KIND, + _build_agentops_exporter, + _fetch_agentops_jwt, + agentops_preset, +) + + +def test_agentops_preset_does_no_network_io(monkeypatch): + # The preset must not fetch the JWT at build time — that would block the + # event loop during callback construction. It only describes the exporter. + def _boom(*_a, **_k): + raise AssertionError("agentops_preset must not fetch the JWT eagerly") + + monkeypatch.setattr(agentops_mod, "_fetch_agentops_jwt", _boom) + monkeypatch.setenv("AGENTOPS_API_KEY", "ak-123") + cfg = agentops_preset() + agentops_exporters = [e for e in cfg.exporters if e.kind == _AGENTOPS_EXPORTER_KIND] + assert len(agentops_exporters) == 1 + spec = agentops_exporters[0] + assert spec.endpoint == _AGENTOPS_ENDPOINT + assert spec.options == {"api_key": "ak-123"} # carried to the lazy exporter + + +def test_agentops_preset_without_key_omits_options(monkeypatch): + monkeypatch.delenv("AGENTOPS_API_KEY", raising=False) + cfg = agentops_preset() + spec = next(e for e in cfg.exporters if e.kind == _AGENTOPS_EXPORTER_KIND) + assert spec.options is None + + +def test_agentops_exporter_factory_is_registered(): + assert _AGENTOPS_EXPORTER_KIND in providers._EXPORTER_FACTORIES + + +def test_agentops_exporter_mints_jwt_lazily(monkeypatch): + pytest.importorskip("opentelemetry.exporter.otlp.proto.http.trace_exporter") + monkeypatch.setattr( + agentops_mod, "_fetch_agentops_jwt", lambda _k: {"token": "jwt-xyz"} + ) + spec = ExporterSpec( + kind=_AGENTOPS_EXPORTER_KIND, + endpoint=_AGENTOPS_ENDPOINT, + options={"api_key": "ak"}, + ) + exporter = _build_agentops_exporter(spec) + + # No auth header until the first export triggers the (off-loop) fetch. + assert "Authorization" not in exporter._session.headers + exporter._ensure_authenticated() + assert exporter._session.headers["Authorization"] == "Bearer jwt-xyz" + + # Cached: a second resolution does not re-fetch. + calls = [] + monkeypatch.setattr( + agentops_mod, + "_fetch_agentops_jwt", + lambda k: calls.append(k) or {"token": "again"}, + ) + exporter._ensure_authenticated() + assert calls == [] + + +def test_agentops_exporter_tolerates_fetch_failure(monkeypatch): + pytest.importorskip("opentelemetry.exporter.otlp.proto.http.trace_exporter") + + def _raise(_k): + raise RuntimeError("auth down") + + monkeypatch.setattr(agentops_mod, "_fetch_agentops_jwt", _raise) + exporter = _build_agentops_exporter( + ExporterSpec( + kind=_AGENTOPS_EXPORTER_KIND, + endpoint=_AGENTOPS_ENDPOINT, + options={"api_key": "ak"}, + ) + ) + exporter._ensure_authenticated() # must not raise + assert "Authorization" not in exporter._session.headers + + +def test_fetch_jwt_uses_owned_client_not_shared_pool(monkeypatch): + """The fetch owns a short-lived client and closes it, rather than closing + the process-wide cached ``_get_httpx_client`` pool shared by other callers.""" + closed = {"n": 0} + + class _FakeResponse: + status_code = 200 + + def json(self): + return {"token": "jwt-123"} + + class _FakeClient: + def __init__(self, *_a, **_k): + pass + + def __enter__(self): + return self + + def __exit__(self, *_a): + closed["n"] += 1 + + def post(self, *_a, **_k): + return _FakeResponse() + + monkeypatch.setattr(httpx, "Client", _FakeClient) + assert not hasattr(agentops_mod, "_get_httpx_client") + + result = _fetch_agentops_jwt("api-key") + assert result == {"token": "jwt-123"} + assert closed["n"] == 1 # the owned client was closed diff --git a/tests/test_litellm/integrations/otel/test_otel_v2_sources_of_truth.py b/tests/test_litellm/integrations/otel/test_otel_v2_sources_of_truth.py new file mode 100644 index 00000000000..e02d33fec17 --- /dev/null +++ b/tests/test_litellm/integrations/otel/test_otel_v2_sources_of_truth.py @@ -0,0 +1,400 @@ +"""Tests for the OTel v2 sources of truth: span registry, semconv keys, config, +and the typed StandardLoggingPayload adapter. These need no OTel SDK.""" + +import pytest + +from litellm.integrations.otel import ( + BAGGAGE_PROMOTED_KEYS, + Error, + GenAI, + GenAIOperation, + HTTP, + LiteLLM, + OpenTelemetryV2Config, + Server, + is_otel_v2_enabled, + promoted_baggage, + resolve_operation, + resolve_provider, +) +from litellm.integrations.otel import spans as spans_mod +from litellm.integrations.otel.payloads import LLMCallSpanData, RequestIdentity +from litellm.integrations.otel.spans import ( + SPAN_REGISTRY, + LiteLLMSpanKind, + SpanRole, + child_roles, + root_roles, + validate_registry, +) + + +def _sample_payload(**overrides): + payload = { + "call_type": "acompletion", + "custom_llm_provider": "openai", + "model": "gpt-4o", + "prompt_tokens": 10, + "completion_tokens": 5, + "total_tokens": 15, + "stream": False, + "model_parameters": { + "temperature": 0.7, + "max_tokens": 256, + "top_p": 0.9, + "top_k": 40, + "frequency_penalty": 0.1, + "presence_penalty": 0.2, + "stop": ["STOP"], + "seed": 42, + }, + "response": { + "id": "resp_1", + "model": "gpt-4o-2024", + "choices": [{"finish_reason": "stop"}], + }, + "metadata": { + "team_id": "t1", + "team_alias": "team one", + "user_api_key_hash": "hsh", + "user_api_key_org_id": "org1", + }, + "api_base": "https://api.openai.com:443/v1", + "status": "success", + "litellm_call_id": "call_1", + "end_user": "u1", + "response_cost": 0.002, + "hidden_params": {}, + } + payload.update(overrides) + return payload + + +# --- span registry (source of truth #2) ------------------------------------- # + + +def test_registry_validates_and_is_complete(): + validate_registry() # raises on inconsistency + assert set(SPAN_REGISTRY) == set(SpanRole) + + +def test_registry_parent_integrity_no_orphans(): + for role, spec in SPAN_REGISTRY.items(): + assert spec.role is role + if spec.parent is not None: + assert spec.parent in SPAN_REGISTRY + + +def test_registry_hierarchy_shape(): + assert set(root_roles()) == {SpanRole.PROXY_REQUEST} + # Guardrails parent to the request span, not the LLM call: a pre-call + # guardrail runs before the LLM call exists, so it's a sibling of it. + assert set(child_roles(SpanRole.PROXY_REQUEST)) == { + SpanRole.LLM_CALL, + SpanRole.GUARDRAIL, + SpanRole.SERVICE, + } + assert SPAN_REGISTRY[SpanRole.LLM_CALL].kind is LiteLLMSpanKind.CLIENT + assert SPAN_REGISTRY[SpanRole.PROXY_REQUEST].kind is LiteLLMSpanKind.SERVER + assert SPAN_REGISTRY[SpanRole.GUARDRAIL].parent is SpanRole.PROXY_REQUEST + + +def test_llm_call_span_name(): + data = LLMCallSpanData.from_standard_logging_payload(_sample_payload()) + assert spans_mod.llm_call_span_name(data) == "chat gpt-4o" + + +# --- semconv (source of truth #1) ------------------------------------------- # + + +def _all_constants(cls): + return { + getattr(cls, name) + for name in vars(cls) + if not name.startswith("__") and isinstance(getattr(cls, name), str) + } + + +def test_attribute_keys_are_unique_across_namespaces(): + # prefixes are allowed to be substrings; exact keys must not collide. + exact = set() + for cls in (GenAI, Error, Server, HTTP): + for key in _all_constants(cls): + assert key not in exact, f"duplicate attribute key {key}" + exact.add(key) + + +def test_provider_resolution(): + assert resolve_provider("openai") == "openai" + assert resolve_provider("bedrock") == "aws.bedrock" + assert resolve_provider("vertex_ai") == "gcp.vertex_ai" + # unknown providers pass through verbatim (semconv allows provider-specific) + assert resolve_provider("my_custom_llm") == "my_custom_llm" + assert resolve_provider(None) == "" + + +def test_operation_resolution(): + assert resolve_operation("acompletion") is GenAIOperation.CHAT + assert resolve_operation("aembedding") is GenAIOperation.EMBEDDINGS + assert resolve_operation("atext_completion") is GenAIOperation.TEXT_COMPLETION + assert resolve_operation(None) is GenAIOperation.CHAT + + +# --- typed adapter (source of truth #3) ------------------------------------- # + + +def test_llm_call_adapter_extracts_all_fields(): + data = LLMCallSpanData.from_standard_logging_payload(_sample_payload()) + assert data.operation is GenAIOperation.CHAT + assert data.provider == "openai" + assert data.request_model == "gpt-4o" + assert data.response_model == "gpt-4o-2024" + assert data.response_id == "resp_1" + assert data.finish_reasons == ("stop",) + assert (data.usage.input_tokens, data.usage.output_tokens) == (10, 5) + assert data.request_params.temperature == 0.7 + assert data.request_params.top_k == 40 + assert data.request_params.stop_sequences == ("STOP",) + assert data.request_params.seed == 42 + assert data.server is not None + assert data.server.address == "api.openai.com" + assert data.server.port == 443 + assert data.response_cost == 0.002 + assert data.error is None + assert data.identity.team_id == "t1" + assert data.identity.key_hash == "hsh" + + +def test_llm_call_adapter_failure_path(): + payload = _sample_payload( + status="failure", + error_information={ + "error_class": "RateLimitError", + "error_message": "429 slow down", + }, + ) + data = LLMCallSpanData.from_standard_logging_payload(payload) + assert data.error is not None + assert data.error.error_type == "RateLimitError" + assert data.error.message == "429 slow down" + + +def test_adapter_is_resilient_to_minimal_payload(): + data = LLMCallSpanData.from_standard_logging_payload({}) + assert data.request_model == "" + assert data.operation is GenAIOperation.CHAT + assert data.server is None + assert data.usage.input_tokens is None + + +def test_content_capture_gated_off_by_default(): + # ``capture_content`` defaults off: prompt/response bodies must not reach the + # span data (and so no vendor mapper can export them) unless explicitly + # opted in. Non-content metadata (finish reasons) is still derived. + payload = _sample_payload( + messages=[{"role": "user", "content": "secret prompt"}], + ) + payload["response"]["choices"] = [ + {"finish_reason": "stop", "message": {"role": "assistant", "content": "secret"}} + ] + data = LLMCallSpanData.from_standard_logging_payload(payload) + assert data.messages_in == () + assert data.choices_out == () + assert data.finish_reasons == ("stop",) + + +def test_request_identity_prefers_canonical_team_keys(): + from litellm.integrations.otel.payloads import RequestIdentity + + payload = _sample_payload( + metadata={ + "user_api_key_team_id": "team-canonical", + "user_api_key_team_alias": "alias-canonical", + "user_api_key_hash": "hsh", + "team_id": "legacy-ignored", # legacy alias loses to the canonical key + } + ) + ident = RequestIdentity.from_payload(payload) + assert ident.team_id == "team-canonical" + assert ident.team_alias == "alias-canonical" + assert ident.key_hash == "hsh" + + +def test_request_identity_falls_back_to_legacy_team_keys(): + from litellm.integrations.otel.payloads import RequestIdentity + + payload = _sample_payload( + metadata={"team_id": "legacy-team", "team_alias": "legacy"} + ) + ident = RequestIdentity.from_payload(payload) + assert ident.team_id == "legacy-team" + assert ident.team_alias == "legacy" + + +def test_guardrail_span_data_block_carries_verdict_and_error(): + from litellm.integrations.otel.payloads import GuardrailSpanData + + entry = { + "guardrail_name": "openai-moderation", + "guardrail_mode": "pre_call", + "guardrail_status": "guardrail_intervened", + "guardrail_provider": "openai", + "guardrail_action": "BLOCKED", + "guardrail_response": {"violated_categories": ["violence"]}, + "violation_categories": ["violence"], + "masked_entity_count": {"EMAIL": 2, "PHONE": 1}, + "duration": 0.05, + } + d = GuardrailSpanData.from_logging_entry(entry) + assert d.guardrail_name == "openai-moderation" + assert d.status == "guardrail_intervened" + assert d.provider == "openai" + assert d.action == "BLOCKED" + assert '"violence"' in (d.response_json or "") + assert d.violation_categories == ("violence",) + assert d.masked_entity_count == 3 # summed across entity types + assert d.duration == 0.05 + assert d.error is not None # intervention → span marked ERROR + + +def test_guardrail_span_data_success_has_no_error(): + from litellm.integrations.otel.payloads import GuardrailSpanData + + d = GuardrailSpanData.from_logging_entry( + { + "guardrail_name": "g", + "guardrail_mode": "pre_call", + "guardrail_status": "success", + } + ) + assert d.error is None + assert d.status == "success" + + +def test_request_identity_from_user_api_key_auth(): + from litellm.integrations.otel.payloads import RequestIdentity + + class _Auth: + team_id = "t9" + team_alias = "team nine" + api_key = "hashed-key" + user_id = "u9" + org_id = "o9" + key_alias = "my-key" + end_user_id = "eu9" + + ident = RequestIdentity.from_user_api_key_auth(_Auth()) + assert (ident.team_id, ident.team_alias, ident.key_hash) == ( + "t9", + "team nine", + "hashed-key", + ) + assert ident.end_user == "eu9" + assert ident.metadata["user_api_key_user_id"] == "u9" + assert ident.metadata["user_api_key_org_id"] == "o9" + assert ident.metadata["user_api_key_alias"] == "my-key" + assert ident.metadata["user_api_key_end_user_id"] == "eu9" + + +def test_content_capture_opt_in_retains_bodies(): + payload = _sample_payload( + messages=[{"role": "user", "content": "secret prompt"}], + ) + payload["response"]["choices"] = [ + {"finish_reason": "stop", "message": {"role": "assistant", "content": "hi"}} + ] + data = LLMCallSpanData.from_standard_logging_payload(payload, capture_content=True) + assert data.messages_in and data.messages_in[0]["content"] == "secret prompt" + assert data.choices_out and data.choices_out[0]["message"]["content"] == "hi" + + +# --- config ----------------------------------------------------------------- # + + +def test_capture_span_content_resolves_modes(): + from litellm.integrations.otel.config import ( + CaptureMessageContent, + OpenTelemetryV2Config, + ) + + # default (no_content) → off + assert OpenTelemetryV2Config().capture_span_content is False + assert ( + OpenTelemetryV2Config( + capture_message_content=CaptureMessageContent.SPAN_ONLY + ).capture_span_content + is True + ) + assert ( + OpenTelemetryV2Config( + capture_message_content=CaptureMessageContent.SPAN_AND_EVENT + ).capture_span_content + is True + ) + # event-only does not authorize span-attribute content + assert ( + OpenTelemetryV2Config( + capture_message_content=CaptureMessageContent.EVENT_ONLY + ).capture_span_content + is False + ) + + +def test_v2_flag_is_off_by_default(monkeypatch): + monkeypatch.delenv("LITELLM_OTEL_V2", raising=False) + assert is_otel_v2_enabled() is False + monkeypatch.setenv("LITELLM_OTEL_V2", "true") + assert is_otel_v2_enabled() is True + + +def test_config_from_env(monkeypatch): + for var in ( + "OTEL_EXPORTER", + "OTEL_EXPORTER_OTLP_PROTOCOL", + "OTEL_ENDPOINT", + "OTEL_EXPORTER_OTLP_ENDPOINT", + "OTEL_HEADERS", + "OTEL_EXPORTER_OTLP_HEADERS", + "OTEL_SERVICE_NAME", + "LITELLM_OTEL_LEGACY_COMPAT", + ): + monkeypatch.delenv(var, raising=False) + + monkeypatch.setenv("OTEL_EXPORTER_OTLP_ENDPOINT", "https://collector:4318") + monkeypatch.setenv("OTEL_SERVICE_NAME", "my-svc") + cfg = OpenTelemetryV2Config.from_env() + # endpoint with no explicit exporter implies OTLP/HTTP + assert cfg.exporter == "otlp_http" + assert cfg.endpoint == "https://collector:4318" + assert cfg.service_name == "my-svc" + assert cfg.legacy_compat is True # dual-emit default during deprecation window + + +def test_config_legacy_compat_env_toggle(monkeypatch): + monkeypatch.setenv("LITELLM_OTEL_LEGACY_COMPAT", "false") + assert OpenTelemetryV2Config.from_env().legacy_compat is False + + +# --- baggage allowlist (the antipattern boundary) --------------------------- # + + +def test_promoted_baggage_is_bounded_allowlist(): + identity = RequestIdentity( + call_id="c1", + team_id="t1", + team_alias="team one", + key_hash="hsh", + end_user="u1", + metadata={"user_api_key_org_id": "org1", "secret_blob": "should-not-promote"}, + ) + promoted = promoted_baggage(identity, "gpt-4o", BAGGAGE_PROMOTED_KEYS) + assert promoted[LiteLLM.TEAM_ID] == "t1" + assert promoted[LiteLLM.TEAM_ALIAS] == "team one" + assert promoted[GenAI.REQUEST_MODEL] == "gpt-4o" + # allowlisted metadata sub-key is promoted under the litellm.metadata.* prefix + assert promoted[f"{LiteLLM.METADATA_PREFIX}user_api_key_org_id"] == "org1" + # full metadata blob is NOT promoted + assert all("secret_blob" not in key for key in promoted) + # http.* is never a promoted key + assert HTTP.ROUTE not in promoted + assert HTTP.REQUEST_METHOD not in promoted diff --git a/tests/test_litellm/integrations/otel/test_otel_v2_vendor_mappers.py b/tests/test_litellm/integrations/otel/test_otel_v2_vendor_mappers.py new file mode 100644 index 00000000000..2494cfe129b --- /dev/null +++ b/tests/test_litellm/integrations/otel/test_otel_v2_vendor_mappers.py @@ -0,0 +1,196 @@ +"""Tests for the vendor mappers (OpenInference, Langfuse, Weave, Langtrace). + +Composition over inheritance: each vendor's vocabulary is a mapper. Layering +mappers on the same span carries multiple naming schemes for different +backends, so one trace lights up every configured destination. +""" + +import json + +import pytest + +from litellm.integrations.otel import GenAIOperation +from litellm.integrations.otel.mappers import ( + GenAIMapper, + LangfuseMapper, + LangtraceMapper, + OpenInferenceMapper, + WeaveMapper, + resolve_mappers, +) +from litellm.integrations.otel.payloads import ( + LLMCallSpanData, + LLMRequestParams, + LLMUsage, + RequestIdentity, + ServerInfo, + ToolDefinition, +) + + +def _llm_call(**overrides): + base = dict( + operation=GenAIOperation.CHAT, + provider="openai", + request_model="gpt-4o", + response_model="gpt-4o-2024", + response_id="resp_1", + request_params=LLMRequestParams(temperature=0.5, top_p=0.9, max_tokens=128), + usage=LLMUsage(input_tokens=12, output_tokens=8, total_tokens=20), + finish_reasons=("stop",), + error=None, + response_cost=0.001, + server=ServerInfo("api.openai.com", 443), + identity=RequestIdentity(call_id="c1", team_id="t1", team_alias="team one"), + is_streaming=False, + tools=( + ToolDefinition( + name="lookup_weather", + description="Get weather", + parameters_json='{"type":"object"}', + ), + ), + messages_in=( + {"role": "system", "content": "Be concise."}, + {"role": "user", "content": "What's the weather?"}, + ), + choices_out=( + { + "finish_reason": "stop", + "message": {"role": "assistant", "content": "Sunny."}, + }, + ), + system_fingerprint="fp_abc", + ) + base.update(overrides) + return LLMCallSpanData(**base) + + +# --------------------------------------------------------------------------- # +# OpenInference (Arize + Phoenix shared vocabulary) +# --------------------------------------------------------------------------- # + + +def test_openinference_mapper_input_output_messages(): + attrs = OpenInferenceMapper().map(_llm_call()) + assert attrs["openinference.span.kind"] == "LLM" + assert attrs["llm.model_name"] == "gpt-4o" + assert attrs["llm.provider"] == "openai" + assert attrs["llm.input_messages.0.message.role"] == "system" + assert attrs["llm.input_messages.0.message.content"] == "Be concise." + assert attrs["llm.input_messages.1.message.role"] == "user" + assert attrs["llm.output_messages.0.message.role"] == "assistant" + assert attrs["llm.output_messages.0.message.content"] == "Sunny." + assert attrs["llm.token_count.prompt"] == 12 + assert attrs["llm.token_count.completion"] == 8 + assert attrs["llm.token_count.total"] == 20 + # tool definitions ride the OpenInference schema + assert attrs["llm.tools.0.tool.name"] == "lookup_weather" + # invocation_parameters is JSON-serialized + params = json.loads(attrs["llm.invocation_parameters"]) + assert params["temperature"] == 0.5 + assert params["max_tokens"] == 128 + + +def test_openinference_mapper_skips_non_llm_roles(): + from litellm.integrations.otel.payloads import GuardrailSpanData + + assert OpenInferenceMapper().map(GuardrailSpanData("presidio")) == {} + + +def test_openinference_multimodal_content_text_only(): + data = _llm_call( + messages_in=( + { + "role": "user", + "content": [ + {"type": "text", "text": "hi "}, + {"type": "image_url", "image_url": {"url": "x"}}, + {"type": "text", "text": "there"}, + ], + }, + ) + ) + attrs = OpenInferenceMapper().map(data) + assert attrs["llm.input_messages.0.message.content"] == "hi there" + + +# --------------------------------------------------------------------------- # +# Langfuse +# --------------------------------------------------------------------------- # + + +def test_langfuse_mapper_observation_attrs(): + attrs = LangfuseMapper().map(_llm_call()) + assert attrs["langfuse.observation.type"] == "generation" + assert attrs["langfuse.observation.model.name"] == "gpt-4o" + assert attrs["langfuse.observation.metadata.provider"] == "openai" + usage = json.loads(attrs["langfuse.observation.usage_details"]) + assert usage["input"] == 12 and usage["output"] == 8 + params = json.loads(attrs["langfuse.observation.model.parameters"]) + assert params["temperature"] == 0.5 + cost = json.loads(attrs["langfuse.observation.cost_details"]) + assert cost["total"] == 0.001 + assert attrs["langfuse.trace.metadata.team_id"] == "t1" + + +def test_langfuse_mapper_skips_when_no_messages(): + data = _llm_call(messages_in=(), choices_out=()) + attrs = LangfuseMapper().map(data) + assert "langfuse.observation.input" not in attrs + assert "langfuse.observation.output" not in attrs + + +# --------------------------------------------------------------------------- # +# Weave +# --------------------------------------------------------------------------- # + + +def test_weave_mapper_display_and_output(): + attrs = WeaveMapper().map(_llm_call()) + assert attrs["weave.display_name"] == "chat gpt-4o" + assert attrs["weave.call_id"] == "c1" + decoded = json.loads(attrs["weave.output"]) + assert decoded[0]["message"]["content"] == "Sunny." + + +# --------------------------------------------------------------------------- # +# Langtrace +# --------------------------------------------------------------------------- # + + +def test_langtrace_mapper_attrs(): + attrs = LangtraceMapper().map(_llm_call()) + assert attrs["gen_ai.operation.name"] == "chat" + assert attrs["langtrace.service.name"] == "openai" + assert attrs["llm.model"] == "gpt-4o" + assert attrs["gen_ai.response.model"] == "gpt-4o-2024" + assert attrs["gen_ai.system_fingerprint"] == "fp_abc" + assert attrs["llm.temperature"] == 0.5 + assert attrs["llm.token.counts.total"] == 20 + + +# --------------------------------------------------------------------------- # +# Composition (the V2 punchline) +# --------------------------------------------------------------------------- # + + +def test_resolve_mappers_composition_layers_vocabularies(): + """One span, three vocabularies — Arize + Langfuse + canonical together.""" + chain = resolve_mappers(["genai", "openinference", "langfuse"]) + data = _llm_call() + union: dict = {} + for mapper in chain: + union.update(mapper.map(data)) + # Canonical + assert union["gen_ai.operation.name"] == "chat" + # OpenInference + assert union["llm.model_name"] == "gpt-4o" + assert union["openinference.span.kind"] == "LLM" + # Langfuse + assert union["langfuse.observation.type"] == "generation" + + +def test_resolve_mappers_rejects_unknown_name(): + with pytest.raises(ValueError, match="unknown mapper name 'nope'"): + resolve_mappers(["genai", "nope"]) diff --git a/uv.lock b/uv.lock index f8e7eeabeeb..93c93fc6f44 100644 --- a/uv.lock +++ b/uv.lock @@ -9,7 +9,7 @@ resolution-markers = [ ] [options] -exclude-newer = "2026-05-25T20:42:18.420988002Z" +exclude-newer = "0001-01-01T00:00:00Z" # This has no effect and is included for backwards compatibility when using relative exclude-newer values. exclude-newer-span = "P3D" [manifest] @@ -294,6 +294,18 @@ wheels = [ { url = "https://files.pythonhosted.org/packages/ee/82/82745642d3c46e7cea25e1885b014b033f4693346ce46b7f47483cf5d448/argon2_cffi_bindings-25.1.0-pp310-pypy310_pp73-win_amd64.whl", hash = "sha256:da0c79c23a63723aa5d782250fbf51b768abca630285262fb5144ba5ae01e520", size = 29187, upload-time = "2025-07-30T10:02:03.674Z" }, ] +[[package]] +name = "asgiref" +version = "3.11.1" +source = { registry = "https://pypi.org/simple" } +dependencies = [ + { name = "typing-extensions", marker = "python_full_version < '3.11'" }, +] +sdist = { url = "https://files.pythonhosted.org/packages/63/40/f03da1264ae8f7cfdbf9146542e5e7e8100a4c66ab48e791df9a03d3f6c0/asgiref-3.11.1.tar.gz", hash = "sha256:5f184dc43b7e763efe848065441eac62229c9f7b0475f41f80e207a114eda4ce", size = 38550, upload-time = "2026-02-03T13:30:14.33Z" } +wheels = [ + { url = "https://files.pythonhosted.org/packages/5c/0a/a72d10ed65068e115044937873362e6e32fab1b7dce0046aeb224682c989/asgiref-3.11.1-py3-none-any.whl", hash = "sha256:e8667a091e69529631969fd45dc268fa79b99c92c5fcdda727757e52146ec133", size = 24345, upload-time = "2026-02-03T13:30:13.039Z" }, +] + [[package]] name = "assemblyai" version = "0.52.4" @@ -3349,6 +3361,7 @@ proxy-runtime = [ { name = "mangum" }, { name = "opentelemetry-api" }, { name = "opentelemetry-exporter-otlp" }, + { name = "opentelemetry-instrumentation-fastapi" }, { name = "opentelemetry-sdk" }, { name = "prometheus-client" }, { name = "pypdf" }, @@ -3411,6 +3424,7 @@ dev = [ { name = "openapi-core" }, { name = "opentelemetry-api" }, { name = "opentelemetry-exporter-otlp" }, + { name = "opentelemetry-instrumentation-fastapi" }, { name = "opentelemetry-sdk" }, { name = "parameterized" }, { name = "psycopg" }, @@ -3444,6 +3458,7 @@ proxy-dev = [ { name = "hypercorn" }, { name = "opentelemetry-api" }, { name = "opentelemetry-exporter-otlp" }, + { name = "opentelemetry-instrumentation-fastapi" }, { name = "opentelemetry-sdk" }, { name = "prisma" }, { name = "prometheus-client" }, @@ -3499,6 +3514,7 @@ requires-dist = [ { name = "openai", specifier = ">=2.20.0,<3.0.0" }, { name = "opentelemetry-api", marker = "extra == 'proxy-runtime'", specifier = "==1.28.0" }, { name = "opentelemetry-exporter-otlp", marker = "extra == 'proxy-runtime'", specifier = "==1.28.0" }, + { name = "opentelemetry-instrumentation-fastapi", marker = "extra == 'proxy-runtime'", specifier = "==0.49b0" }, { name = "opentelemetry-sdk", marker = "extra == 'proxy-runtime'", specifier = "==1.28.0" }, { name = "orjson", marker = "extra == 'proxy'", specifier = ">=3.11.6,<4.0" }, { name = "polars", marker = "extra == 'proxy'", specifier = ">=1.38.1,<2.0" }, @@ -3573,6 +3589,7 @@ dev = [ { name = "openapi-core", marker = "python_full_version < '3.14'", specifier = "==0.22.0" }, { name = "opentelemetry-api", specifier = "==1.28.0" }, { name = "opentelemetry-exporter-otlp", specifier = "==1.28.0" }, + { name = "opentelemetry-instrumentation-fastapi", specifier = "==0.49b0" }, { name = "opentelemetry-sdk", specifier = "==1.28.0" }, { name = "parameterized", specifier = "==0.9.0" }, { name = "psycopg", specifier = "==3.3.3" }, @@ -3606,6 +3623,7 @@ proxy-dev = [ { name = "hypercorn", specifier = "==0.17.3" }, { name = "opentelemetry-api", specifier = "==1.28.0" }, { name = "opentelemetry-exporter-otlp", specifier = "==1.28.0" }, + { name = "opentelemetry-instrumentation-fastapi", specifier = "==0.49b0" }, { name = "opentelemetry-sdk", specifier = "==1.28.0" }, { name = "prisma", specifier = "==0.11.0" }, { name = "prometheus-client", specifier = "==0.20.0" }, @@ -4561,6 +4579,22 @@ wheels = [ { url = "https://files.pythonhosted.org/packages/ba/46/ba2dc8d18b04acae3d34facd8fe1e5e0cdc9fe64292d45eca9d1d4a8a298/opentelemetry_instrumentation_anthropic-0.33.12-py3-none-any.whl", hash = "sha256:b31618d12a429045db14ed982a142a25df0f0f1dbf03d756e8d597f25b9a053d", size = 11024, upload-time = "2024-11-13T20:27:14.622Z" }, ] +[[package]] +name = "opentelemetry-instrumentation-asgi" +version = "0.49b0" +source = { registry = "https://pypi.org/simple" } +dependencies = [ + { name = "asgiref" }, + { name = "opentelemetry-api" }, + { name = "opentelemetry-instrumentation" }, + { name = "opentelemetry-semantic-conventions" }, + { name = "opentelemetry-util-http" }, +] +sdist = { url = "https://files.pythonhosted.org/packages/e8/55/693c3d0938ba5fead5c3aa4ac7022a992b4ff99a8e9979800d0feb843ff4/opentelemetry_instrumentation_asgi-0.49b0.tar.gz", hash = "sha256:959fd9b1345c92f20c6ef1d42f92ef6a76b3c3083fbc4104d59da6859b15b083", size = 24117, upload-time = "2024-11-05T19:21:46.769Z" } +wheels = [ + { url = "https://files.pythonhosted.org/packages/2c/0b/7900c782a1dfaa584588d724bc3bbdf8405a32497537dd96b3fcbf8461b9/opentelemetry_instrumentation_asgi-0.49b0-py3-none-any.whl", hash = "sha256:722a90856457c81956c88f35a6db606cc7db3231046b708aae2ddde065723dbe", size = 16326, upload-time = "2024-11-05T19:20:46.176Z" }, +] + [[package]] name = "opentelemetry-instrumentation-bedrock" version = "0.33.12" @@ -4607,6 +4641,22 @@ wheels = [ { url = "https://files.pythonhosted.org/packages/bf/08/dce2b7926ace0204ce7946563348e1ff755873e387833484791e4ed391c8/opentelemetry_instrumentation_cohere-0.33.12-py3-none-any.whl", hash = "sha256:3bee3f7f7105259c85145be8c3b68612421860c95ad170f4d03144a3b8c07418", size = 5589, upload-time = "2024-11-13T20:27:21.317Z" }, ] +[[package]] +name = "opentelemetry-instrumentation-fastapi" +version = "0.49b0" +source = { registry = "https://pypi.org/simple" } +dependencies = [ + { name = "opentelemetry-api" }, + { name = "opentelemetry-instrumentation" }, + { name = "opentelemetry-instrumentation-asgi" }, + { name = "opentelemetry-semantic-conventions" }, + { name = "opentelemetry-util-http" }, +] +sdist = { url = "https://files.pythonhosted.org/packages/fe/bf/8e6d2a4807360f2203192017eb4845f5628dbeaf0597adf3d141cc5c24e1/opentelemetry_instrumentation_fastapi-0.49b0.tar.gz", hash = "sha256:6d14935c41fd3e49328188b6a59dd4c37bd17a66b01c15b0c64afa9714a1f905", size = 19230, upload-time = "2024-11-05T19:21:59.361Z" } +wheels = [ + { url = "https://files.pythonhosted.org/packages/b1/f4/0895b9410c10abf987c90dee1b7688a8f2214a284fe15e575648f6a1473a/opentelemetry_instrumentation_fastapi-0.49b0-py3-none-any.whl", hash = "sha256:646e1b18523cbe6860ae9711eb2c7b9c85466c3c7697cd6b8fb5180d85d3fe6e", size = 12101, upload-time = "2024-11-05T19:21:01.805Z" }, +] + [[package]] name = "opentelemetry-instrumentation-google-generativeai" version = "0.33.12"