mirror of
https://github.com/BerriAI/litellm.git
synced 2026-08-28 05:25:59 +00:00
* test(e2e): add other suite covering master-key auth and health lifecycle Covers the other.* holding-pen cells that were uncovered: master-key valid_allows/invalid_denied on the admin /user/list gate, and the lifecycle probes liveness.ping, readiness.public_probe, readiness.reports_db_status, and readiness_details.authenticated_diagnostics. New tests/e2e/other/ suite on the shared ProxyClient; the health probes send no auth header to prove the public routes need no credential, and the details route is asserted to reject an anonymous caller while exposing version/db diagnostics to the master key. * test(e2e): cover block_code_execution and openai_moderation guardrails Extends the guardrails suite with two built-in guardrails registered per request (default_on=False, opted in via the chat body's guardrails selector) so neither intercepts unrelated traffic on the shared proxy. block_code_execution.pre_call.blocks: a python code block plus a run-this request is intercepted with the canned content-blocked message and the model never runs, while the same code block asked about with don't-run-it reaches the model. Verified live. openai_moderations.pre_call.blocks: a flagged prompt is rejected 400 naming the moderation policy while a benign prompt passes. The guardrail calls OpenAI's moderation API; verifying it needs an OpenAI key with moderation quota (this account currently 429s the moderation endpoint). Adds a shared create_backend_model helper and a generic register() plus per-request guardrails/max_tokens on the client so more built-ins can reuse the same path. * test(e2e): cover presidio PII masking (pre_call + post_call) Registers a presidio guardrail per request (default_on=False) with the analyzer/anonymizer bases supplied in the registration params, so the test controls its own dependency and needs no proxy restart. presidio.pre_call.masks: a repeat-verbatim request comes back with the <EMAIL_ADDRESS> placeholder and never the raw email, proving the prompt was anonymized before the model saw it. presidio.post_call.masks: with apply_to_output the model's own emitted email is masked on the way out, so the caller never receives the raw value. Both verified live against real presidio analyzer + anonymizer containers. logging_only is intentionally not covered: /spend/logs exposes no prompt messages to read back the masked log, and a logging_only run also masked the response, contradicting its contract; noted in the module docstring for a follow-up. * test(e2e): cover presidio logging_only masking via OTEL read-back Adds the third presidio cell, guardrail.presidio.logging_only.masks. The logging_only contract (mask what is logged, do not block) is verified by reading the request's gen-AI span back from the real OTEL destination: the span's gen_ai.input.messages attribute carries the <EMAIL_ADDRESS> placeholder, never the raw email, and the call itself is not blocked. Reads the trace via the shared OtelReader, promoted from logging/ to the suite root so both suites use it. The masked prompt is polled to a deadline because logging_only masks the payload asynchronously and the span can briefly export before the mask lands. Drops the throwaway chat_send in favor of the existing transport.send for the call-id capture. * fix(e2e): tolerate cross-pod guardrail sync delay in team-opt-out test Stage runs multiple gateway pods behind the shared key. POST /guardrails registers a new default-on guardrail in-process immediately only on the pod that served the create call; every other pod picks it up on its next periodic DB sync (proxy_server.py, every 30s), so the very next chat call can race a pod that has not synced yet. Poll to a 40s deadline instead of asserting on the first response, matching the existing pattern in test_budget_reset_advances_e2e.py. * test(e2e): cover a guardrail on the MCP tool-call path (content_filter pre_mcp_call) Adds guardrail.litellm_content_filter.pre_mcp_call.blocks: against the real Datadog MCP server, a content_filter guardrail configured mode=pre_mcp_call blocks a banned keyword in an MCP tool call's arguments with HTTP 400 attributed to the pre_mcp_call hook, and lets a clean argument reach the upstream server. The guardrail attaches with default_on because per-key/request guardrail selection is dropped from the synthetic MCP request the hook sees; the banned keyword is unique per run so default_on only intercepts this test's own call. mode must be pre_mcp_call - a pre_call config silently no-ops on tools/call because the event type is rewritten for call_mcp_tool. Drives the tool directly via /mcp-rest/tools/call for a deterministic check of the same pre_mcp_call enforcement the OpenAI-SDK chat path hits when a model invokes an MCP tool. * fix(e2e): mid-conversation messages test uses client.proxy not client.gateway EndpointsClient exposes .proxy after the Gateway->ProxyClient rename; the mid-conversation system test still referenced .gateway, which fails the e2e basedpyright gate. Aligns it with the rest of the harness. * test(e2e): address review on the guardrail coverage MCP tool-call guardrail: poll the banned call until the guardrail is enforced instead of asserting on the first call, so the control-plane -> data-plane guardrail sync cannot race the check into a false pass-through; add a repeat banned call after enforcement to guard against a partial-propagation state. OpenAI moderation: distinguish a moderation-endpoint 429 (rate limit / no moderation quota) from a guardrail failure, so an account-capability gap reads as such rather than as "did not block". Runs green with a moderation-capable key. * test(e2e): close partial-propagation false-pass in MCP guardrail block test The single post-block repeat call could be load-balanced back to the same already-synced data-plane pod, so the test could pass while another pod still lacked the guardrail and let the banned MCP call reach Datadog. Anchor a wait to the guardrail create time (every pod is guaranteed to have DB-synced only after a full ~30s sync interval), then require the banned call to stay blocked across several attempts; a pass-through after that window is a real leak, not a race. * test(e2e): drop xfail-style rate-limit branch from openai_moderation test OpenAI's /v1/moderations is free and returns 200 with the env key (verified directly), so the RateLimitedError branch mislabeled the failure: a 429 there is insufficient_quota (no account billing), not throttling. The branch also only printed a softer message before failing anyway, an xfail-in-disguise the e2e rules forbid. A 429 now falls through and fails loudly with the full result.
140 lines
4.8 KiB
Python
140 lines
4.8 KiB
Python
"""Jaeger read-back for the OTEL trace-completeness tests: typed models over the
|
|
Jaeger query API (the destination's own API - completeness is judged on what the
|
|
backend actually holds, never on "export succeeded" proxy-side).
|
|
|
|
Traces are fetched server-side by the ``litellm.call_id`` tag the gen-AI span
|
|
carries (the request's x-litellm-call-id response header), so read-back is
|
|
immune to the query page filling up with unrelated traffic (background jobs,
|
|
other suites sharing the stack). Jaeger returns every span of a matching trace,
|
|
so the completeness assertions see the whole tree. A failed query is a hard
|
|
failure, never an empty result - an unreachable destination must not read as
|
|
"the trace never arrived".
|
|
|
|
External reads go through ``e2e_http`` (the only module allowed to call
|
|
``requests.*``).
|
|
"""
|
|
|
|
from __future__ import annotations
|
|
|
|
import json
|
|
import time
|
|
from dataclasses import dataclass
|
|
|
|
import pytest
|
|
from pydantic import BaseModel, ConfigDict, Field
|
|
|
|
from e2e_config import OTEL_QUERY_URL, POLL_INTERVAL, POLL_TIMEOUT
|
|
from e2e_http import URL, NoBody, Success, get
|
|
|
|
#: OTEL resource service.name the proxy exports under (OTEL_SERVICE_NAME default).
|
|
JAEGER_SERVICE = "litellm"
|
|
#: Span tag carrying the request's x-litellm-call-id (stamped on the gen-AI span).
|
|
CALL_ID_TAG = "litellm.call_id"
|
|
|
|
|
|
class JaegerTag(BaseModel):
|
|
model_config = ConfigDict(extra="ignore")
|
|
|
|
key: str
|
|
value: str | int | float | bool | None = None
|
|
|
|
|
|
class JaegerReference(BaseModel):
|
|
model_config = ConfigDict(extra="ignore", populate_by_name=True)
|
|
|
|
ref_type: str = Field(alias="refType")
|
|
trace_id: str = Field(alias="traceID")
|
|
span_id: str = Field(alias="spanID")
|
|
|
|
|
|
class JaegerSpan(BaseModel):
|
|
model_config = ConfigDict(extra="ignore", populate_by_name=True)
|
|
|
|
span_id: str = Field(alias="spanID")
|
|
operation_name: str = Field(alias="operationName")
|
|
start_time: int = Field(default=0, alias="startTime")
|
|
#: Span duration in microseconds, as reported by the Jaeger query API.
|
|
duration: int = 0
|
|
references: list[JaegerReference] = []
|
|
tags: list[JaegerTag] = []
|
|
|
|
@property
|
|
def kind(self) -> str:
|
|
for tag in self.tags:
|
|
if tag.key == "span.kind":
|
|
return str(tag.value)
|
|
return ""
|
|
|
|
|
|
class JaegerTrace(BaseModel):
|
|
model_config = ConfigDict(extra="ignore", populate_by_name=True)
|
|
|
|
trace_id: str = Field(alias="traceID")
|
|
spans: list[JaegerSpan] = []
|
|
|
|
def span_names(self) -> list[str]:
|
|
return sorted(span.operation_name for span in self.spans)
|
|
|
|
|
|
class JaegerTracesPage(BaseModel):
|
|
model_config = ConfigDict(extra="ignore")
|
|
|
|
data: list[JaegerTrace] = []
|
|
|
|
|
|
class _TracesQuery(BaseModel):
|
|
service: str
|
|
tags: str
|
|
limit: int = 20
|
|
lookback: str = "1h"
|
|
|
|
|
|
def _settled(trace: JaegerTrace, names: set[str], prefixes: set[str]) -> bool:
|
|
present = set(trace.span_names())
|
|
return names.issubset(present) and all(
|
|
any(name.startswith(prefix) for name in present) for prefix in prefixes
|
|
)
|
|
|
|
|
|
@dataclass(frozen=True, slots=True)
|
|
class OtelReader:
|
|
query_url: str
|
|
|
|
def traces_for_call(self, call_id: str) -> list[JaegerTrace]:
|
|
"""Every trace holding a span tagged with this call id. Jaeger matches
|
|
spans server-side and returns their full traces; more than one hit for
|
|
one call IS the split-trace bug, so this never collapses to one."""
|
|
result = get(
|
|
URL(f"{self.query_url}/api/traces"),
|
|
headers=NoBody(),
|
|
params=_TracesQuery(service=JAEGER_SERVICE, tags=json.dumps({CALL_ID_TAG: call_id})),
|
|
response_type=JaegerTracesPage,
|
|
timeout=30.0,
|
|
)
|
|
match result:
|
|
case Success(data=page):
|
|
return page.data
|
|
case failure:
|
|
pytest.fail(f"Jaeger query API at {self.query_url} failed: {failure}")
|
|
|
|
def poll_traces_for_call(
|
|
self, *, call_id: str, settled_names: set[str], settled_prefixes: set[str]
|
|
) -> list[JaegerTrace]:
|
|
"""Poll until exactly one trace holds the call and it carries every span
|
|
name in ``settled_names`` plus at least one name per prefix in
|
|
``settled_prefixes`` (spans flush in batches, the cost write lands after
|
|
the response), then return the hits. At the deadline the last hits are
|
|
returned as-is so the caller's assertions report the real final state -
|
|
on a split trace this never settles and the orphan comes back."""
|
|
deadline = time.monotonic() + POLL_TIMEOUT
|
|
hits: list[JaegerTrace] = []
|
|
while time.monotonic() < deadline:
|
|
hits = self.traces_for_call(call_id)
|
|
if len(hits) == 1 and _settled(hits[0], settled_names, settled_prefixes):
|
|
return hits
|
|
time.sleep(POLL_INTERVAL)
|
|
return hits
|
|
|
|
|
|
def build_otel_reader() -> OtelReader:
|
|
return OtelReader(query_url=OTEL_QUERY_URL)
|