mirror of
https://github.com/BerriAI/litellm.git
synced 2026-09-13 23:11:40 +00:00
* fix(e2e): stop the cache-settings test from persisting a degraded Redis config TestCacheSettings.test_update_persists_cache_backend_to_get read the live cache settings and wrote them back, intending a no-op. Its capture modelled only type/host/port, so on a TLS cluster the write-back silently dropped `ssl` and `redis_startup_nodes`. That is not recoverable on its own. `/cache/settings` persists what it receives into LiteLLM_CacheConfig, that row outranks the YAML `cache_params`, and init_cache_settings_in_db re-applies it on a timer, so a restart does not clear it. The proxy ends up driving a TLS-only cluster endpoint as a plaintext standalone node and every Redis call blocks to socket timeout. On the affected deployment that took out rate limiting entirely (the v3 limiter is a Lua script on Redis with no DB fallback), Redis-only budget levels (tag, per-model, team-member, per-window), spend tracking, `ResetBudgetJob` (which self-starved at 54 skipped runs per 15 min), and `ProxyConfig.add_deployment`, whose last statement syncs guardrails and never ran. 60 of 72 failures in one run traced back here. The settings blob is now round-tripped verbatim via a RootModel over an exhaustive value union, so a subset cannot be written. Two guards make a regression fail loudly at this test instead of silently downstream: - refuse to write when GET reports redis_type=cluster but omits redis_startup_nodes, which is the exact precondition for persisting a downgrade. GET resolves the stored row overlaid with REDIS_* env and never reads YAML, so a cluster configured only in YAML cannot round-trip here - compare /cache/ping before and after, so a write that breaks connectivity fails this test rather than every suite that follows The underlying product defect is filed as LIT-4816: GET cannot express the effective config, and a partial POST is allowed to downgrade transport. This change only stops the suite from triggering it; the Admin UI can still do so. basedpyright clean (0 errors) under the e2e gate. * fix(e2e): scope the bedrock guardrail per request and send OpenAI's current token param Two failures that had nothing to do with the guardrail or route under test. create_bedrock_guardrail registered with default_on=True, which applies the guardrail to every request the proxy serves. The upstream ApplyGuardrail call was answering 403, and that came back to unrelated traffic as `403 Bedrock guardrail request failed`, failing three a2a tests and a passthrough headers test alongside the bedrock one. The harness already supports the per-request `guardrails` selector, so the guardrail is now registered opted out of default_on and selected by the test that wants it. A broken upstream guardrail fails its own test instead of whatever else is running. Note this only contains the blast radius; the 403 itself still needs the bedrock:ApplyGuardrail permission (or a valid guardrail identifier) on the deployment, so test_bedrock_pre_call_blocks_harmful_prompt can still fail on its own until that is sorted. The OpenAI passthrough body sent `max_tokens`, which newer models reject with "Unsupported parameter: 'max_tokens' is not supported with this model. Use 'max_completion_tokens' instead." Passthrough forwards the body untranslated, so drop_params does not apply and the body has to satisfy OpenAI's contract directly. vllm_chat keeps max_tokens, which vLLM accepts. basedpyright clean (0 errors) under the e2e gate. * fix(e2e): drop the pinned a2a api_key that broke every message/send #34512 pinned `api_key="os.environ/ANTHROPIC_API_KEY"` on the a2a bridge agent. The a2a bridge forwards the agent's litellm_params straight into litellm.acompletion() without expanding "os.environ/" indirection, so that literal string was sent upstream as x-api-key and every message/send failed with `AnthropicException - {"type":"authentication_error","message":"invalid x-api-key"}`. Omitting api_key restores the normal provider resolution: litellm reads ANTHROPIC_API_KEY from the proxy's own environment for this provider, which is what the agent-owner flow depends on and what the suite did before #34512. Verified against a live proxy, same agent shape each time: api_key omitted -> message/send 200 api_key "os.environ/ANTHROPIC_API_KEY" -> message/send 500 invalid x-api-key api_key <literal key> -> message/send 200 and the key itself is valid (direct call to api.anthropic.com returns 200), so this was indirection that never got expanded rather than a bad credential. This accounts for four failures (test_semver_protocol_version_registers_and_serves, test_message_send_runs_completion_bridge, test_pinned_v0_3_serves_flat_message_shape, test_pinned_v1_0_serves_nested_message_shape). They were previously reported as `403 Bedrock guardrail request failed`, because a default_on Bedrock guardrail short-circuited the request before it ever reached the bridge and hid this. The bridge silently ignoring "os.environ/" in agent params is a product defect in its own right, filed separately; anyone configuring an agent credential that way through the UI hits the same wall. basedpyright clean (0 errors) under the e2e gate. * test(e2e): make the load suite less aggressive against a shared proxy 750 users at spawn rate 50 saturated the request path hard enough to distort the latency-sensitive suites sharing the same proxy, and it spends real provider money at that rate. Drop to 200 users at spawn rate 20. The RPS floor moves with the user count rather than staying put, so the assertion keeps its meaning instead of becoming a formality: 355 RPS over 750 users is ~0.47 RPS/user, and 90 over 200 holds that same per-user expectation with a similar pass margin. A request-path regression still trips it. All four knobs stay env-overridable (E2E_LOAD_USERS, E2E_LOAD_SPAWN_RATE, E2E_LOAD_DURATION_SECONDS, E2E_LOAD_MIN_RPS) for a deliberate load run. Note the recorded failure for this test was "no requests completed in 60s", which was the gateway wedged on unreachable Redis rather than a throughput regression; this change is about not perturbing its neighbours, not about that failure. * fix(e2e): make the reasoning-tokens assertion exercise a request that reasons test_openai_chat_reasoning_reports_reasoning_tokens asked "A train travels 60 miles in 1.5 hours. What is its average speed in mph?" at reasoning_effort="low", then asserted reasoning_tokens > 0. The model answers that directly without reasoning, so 0 is correct behavior and the assertion was testing the model's discretion rather than litellm's reporting. Verified against a live proxy on a dedicated openai/gpt-5.6 deployment, matching how the test provisions its model: reasoning_effort=low, one-step arithmetic -> reasoning_tokens=0 reasoning_effort=high, the prompt used here -> reasoning_tokens=114 Raised to high effort with a prompt that requires a proof plus a search, so the field under test is actually populated and the assertion fails only if litellm stops surfacing it. While confirming this I also checked prompt caching, which needed no change: cached_tokens comes back 3615 of 3618 prompt tokens on a repeated large prefix against a dedicated deployment. An earlier reading of 0 was an artifact of probing a fan-out alias whose requests land on different deployments, not a caching defect. * test(e2e): skip the files-list test while LIT-4820 is open GET /v1/files does not include a just-uploaded file. The upload returns 200 and GET /v1/files/{id} resolves it, but the listing never contains it: the returned set stays fixed at 27 entries whose newest created_at is roughly ten hours older than the upload, on both the managed (/v1/files?model=) and provider-scoped (/openai/v1/files) routes. Polled for 40s, so not an eventual-consistency window. Filed as LIT-4820. Skipping keeps a known, ticketed product bug from holding the suite red and masking a new regression somewhere else in the same test. The assertion is left exactly as it was on purpose. It encodes the contract we actually want, that a file retrievable by id is also enumerable, and anything that lists files (a UI picker, cleanup tooling that lists then deletes and would therefore leak provider-side files) depends on it. Relaxing it to get green would delete the signal. The skip reason says so and links the ticket, and the ticket records that removing this marker is part of its definition of done. Matches the existing pattern in this file, where test_unified_file_and_batch_create skips with a reason citing LIT-3266. While skipped, the registry cell llm.files.openai.list.nonstream.works has no passing covering test, so files-list coverage reports as uncovered rather than passing, which is the honest state. * fix(e2e): parse Sentinel node lists in the cache-settings model The value union covered scalar lists and lists of mappings, but not lists of lists. `redis_startup_nodes` holds host/port mappings while `sentinel_nodes` holds positional pairs (CACHE_SETTINGS_FIELDS documents `[['localhost', 26379]]`), so on a Sentinel deployment pydantic rejected the response: sentinel_nodes.list[dict[str,...]].1 Input should be a valid dictionary [input_value=['localhost', 26380]] The round-trip test reads GET /cache/settings before it writes anything, so that rejection failed the test at the read, before any assertion ran. A Sentinel deployment would have looked like a broken cache-settings route rather than a model too narrow to parse a documented shape. A list element may now be a scalar, a list or a mapping, which covers both node shapes without special-casing either and tolerates a heterogeneous list instead of rejecting the whole response. Adds TestCacheSettingsModel, harness-level with no `e2e` marker so it runs without a proxy, covering all four backend shapes (cluster mappings, sentinel pairs, plain node, url mode with a null discrete field) plus transport() key selection. Confirmed it fails on the previous union and passes on this one: old union -> 1 failed, 4 passed (the sentinel case) new union -> 5 passed * test(e2e): remove the cache-settings round-trip test The test could not fail for the thing it claimed to test, and could break the deployment it ran against. Both halves of that are worth stating. It read the live settings, wrote back identical values, and asserted the read-back matched. If POST /cache/settings were a complete no-op that returned 200 and touched nothing, GET would still return the values read a moment earlier and the test would pass. It verified that GET is stable, not that the route persists anything. Against that, /cache/settings persists what it receives into LiteLLM_CacheConfig, that row outranks YAML cache_params, and init_cache_settings_in_db re-applies it on a timer. A write that omits ssl or redis_startup_nodes converts a TLS cluster into a plaintext standalone client and every later Redis call blocks to socket timeout. On 2026-07-25 that failed 60 of 72 tests in one run: rate limiting stopped enforcing, Redis-only budgets admitted billable over-budget spend, ResetBudgetJob self-starved, and guardrail sync never ran. Guarding the previous shape was not sufficient. Writing the blob verbatim plus a cluster precondition and a /cache/ping check narrowed the hazard but did not remove it, because GET cannot express the effective config: it resolves the stored row overlaid with REDIS_* env and never reads YAML. On a fresh deploy it cannot see YAML's ssl to echo back, so a TLS non-cluster deployment could still have a row written that drops it. No round-trip through this route is safe on a shared proxy. Removed with the models and helpers it owned, and TestCacheSettingsModel with them since it existed only to protect that parsing. The registry row mgmt.cache_settings.update.happy_path stays, now carrying the rationale for why it is deliberately uncovered and what a safe test would require (an isolated proxy, or LIT-4816 fixed so a partial write cannot downgrade transport). Coverage therefore reports this cell as a gap, which is the honest state. Collector passes --strict; the module still collects 11 tests.
131 lines
6.6 KiB
Python
131 lines
6.6 KiB
Python
"""Generic configuration for live e2e tests against a running LiteLLM proxy.
|
|
|
|
Shared by every e2e suite under tests/e2e/. Values come from the
|
|
environment so the same tests run against localhost or a deployed proxy.
|
|
"""
|
|
|
|
from __future__ import annotations
|
|
|
|
import os
|
|
import uuid
|
|
from pathlib import Path
|
|
|
|
from dotenv import load_dotenv
|
|
|
|
# Local runs keep provider / DataDog keys in tests/e2e/.env (see CONTRIBUTING.md).
|
|
# Compose injects them into the proxy container, but pytest on the host does not
|
|
# inherit that file unless we load it. override=False so a real shell export wins.
|
|
load_dotenv(Path(__file__).resolve().parent / ".env", override=False)
|
|
|
|
PROXY_BASE_URL = os.environ.get("LITELLM_PROXY_URL", "http://localhost:4000").rstrip("/")
|
|
MASTER_KEY = os.environ.get("LITELLM_MASTER_KEY", "sk-1234")
|
|
|
|
# Control-plane (management/admin) base URL. Defaults to PROXY_BASE_URL so a
|
|
# single path-routing host (stage ALB, compose monolith) works for both planes.
|
|
# Set LITELLM_CONTROL_PLANE_URL only when management is a different base than
|
|
# the LLM host and you are not going through an ingress that path-routes.
|
|
CONTROL_PLANE_BASE_URL = os.environ.get(
|
|
"LITELLM_CONTROL_PLANE_URL", PROXY_BASE_URL
|
|
).rstrip("/")
|
|
|
|
UI_USERNAME = os.environ.get("E2E_UI_USERNAME", "admin")
|
|
UI_PASSWORD = os.environ.get("E2E_UI_PASSWORD", MASTER_KEY)
|
|
|
|
# Dashboard base for playwright. Defaults to PROXY_BASE_URL so one ALB/monolith
|
|
# host covers /ui as well. Override E2E_UI_BASE_URL only if the UI is elsewhere.
|
|
UI_BASE_URL = os.environ.get("E2E_UI_BASE_URL", PROXY_BASE_URL).rstrip("/")
|
|
|
|
CHEAP_ANTHROPIC_MODEL = os.environ.get("E2E_CHEAP_ANTHROPIC_MODEL", "claude-haiku-4-5")
|
|
CHEAP_OPENAI_MODEL = os.environ.get("E2E_CHEAP_OPENAI_MODEL", "gpt-5.5")
|
|
|
|
LINEAR_MCP_URL = os.environ.get("E2E_LINEAR_MCP_URL", "https://mcp.linear.app/mcp")
|
|
LINEAR_STORAGE_STATE = os.environ.get("E2E_LINEAR_STORAGE_STATE", "")
|
|
|
|
# Jaeger query API of the compose stack's OTEL trace destination (the `jaeger`
|
|
# service in docker-compose.yml maps it to host 16686). Trace-completeness tests
|
|
# read exported spans back through it.
|
|
OTEL_QUERY_URL = os.environ.get("E2E_OTEL_QUERY_URL", "http://localhost:16686").rstrip("/")
|
|
|
|
# Real-DataDog read-back (no local sink - destination fakes cannot be deployed
|
|
# on the cluster): the proxy delivers with DD_API_KEY as in production, and the
|
|
# tests read ingested events back through the DataDog Logs Search API, which
|
|
# additionally needs an application key. On the cluster the secret manager
|
|
# injects both; locally tests/e2e/.env provides them.
|
|
DD_SITE = os.environ.get("DD_SITE", "datadoghq.com").strip()
|
|
DD_API_KEY = os.environ.get("DD_API_KEY", "").strip()
|
|
DD_APP_KEY = os.environ.get("DD_APP_KEY", "").strip()
|
|
|
|
# After the first event is searchable, keep watching this long for a late
|
|
# duplicate before the exactly-one assertion: real-DataDog ingestion jitter can
|
|
# make one call's two events searchable tens of seconds apart, and a duplicate
|
|
# that surfaces late IS the bug (LIT-4447), so one poll interval is not enough.
|
|
DD_SETTLE_SECONDS = float(os.environ.get("E2E_DD_SETTLE_SECONDS", "30"))
|
|
# DataDog Logs Search `from` window (relative to now). Wide enough for a suite
|
|
# run plus ingestion lag; override if a long CI queue needs a wider lookback.
|
|
DD_SEARCH_FROM = os.environ.get("E2E_DD_SEARCH_FROM", "now-30m").strip() or "now-30m"
|
|
# The Logs Search API budget is tight - 2 requests per 10s org-wide
|
|
# (x-ratelimit-name logs_public_search_api) - so read-backs pace their search
|
|
# calls at this interval instead of POLL_INTERVAL, and back off when a 429
|
|
# still slips through (the budget is shared with anything else searching).
|
|
DD_SEARCH_INTERVAL = float(os.environ.get("E2E_DD_SEARCH_INTERVAL", "10"))
|
|
|
|
# Writes on the proxy are eventually consistent (e.g. spend rows flush on
|
|
# proxy_batch_write_at, ~60s). Read-backs poll to this deadline, never sleep-once.
|
|
POLL_TIMEOUT = float(os.environ.get("E2E_POLL_TIMEOUT", "120"))
|
|
POLL_INTERVAL = float(os.environ.get("E2E_POLL_INTERVAL", "5"))
|
|
REQUEST_TIMEOUT = float(os.environ.get("E2E_REQUEST_TIMEOUT", "60"))
|
|
|
|
EXPECT_RUST = os.environ.get("E2E_EXPECT_RUST", "").strip().lower() in ("1", "true", "yes")
|
|
|
|
# Deliberately modest concurrency. The suite shares its proxy with every other
|
|
# suite in the run, and 750 users at spawn rate 50 saturated the request path hard
|
|
# enough to distort latency-sensitive neighbours (and to spend real provider money
|
|
# fast). The SLO is scaled with the user count to keep the same per-user throughput
|
|
# expectation (~0.47 RPS/user), so this still catches a request-path regression
|
|
# rather than becoming a formality. Raise all four via env for a real load run.
|
|
LOAD_USERS = int(os.environ.get("E2E_LOAD_USERS", "200"))
|
|
LOAD_SPAWN_RATE = float(os.environ.get("E2E_LOAD_SPAWN_RATE", "20"))
|
|
LOAD_DURATION_SECONDS = float(os.environ.get("E2E_LOAD_DURATION_SECONDS", "60"))
|
|
LOAD_MIN_RPS = float(os.environ.get("E2E_LOAD_MIN_RPS", "90"))
|
|
LOAD_MAX_FAILURE_RATIO = float(os.environ.get("E2E_LOAD_MAX_FAILURE_RATIO", "0.01"))
|
|
|
|
WEEKLY_ANOMALY_OPT_IN_ENV = "E2E_WEEKLY_ANOMALY"
|
|
ANOMALY_SESSIONS = int(os.environ.get("E2E_ANOMALY_SESSIONS", "6"))
|
|
ANOMALY_TURNS_PER_SESSION = int(os.environ.get("E2E_ANOMALY_TURNS_PER_SESSION", "6"))
|
|
ANOMALY_TURN_ATTEMPTS = int(os.environ.get("E2E_ANOMALY_TURN_ATTEMPTS", "3"))
|
|
ANOMALY_MAX_ERROR_RATIO = float(os.environ.get("E2E_ANOMALY_MAX_ERROR_RATIO", "0.05"))
|
|
ANOMALY_MIN_WARM_CACHE_READ_SHARE = float(
|
|
os.environ.get("E2E_ANOMALY_MIN_WARM_CACHE_READ_SHARE", "0.65")
|
|
)
|
|
ANOMALY_MAX_P95_TURN_SECONDS = float(
|
|
os.environ.get("E2E_ANOMALY_MAX_P95_TURN_SECONDS", "30")
|
|
)
|
|
ANOMALY_MAX_KEY_SPEND_USD = float(
|
|
os.environ.get("E2E_ANOMALY_MAX_KEY_SPEND_USD", "0.60")
|
|
)
|
|
ANOMALY_SPEND_SETTLE_SECONDS = float(
|
|
os.environ.get("E2E_ANOMALY_SPEND_SETTLE_SECONDS", "75")
|
|
)
|
|
|
|
|
|
def datadog_mcp_url(*, toolsets: str = "core") -> str:
|
|
"""Regional Datadog remote MCP endpoint for this process's DD_SITE.
|
|
|
|
US1 is mcp.datadoghq.com; every other site is mcp.<site> (e.g. us5 ->
|
|
mcp.us5.datadoghq.com). A fixed mcp.datadoghq.com URL 403s when the keys
|
|
belong to a non-US1 org.
|
|
"""
|
|
site = (
|
|
os.environ.get("DD_SITE", DD_SITE) or "datadoghq.com"
|
|
).strip().removeprefix("https://").removeprefix("http://").rstrip("/")
|
|
if site.startswith("app."):
|
|
site = site[len("app.") :]
|
|
host = "mcp.datadoghq.com" if site in ("", "datadoghq.com") else f"mcp.{site}"
|
|
base = f"https://{host}/v1/mcp"
|
|
return f"{base}?toolsets={toolsets}" if toolsets else base
|
|
|
|
|
|
def unique_marker() -> str:
|
|
"""A short unique token per call/run, so concurrent runs and the shared
|
|
response cache never collide on prompts, tags, or customer ids."""
|
|
return uuid.uuid4().hex[:12]
|