* test(e2e): add other suite covering master-key auth and health lifecycle Covers the other.* holding-pen cells that were uncovered: master-key valid_allows/invalid_denied on the admin /user/list gate, and the lifecycle probes liveness.ping, readiness.public_probe, readiness.reports_db_status, and readiness_details.authenticated_diagnostics. New tests/e2e/other/ suite on the shared ProxyClient; the health probes send no auth header to prove the public routes need no credential, and the details route is asserted to reject an anonymous caller while exposing version/db diagnostics to the master key. * test(e2e): cover block_code_execution and openai_moderation guardrails Extends the guardrails suite with two built-in guardrails registered per request (default_on=False, opted in via the chat body's guardrails selector) so neither intercepts unrelated traffic on the shared proxy. block_code_execution.pre_call.blocks: a python code block plus a run-this request is intercepted with the canned content-blocked message and the model never runs, while the same code block asked about with don't-run-it reaches the model. Verified live. openai_moderations.pre_call.blocks: a flagged prompt is rejected 400 naming the moderation policy while a benign prompt passes. The guardrail calls OpenAI's moderation API; verifying it needs an OpenAI key with moderation quota (this account currently 429s the moderation endpoint). Adds a shared create_backend_model helper and a generic register() plus per-request guardrails/max_tokens on the client so more built-ins can reuse the same path. * test(e2e): cover presidio PII masking (pre_call + post_call) Registers a presidio guardrail per request (default_on=False) with the analyzer/anonymizer bases supplied in the registration params, so the test controls its own dependency and needs no proxy restart. presidio.pre_call.masks: a repeat-verbatim request comes back with the <EMAIL_ADDRESS> placeholder and never the raw email, proving the prompt was anonymized before the model saw it. presidio.post_call.masks: with apply_to_output the model's own emitted email is masked on the way out, so the caller never receives the raw value. Both verified live against real presidio analyzer + anonymizer containers. logging_only is intentionally not covered: /spend/logs exposes no prompt messages to read back the masked log, and a logging_only run also masked the response, contradicting its contract; noted in the module docstring for a follow-up. * test(e2e): cover presidio logging_only masking via OTEL read-back Adds the third presidio cell, guardrail.presidio.logging_only.masks. The logging_only contract (mask what is logged, do not block) is verified by reading the request's gen-AI span back from the real OTEL destination: the span's gen_ai.input.messages attribute carries the <EMAIL_ADDRESS> placeholder, never the raw email, and the call itself is not blocked. Reads the trace via the shared OtelReader, promoted from logging/ to the suite root so both suites use it. The masked prompt is polled to a deadline because logging_only masks the payload asynchronously and the span can briefly export before the mask lands. Drops the throwaway chat_send in favor of the existing transport.send for the call-id capture. * fix(e2e): tolerate cross-pod guardrail sync delay in team-opt-out test Stage runs multiple gateway pods behind the shared key. POST /guardrails registers a new default-on guardrail in-process immediately only on the pod that served the create call; every other pod picks it up on its next periodic DB sync (proxy_server.py, every 30s), so the very next chat call can race a pod that has not synced yet. Poll to a 40s deadline instead of asserting on the first response, matching the existing pattern in test_budget_reset_advances_e2e.py. * test(e2e): cover a guardrail on the MCP tool-call path (content_filter pre_mcp_call) Adds guardrail.litellm_content_filter.pre_mcp_call.blocks: against the real Datadog MCP server, a content_filter guardrail configured mode=pre_mcp_call blocks a banned keyword in an MCP tool call's arguments with HTTP 400 attributed to the pre_mcp_call hook, and lets a clean argument reach the upstream server. The guardrail attaches with default_on because per-key/request guardrail selection is dropped from the synthetic MCP request the hook sees; the banned keyword is unique per run so default_on only intercepts this test's own call. mode must be pre_mcp_call - a pre_call config silently no-ops on tools/call because the event type is rewritten for call_mcp_tool. Drives the tool directly via /mcp-rest/tools/call for a deterministic check of the same pre_mcp_call enforcement the OpenAI-SDK chat path hits when a model invokes an MCP tool. * fix(e2e): mid-conversation messages test uses client.proxy not client.gateway EndpointsClient exposes .proxy after the Gateway->ProxyClient rename; the mid-conversation system test still referenced .gateway, which fails the e2e basedpyright gate. Aligns it with the rest of the harness. * test(e2e): address review on the guardrail coverage MCP tool-call guardrail: poll the banned call until the guardrail is enforced instead of asserting on the first call, so the control-plane -> data-plane guardrail sync cannot race the check into a false pass-through; add a repeat banned call after enforcement to guard against a partial-propagation state. OpenAI moderation: distinguish a moderation-endpoint 429 (rate limit / no moderation quota) from a guardrail failure, so an account-capability gap reads as such rather than as "did not block". Runs green with a moderation-capable key. * test(e2e): close partial-propagation false-pass in MCP guardrail block test The single post-block repeat call could be load-balanced back to the same already-synced data-plane pod, so the test could pass while another pod still lacked the guardrail and let the banned MCP call reach Datadog. Anchor a wait to the guardrail create time (every pod is guaranteed to have DB-synced only after a full ~30s sync interval), then require the banned call to stay blocked across several attempts; a pass-through after that window is a real leak, not a race. * test(e2e): drop xfail-style rate-limit branch from openai_moderation test OpenAI's /v1/moderations is free and returns 200 with the env key (verified directly), so the RateLimitedError branch mislabeled the failure: a 429 there is insufficient_quota (no account billing), not throttling. The branch also only printed a softer message before failing anyway, an xfail-in-disguise the e2e rules forbid. A 429 now falls through and fails loudly with the full result.
18 KiB
e2e harness conventions
Code-style rules for writing tests under tests/e2e/. The harness already encodes the plumbing; your job is the feature-specific behavior, not reinventing it. For what a complete test must do (the lifecycle contract, asserting both recorded state and enforced behavior) and how to run a suite, see CONTRIBUTING.md in this directory. Repo-wide conventions live in the root CLAUDE.md
Suite folders
Each subdirectory under tests/e2e/ is one suite, scoped to an endpoint family or behavior area. If you add a new folder, you must add a line here describing what kind of tests belong in it, so the layout stays self-describing. gateway/ is the exception: it holds proxy configuration only and never tests
llm_translation/- LLM endpoint and provider-translation behavior: passthrough, custom pricing, OCR, and the non-chat inference endpoints (/v1/responses,/v1/messages,/embeddings,/v1/rerank,/v1/audio/speech,/v1/images/generations), each against a deployment the test creates via/model/newand deletes on teardownaccess_control/- the gateway's authorization and error-shape contract: per-key model allow-lists, route-group permissions (allowed_routes), and unknown-model validationembeddings/- the/embeddingsendpoint across providersbatches/- the/batchesendpoint (placeholder until the first test lands)realtime/- realtime websocket sessions, including the pipecat audio pathquota_management/- quota enforcement and accounting, one subfolder per behavior:ratelimit/(rpm/tpm blocks, window reset, pacing headers on live traffic),budgets/(budget definition, enforcement, and reset windows: key, team, tag, soft, multi-window), andspend_tracking/(spend logging and cost attribution on/spend/*)management/- key/team/user/organization management routes: create/update/delete persistence via the info routes, team membership, and llm-only-key route denials (API surface; not Playwright)mcp/- the MCP server surface over api_key auth against the real Datadog remote MCP server only (see "MCP suite: real Datadog only" below)logging/- logging-integration delivery (datadog and friends)security/- secret handling and log-leak protectionrouter/- routing and reliability behavior (fallbacks, cooldowns)load/- throughput/performance under concurrency: drives real concurrent traffic through the whole stack with Locust and asserts a throughput SLO; markedloadso the parent conftest collects it last and it never perturbs latency-sensitive suitesother/- the holding-pen suite for theother.*registry cluster with no home of its own yet: the master-key auth gate and the process-lifecycle health probes (liveness, public readiness, authenticated readiness diagnostics). Promote a cluster out once it is large/stable enough for its own suitegateway/- proxy configuration only (litellm-config.yml); no testsclaude_code/- the Claude Code compatibility matrix: drives the realclaudeCLI (and HTTP probes) against a proxy for each feature x provider cell, reporting tagged-union outcomes via thecompat_resultfixture; ships its own driver/builder/publisher plus_*_unit_tests/trees. The HTTP probes ride the shared transport (ProxyClient.count_tokens/ProxyClient.messages); the CLI-driving path stays bespoke
MCP suite: real Datadog only
Every test under tests/e2e/mcp/ must exercise the proxy against the real Datadog remote MCP server. Do not add a compose service, FastMCP fixture, mock upstream, or any other fake MCP host for this suite
- Register via
register_datadog_mcpintests/e2e/mcp/datadog_mcp.py(or extend that helper if you need a differenttoolsets=/allowed_toolsslice of the same Datadog endpoint). That posts/v1/mcp/serverwithurl=datadog_mcp_url(...)and static headersDD-API-KEY/DD-APPLICATION-KEYfrom the process env - Auth is Datadog's documented CI/header path, not a browser OAuth authorize/token dance. Hard-fail when
DD_API_KEYorDD_APP_KEYis missing (assert_dd_mcp_creds); never skip for a missing fake upstream - Prefer calling real Datadog tools that prove the product path (e.g.
search_datadog_logsfor list/call and permission denials). Seed a unique marker (e2e-datadog-mcp-*) in a chat completion when you need a log the tool can find; dual-read withdd_logsfrom conftest when delivery matters - Delete the MCP server (and any keys) through
resources.deferthe same way every other suite tears down - If a new MCP behavior cannot be covered with Datadog's tool surface, say so in the PR and get agreement before inventing another upstream; the default is always Datadog
Lay the pattern down in a class
Keep the cases for one feature inside a class so the file reads as a spec for how that feature behaves in production. The class name says what is under test; each method is one behavior. Think of it as documenting the contract, with the rough intent being
# pseudo-code to convey intent, not the real API
class TestPromptCompression:
def test_prompt_compression_add_to_virtual_key(self):
new_key = self.resources.create_key(user_id, compression=True) # turn the feature on
resources._defer(new_key) # queue key deletion
def test_prompt_compression_accumulate_spend(self, key_id, user_id):
for _ in range(10):
response = self.resources.proxy.post("gemini-2.5-flash", key_id, user_id)
compressed_value = ...
assert response.cost == compressed_value # the cost was actually reduced
That snippet only conveys intent. What you actually write uses the real harness: the client fixture for your suite, the scoped_key fixture for an auto-deleted key, typed pydantic bodies from models.py, and unwrap(...) on the tagged-union result. tests/e2e/llm_translation/test_custom_pricing_e2e.py is the reference to copy from; it creates a scoped key, drives a real gemini call, polls /spend/logs to a deadline for the cost-breakdown row, then asserts the input and output costs match the configured custom rates and that a sibling deployment kept its own price. Read it before writing yours
Use the shared transport; never touch requests directly
Every HTTP call goes through the shared transport, never through requests.* in a test. e2e_http.py is the only module permitted to call requests.*, and that is enforced in CI by tests/code_coverage_tests/check_e2e_no_raw_requests.py. A test that imports requests will fail the check
The shape is layered so tests stay declarative
transport.py exposes a Transport Protocol with post, get, delete, send, stream, probe, plus bearer(key) and the master header. HttpTransport fulfils it, and SplitTransport routes each call by path to the data plane or the control plane so a split control-plane/data-plane deployment works without any change in the test
proxy_client.py holds ProxyClient, a frozen dataclass that wraps a Transport and adds the operations tests reuse: generate_key / delete_key / key_info, model_info, the LLM calls chat / chat_stream / embed / ocr, the spend read-back spend_logs, and the poll helpers poll_logs_for_key / poll_logs_for_request_id that loop to poll_timeout instead of sleeping once. It is exposed as the session-scoped proxy fixture (see tests/e2e/conftest.py), which each suite's client fixture depends on and injects. Add a new route as a method here so other suites get it for free
Each suite provides its own client fixture (see llm_translation/passthrough_client.py), a frozen dataclass that holds the shared ProxyClient (as .proxy) and adds suite-specific routes. Cleanup runs through that same ProxyClient, so whatever keys or customers your test creates get torn down by the resources fixture
Request and response bodies are typed pydantic models in models.py; only the fields a test reads are modelled, and nothing passes raw dicts. Outcomes come back as a Result[R] tagged union (Success, NetworkError, UnauthorizedError, RateLimitedError, ValidationError, UnknownApiError). Handle them with match, or call unwrap(...) when a non-success should fail the test. The harness hard-fails and never skips: a test marked e2e fails when no proxy answers its liveness probe, and once a request reaches the proxy any wrong behavior is likewise a hard failure, so a missing proxy turns the run red instead of being mistaken for a pass
Mark live tests with @pytest.mark.e2e (on the class or the module). Pure coverage of the harness itself carries no marker and runs regardless. Use scoped_key for a fresh all-models key that auto-deletes, resources when you need to create and tear down more than a key, and unique_marker() from e2e_config to keep prompts, tags, and customer ids from colliding across concurrent runs and the shared response cache
Typing
The harness is fully typed with no error budget: make lint-e2e-basedpyright must report zero basedpyright errors, and CI enforces that on any PR touching tests/e2e/**/*.py. When a response field is untyped, model it in models.py (just the fields you read) and let pydantic validate it, rather than threading a dict or Any through the test
Coverage registry
The set of tests we want is a registry checked into this repo, one row per behavior; that file is the definition of done and the denominator. Each e2e test declares what it covers with @pytest.mark.covers("..."), and a small collector diffs the registry against the tests and ships coverage to the existing Grafana. No Allure, no new dependencies
Coverage is organized as module > feature > test. Dashboard modules are Core LLMs, Non-Core LLMs, MCPs, Management/UI, Reliability & Performance, Quota Management, Logging & Guardrails, and Other. The Loki stdout formatter maps those display modules to log-safe labels (core_llms, non_core_llms, mcp, management_ui, reliability_performance, quota_management, logging_guardrails, and other) without changing JSON or Prometheus labels. A feature is either an endpoint (/chat/completions) or a behavior (fallbacks, rate limits; config-driven, with no route of its own). A cell reads like llm.chat_completions.bedrock_converse.tool_use.stream.works
The metric is coverage: the share of registry rows that have a passing covering test, reported to Grafana per module so a gap surfaces as an uncovered row rather than a silent absence
Tests do not declare a dashboard module directly. They only declare the registry cell id with @pytest.mark.covers("..."); the registry row decides the module, tier, endpoint, and dashboard rollup. Run python -m coverage_registry.collector --strict when you want CI to reject unknown marker ids. Add --fail-on-collection-errors when the job should also fail on pytest collection errors.
Naming grammar per module
LLMs - endpoint features (subject = the route), seeded from the Claude Code compat matrix. chat_completions, messages, and responses roll up to Core LLMs. Other LLM endpoints, including batches and realtime, roll up to Non-Core LLMs.
llm.<endpoint>.<route>.<capability>.<streaming>.<assertion>
endpoint : chat_completions | messages | responses | embeddings | batches | files
| rerank | images_generations | audio_speech | audio_transcriptions | moderations
| realtime
route : openai | azure_openai | anthropic | bedrock_converse | bedrock_invoke | vertex
| azure_foundry | cohere | together_ai
(vocab varies per endpoint; messages is anthropic-format only)
capability : basic | tool_use | prompt_cache_5m | vision | thinking | structured_output
| service_tier | mid_conversation_system
streaming : stream | nonstream (omit where n/a)
assertion : works | cost_logged | cache_hit
label (not in id): model = haiku-4.5 | sonnet-4.6 | opus-4.7 | gpt-*
e.g. llm.chat_completions.bedrock_converse.tool_use.stream.works
llm.messages.anthropic.prompt_cache_1h.nonstream.cache_hit
Management / UI - endpoint features (surface tag: api | ui)
mgmt.<endpoint>.<assertion>
endpoint : key.generate | key.update | key.delete | team.new | user.new
| budget.new | model.add | ... (one per management route)
assertion : persists | member_forbidden | admin_only | happy_path
e.g. mgmt.key.generate.persists (surface=api)
mgmt.key.generate.happy_path (surface=ui)
MCPs - endpoint features with the protocol op as the variant
mcp.<operation>.<auth_family>.<assertion>
operation : list_tools | call_tool | list_resources | read_resource | list_prompts | get_prompt
auth_family : none | api_key | bearer | oauth
assertion : succeeds | denied_without_permission
e.g. mcp.call_tool.oauth.succeeds
Reliability & Performance - behavior features (no route; endpoint is exercised_on)
reliability.<behavior>.<variant>.<assertion>
behavior : fallback | retry | cooldown | timeout | routing | cache | circuit_breaker | perf
variant : <trigger> 5xx | context_window | content_policy | 429 | timeout
<strategy> simple_shuffle | usage_based | latency_based | cost_based | least_busy
<dimension> latency | throughput (perf only; SLO/threshold assertion, not binary)
assertion : routes_to_fallback | succeeds_within_retries | picks_under_tpm | returns_cached
| trips_then_recovers | under_slo
e.g. reliability.fallback.context_window.routes_to_fallback exercised_on=[chat_completions]
reliability.cooldown.429.trips_then_recovers exercised_on=[chat_completions, messages]
Quota Management - behavior features (entity- or config-driven caps and their accounting; endpoint is exercised_on)
quota_management.<behavior>.<variant>.<assertion>
behavior : ratelimit | budget | spend_tracking
variant : <ratelimit> rpm | tpm | priority_generous | priority_strict
<budget> key | internal_user | end_user | organization | team | team_member | tag
| model_max | soft | key_multi_window | team_multi_window
| fallback | spend_counter
<spend_tracking> chat_completions | stream | messages_bridge | embeddings
| cache_hit | key_rollup | concurrent_burst | tags | end_user
| per_model | failure | spend_calculate | pagination
assertion : blocks_over_limit | resets_after_window | headers_report_remaining | picks_under_tpm
| blocks_then_resets | resets_windows_independently | alerts_without_blocking
| isolates_per_model | isolates_per_member | enforced_across_keys | routes_to_fallback
| reseed_matches_db | logs_cost | zero_cost
| matches_sum_of_logs | loses_no_spend | attributes_spend | writes_own_rows
| writes_failure_row | returns_cost | keeps_total
e.g. quota_management.ratelimit.rpm.blocks_over_limit exercised_on=[chat_completions, messages]
quota_management.budget.key.blocks_over_limit exercised_on=[chat_completions]
Logging & Guardrails - behavior features (config-driven; endpoint is exercised_on)
logging.<integration>.<event>.<assertion>
integration : langfuse | s3 | otel | prometheus | datadog | ...
event : success | failure | stream
assertion : logs_spend | writes_object | exports_metric
e.g. logging.langfuse.success.logs_spend exercised_on=[chat_completions]
guardrail.<provider>.<hook_point>.<assertion>
provider : presidio | lakera | bedrock | aporia | ...
hook_point : pre_call | post_call | during | logging_only
assertion : blocks | masks | allows
e.g. guardrail.presidio.pre_call.masks exercised_on=[chat_completions]
Other - holding pen (endpoint or behavior)
other.<area>.<case>.<assertion>
area : auth | lifecycle | config | ...
rule : audited periodically; a cluster here promotes to a new component
e.g. other.auth.jwt.valid_token_allows
other.lifecycle.readiness.reports_db
Hard Rules
-
no monkeypatching or mock tests, and never substitute a unit test for e2e feature coverage: a product feature is proven end to end against a live proxy, not with a unit test. if a contributor asks you to write an end to end test, do NOT stage a unit test of the feature with it; if you find a product gap, call it out in the PR description. tests that cover the harness itself are the exception and are allowed (for example
coverage_registry/test_collector.py, which unit-tests the coverage collector): they carry noe2emarker, exercise harness plumbing rather than a product feature, and run whether or not a proxy is up -
use model management endpoints to create new models for a test. this could be in a conftest / inline for each test. ask the user what they want.
-
do not overengineer a test, i need you to write readable, clean code of what would look like a natural user scenario
-
when it comes to typing an input schema for an api endpoint, have it type X = A | B | C ... where X = exhaustive union of all supported input schemas and A, B, C typically are composed by a base type. types are only pretty for a api request / response body. make sure to compose types instead of repeating the same base attributes over and over again.
-
spin up a local proxy by running the litellm proxy locally (
litellm --config <your-e2e-config>.yml --port 4000; see CONTRIBUTING.md), make sure all tests pass. if a test fails due to an internally found issue, let users know to create a linear ticket for it. -
do not use xfail markers, tests should be written in a form that the end user expects it to pass