litellm/tests/integration
yuneng-jiang 9b5562f89b
test: repair stale and polluting tests red on scheduled main CI (#44229)
* test(proxy): stop the proxy_server app fixture leaking LITELLM_LOG

The session app fixture set LITELLM_LOG=ERROR with os.environ.setdefault and never removed it, so later tests on the same xdist worker inherited it. test_drop_params_env_var spawns a subprocess with os.environ and lost the warning it asserts on. Scope the variable to the import with a MonkeyPatch context

* test(secret-detection): give the hand-built redaction request an ASGI path

Since #43975 _read_request_body checks the route path via request.scope, and a scope without path raised KeyError that was swallowed into an empty body, so chat_completion failed with a missing messages parameter. Real ASGI scopes always carry path

* test(integration): isolate litellm callback lists per sdk test

usage-based-routing-v2 Routers register their selector in litellm.callbacks and nothing removes it, not even Router.reset(). The counter TTL and Redis service metrics tests left their selectors behind, and the next usage routing test ran their pre-call checks against its own rpm=1 deployments, raising "Deployment over defined rpm limit". An autouse fixture now gives each sdk test copies of the callback lists and restores the originals afterwards

* test(integration): keep the owner-lookup fault proxy off the shared read replica

The owned proxy points DATABASE_URL at a scratch database but inherited
DATABASE_URL_READ_REPLICA from the replica job, so auth read the shared
database and rejected the freshly created key with token_not_found_in_db.
Drop the replica variable like the other scratch-database owned proxies

* test(integration): request every seeded key in the team owner breakdown

The aggregated team activity endpoint now caps breakdown.api_keys at the top
100 keys by default (#43398), so the 300 seeded keys came back as 100 rows.
The test guarantees each key is reported with its own owner, so ask for an
api_key_limit that covers all seeded keys

* test(integration): give every owned Redis its own port in the redis-cache container

On CircleCI every owned Redis ran on the fixed port 16379 inside the shared
redis-cache container. When an earlier server still held that port, the new
one failed to bind, readiness pinged the old server, the pidfile read failed
and cleanup then reported "Owned Redis still serves after shutdown"

Reserve an ephemeral port for the docker-exec path the same way the local
binary path already does, and refuse to start when something already serves
the chosen port so the failure names the real cause

* test(e2e): skip the Vertex Mistral partner case the e2e project cannot reach

The e2e Vertex project gets a 404 publisher model not found for vertex_ai/mistral-small-2503, so the case can only fail

* test(e2e): skip the Vertex gpt-oss partner case the e2e project never serves

vertex_ai/openai/gpt-oss-120b-maas has hit a 60s read timeout with no response headers on every run in the e2e Vertex project since the case was ported, and no other Vertex partner chat model passes there to switch to

* test(e2e): check only stored message content for a leaked card number

The Presidio spend-log check ran the card-number pattern over the whole serialized response, so a Luhn-valid usage.cost float (0.0003466000000000001) failed the streaming /v1/messages case although the stored content was <CREDIT_CARD>. The check now reads the content and text strings of the stored response, which is where a raw card would land, and still requires the placeholder there

* test(e2e): assert the proxy decodes token-array embeddings for titan

The port in #44120 carried over a legacy SDK-direct test that expected Bedrock to reject token ids with a 400. Through the proxy, /embeddings decodes token arrays to text for providers that cannot embed tokens, so titan answers 200. The test now sends a token array and its decoded sentence and requires the two vectors to match, which fails if the proxy stops decoding or decodes with the wrong tokenizer

* test(e2e): run the Bedrock extended-thinking round trip on a model that honors enabled thinking

us.anthropic.claude-sonnet-5-5 is adaptive-only, so litellm sends thinking.type=enabled with a 1024 budget as adaptive with low effort, and Bedrock returned no reasoning blocks on 5 of 5 identical Converse calls (boto3 direct agreed). us.anthropic.claude-sonnet-4-6 accepts the legacy shape verbatim and returned reasoning on 5 of 5. The non-thinking Bedrock case stays on sonnet-5-5

* test(proxy): stop unit modules forcing DEBUG logging into the event-loop lag tests

Five tests/unit modules set verbose_proxy_logger to DEBUG at import, so every xdist worker that collected them logged the 2.4MB pass-through response from a worker thread, and secret redaction of that line held the GIL for ~0.8s+ inside the timed window. The lag tests now pin the LiteLLM loggers to WARNING and freeze gc while timing, and the module-level DEBUG overrides are removed

* test(e2e): cite the tokenizer and date behind the titan token-array fixture

* test(e2e): let migration seed replicas finish their request-log indexes before cloning

Since #43948 a serving proxy builds the two LiteLLM_SpendLogs indexes on a background thread after it reports ready. The seed fixtures stopped the replica at readiness, so every cloned legacy database lacked an index no real deployment would be missing, and the v2 baseline diff refused it. Seeds now wait until both indexes exist and are valid in the database's schema

* test(passthrough): give the pass-through MockRequest an httpx URL and ASGI scope

#43626 made get_request_route read request.scope during pass-through kwarg setup; the MockRequest in tests/unit/passthrough had neither a scope nor a URL object, so both stream-param tests raised before reaching the code they check. Mirrors the repair #43626 made to the tests/pass_through_unit_tests fake

* test(integration): ignore foreign allow_all_keys MCP servers in the access matrix tool list

test_toolset_gateway_url_serves_a_team_granted_toolset_to_a_key_without_its_own_grant (#43908) registers an allow_all_keys server on the shared gateway, and allow_all_keys servers are listed to every key by design, so a matrix case running on another xdist worker at the same time saw its tools. The matrix now drops tools of allow_all_keys servers it did not create, read from LiteLLM_MCPServerTable before and after listing, and still compares everything else exactly
2026-10-02 21:32:17 +00:00
..
_support test: repair stale and polluting tests red on scheduled main CI (#44229) 2026-10-02 21:32:17 +00:00
authorization feat(jwt): auto_register_map_existing_key maps JWT to the user's existing virtual key (#42375) 2026-10-02 12:11:33 -07:00
compatibility fix(proxy): document request body and response schemas for the Responses API in OpenAPI (#42802) 2026-09-24 01:28:34 +00:00
configuration feat(proxy): add LITELLM_DISABLE_LAZY_ROUTES to register optional routers at startup (#43911) 2026-10-01 22:50:41 +00:00
cost_calculation test(integration): chain a proxy-issued previous_response_id in the cost suite (#42396) 2026-09-21 20:36:53 -07:00
database fix(proxy): always exit when database setup fails at boot (#44141) 2026-10-01 21:51:06 -07:00
management test: repair stale and polluting tests red on scheduled main CI (#44229) 2026-10-02 21:32:17 +00:00
mcp test: repair stale and polluting tests red on scheduled main CI (#44229) 2026-10-02 21:32:17 +00:00
messages_endpoint test(anthropic): native /v1/messages reasoning integration tests built on a captured Claude Code request (#43361) 2026-10-01 09:12:53 -07:00
observability fix(otel): tolerate non-dict callback_settings.otel and ignore bare EXCLUDED_SERVICES env (#44086) 2026-10-02 12:36:46 -07:00
pricing fix(router): bill service tiers at catalog rates for custom-priced deployments (#43890) 2026-09-30 16:06:37 -07:00
providers test(integration): move legacy proxy, router and Redis tests into tests/integration (#44128) 2026-10-01 23:00:54 -07:00
routing refactor(types): replace Any with proven types in 7 files (#43844) 2026-10-02 02:47:47 -07:00
sandbox test: finish the non-proxy half of tests/test_litellm (#43281) 2026-09-25 22:43:41 -07:00
sdk test: repair stale and polluting tests red on scheduled main CI (#44229) 2026-10-02 21:32:17 +00:00
security test(ci): repair stale tests and move retired OpenAI text-completion fixtures (#43958) 2026-09-30 19:19:59 -07:00
spend test: repair stale and polluting tests red on scheduled main CI (#44229) 2026-10-02 21:32:17 +00:00
streaming test(integration): regression tests for July provider translation, routing and streaming bugs (#42693) 2026-09-23 09:51:06 -07:00
__init__.py test: add CircleCI integration contract foundation 2026-09-14 03:30:52 -07:00
AGENTS.md test(integration): add MCP gateway coverage wave 1 with a dedicated mcp shard and proxy coverage artifact (#42711) 2026-09-23 07:48:46 -07:00
conftest.py ci: cut CircleCI wall time without loosening test isolation (#43347) 2026-09-26 15:34:53 -07:00
coordination_redis_proxy_config.yaml fix(proxy): publish auth cache invalidations in the background so a wedged coordination Redis cannot stall user updates (#42534) 2026-09-22 23:33:32 -07:00
mcp_coverage.toml test(integration): add MCP gateway coverage wave 1 with a dedicated mcp shard and proxy coverage artifact (#42711) 2026-09-23 07:48:46 -07:00
oci_proxy_test_config.yaml CI: copy of #25177 (OCI GenAI: embeddings, streaming/reasoning fixes, model catalog) (#28223) 2026-05-23 12:15:41 -07:00
proxy_config.yaml test(integration): regression tests for August cost tracking and budgeting bugs (#42622) 2026-09-23 04:13:03 +00:00
README.md ci(circleci): test Redis behavior against local Redis and print short tracebacks (#44062) 2026-10-01 13:00:54 -07:00
run.py ci(circleci): test Redis behavior against local Redis and print short tracebacks (#44062) 2026-10-01 13:00:54 -07:00
test_oci_integration.py test: drop the cwd-relative sys.path.insert calls from the test suite (#37802) 2026-08-22 09:25:58 -07:00
test_oci_proxy_integration.py feat: litellm oss 110626 (#30202) 2026-06-11 22:30:26 -07:00

Integration contracts

These tests exercise a running gateway, PostgreSQL and Redis with an owned local upstream. CircleCI owns this suite. Tests are grouped by behavior, with no automatic test retries or fallback to paid provider calls

The cost group is driven by cost_tracking_cases.json, which contains the cost map, literal requests, literal provider responses and expected accounting values. Each case has a name, contract ID, cost-map model, optional deployment overrides, request body, tagged response and exact or recount expectations. Request bodies use $MODEL for the registered proxy model, while responses use $REQUEST_ID for the per-run scenario ID. To add a case, add a cost-map entry when the model is new, add the request body and exact provider response data, and add hand-computed expected values. The upstream serves each stored response for any path under /<scenario_id>, while the test-owned cost map is served over loopback through LITELLM_MODEL_COST_MAP_URL

Use tests/integration/run.py management, accounting, database, providers, extensions, mcp, sdk or cost to run a selected group. The group to directory mapping is the GROUPS literal at the top of run.py; a new directory needs a GROUPS entry and an OWNED_DIRECTORIES entry in _support/manifest.py. Set INTEGRATION_WORKERS above 1 to run a group under pytest-xdist; the mcp job does this in CI, so MCP tests must own their resources per scenario. Set INTEGRATION_PROXY_URL, INTEGRATION_UPSTREAM_URL, INTEGRATION_MASTER_KEY and DATABASE_URL to an isolated test deployment. The runner selects the new domain directories explicitly; the legacy OCI and sandbox selections remain separate

Management also requires INTEGRATION_PEER_URL, REDIS_HOST and REDIS_PORT. CircleCI starts two directly addressed proxy processes sharing only that job's stores. The test-only CLI wrapper supplies enterprise route entitlement, following the existing behavior suite's convention. It does not qualify license validation; run it with one worker and no reload

The generated lifecycle models use 20 examples, eight steps, generation and shrinking, with isolated resources per example. HTTP operation caps include generation and shrinking and exempt cleanup. Local qualification defaults to seed 4106601 and canonical order; CircleCI derives exploration and ordering seeds from the checked-out revision and workflow ID. The ordering seed shuffles the file order and the test order inside each file but keeps each file's tests together, so module fixtures are built once per file. Use --seed and --order-seed to reproduce a run. Actual installed Hypothesis version, settings, seeds and collected order are written beside the execution manifest

Reuse the existing canned provider handlers through _support/upstream.py. It rejects internal request fields and exposes actual received requests for independent assertions. Register every created resource for cleanup immediately, keep expected values independent of production calculations, and assert readback plus the runtime effect of a change

The CircleCI workflow starts its own database and Redis, restricts test-phase egress to its owned services and writes JUnit plus an executed-node manifest. Missing setup, failed cleanup or a selected test with neither a passed call nor a skip fail qualification. Skipped nodes are listed under skipped in execution.json, so the skip reasons double as the open bug list. Existing GitHub Actions jobs do not own these tests

There is no per-node manifest. The runner fails only when pytest fails, when collection errors, or when a selected file collects zero tests. Older tests still carry @pytest.mark.covers(...) decorators; the marker stays registered so they collect, but the IDs are not checked against anything and new tests should not use it. The GitHub Actions coverage census reads the GROUPS literal in run.py and treats every tests/integration/<directory>/test_*.py file in a scheduled group as owned by CircleCI

Provider sentinels currently use the controlled server, not live recordings. The provider shard also runs the existing strict replay controls for changed requests, exhausted interactions, leftover interactions and no provider connection. Future recorded scenarios must use that replay-only implementation; missing recordings cannot fall back to a real provider. The observation endpoint is destructive and the current selection runs serially against one owned upstream

Fixtures must contain synthetic data only. Keep private incident records and source documents out of code, fixtures, logs and PR descriptions

Database cases own their temporary schemas, roles, constraints and proxy processes. They prove reader-versus-writer execution with PostgreSQL lock observations, exercise real transaction wait limits and verify rollback after a reached database failure

Accounting cases compare persisted input and output cost components against literal rates, including zero and default prices. Cache state models assert actual upstream calls, response identity and every persisted charge. Generated accounting tests have a 180-second test limit to accommodate the asynchronous spend writer

Provider contracts exercise actual TCP requests with synthetic credentials and local protocol peers. The S3 verifier uses independently implemented equations, a published known-answer vector, a fixed signing clock and deliberately invalid signed requests. Bedrock cases clear ambient AWS credential sources and check the literal model path, loaded role references, STS requests and bearer-only behavior

Streaming checks send real HTTP transfer chunks, including one-byte partitions, fragmented tools, incomplete transfers and a cancellation barrier. They assert meaningful text, tool arguments, final usage and persisted cost. The Redis recovery case owns a separate database and Redis process, uses the supported one-second circuit-breaker recovery setting, waits for the real subscriber and verifies response data in Redis after restart. CircleCI reuses its existing Redis image for that extra process; it never pulls an image during tests

The messages_endpoint/ directory holds /v1/messages endpoint contracts: native-provider backends under providers/ (anthropic, bedrock, gemini) and the translation bridges (responses_bridge, chat_bridge) at the top level. It runs in the providers shard; run.py selects test files recursively under each scheduled directory

The sdk shard exercises the SDK's own HTTP clients against local protocol peers with no gateway in the path, so a case here fails only when the client library or its wire behavior changes. The HTTP/2 case runs a hypercorn TLS peer offering h2 and http/1.1 over ALPN, drives the sync and async httpx handlers at it with LITELLM_HTTP2 off and on, and asserts the version both the client and the peer observed on the wire. Put a test here only when it needs no proxy or database. CircleCI starts a local Redis for this shard like the others, so SDK-side caching cases that need a real Redis server belong here too; a case that reaches the gateway belongs in one of the other shards

The extensions shard uses the built-in generic callback and guardrail transports. It checks callback correlation and credential exclusion, guardrail rewriting and denial, retained OpenAI consumers and A2A wire versions. CircleCI runs it on parallel nodes, and each node starts its own database, Redis, upstream and proxy and runs its share of the group's files serially, split by recorded timings with circleci tests split. Tests keep the isolation of a serial run; they still must not assume a particular set of sibling files. run.py <group> --list prints a group's files and run.py <group> <file>... runs a subset of them

The mcp shard runs the MCP gateway against SDK peers owned by each test (_support/mcp.py): streamable HTTP, SSE and stdio peers, an OpenAPI-spec app, and an OAuth 2.1 authorization-server double. Every peer records the requests it receives so a test can assert what reached the peer, not only what the proxy answered. The shard runs with INTEGRATION_WORKERS set and with INTEGRATION_COVERAGE=1, which starts the proxy under coverage run --parallel-mode limited to the MCP modules and stores coverage.txt plus an HTML report with the job artifacts. A test that fails because the product is wrong is skipped with pytest.skip("BUG: <symptom>") so the skip list in execution.json is the open MCP bug list

Browser contracts live in tests/e2e/ui/tests/integrationCritical and run only through tests/e2e/ui/integration.config.ts. The expected browser results are listed in expected.json in that directory and checked by .circleci/scripts/verify_integration_browser.py. The CircleCI browser shard builds the checked-out dashboard, starts the owned proxy with that build, and verifies one exact browser result without retries or skips. The default Playwright selection excludes this directory. The focused project flow asserts the submitted create and clear values, fresh SQL state and actual blocked/restored serving while preserving model restrictions

Two always-on -replica CircleCI jobs (management, database) run their groups in replica mode, where every proxy connects through a real litellm_writer role and a real read-only litellm_reader role against the same PostgreSQL. Nothing is captured there: the job passes when the tests pass, and a write routed to the read-only reader fails the test that issued it. A deeper check runs on demand as the routing_parity workflow, triggered through the CircleCI API v2 pipeline endpoint on the PR branch with {"parameters": {"routing_parity_base": "<40-hex merge-base sha>"}}. The workflow fans out over the seven groups, and each routing-parity-<group> job runs its own group twice against the same test harness, once with litellm/, enterprise/, and litellm-proxy-extras/ checked out from the base revision and once from the head, with a pytest plugin snapshotting pg_stat_statements into routing-observed.json per side. The check step then compares the two observations and writes routing-diff.txt: a statement seen on both sides fails when its role set changed, globally or for the same test (per-test capture is skipped under xdist), unless it is listed in tests/integration/routing/either_role.json, where each entry names the statement and a one-line reason it legitimately runs on whichever role asks for it, printed under == either role ==. Queries seen on only one side are listed, never failed, pg_stat_statements evictions and a role that never ran a statement are failures