Commit graph

38229 commits

Author SHA1 Message Date
mateo-berri
53f71fbf4d tests(vcr): emit per-test verdicts via xdist controller's terminalreporter
Previous attempt wrote to sys.__stderr__ from the test fixture. Under
xdist, fixtures run inside worker subprocesses whose stderr is captured
by the controller and only released to the live log on test failure —
so passing tests' verdicts were silently swallowed.

Round-trip via report.user_properties: the worker-side fixture stashes
the verdict on user_properties, xdist serializes it onto the report,
and a controller-side pytest_runtest_logreport hook writes it via the
TerminalReporter (the same plugin that emits PASSED/FAILED markers).
TerminalReporter is resolved lazily on first hook call because it's
not yet registered when conftest's pytest_configure runs.

Verified locally in both serial and xdist modes.
2026-05-01 14:29:06 -07:00
mateo-berri
965185c106 fix(tests): host PDF fixture via jsDelivr with proper application/pdf MIME
Raw github serves application/octet-stream which OpenAI/Gemini reject
when LiteLLM fetches the URL client-side. jsDelivr serves the same
file with content-type: application/pdf. Pin to a commit SHA so the
asset is immutable and jsDelivr can cache it for a year.
2026-05-01 14:16:57 -07:00
mateo-berri
c05c865a1c tests(vcr): emit verbose verdicts to un-redirected stderr
Previously, the per-test [VCR HIT/MISS/...] line was written via
TerminalReporter.write_line from inside fixture teardown. Pytest
captures that stream by default and only surfaces it on FAILED tests
(under 'Captured stdout teardown'), so passing tests' verdicts were
invisible in CI logs and the user couldn't tell whether the cache
was working.

Write directly to sys.__stderr__ so the line bypasses pytest's
capture entirely. Under xdist each worker has its own __stderr__
which CircleCI aggregates into the live job log alongside the
PASSED/FAILED markers.
2026-05-01 14:14:44 -07:00
mateo-berri
225d01cb4f fix(tests): use github-hosted PDF fixture for Anthropic Files API test
Anthropic's URL fetcher intermittently returns 400 'Unable to download
the file' for the Wikipedia URL the test was using. Point it at the
repo's existing tests/llm_translation/fixtures/dummy.pdf via raw
GitHub instead — small, deterministic, reliably fetchable.

With a stable URL the test no longer needs to be opted out of VCR;
remove it from the incompatible list so it can replay from cassette.
2026-05-01 14:07:32 -07:00
mateo-berri
a47c4e7d1b tests(vcr): refuse to persist cassettes past 50 episodes
A test that produces non-deterministic request bodies (e.g. uuid in
the prompt) under record_mode=new_episodes never replays — every CI
run appends fresh unmatched episodes. The cassette grows unbounded
over time and silently inflates Redis (we observed one cassette at
22 episodes / ~860KB after ~5 CI runs).

Refuse the save when episode count exceeds MAX_EPISODES_PER_CASSETTE
so the pathology surfaces with a loud warning that points to the
opt-out fix instead of festering invisibly.
2026-05-01 13:56:20 -07:00
mateo-berri
8c01b02779 tests(vcr): opt out tests that observe live cross-call provider state
Some tests can't benefit from cassette replay because they assert on
state that only exists in the live provider between two calls (e.g.
prompt-cache propagation, intermittent provider quirks). Marking them
with @pytest.mark.vcr just wastes cycles trying to record cassettes
they will never replay against successfully.

Opt-out by nodeid suffix so subclassed/parametrized variants are
covered:

- ::test_prompt_caching — Anthropic/Bedrock prompt-cache propagation
  isn't deterministic in the 0–1s window the test gives it.
- ::test_async_pdf_handling_with_file_id — flaky upstream Wikipedia
  fetch through the Anthropic Files API.
- TestBedrockInvokeNovaJson::test_json_response_pydantic_obj —
  Bedrock Nova returns tool_call vs JSON nondeterministically (other
  providers' subclasses are healthy).
- ::test_bedrock_converse__streaming_passthrough — Bedrock streaming
  response_cost calc returns None intermittently.

These tests keep their existing @pytest.mark.flaky retry behavior.
2026-05-01 13:50:34 -07:00
mateo-berri
ff63bdb984 tests(vcr): only persist cassette on test pass to avoid poisoning cache
A test that fails (incl. all the failing retries before a passing one)
can otherwise overwrite a known-good cassette with a 'bad luck'
recording. Tests like test_prompt_caching, which assert on provider
state across two calls, can produce a 200 response that semantically
fails the assertion — the 2xx filter doesn't catch this because the
HTTP layer is fine.

- pytest_runtest_makereport hook attaches each phase report to the
  pytest item.
- _vcr_outcome_gate fixture (combining the verbose-mode reporter)
  reads the call-phase outcome at teardown and informs the persister
  via mark_test_outcome_for_cassette before vcrpy's Cassette.__exit__
  triggers save_cassette.
- save_cassette consults the per-key 'did the test pass?' flag and
  short-circuits when False, leaving any prior good recording intact.
- Defaults to passed=True when no marker is present so non-test
  usage of the persister still works.
2026-05-01 13:35:29 -07:00
mateo-berri
cf4c9ede61 tests(vcr): add LITELLM_VCR_VERBOSE per-test hit/miss reporting
Set LITELLM_VCR_VERBOSE=1 to print a one-line cassette verdict per
test (HIT / MISS / PARTIAL / NOOP) showing replay vs new-recording
counts. Useful for local QA to confirm which tests actually exercised
the cache and which fell through to the live provider.
2026-05-01 13:00:20 -07:00
mateo-berri
a7e8189b17 tests(vcr): make redis persister resilient to transient outages
Managed Redis (e.g. Upstash) drops idle TLS connections, which surfaced
in CI as a teardown ERROR on test_gemini_image_size_limit_exceeded:

  redis.exceptions.ConnectionError: EOF occurred in violation of
  protocol (_ssl.c:2427)

Cassette persistence is a cache, not test correctness, so:

- Configure the redis client with Retry(ExponentialBackoff, retries=2)
  on ConnectionError/TimeoutError to absorb single-socket drops.
- Wrap save_cassette so a final failure logs a warning instead of
  failing teardown — the next run re-records.
- Wrap load_cassette so an outage on read becomes a cache miss
  (CassetteNotFoundError) instead of erroring in setup.
2026-05-01 12:53:36 -07:00
mateo-berri
4c69557621 tests(vcr): isolate cassette redis to CASSETTE_REDIS_URL
Stop falling back to REDIS_URL/REDIS_SSL_URL/REDIS_HOST for the VCR
persister. Sharing a Redis with the application cache risks cassettes
being wiped by tests that flush the app Redis.
2026-05-01 12:32:59 -07:00
mateo-berri
67287460e5 tests(vcr): drop redundant comments and docstrings
Some checks failed
Unit Tests: Caching (Redis) / caching-redis (push) Has been cancelled
Unit Tests: Proxy DB Operations / assert-shard-coverage (push) Has been cancelled
Unit Tests: Security / security (push) Has been cancelled
Unit Tests: Proxy DB Operations / auth-checks (push) Has been cancelled
Unit Tests: Proxy DB Operations / budgets (push) Has been cancelled
Unit Tests: Proxy DB Operations / custom-logging (push) Has been cancelled
Unit Tests: Proxy DB Operations / db-and-spend (push) Has been cancelled
Unit Tests: Proxy DB Operations / proxy-runtime (push) Has been cancelled
Unit Tests: Proxy DB Operations / endpoints-and-responses (push) Has been cancelled
Unit Tests: Proxy DB Operations / guardrails-hooks (push) Has been cancelled
Unit Tests: Proxy DB Operations / jwt-and-keys (push) Has been cancelled
Unit Tests: Proxy DB Operations / key-generation (push) Has been cancelled
Unit Tests: Proxy DB Operations / logging-misc (push) Has been cancelled
Unit Tests: Proxy DB Operations / proxy-server-core (push) Has been cancelled
Unit Tests: Proxy DB Operations / schema-migration (push) Has been cancelled
Unit Tests: Proxy DB Operations / proxy-utils (push) Has been cancelled
2026-04-30 18:45:56 -07:00
mateo-berri
687ff32616 tests(vcr): patch vcrpy aiohttp record path instead of forcing httpx transport
vcrpy's aiohttp stub captures response bodies via 'await response.read()',
which drains aiohttp's StreamReader. Downstream consumers of the same
ClientResponse (litellm's AiohttpResponseStream, which iterates
response.content.iter_chunked) then see an empty body and surface as
JSON 'Expecting value: line 1 column 1 (char 0)' errors on every
record-path call.

The previous workaround set litellm.disable_aiohttp_transport=True for
the whole VCR-active session, which made the tests exercise pure httpx
instead of the production aiohttp transport. That hid the production
transport from coverage and surfaced its own bugs (e.g. the Azure
DELETE-with-empty-body case fixed in upstream staging).

Replace the workaround with a targeted monkey-patch that re-feeds the
captured body into the StreamReader via unread_data after vcrpy records
it. Tests now run through the same transport customers do, both on
first record and on replay, for both unary and streaming endpoints.

Verified locally against api.anthropic.com with the production
LiteLLMAiohttpTransport: record path passes (real network, 4.2s),
replay path passes (Redis cache, 1.8s).
2026-04-30 18:43:06 -07:00
mateo-berri
95bce9a72e tests(vcr): assert response shape, not exact bytes, in replay tests
The Anthropic replay tests hardcoded specific token counts and content
strings ('Hello! How can I help you today?', prompt_tokens == 12). On a
fresh CI Redis those values must match a pre-recorded cassette that
doesn't exist, so the first run hits the live API and gets different
real bytes back.

Assert on shape instead: non-empty content, positive token counts,
finish_reason in the known set, and (for streaming) more than one chunk.
The tests still exercise the full transformation pipeline end-to-end and
catch shape regressions; drift in the exact text/token counts is
expected and now tolerated.
2026-04-30 18:35:01 -07:00
mateo-berri
722a1a9f8f Merge remote-tracking branch 'origin/litellm_internal_staging' into litellm_vcr-cassette-llm-tests-af37
# Conflicts:
#	litellm/llms/custom_httpx/llm_http_handler.py
2026-04-30 17:56:02 -07:00
mateo-berri
8bdc46ea74 fix(responses): omit empty JSON body on DELETE response API requests
Azure OpenAI's responses-API DELETE endpoint rejects requests that carry
a JSON body with: "Unexpected body with size 2. This API method does
not accept a request body.". The default LiteLLMAiohttpTransport silently
elides empty-dict bodies on DELETE so this was masked, but the pure-httpx
transport (used when DISABLE_AIOHTTP_TRANSPORT=True or under vcrpy/respx
patching) sends literal '{}' (2 bytes), which Azure rejects.

Only attach json= when the provider's transform actually returned a
non-empty dict; otherwise issue a bodyless DELETE.
2026-04-30 17:55:27 -07:00
Michael-RZ-Berri
e4fb325a3a
Merge pull request #26914 from BerriAI/litellm_googleGenContentHooks
Run pre_call_hook on Google generateContent endpoints
2026-04-30 17:53:38 -07:00
ryan-crabbe-berri
76e43b7bb2
Merge pull request #26949 from BerriAI/litellm_/condescending-hawking-19bdeb
[Fix] Responses API: Omit Empty Body On DELETE
2026-04-30 17:49:30 -07:00
mateo-berri
265a94cd60 tests(vcr): force pure-httpx transport when VCR is active
litellm's default LiteLLMAiohttpTransport routes requests through aiohttp,
which sits below httpx and is invisible to vcrpy's httpx-stub interception.
Under vcrpy + aiohttp, requests reach the real network but responses come
back through the stubbed httpx transport as empty 200s, surfacing as
'Unable to get json response - Expecting value: line 1 column 1 (char 0)'
in providers like Anthropic, Gemini, and any other path that exercises the
aiohttp transport.

Disabling the aiohttp transport when the VCR persister is registered
forces all calls through pure httpx, which vcrpy can record and replay
correctly.
2026-04-30 17:42:17 -07:00
Yuneng Jiang
bd638245e8
[Fix] Responses API: Omit Empty Body On DELETE
The async/sync delete_response_api_handler always passed json=data into
httpx.delete, where data is {} from the transformer. httpx serializes that
to a 2-byte body. The Azure Responses DELETE endpoint now rejects any
request body with code: unexpected_body, breaking
test_basic_openai_responses_delete_endpoint on the llm_responses_api_testing
job. Build the kwargs dict and only set json= when data is truthy.

Add unit tests that patch httpx.delete and assert json/data are not in the
captured kwargs for the Azure DELETE path (sync and async).
2026-04-30 17:39:55 -07:00
yuneng-jiang
326bcd6cec
Merge pull request #26941 from BerriAI/litellm_/stoic-jemison-cbb6cf
[Test] Proxy E2E: Opt In To Client Mock Response For Model Access Tests
2026-04-30 17:35:16 -07:00
mateo-berri
68db1c5e9e tests(vcr): switch to record_mode=new_episodes to avoid partial-cassette poisoning
record_mode='once' refused to add new requests once any cassette
existed in Redis. Combined with filter_non_2xx_response (which drops
non-2xx responses from the saved cassette) and a 24h shared-Redis TTL,
a single transient API failure mid-test left the cassette stuck with
only the leading non-API requests (e.g. the model_prices fetch from
raw.githubusercontent.com), and every subsequent run for the next 24h
errored with 'Can't overwrite existing cassette'.

new_episodes records anything not already present, so partially
populated cassettes recover on the next run instead of poisoning the
suite for a full TTL window.
2026-04-30 17:22:21 -07:00
yuneng-jiang
bdcc23853c
Merge pull request #26835 from stuxf/codex/cli-sso-flow-binding
chore(cli): tighten CLI SSO session flow
2026-04-30 17:10:27 -07:00
yuneng-jiang
15b7386859
Merge pull request #26815 from stuxf/fix/get-image-lfi-ssrf
chore(proxy): contain UI_LOGO_PATH / LITELLM_FAVICON_URL on unauthenticated asset endpoints
2026-04-30 17:10:15 -07:00
yuneng-jiang
71d5015975
Merge pull request #26827 from stuxf/fix/passthrough-auth-default
chore(passthrough): default auth=True and drop enterprise gate on the safe option
2026-04-30 17:06:37 -07:00
Yuneng Jiang
be0e9914dc
[Test] Proxy E2E: Opt In To Client Mock Response For Model Access Tests
The proxy's ingress hardening (commit 842eea0131) now strips client-supplied
`mock_response` from the request body unless the calling key or team has the
`allow_client_mock_response: true` admin-metadata flag set. The e2e model
access tests rely on `mock_response` to short-circuit the LLM call, so without
the flag they hit real backends — the bedrock wildcard route fakes out to a
shared example endpoint that now 404s on unsupported paths, causing
`test_model_access_patterns[key_models2-bedrock/anthropic.claude-3-True]`
(and the bedrock/anthropic.* row that pytest -x never reaches) to fail.

Set `allow_client_mock_response: true` on every key and team this test file
provisions so `mock_response` is preserved end-to-end.
2026-04-30 17:05:31 -07:00
mateo-berri
f6a37a6a15 style: reformat to pass ci 2026-04-30 17:01:44 -07:00
Michael Riad Zaky
053e040171 run pre_call_hook on Google generateContent endpoints 2026-04-30 16:43:42 -07:00
Michael-RZ-Berri
e810d8735d
Merge pull request #26934 from BerriAI/litellm_lazyStartupTestFix
[Fix] Replace subprocess startup-import diff with static source scan
2026-04-30 16:42:52 -07:00
mateo-berri
468b849072
tests(vcr): drop YAML/cassettes-directory metaphor from Redis keys 2026-04-30 23:37:17 +00:00
Michael Riad Zaky
47b2832d6f test: replace subprocess startup-import diff with static source scan 2026-04-30 16:15:46 -07:00
mateo-berri
f55a710e92
tests(vcr): accept REDIS_URL / REDIS_SSL_URL for managed Redis with TLS 2026-04-30 23:08:01 +00:00
mateo-berri
efdeff89d8
fix(llm_request_utils): handle None proxy_server_request without AttributeError 2026-04-30 23:01:10 +00:00
mateo-berri
59d5901766
tests(vcr): allow playback repeats so duplicate intra-test requests serve from cache 2026-04-30 22:57:21 +00:00
mateo-berri
73594262ee
tests(vcr): drop redundant num_retries=3 layer for vcr-marked tests
Provider SDKs already retry transient 5xx/429 with exponential backoff
(default max_retries=2), and pytest.mark.flaky covers test-level
retries on top of that. Setting litellm.num_retries=3 here just
multiplied the existing layers — worst case 6 (flaky) x 3 (this) x
2 (CI rerunfailures) = 36 attempts on a single test.

Removing it keeps SDK-level network-blip protection intact and
shortens worst-case latency on cache-miss runs.
2026-04-30 22:42:30 +00:00
mateo-berri
e1f2b4b818
tests(vcr): trim non-load-bearing comments and docstrings
Removes commentary that restated the code, including:

- module-level banners explaining what the conftest does (covered by
  Readme.md and the function bodies)
- docstrings on _scrub_response, _before_record_response, vcr_config,
  _vcr_disabled, pytest_recording_configure (function names + bodies
  are self-evident)
- inline notes about header filtering, match_on, etc.
- per-test docstrings restating the test name

Keeps the two non-obvious notes that aren't recoverable from the code:
the vcrpy/respx httpx-transport collision rationale on
_RESPX_CONFLICTING_FILES, the vcrpy "return None to skip persisting"
contract on filter_non_2xx_response, and the fixture-ordering
dependency on _vcr_record_retries.
2026-04-30 21:48:48 +00:00
mateo-berri
c7d647b567
tests: drop YAML cassettes, make Redis-backed VCR the default
Removes the YAML cassette feature entirely and replaces it with a
Redis-only flow. Every test in tests/llm_translation/ and
tests/llm_responses_api_testing/ is auto-marked @pytest.mark.vcr via
conftest.pytest_collection_modifyitems, so any provider call lands in
the Redis cache (litellm:vcr:cassette:<rel_path>, 24h TTL). First run
records, runs within the day replay, day rollover re-records and
surfaces upstream API drift within 24h.

VCR is on by default. Set LITELLM_VCR_DISABLE=1, or simply leave
REDIS_HOST unset, to opt out — both bypass the auto-marker entirely so
nothing about cassettes runs. record_mode is "once" so cache-miss
records and cache-hit replays.

The 8 existing respx-using files in tests/llm_translation are excluded
from the auto-marker (vcrpy and respx both patch the httpx transport;
applying both makes one silently win). The persister's own unit-test
file is also excluded so it doesn't recursively run inside a cassette.

The persister moved from tests/llm_translation/_vcr_redis_persister.py
to tests/_vcr_redis_persister.py so both conftests share it. The two
demo tests in test_anthropic_completion_vcr.py were ported into
test_anthropic_completion.py and the demo file was deleted.

Adds tests/_flush_vcr_cache.py + a Make target
(test-llm-translation-flush-vcr-cache) that scans
litellm:vcr:cassette:* and pipelines DELETEs, for the
"I want the next CI run to re-record now" workflow. Drops the now-dead
test-llm-translation-record target.

Provider keys are still required on cache-miss (which happens on first
run and once a day after that). Replay-mode runs need only Redis.
2026-04-30 21:40:58 +00:00
yuneng-jiang
256e05e474
Merge pull request #26849 from stuxf/fix/mcp-oauth-discovery-ssrf
chore(mcp): SSRF guard on OAuth metadata discovery follow-up fetches
2026-04-30 13:44:16 -07:00
yuneng-jiang
174c770b07
Merge pull request #26836 from stuxf/fix/byok-credential-encryption
chore(mcp): encrypt user-scoped MCP credentials at rest
2026-04-30 13:42:57 -07:00
mateo-berri
33a051636d
tests(llm_translation): add Redis cassette persister with 24h TTL
Stores VCR cassettes in Redis under litellm:vcr:cassette:<rel_path> with
a 24h expiry instead of YAML on disk. The TTL means each daily CI run
starts with an aged-out cache, naturally re-records against live providers,
and surfaces upstream API drift within a day without a manual `make`
re-record sweep. Opt-in via LITELLM_VCR_REDIS=1; default behaviour is
unchanged so local dev keeps the on-disk cassettes.

before_record_response now drops non-2xx responses so a transient 5xx or
429 from a provider can't poison the cache for the rest of the TTL window.
Vcr-marked tests bump litellm.num_retries to 3 during recording so
provider-SDK exponential backoff kicks in on the cache-miss path.

Tests cover the three surfaces we depend on in CI: serialize/deserialize
roundtrip via the real vcrpy serializer, TTL is actually applied to saved
keys, cache miss raises CassetteNotFoundError so vcrpy falls through to
record mode, and 2xx-only filtering across the status-code matrix
(2xx kept, 3xx/4xx/5xx dropped, with 429 and 503 explicitly pinned).
2026-04-30 20:28:11 +00:00
yuneng-jiang
4ff8f0e901
Merge pull request #26851 from stuxf/codex/fix-callback-env-secret-resolution
chore(proxy): block env callback refs in key metadata
2026-04-30 13:11:32 -07:00
yuneng-jiang
aa76ab2df7
Merge pull request #26862 from stuxf/codex/control-field-sanitization
chore(proxy): harden request control fields
2026-04-30 13:10:58 -07:00
Michael-RZ-Berri
9637d8c17b
Merge pull request #26802 from BerriAI/litellm_lazyLoadedFrontPage
[Feat / Fix] Lazy loaded imports, lazy loaded front page
2026-04-30 13:04:42 -07:00
yuneng-jiang
08541f49ee
Merge pull request #26910 from BerriAI/litellm_fix/drop-milvus-db-params
fix: drop milvus dbName and partitionNames from MILVUS_OPTIONAL_PARAMS
2026-04-30 12:53:18 -07:00
Yassin Kortam
d84b35cc40
Merge pull request #26906 from BerriAI/litellm_fix/validate-aws-region
fix: validate aws region name
2026-04-30 12:52:31 -07:00
user
51a3e90451 fix(mcp): reuse safe URL fetch for OAuth discovery 2026-04-30 12:42:52 -07:00
yuneng-jiang
3c060364fb
Merge pull request #26840 from stuxf/codex/mcp-oauth-root-visibility
chore(mcp): tighten OAuth root endpoint resolution
2026-04-30 11:59:03 -07:00
yuneng-jiang
a9db887bdd
Merge pull request #26843 from stuxf/codex/fix-onboarding-invite-token
chore(auth): harden invite-link onboarding token flow
2026-04-30 11:56:26 -07:00
Yassin Kortam
dfc080f580 fix: drop milvus dbName and partitionNames from MILVUS_OPTIONAL_PARAMS 2026-04-30 11:51:32 -07:00
yuneng-jiang
0efa8b8828
Merge pull request #26854 from stuxf/fix/team-authz-available-team-bypass
chore(team): close authz bypass via the available-team check
2026-04-30 11:47:19 -07:00
user
b67a81da47 test(proxy): align favicon remote asset expectations 2026-04-30 11:46:45 -07:00