mitmproxy serves its current-instance CA certificate at the magic
host ``mitm.it/cert/pem``. The CI pipeline downloads that CA in the
``Fetch CA from cassette-proxy`` step and trusts it (certifi append +
system trust store) so subsequent in-process tests can validate
mitmproxy's MITM leaf certs.
Once the previous commit fixed the replay path so cache hits actually
serve, the very first ``mitm.it/cert/pem`` lookup hit a stale entry
from a prior CI run — the proxy handed back a CA generated by a *previous*
mitmdump instance with a different private key. The CI then trusted that
stale CA, mitmproxy in the *current* run signed leaf certs with its fresh
CA, and every TLS handshake failed:
SSL: CERTIFICATE_VERIFY_FAILED:
certificate verify failed: authority and subject key identifier mismatch
This took down langfuse_logging_unit_tests (at ``test_embedding.py``
collection), search_testing, and image_gen_testing.
Adding ``mitm.it`` to the passthrough host list short-circuits in
``_should_skip`` before any Redis lookup, so the CA is always served
fresh from the running mitmdump instance. The poisoned key in the
shared dev Redis was deleted manually before pushing this commit.
Adds a regression test that exercises both hooks: ``request()`` must
return without setting ``flow.response``, and ``response()`` must not
persist anything to Redis.
Every cache hit was silently falling through to the upstream because
the addon's replay path called http.Response.make(status, body, list[(str,str)]),
and mitmproxy 11's Headers constructor demands bytes — it raised
``TypeError: Header fields must be bytes.`` inside the addon, mitmdump
logged ``Addon error: Header fields must be bytes.`` to its background
log, and the request continued out to the real provider as if it had
been a miss. (Stats counted it as a hit because the log line was
emitted before the raise.) Net effect: the proxy was record-only.
Three fixes:
1. _build_replay_response constructs http.Response directly,
encoding the cached headers back to bytes (latin-1, RFC 7230) and
handing mitmproxy the raw on-the-wire body. Going through
Response.make/set_content would also have re-encoded the
body (e.g. double-gzip), so we bypass that codepath entirely.
2. The recording side now calls flow.response.headers.items(multi=True)
so repeated headers (notably multiple Set-Cookie) are preserved
as distinct entries instead of being silently merged.
3. Adds a contract test file that runs against *real* mitmproxy
(skipped automatically when only the test stub is loaded). This is
what would have caught the original bug — the existing fake stubs
don't model Headers's bytes-strictness, which is precisely why
the issue hid for the whole record-only run on the previous commit.
The existing addon unit tests now prefer real mitmproxy when it's
installed, so the contract tests run in the same process when both
are available.
Introduces a mitmproxy-based recording HTTP/HTTPS sidecar that any CI
job can opt into to cache LLM-provider responses across runs. Unlike
the in-process VCR persister at tests/_vcr_redis_persister.py — which
can only intercept HTTP traffic from the same Python process where it
was loaded — this sidecar operates at the network layer, so it works
for any e2e job whose system-under-test runs in a Docker container
(every job under e2e_*, proxy_*, etc.).
Components
- tests/e2e_cassette_proxy/cache_key.py: pure-function cache-key
derivation. Hashes (method, scheme, host, path, sorted query,
allowlisted headers, canonical-JSON body); strips auth, tracing,
and SDK-metadata headers so equivalent requests collide regardless
of run-to-run noise.
- tests/e2e_cassette_proxy/redis_store.py: thin Redis wrapper that
stores one (request, response) pair per key as MessagePack
(JSON+base64 fallback). Caps per-key payload size, drops oversize
responses with a log line, and never blocks the request path on
Redis errors.
- tests/e2e_cassette_proxy/addon.py: mitmproxy addon that ties the
two together. Hosts on the passthrough list (localhost, the proxy
itself) are never cached; non-2xx upstream responses are not
persisted.
- tests/e2e_cassette_proxy/Dockerfile: pinned python:3.12-slim base +
pinned mitmproxy 11.0.2 + pinned redis-py + pinned msgpack.
- tests/e2e_cassette_proxy/trust_ca.sh: helper for SUT containers to
trust the proxy CA in every Python / curl / boto3 / node trust
store at once.
- tests/e2e_cassette_proxy/README.md: usage guide + opt-in checklist
for other e2e jobs.
CI integration
- New reusable command 'start_cassette_proxy' in .circleci/config.yml.
Builds the image, runs the sidecar wired to the project Redis, fetches
the proxy CA, and exports CASSETTE_PROXY_URL / CASSETTE_PROXY_CA into
$BASH_ENV for downstream steps.
- e2e_openai_endpoints is wired up as the canonical demo: two-line opt-in
pattern documented in the README.
- Job logs include a 'Cassette-proxy stats' step that dumps the
hit/miss/store summary via 'docker logs cassette-proxy | grep
[E2ECASS]'.
Tests
- 31 hermetic unit tests under tests/test_litellm/e2e_cassette_proxy:
- test_cache_key.py: 14 tests pinning equivalence-class behavior of
the key derivation (auth header, tracing header, JSON key order,
query order, host case all collapse; method/path/body/allowlisted
headers / query-param values do not).
- test_redis_store.py: 9 tests covering set/get round-trip, default
TTL, binary body round-trip, oversize-payload rejection, corrupt-
blob eviction, and graceful behavior when the Redis client raises.
- test_addon.py: 8 tests using a fake-mitmproxy flow to exercise
the addon end-to-end (passthrough miss, persist on 2xx, hit on
canonicalized-equivalent re-request, no-persist on 5xx, host
passthrough, replay-only 599-on-miss, record-only never serves
cache).
- 31/31 pass.
Co-authored-by: Mateo Wang <mateo-berri@users.noreply.github.com>
The async/sync delete_response_api_handler always passed json=data into
httpx.delete, where data is {} from the transformer. httpx serializes that
to a 2-byte body. The Azure Responses DELETE endpoint now rejects any
request body with code: unexpected_body, breaking
test_basic_openai_responses_delete_endpoint on the llm_responses_api_testing
job. Build the kwargs dict and only set json= when data is truthy.
Add unit tests that patch httpx.delete and assert json/data are not in the
captured kwargs for the Azure DELETE path (sync and async).
Cover the full litellm.rerank()/arerank() path with HTTP mocked, asserting
metadata.requester_metadata reaches the Discovery Engine :rank body as
userLabels (and stays absent when no metadata is set). Catches plumbing
regressions that unit tests on transform_rerank_request alone would miss.
Two Greptile P2s addressed:
1. (security) The audit-log row for an ``add_team_callbacks`` call would
serialize the entire ``callback_vars`` block — including
``langfuse_secret_key``, ``langsmith_api_key``, and the GCS service
account path — verbatim into ``LiteLLM_AuditLogs``. Anyone with read
access to the audit table could harvest team callback credentials.
Same risk for ``disable_team_logging`` when the team's existing row
has populated ``callback_settings.callback_vars``.
Add ``_redact_callback_secrets``: deep-copies the metadata snapshot
and replaces every ``callback_vars`` value with ``***REDACTED***``.
The keys are kept so an auditor can still see *which* fields
changed. Applied to both before and after snapshots.
2. ``asyncio.create_task`` is fire-and-forget; if the audit-log write
raises (transient DB error etc.) the exception is silently
discarded by the event loop and the audit row is just missing —
the exact gap this PR is closing. Attach a ``done_callback`` that
logs the exception at warning level via ``verbose_proxy_logger`` so
the operator sees there's a gap.
Tests assert that callback values are not present in the serialized
audit payload (both for ``add_team_callbacks`` and for
``disable_team_logging`` when the team's existing row has
populated secrets).
The two mutating endpoints in team_callback_endpoints.py
(``/team/{id}/callback`` POST and ``/team/{id}/disable_logging``) wrote
team metadata without emitting an audit-log row. The disable variant
is the worst case: a logging-control action that itself isn't logged,
so an admin (or compromised admin) could zero out a team's
observability with no forensic trail.
Add ``_emit_team_callback_audit_log`` mirroring the
``store_audit_logs``-gated pattern already used in team_endpoints.py
for /team/new and /team/update. When ``litellm.store_audit_logs`` is
True, both endpoints now emit an ``LiteLLM_AuditLogs`` row capturing
the calling user, the API key, and the before/after team metadata.
When the flag is False the helper is a no-op, so non-Enterprise
deployments are unaffected.
The ``litellm_changed_by`` header is now also accepted on
``/team/{id}/disable_logging`` to match the existing
``add_team_callbacks`` shape; the header is optional so existing
callers are unaffected.
Variant scope: the file has three endpoints — both mutating variants
are now logged. The read-only ``GET /team/{id}/callback`` is unchanged.
Other unlogged callback / logging-control admin endpoints elsewhere in
the proxy (e.g. ``/cache/settings``, ``/config_overrides/hashicorp_vault``)
are out of scope here and would be addressed in a separate PR.
Tests cover both endpoints in both ``store_audit_logs`` states and
verify that the captured before/after metadata reflects the actual
mutation, plus that the ``litellm-changed-by`` header overrides the
auth user_id when supplied.
Greptile P2: the bypass removal in update_team_member_permissions had
no dedicated regression test. Adds an integration-style test that
posts to /team/permissions_update as a non-admin caller while
``_is_available_team`` is mocked True, and asserts a 403 — pinning
the bypass-removal against future regressions in the same way the
new member-add unit tests pin the self-join enforcement.
Two paths previously treated ``_is_available_team`` as a blanket
authorization bypass — the function was meant to let standard users
self-join a public team but was wired into the broader admin gate
without bounding the action being performed. Three concrete
exposures resulted:
1. ``/team/member_add``: the bypass let an unprivileged caller add
themselves as a Team Admin, or add an arbitrary other ``user_id``
into the team.
2. ``/team/permissions_update``: the same bypass let any authenticated
user overwrite a team's ``team_member_permissions`` array, mutating
the access policy for every member.
3. (Read endpoint ``/team/permissions_list`` is unchanged — it leaks
read-only policy state to non-members but is out of scope of the
advisory's recommendation; tracking separately.)
This commit:
- Splits ``_validate_team_member_add_permissions`` into early-return
admin checks followed by an available-team self-join branch that
enforces ``member.user_id == caller.user_id`` AND
``member.role == "user"`` for every member entry in the request.
The bulk shape (``member: List[Member]``) is checked the same way,
so a list with one valid self-entry plus one ``role=admin`` entry
is rejected. Email-only members are rejected on the self-join
path: matching by ``user_id`` is the only safe primitive at
pre-validation time (resolving email→user_id earlier would let
unauthenticated callers probe user existence).
- Removes the ``_is_available_team`` clause from
``update_team_member_permissions`` entirely. Only proxy / team /
org admins can update permission policies.
Tests:
- Update the two existing ``_validate_team_member_add_permissions``
unit tests to pass the new ``data`` argument.
- Add six regression tests covering the privesc shape (role=admin),
the cross-user-injection shape (other user_id), the no-caller-uid
fail-closed case, the email-only rejection, and the bulk shape.
- ``test_team_endpoints.py`` 133/133 pass.