Two Greptile P2s addressed:
1. (security) The audit-log row for an ``add_team_callbacks`` call would
serialize the entire ``callback_vars`` block — including
``langfuse_secret_key``, ``langsmith_api_key``, and the GCS service
account path — verbatim into ``LiteLLM_AuditLogs``. Anyone with read
access to the audit table could harvest team callback credentials.
Same risk for ``disable_team_logging`` when the team's existing row
has populated ``callback_settings.callback_vars``.
Add ``_redact_callback_secrets``: deep-copies the metadata snapshot
and replaces every ``callback_vars`` value with ``***REDACTED***``.
The keys are kept so an auditor can still see *which* fields
changed. Applied to both before and after snapshots.
2. ``asyncio.create_task`` is fire-and-forget; if the audit-log write
raises (transient DB error etc.) the exception is silently
discarded by the event loop and the audit row is just missing —
the exact gap this PR is closing. Attach a ``done_callback`` that
logs the exception at warning level via ``verbose_proxy_logger`` so
the operator sees there's a gap.
Tests assert that callback values are not present in the serialized
audit payload (both for ``add_team_callbacks`` and for
``disable_team_logging`` when the team's existing row has
populated secrets).
The two mutating endpoints in team_callback_endpoints.py
(``/team/{id}/callback`` POST and ``/team/{id}/disable_logging``) wrote
team metadata without emitting an audit-log row. The disable variant
is the worst case: a logging-control action that itself isn't logged,
so an admin (or compromised admin) could zero out a team's
observability with no forensic trail.
Add ``_emit_team_callback_audit_log`` mirroring the
``store_audit_logs``-gated pattern already used in team_endpoints.py
for /team/new and /team/update. When ``litellm.store_audit_logs`` is
True, both endpoints now emit an ``LiteLLM_AuditLogs`` row capturing
the calling user, the API key, and the before/after team metadata.
When the flag is False the helper is a no-op, so non-Enterprise
deployments are unaffected.
The ``litellm_changed_by`` header is now also accepted on
``/team/{id}/disable_logging`` to match the existing
``add_team_callbacks`` shape; the header is optional so existing
callers are unaffected.
Variant scope: the file has three endpoints — both mutating variants
are now logged. The read-only ``GET /team/{id}/callback`` is unchanged.
Other unlogged callback / logging-control admin endpoints elsewhere in
the proxy (e.g. ``/cache/settings``, ``/config_overrides/hashicorp_vault``)
are out of scope here and would be addressed in a separate PR.
Tests cover both endpoints in both ``store_audit_logs`` states and
verify that the captured before/after metadata reflects the actual
mutation, plus that the ``litellm-changed-by`` header overrides the
auth user_id when supplied.
Greptile P2: the bypass removal in update_team_member_permissions had
no dedicated regression test. Adds an integration-style test that
posts to /team/permissions_update as a non-admin caller while
``_is_available_team`` is mocked True, and asserts a 403 — pinning
the bypass-removal against future regressions in the same way the
new member-add unit tests pin the self-join enforcement.
Two paths previously treated ``_is_available_team`` as a blanket
authorization bypass — the function was meant to let standard users
self-join a public team but was wired into the broader admin gate
without bounding the action being performed. Three concrete
exposures resulted:
1. ``/team/member_add``: the bypass let an unprivileged caller add
themselves as a Team Admin, or add an arbitrary other ``user_id``
into the team.
2. ``/team/permissions_update``: the same bypass let any authenticated
user overwrite a team's ``team_member_permissions`` array, mutating
the access policy for every member.
3. (Read endpoint ``/team/permissions_list`` is unchanged — it leaks
read-only policy state to non-members but is out of scope of the
advisory's recommendation; tracking separately.)
This commit:
- Splits ``_validate_team_member_add_permissions`` into early-return
admin checks followed by an available-team self-join branch that
enforces ``member.user_id == caller.user_id`` AND
``member.role == "user"`` for every member entry in the request.
The bulk shape (``member: List[Member]``) is checked the same way,
so a list with one valid self-entry plus one ``role=admin`` entry
is rejected. Email-only members are rejected on the self-join
path: matching by ``user_id`` is the only safe primitive at
pre-validation time (resolving email→user_id earlier would let
unauthenticated callers probe user existence).
- Removes the ``_is_available_team`` clause from
``update_team_member_permissions`` entirely. Only proxy / team /
org admins can update permission policies.
Tests:
- Update the two existing ``_validate_team_member_add_permissions``
unit tests to pass the new ``data`` argument.
- Add six regression tests covering the privesc shape (role=admin),
the cross-user-injection shape (other user_id), the no-caller-uid
fail-closed case, the email-only rejection, and the bulk shape.
- ``test_team_endpoints.py`` 133/133 pass.
PR #26484 substitutes LITELLM_PROXY_MASTER_KEY_ALIAS for
hash_token(master_key) in UserAPIKeyAuth so the master key (or its
hash) never reaches spend logs / metrics. The otel prometheus tests
still hardcoded the SHA-256 of "sk-1234"
("88dc28d0f030c55ed4ab77ed8faf098196cb1c05df778539800c9f1243fe6b4b"),
so the metric labels no longer matched and test_proxy_failure_metrics
failed. Reference the alias constant directly.
https://claude.ai/code/session_01UkzyZKiADEkZDbZFwB98yV
Co-authored-by: Mateo Wang <mateo-berri@users.noreply.github.com>
npm's `min-release-age` config has type `[null, Number]`. The value `3d`
parses to NaN, which propagates into `before = new Date(NaN)` (Invalid
Date). Pacote then calls `.toISOString()` on it and throws
`RangeError: Invalid time value`, breaking every local `npm install`.
Drop the `d` suffix in all six `.npmrc` files. The `<days>` in npm's
type hint is a label, not part of the value.
This is a no-op for CI (`npm ci` ignores this setting per the comment
in the file) but unblocks local `npm install`.
Address Greptile feedback: the bare `except Exception: pass` in the
finally blocks of _sync_streaming / _async_streaming silently dropped
errors from executor.submit() / asyncio.create_task() (e.g. saturated
thread pool, closed event loop). Since the entire point of the fix is
that spend tracking should not silently lose data, mirror the peer
streaming_handler.py logging pattern so any scheduling failure is
diagnosable in production.
Co-authored-by: Mateo Wang <mateo-berri@users.noreply.github.com>
Strip module-level docstrings and per-test/per-block prose from the
LIT-2642 fix and tests. Keep one short comment in each streaming site
that flags the GeneratorExit-vs-Exception subtlety, since that's the
non-obvious reason the flush lives in finally rather than after the loop.
Pure cleanup; no behavior change. All 12 regression tests still pass.
Co-authored-by: Mateo Wang <mateo-berri@users.noreply.github.com>
When a client disconnects mid-stream from a Bedrock pass-through endpoint,
Starlette calls aclose() on the async generator, raising GeneratorExit
(a BaseException, not Exception) at the suspended yield. The previous
`except Exception` blocks in _async_streaming/_sync_streaming
(litellm/passthrough/main.py) and PassThroughStreamingHandler.chunk_processor
did not catch GeneratorExit, so the post-loop flush that hands collected
raw bytes to async_flush_passthrough_collected_chunks /
_route_streaming_logging_to_handler never ran. All per-chunk usage data
was silently dropped, undercounting spend for interrupted Bedrock invoke
and converse streams.
Move the flush into a finally block in all three sites and guard with a
`flush_scheduled` flag so the success path still flushes exactly once.
Also pull raise_for_status() out of the chunk-collection try block in
_async_streaming so 4xx/5xx responses still raise and don't enter the
flush path with zero bytes (preserving the behavior tested by
test_async_streaming_error_propagation.py).
Add regression coverage:
- test_async_streaming_flushes_on_client_disconnect
- test_async_streaming_flushes_on_upstream_exception_with_partial_data
- test_sync_streaming_flushes_on_early_close
- test_chunk_processor_logs_on_client_disconnect
plus baseline tests for normal completion and the 4xx no-flush path.
Fixes LIT-2642.
Co-authored-by: Mateo Wang <mateo-berri@users.noreply.github.com>
After dedaf74a5e, _async_retrieve_batch wraps the GET in async_safe_get,
which inspects response.is_redirect. test_avertex_batch_prediction's
MagicMock response left is_redirect unset, so it auto-generated a truthy
mock, sent the redirect-follow loop into _extract_redirect_url, and
httpx.URL().join(<MagicMock>) raised TypeError. Set is_redirect=False
so the response is treated as terminal.
* chore(auth): validate clientside api_base against SSRF guard; clear admin secrets on base override
Two related issues with how the proxy handles client-supplied
``api_base`` / ``base_url`` overrides on chat-completion requests:
1. **SSRF gate bypass** — ``check_complete_credentials()`` returned
``True`` for any non-empty ``api_key``, allowing the
``is_request_body_safe`` ``banned_params`` loop to admit ``api_base``
/ ``base_url`` values that point at private (RFC 1918), loopback,
link-local, or cloud-metadata addresses. Now: when the gate sees a
client-supplied ``api_base`` / ``base_url``, it runs the URL through
``litellm_core_utils.url_utils.validate_url`` (DNS-resolves, blocks
internal/IMDS/LL networks, defends against rebinding). Rejection
raises with a clear message.
2. **Admin-config leak on base override** —
``get_dynamic_litellm_params`` only carried the three clientside keys
(``api_key``, ``api_base``, ``base_url``) from request to upstream
call. Other admin-configured fields on ``litellm_params`` —
``organization``, ``extra_body``, ``extra_headers``, ``api_version``,
``azure_ad_token``, AWS / Vertex creds, etc. — flowed through
unchanged. With base redirected to a client-controlled server, those
admin secrets were sent to the attacker. Now: when ``api_base`` /
``base_url`` is in ``request_kwargs``, drop those admin-config
fields from ``litellm_params`` unless the caller re-supplied them.
Tests cover the SSRF-target rejection per URL field, the admin-secret
clearing on base override, the don't-clear case when only ``api_key``
is overridden (BYOK pattern), and the don't-overwrite case when the
caller resupplies fields like ``organization`` themselves.
Co-Authored-By: Claude Opus 4.7 (1M context) <noreply@anthropic.com>
* chore(vertex-batches): wrap api_base GET in safe_get for defense-in-depth
The vertex batches status-poll fetches an attacker-influenceable
``api_base`` URL with a raw ``sync_handler.get()``. The proxy auth gate
already validates clientside ``api_base`` before reaching this sink, so
the proxy flow is covered. This adds the per-sink wrap so SDK callers
and any future code path that bypasses the proxy gate pick up the same
SSRF defense from ``url_utils.safe_get``.
Operators with a legitimate private Vertex base can either allowlist
the host via ``litellm.user_url_allowed_hosts`` or disable validation
with ``litellm.user_url_validation = False``.
Co-Authored-By: Claude Opus 4.7 (1M context) <noreply@anthropic.com>
* refactor(auth): hoist url_utils import; derive admin-config field list from CredentialLiteLLMParams
/simplify pass:
- Move ``from litellm.litellm_core_utils.url_utils import SSRFError, validate_url``
to module top in ``proxy/auth/auth_utils.py``. CLAUDE.md prefers
module-level imports unless avoiding a circular dependency, and
there's no cycle here (``url_utils`` doesn't depend on ``proxy.auth``).
- Replace the hardcoded ``_ADMIN_CONFIG_FIELDS_TO_CLEAR_ON_BASE_OVERRIDE``
literal with ``_admin_config_fields_to_clear_on_base_override()`` that
derives the typed-field portion from
``CredentialLiteLLMParams.model_fields``. Adds three fields the
hardcoded list missed (``aws_bedrock_runtime_endpoint``,
``watsonx_region_name``, ``region_name``) and stays in sync as new
provider fields are declared on the model. The kwargs-only set
(``organization``, ``extra_body``, ``azure_ad_token``, ``aws_session_token``,
``aws_sts_endpoint``, ``aws_web_identity_token``, ``aws_role_name``, …)
remains explicit since those fields aren't on the typed model.
Co-Authored-By: Claude Opus 4.7 (1M context) <noreply@anthropic.com>
* fix(auth): close field-echo bypass; gate URL check on toggle; cover async batch path
Three issues from review:
1. ``get_dynamic_litellm_params`` used ``if field not in request_kwargs:
pop`` to clear admin-set provider config when the caller redirected
``api_base``. A caller could *echo* any clear-list field name (with any
value, including an empty string) to skip the pop, leaving the admin's
value in ``litellm_params`` to be forwarded to the redirected upstream.
Fix: always pop, then write the caller's value back if they resupplied
the field.
2. ``check_complete_credentials`` called ``validate_url`` directly. That
helper doesn't itself consult ``litellm.user_url_validation``; the
toggle is honoured by ``safe_get`` / ``async_safe_get``. Mirror that
here so admins who explicitly disabled URL validation aren't blocked
at the proxy boundary.
3. ``VertexAIBatchesHandler._async_retrieve_batch`` still used a bare
``await client.get(api_base, ...)`` while the sync sibling was wrapped
in ``safe_get``. Wrap the async call in ``async_safe_get`` so SDK
callers on the async path get the same DNS-rebind / private /
cloud-metadata defenses as the sync path.
Tests:
- ``TestCheckCompleteCredentialsBlocksSSRF`` is now mock-only; an autouse
fixture flips the toggle on, ``validate_url`` is patched in the
parametrized blocking tests, and the positive path no longer makes a
real DNS call to api.openai.com.
- ``test_skips_url_validation_when_toggle_is_off`` documents the new
toggle-off behaviour and asserts ``validate_url`` is not called.
- ``test_caller_resupplied_value_overrides_admin_value_on_base_override``
replaces the prior test that asserted the buggy
preserve-admin-value-on-echo behaviour.
- ``test_field_echo_does_not_preserve_admin_value`` is a focused
regression test for the empty-string echo vector.
Co-Authored-By: Claude Opus 4.7 (1M context) <noreply@anthropic.com>
* fix(auth): close provider-confusion credential exfil; expand banned-params; cover OCI
Three additions on top of the entry-point URL gate so the cluster is
fully closed against caller-supplied ``api_base`` redirection:
1. ``get_llm_provider_logic.py`` matched registered openai-compatible
endpoints against ``api_base`` with an unanchored substring search
(``if endpoint in api_base:``). A caller could pass an api_base like
``https://attacker.com/api.groq.com/openai/v1`` to coerce the proxy
into reading ``GROQ_API_KEY`` from the environment and forwarding it
as a Bearer credential to the attacker's host. Replaced with parsed-
URL semantics (hostname exact-match plus segment-bounded path-prefix)
in a new ``_endpoint_matches_api_base`` helper.
2. ``is_request_body_safe`` rejects ``api_base`` / ``base_url`` /
``user_config`` / a handful of AWS / vertex fields, but the list
omitted three other endpoint-targeting fields:
* ``aws_bedrock_runtime_endpoint`` — Bedrock endpoint redirect
* ``langsmith_base_url`` / ``langfuse_host`` — observability callback
hostnames; attacker-controlled values exfiltrate the entire request
payload (incl. message content) via the logging hook.
Added all three to the blocklist.
3. ``_admin_config_fields_to_clear_on_base_override`` derives its typed-
field list from ``CredentialLiteLLMParams.model_fields``, which does
not declare any of the OCI provider's auth fields. Added
``oci_signer``, ``oci_user``, ``oci_fingerprint``, ``oci_tenancy``,
``oci_key``, and ``oci_key_file`` to the kwargs-only fixed list so
they are cleared on caller-redirected ``api_base`` like the AWS /
Azure / Vertex equivalents.
Tests:
- ``TestEndpointMatchesApiBase`` — direct unit tests on the new
matcher: legitimate provider URLs (5 shapes) match; attacker
smuggling via path injection, suffix label, prefix label, userinfo
``@`` injection, and path-segment lookalikes (7 shapes) do not.
- ``TestGetLlmProviderRejectsAttackerSmuggledApiBase`` — end-to-end
invariant that ``GROQ_API_KEY`` is never read against an attacker-
controlled host while the legitimate ``api.groq.com`` path still
resolves the provider correctly.
- ``TestIsRequestBodySafeBlocksEndpointTargetingFields`` — parametrized
coverage that each of the three new banned-params raises a clear
rejection naming the offending field.
Co-Authored-By: Claude Opus 4.7 (1M context) <noreply@anthropic.com>
* fix(auth): remove implicit api-key bypass + add posthog/braintrust/slack to blocklist
The historical ``check_complete_credentials`` clause inside
``is_request_body_safe`` was a third, *implicit*, *caller-controlled*
BYOK path: any caller that supplied a non-empty ``api_key`` caused the
entire banned-params blocklist to be skipped. That turned every missing
entry on the blocklist into an exploitable SSRF / credential-exfil hole
and is the root cause of the chain of api_base advisories that have
been re-discovered with each new integration:
* GHSA-jh89-88fc-qrfp (critical, triage) — env-var exfil via api_base
* GHSA-3frq-6r6h-7j64 (high, triage) — admin org / extra_body leak
* veria-admin Dv_m860l, b_yRJeQ5, stN90yjP, LBlyOAc8, U2TD78kg —
variations on "list X is missing field Y"
Two explicit, admin-controlled BYOK paths already exist and remain:
``general_settings.allow_client_side_credentials = true`` (proxy-wide)
and ``configurable_clientside_auth_params: [...]`` per deployment.
Removing the implicit bypass converts the failure mode of a missing
blocklist entry from "live credential leak" to "predictable 400 with
a clear remediation message," which is the structural fix.
Also adds the three remaining endpoint-targeting fields the dynamic
callback layer reads from request body: ``posthog_host``,
``braintrust_host``, ``slack_webhook_url``. ``slack_webhook_url`` in
particular was a direct exfil channel (caller-set webhook → proxy
mirrors every request to attacker's Slack).
Tests:
- ``test_api_key_does_not_bypass_blocklist`` — parametrized regression
asserting api_key=anything no longer skips the gate for any of the
five highest-risk fields.
- ``test_admin_opt_in_proxy_wide_still_allows`` — confirms the
documented BYOK opt-in still works.
- Extends ``test_endpoint_targeting_field_in_request_body_is_rejected``
to cover posthog / braintrust / slack.
Co-Authored-By: Claude Opus 4.7 (1M context) <noreply@anthropic.com>
* fix(auth): block sagemaker_base_url, s3_endpoint_url, deployment_url
Provider-specific endpoint overrides surfaced by a wider audit of
``optional_params`` consumers in ``litellm/llms/``. Same threat as
``api_base``: a caller-supplied value redirects the outbound request
to an attacker host.
* ``s3_endpoint_url`` — read in ``litellm/llms/bedrock/files/transformation.py``
to build the S3 upload URL for Bedrock files. Caller redirects file
uploads to attacker-controlled S3.
* ``sagemaker_base_url`` — read in ``litellm/llms/sagemaker/{chat,completion}/*``.
Caller redirects SageMaker traffic. This is the primary vector
described in veria-admin mNqEBBtG.
* ``deployment_url`` — popped in ``litellm/llms/sap/chat/transformation.py``.
Caller redirects SAP deployment requests.
Tests parametrize ``test_endpoint_targeting_field_in_request_body_is_rejected``
to cover the three new fields.
Co-Authored-By: Claude Opus 4.7 (1M context) <noreply@anthropic.com>
---------
Co-authored-by: Claude Opus 4.7 (1M context) <noreply@anthropic.com>
* feat(schema): add workflow run tracking tables (LiteLLM_WorkflowRun, LiteLLM_WorkflowEvent, LiteLLM_WorkflowMessage)
* feat(proxy): add /v1/workflows/runs endpoints for durable agent workflow tracking
* feat(proxy): register workflow management router in proxy_server
* docs(workflows): add README for workflow run tracking API
* test(workflows): add unit tests for /v1/workflows/runs endpoints
* fix(workflows): atomic event+status update via tx(), run_id 404 guard, sequence retry on collision
* test(workflows): add tx mock, 404 on unknown run_id, retry-on-collision tests
* fix(workflows): constrain status to Literal enum, rename total→count in list responses
* add tenant isolation and bounded limits to workflow endpoints
* add created_by column and index to LiteLLM_WorkflowRun
* add ownership and bounded-limit tests for workflow endpoints
* Fix workflow run ownership for null owners
* guard prisma import in workflow_management_endpoints
* sync schema.prisma copies with workflow run models
* black: format workflow_management_endpoints.py
---------
Co-authored-by: Cursor Agent <cursoragent@cursor.com>
Read user_id and team_id from the request's litellm_params metadata when
fabricating the UserAPIKeyAuth handed to the managed_files hook, so
batches created via passthrough are attributed to the real requester
instead of a hardcoded fallback. Adds parametrized regression coverage
for both the populated-metadata and empty-kwargs cases.
Drop prompt_variables and client_messages from the re-raised error so
callers cannot leak secrets, tokens, or PII embedded in those payloads
through HTTP error responses. Both sync and async variants.
Remove parameters that may contain credentials from the messages built
inside broad except handlers. These messages can surface in HTTP error
responses, so caller-supplied secrets and integration tokens shouldn't
be interpolated into them.
Use key-in-dict membership instead of truthy value lookup so explicitly
supplied empty/falsy payloads still trigger the permission check.
Adds parametrized regression coverage across all gated keys.
Resolves merge conflict in tests/test_litellm/llms/bedrock/chat/test_converse_transformation.py
by keeping both the new bedrock tool-result file/document tests and the
transform_response body-leak regression test.
Also addresses Greptile P2 comment: when BedrockImageProcessor returns a
block with neither 'image' nor 'document' keys on the tool-result path
(image_url and file content types), log a warning instead of silently
dropping the block.
Co-authored-by: Mateo Wang <mateo-berri@users.noreply.github.com>