Several guardrail hooks short-circuit when ``message.content`` is a list
or when the request uses the Responses-API ``input`` field instead of
``messages``. Centralise the content-walking logic in a shared helper and
update the affected hooks so list-format and Responses-API payloads no
longer skip inspection.
Also: Aim's ``async_post_call_success_hook`` now inspects every choice
(via ``asyncio.gather``) instead of only ``choices[0]`` — the prior
behaviour let ``n>1`` callers hide content in subsequent completions.
Hooks updated to use the new helper:
- aim, lakera_ai_v2, lasso (post a synthesised messages list to a remote
guardrail service)
- azure_content_safety, ibm_detector, banned_keywords, openai_moderation,
google_text_moderation (iterate text fragments locally)
- secret_detection (walk-and-rewrite to redact in place)
Drive-by fix: the legacy ``data["prompt"]`` list-handling path in
secret_detection rebound the loop variable instead of mutating the list,
leaving secrets unredacted on text-completion calls; corrected to index
back into the list.
Co-Authored-By: Claude Opus 4.7 (1M context) <noreply@anthropic.com>
Greptile P2 follow-up: when a litellm_configoverrides row exists with a
NULL config_value (e.g. an earlier failed write left a stub), the audit
action was mislabeled "created" because we keyed off existing_decrypted
(which is only set when config_value is non-null). Key off existing_record
instead — a row is a row regardless of its value.
Also hoist asyncio + patch imports to module top in the test file.
Co-Authored-By: Claude Opus 4.7 (1M context) <noreply@anthropic.com>
The previous fix for the TOCTOU bypass relied on a per-instance asyncio.Lock,
which closed the window only within a single proxy worker. Multi-replica
deployments still raced across processes — A and B both read counter=99,
both passed validation, both incremented to 100/100 → effective limit doubled.
Add `CHECK_AND_INCREMENT_BY_N_SCRIPT` Lua script that processes any number of
(window_key, counter_key, limit, increment, ttl) descriptors atomically with
all-or-nothing semantics: if any descriptor would exceed its limit, no counter
is modified and the script returns OVER_LIMIT with the offending descriptor's
state. When Redis isn't configured, the in-memory fallback uses the existing
asyncio.Lock for single-process atomicity.
Expose this as `_PROXY_MaxParallelRequestsHandler_v3.atomic_check_and_increment_by_n`
and rewire both call sites:
- batch_rate_limiter._check_and_increment_batch_counters: replace the
read_only=True check + separate async_increment_tokens_with_ttl_preservation
with a single atomic call passing the batch's (request_count, total_tokens)
as the increment.
- dynamic_rate_limiter_v3._check_rate_limits: bundle model_saturation_check
(always enforced) and priority_model (enforced only when saturated) into
one atomic call. When priority is unenforced, increment its counter via
the existing should_rate_limit(read_only=False) path for tracking only.
Update structural regression tests to assert the new atomic path is used
rather than the legacy two-phase pattern.
Tests: 4/4 TOCTOU tests pass, 59 existing rate-limiter tests pass, no
regressions.
Co-Authored-By: Claude Opus 4.7 (1M context) <noreply@anthropic.com>
Three Greptile/CI findings on the prior commit:
1. **P1 (real bug):** ``before_config = existing_decrypted if existing_record
is not None else env_values`` would NameError when ``existing_record``
exists but its ``config_value`` is null (a valid nullable DB state) —
the upper branch only defines ``existing_decrypted`` when *both*
conditions are met, but the ternary only checked the first. Pre-bind
``existing_decrypted: Optional[Dict] = None`` and ``env_values = {}``
above the if/else so both names are always in scope, and key the
audit-log decision off ``existing_decrypted is not None`` instead.
2. **mypy lint:** ``action: str`` rejected — the field is typed
``AUDIT_ACTIONS = Literal[...]``. Annotate both helper signatures
with ``AUDIT_ACTIONS`` and pre-bind the call-site ternary so mypy
infers the literal correctly.
3. **P2:** ``import asyncio`` was at the bottom of the test file.
Moved to the stdlib import block at top.
The batch rate limiter (`_check_and_increment_batch_counters`) and the
dynamic rate limiter (`_check_rate_limits`) implemented rate limiting in
two disjoint awaits: a `should_rate_limit(read_only=True)` check followed
by a separate increment. Concurrent requests could all observe the same
pre-increment state, all pass enforcement, and all then increment —
multiplying the effective quota by the concurrency level.
Demonstrated bypass (see new test):
- Batch: 5 concurrent batches of 40 tokens each against TPM=100 consumed
200 tokens (100% over).
- Dynamic: 5 concurrent priority="high" requests against RPM=2 all
passed Phase 1 + Phase 3.
Wrap both critical sections in a per-instance asyncio.Lock so the read
and increment execute atomically within a process. Multi-replica
deployments still rely on Redis Lua atomicity for cross-process safety;
that is a follow-up.
Co-Authored-By: Claude Opus 4.7 (1M context) <noreply@anthropic.com>
Follow-up to merged #26859 (team-callback audit log). The variant
scan from that PR flagged two more high-risk admin endpoints whose
mutations weren't audit-logged:
- ``POST /cache/settings`` (cache_settings_endpoints.py) — writes the
team / global Redis cache configuration, including credentials. An
admin (or compromised admin) flipping the cache backend is a
data-routing pivot — every subsequent LLM response cache write goes
to the new destination — so the action needs to be traceable.
- ``POST /config_overrides/hashicorp_vault`` and
``DELETE /config_overrides/hashicorp_vault``
(config_override_endpoints.py) — control the proxy's KMS config.
Mutating these affects every secret retrieval going forward.
Each endpoint now emits an ``LiteLLM_AuditLogs`` row gated on
``litellm.store_audit_logs`` (Enterprise feature), mirroring the
shape merged in #26859. Both helpers redact every field value
before serialization (replacing them with ``***REDACTED***`` while
keeping field names) so the audit table cannot itself become a
credential-harvest sink — Redis passwords / vault tokens /
``approle_secret_id`` / ``client_key`` would otherwise be persisted
in plaintext JSON.
Both helpers also attach a ``done_callback`` that surfaces a
``verbose_proxy_logger.warning`` when the fire-and-forget audit-log
task fails, so a transient DB error doesn't silently lose the row.
Adds ``CACHE_CONFIG_TABLE_NAME`` / ``CONFIG_OVERRIDES_TABLE_NAME`` to
``LitellmTableNames`` so the audit rows co-locate with the table they
mutate.
Tests:
- /cache/settings: emits when ``store_audit_logs=True``, no emission
when off, plaintext credentials don't appear in the serialized row.
- /config_overrides/hashicorp_vault POST: same, plus checks the
redaction of ``vault_token`` and ``vault_addr``.
- /config_overrides/hashicorp_vault DELETE: emits with ``action="deleted"``
only when an actual row was removed; idempotent delete of a non-existent
row produces no audit log.
Two Greptile review findings addressed:
1. (P1, security) The ``litellm_oauth_state`` cookie is the sole
guard against Login-CSRF in the PKCE flow but was set without the
``Secure`` attribute, so a network observer on plain HTTP could
read and replay it — bypassing the protection this PR adds.
Thread the originating ``Request`` down through
``get_sso_login_redirect`` and ``get_generic_sso_redirect_response``
and set ``Secure`` based on ``request.url.scheme == "https"``.
When no request is supplied (programmatic callers / tests) default
to ``Secure=True`` — production-safe. Local HTTP dev still works
because the request scheme is observed at runtime.
2. (P2) The cookie was set unconditionally, but the callback only
validates it inside the PKCE branch. Two concurrent SSO sessions
(one PKCE, one plain) could overwrite each other's state cookie
and produce spurious 400s for the plain-flow user.
Move the ``set_cookie`` call inside the existing
``if code_verifier and "state" in redirect_params`` block so the
cookie is only written when PKCE is active and the validation
will actually fire.
Tests cover both paths: PKCE-on (cookie set with Secure default),
PKCE-off (cookie not set), and HTTP dev request (Secure dropped so
the browser will actually attach the cookie on the callback hop).
The Generic SSO PKCE flow used the URL ``state`` parameter as the
cache key for the PKCE ``code_verifier`` without binding the state
to the caller's browser. An attacker who pre-minted a state and
cached a verifier under it could hand the resulting login link to a
victim; the victim's auth code would then be exchanged with the
attacker's verifier on the callback, producing an access token
under the attacker's control (Login CSRF / token theft).
The non-PKCE branch is unaffected because it delegates to
fastapi-sso's ``verify_and_process``, which performs its own
session-cookie check. The PKCE branch bypasses that helper, which
is exactly the gap this commit closes.
Two-part fix in ``ui_sso.py``:
- ``get_generic_sso_redirect_response`` now sets a
``litellm_oauth_state`` cookie (HttpOnly, SameSite=Lax, 10-min TTL)
carrying the state value used in the redirect URL. The cookie is
set on the redirect response just like the existing
``litellm_cp_return_to`` cookie a few lines earlier in the file.
- ``get_generic_sso_response`` validates ``request.cookies.get(
"litellm_oauth_state")`` against ``request.query_params.get(
"state")`` via ``secrets.compare_digest`` before invoking the
PKCE token exchange. Mismatch (or either being missing) raises a
``ProxyException`` with HTTP 400.
The pre-existing TODO above the redirect logic ("state should be a
random string and added to the user session with cookie") is now
addressed and removed.
Tests cover the redirect-side cookie set, the missing-cookie reject
shape, the URL/cookie-mismatch reject shape, and the matching-cookie
happy path.