Greptile P1: Aim's ``_anonymize_request`` and Lakera v2's mask-PII path
both wrote redacted content only to ``data["messages"]``. The Responses
API backend reads ``data["input"]``, so when a request arrived via
``/v1/responses`` with a plain string ``input`` the hook would update
``messages`` (which the backend ignores) and leave ``input`` carrying
the original unredacted text. Net effect: anonymize/mask silently passed
PII through to the LLM.
Add ``apply_redacted_messages_back`` to ``_content_utils`` — it writes
the redacted messages back to ``data["messages"]`` AND, when present,
re-flattens the redacted content into ``data["input"]``. Aim and
Lakera v2 now route their mask writeback through this helper. List
``input`` (multimodal) is still handled by the upstream
block-on-multimodal guard.
Adds unit tests for the helper and regression tests asserting
``data["input"]`` is redacted for both hooks on Responses-API string
input.
Co-Authored-By: Claude Opus 4.7 (1M context) <noreply@anthropic.com>
Apply organization object_permission as a ceiling on allowed MCP servers
and tool permissions, consistent with vector store org checks.
Includes unit tests for org ceiling, intersection, and tool filtering.
Made-with: Cursor
Two more in-place rewrite paths exhibit the same regression as Lakera v2:
overwriting ``data["messages"]`` with text-only redacted versions silently
strips image/audio parts from multimodal requests.
- ``LassoGuardrail._run_lasso_guardrail``: when ``mask=True`` AND input
is multimodal/Responses-API list, fall back to the classify endpoint
(which raises on BLOCK actions but never overwrites the payload).
- ``AimGuardrail._anonymize_request``: when input is multimodal, raise
the standard 400 instead of replacing ``data["messages"]`` with the
text-only ``redacted_chat`` from Aim. The error message tells the
user to either send plain string content or rely on block-mode.
Co-Authored-By: Claude Opus 4.7 (1M context) <noreply@anthropic.com>
Mask-in-place uses the offsets that Lakera returns for the inspection
payload. ``build_inspection_messages`` flattens multimodal content into
joined text before sending to Lakera, so the offsets refer to the
flattened representation. Writing those offsets back via
``_mask_pii_in_messages`` and overwriting ``data["messages"]`` would
silently strip image/audio parts from the original request — that is a
real functional regression for Lakera + mask mode + multimodal input.
Detect multimodal input (any list-format ``content`` or non-string
``data["input"]``) up front and skip the mask-in-place branch in that
case. The hook then falls into the standard block-on-detect path so PII
is still blocked but the multimodal payload is never silently rewritten.
Per-part masking that preserves multimodal structure is the right
long-term fix; tracking that as a follow-up.
Also: add ``has_non_string_content`` to ``_content_utils`` (with tests)
and a regression test that asserts multimodal+PII raises an HTTPException
instead of returning a flattened request body.
Co-Authored-By: Claude Opus 4.7 (1M context) <noreply@anthropic.com>
Greptile P2 follow-ups on _content_utils.py:
- Drop unreachable ``_resolve_messages``. The new
``_iter_inspection_messages`` walks ``messages`` AND ``input``
independently; leaving the old fallback-only variant around invited a
future maintainer to wire it back up and silently narrow coverage.
- Rename ``iter_user_text`` → ``iter_message_text``. The helper walks
every role (user, assistant, system); the old name implied user-turn
content only. Callers and tests updated.
- Close mixed-list coverage gap. When ``data["input"]`` was a list
mixing content-part dicts and bare strings, ``iter_message_text`` and
``build_inspection_messages`` only saw the dict parts while
``walk_user_text`` already inspected both. ``_iter_text_parts_in_content``
now treats bare strings inside a content list as text fragments, so
read and write helpers agree on coverage.
Adds two regression tests for the mixed-list shape.
Co-Authored-By: Claude Opus 4.7 (1M context) <noreply@anthropic.com>
Several guardrail hooks short-circuit when ``message.content`` is a list
or when the request uses the Responses-API ``input`` field instead of
``messages``. Centralise the content-walking logic in a shared helper and
update the affected hooks so list-format and Responses-API payloads no
longer skip inspection.
Also: Aim's ``async_post_call_success_hook`` now inspects every choice
(via ``asyncio.gather``) instead of only ``choices[0]`` — the prior
behaviour let ``n>1`` callers hide content in subsequent completions.
Hooks updated to use the new helper:
- aim, lakera_ai_v2, lasso (post a synthesised messages list to a remote
guardrail service)
- azure_content_safety, ibm_detector, banned_keywords, openai_moderation,
google_text_moderation (iterate text fragments locally)
- secret_detection (walk-and-rewrite to redact in place)
Drive-by fix: the legacy ``data["prompt"]`` list-handling path in
secret_detection rebound the loop variable instead of mutating the list,
leaving secrets unredacted on text-completion calls; corrected to index
back into the list.
Co-Authored-By: Claude Opus 4.7 (1M context) <noreply@anthropic.com>
Greptile P2 follow-up: when a litellm_configoverrides row exists with a
NULL config_value (e.g. an earlier failed write left a stub), the audit
action was mislabeled "created" because we keyed off existing_decrypted
(which is only set when config_value is non-null). Key off existing_record
instead — a row is a row regardless of its value.
Also hoist asyncio + patch imports to module top in the test file.
Co-Authored-By: Claude Opus 4.7 (1M context) <noreply@anthropic.com>
The previous fix for the TOCTOU bypass relied on a per-instance asyncio.Lock,
which closed the window only within a single proxy worker. Multi-replica
deployments still raced across processes — A and B both read counter=99,
both passed validation, both incremented to 100/100 → effective limit doubled.
Add `CHECK_AND_INCREMENT_BY_N_SCRIPT` Lua script that processes any number of
(window_key, counter_key, limit, increment, ttl) descriptors atomically with
all-or-nothing semantics: if any descriptor would exceed its limit, no counter
is modified and the script returns OVER_LIMIT with the offending descriptor's
state. When Redis isn't configured, the in-memory fallback uses the existing
asyncio.Lock for single-process atomicity.
Expose this as `_PROXY_MaxParallelRequestsHandler_v3.atomic_check_and_increment_by_n`
and rewire both call sites:
- batch_rate_limiter._check_and_increment_batch_counters: replace the
read_only=True check + separate async_increment_tokens_with_ttl_preservation
with a single atomic call passing the batch's (request_count, total_tokens)
as the increment.
- dynamic_rate_limiter_v3._check_rate_limits: bundle model_saturation_check
(always enforced) and priority_model (enforced only when saturated) into
one atomic call. When priority is unenforced, increment its counter via
the existing should_rate_limit(read_only=False) path for tracking only.
Update structural regression tests to assert the new atomic path is used
rather than the legacy two-phase pattern.
Tests: 4/4 TOCTOU tests pass, 59 existing rate-limiter tests pass, no
regressions.
Co-Authored-By: Claude Opus 4.7 (1M context) <noreply@anthropic.com>
Three Greptile/CI findings on the prior commit:
1. **P1 (real bug):** ``before_config = existing_decrypted if existing_record
is not None else env_values`` would NameError when ``existing_record``
exists but its ``config_value`` is null (a valid nullable DB state) —
the upper branch only defines ``existing_decrypted`` when *both*
conditions are met, but the ternary only checked the first. Pre-bind
``existing_decrypted: Optional[Dict] = None`` and ``env_values = {}``
above the if/else so both names are always in scope, and key the
audit-log decision off ``existing_decrypted is not None`` instead.
2. **mypy lint:** ``action: str`` rejected — the field is typed
``AUDIT_ACTIONS = Literal[...]``. Annotate both helper signatures
with ``AUDIT_ACTIONS`` and pre-bind the call-site ternary so mypy
infers the literal correctly.
3. **P2:** ``import asyncio`` was at the bottom of the test file.
Moved to the stdlib import block at top.
The batch rate limiter (`_check_and_increment_batch_counters`) and the
dynamic rate limiter (`_check_rate_limits`) implemented rate limiting in
two disjoint awaits: a `should_rate_limit(read_only=True)` check followed
by a separate increment. Concurrent requests could all observe the same
pre-increment state, all pass enforcement, and all then increment —
multiplying the effective quota by the concurrency level.
Demonstrated bypass (see new test):
- Batch: 5 concurrent batches of 40 tokens each against TPM=100 consumed
200 tokens (100% over).
- Dynamic: 5 concurrent priority="high" requests against RPM=2 all
passed Phase 1 + Phase 3.
Wrap both critical sections in a per-instance asyncio.Lock so the read
and increment execute atomically within a process. Multi-replica
deployments still rely on Redis Lua atomicity for cross-process safety;
that is a follow-up.
Co-Authored-By: Claude Opus 4.7 (1M context) <noreply@anthropic.com>
Follow-up to merged #26859 (team-callback audit log). The variant
scan from that PR flagged two more high-risk admin endpoints whose
mutations weren't audit-logged:
- ``POST /cache/settings`` (cache_settings_endpoints.py) — writes the
team / global Redis cache configuration, including credentials. An
admin (or compromised admin) flipping the cache backend is a
data-routing pivot — every subsequent LLM response cache write goes
to the new destination — so the action needs to be traceable.
- ``POST /config_overrides/hashicorp_vault`` and
``DELETE /config_overrides/hashicorp_vault``
(config_override_endpoints.py) — control the proxy's KMS config.
Mutating these affects every secret retrieval going forward.
Each endpoint now emits an ``LiteLLM_AuditLogs`` row gated on
``litellm.store_audit_logs`` (Enterprise feature), mirroring the
shape merged in #26859. Both helpers redact every field value
before serialization (replacing them with ``***REDACTED***`` while
keeping field names) so the audit table cannot itself become a
credential-harvest sink — Redis passwords / vault tokens /
``approle_secret_id`` / ``client_key`` would otherwise be persisted
in plaintext JSON.
Both helpers also attach a ``done_callback`` that surfaces a
``verbose_proxy_logger.warning`` when the fire-and-forget audit-log
task fails, so a transient DB error doesn't silently lose the row.
Adds ``CACHE_CONFIG_TABLE_NAME`` / ``CONFIG_OVERRIDES_TABLE_NAME`` to
``LitellmTableNames`` so the audit rows co-locate with the table they
mutate.
Tests:
- /cache/settings: emits when ``store_audit_logs=True``, no emission
when off, plaintext credentials don't appear in the serialized row.
- /config_overrides/hashicorp_vault POST: same, plus checks the
redaction of ``vault_token`` and ``vault_addr``.
- /config_overrides/hashicorp_vault DELETE: emits with ``action="deleted"``
only when an actual row was removed; idempotent delete of a non-existent
row produces no audit log.
When model_info.id equals model_name (common for batch models), the router
resolves via has_model_id and returns one deployment dict instead of a list.
The dict branch incorrectly iterated deployment keys (model_name,
litellm_params, model_info), producing non-string values that broke
LiteLLM_ManagedFileTable validation on managed file upload.
Normalize list vs dict by wrapping single deployments and extracting
model_info.id for each response pair.
Add regression tests including the batch model id == model_name case.
Made-with: Cursor