Success handlers already run when CustomStreamWrapper or
CachedResponsesAPIStreamingIterator finishes replay. Logging at
cache-hit time for acompletion/completion streaming duplicated spend
and callbacks. Align tests with deferred behavior.
Made-with: Cursor
_base_process_chunk only encodes response IDs when parsed_chunk contains
a top-level "response" key. Align test_process_chunk_completed_response_
updates_id_and_usage_cost with that contract and test_base_responses_api_streaming_iterator.
Made-with: Cursor
The test supplies a minimal PDF base64 payload but expected the wrong
constant (base64 for "test"). Assert against the same pdf_b64 value
and drop the unused import.
Made-with: Cursor
CodeQL flagged the previous ``from litellm.types.router import
MockRouterTestingParams`` at module top-level — ``litellm.types.router``
indirectly imports back into proxy modules, so the dataclass may not
exist yet when ``route_llm_request`` is being imported.
Hardcode the three flag names instead, with a guard test
(``test_mock_testing_kwarg_names_matches_dataclass``) that asserts the
hardcoded list matches ``MockRouterTestingParams.fields`` so drift is
caught at test time rather than missed in production.
Co-Authored-By: Claude Opus 4.7 (1M context) <noreply@anthropic.com>
Two changes that together prevent a caller from smuggling unauthorized
models past the API key's allowlist via per-request router overrides.
1. ``_enforce_key_and_fallback_model_access``: also walk fallback models
nested inside ``router_settings_override.fallbacks`` /
``context_window_fallbacks`` / ``content_policy_fallbacks``.
``route_llm_request.py`` promotes those to per-request kwargs after
auth, so without this they bypassed the model allowlist entirely.
New ``iter_router_fallback_model_names`` helper extracts leaf names
from both the simple top-level shape (str | {"model": str}) and the
nested router-config shape ({primary: [fallbacks]}). The two fallback
validation loops are unified — every name (top-level + override) is
deduplicated and validated once via ``can_key_call_model`` +
``is_valid_fallback_model``.
2. ``route_request``: strip router-internal ``mock_testing_*`` flags
from user-supplied data. These are testing-only flags that
deterministically force the router into fallback logic by raising a
synthetic ``InternalServerError`` etc. Combined with override
fallbacks they made the smuggling path trivially exploitable. Test
code that calls the router directly bypasses the strip and is
unaffected. The strip list is derived from ``MockRouterTestingParams``
so a new ``mock_testing_*`` flag added to that dataclass is
automatically covered.
Co-Authored-By: Claude Opus 4.7 (1M context) <noreply@anthropic.com>
The pre-release detector in create-release.yml uses `\.dev` (literal dot
before `dev`), which matches PEP 440 canonical tags like `1.84.0.dev2`
but misses the SemVer/Docker form `1.84.0-dev.2` (hyphen-dev). Per the
release design doc's PyPI<->Docker mapping rule, both forms are valid
production-track release tags and both are pre-releases (opt-in via
`pip install --pre litellm`), so the workflow should mark them as
GitHub pre-releases either way.
Change the regex to `[-.]dev` so it accepts `.dev` and `-dev`.
* fix(bedrock): handle document content blocks in Converse API message conversion
Document content blocks (used for PDF support) were silently dropped
during message conversion for Bedrock's Converse API. The content block
processing loop only handled text, image_url, and file types — document
blocks were skipped without warning, causing the model to respond as if
no document was provided.
Adds document block handling in three locations:
- Sync user message processing (_bedrock_converse_messages_pt)
- Async user message processing (_bedrock_converse_messages_pt_async)
- Tool result conversion (_convert_to_bedrock_tool_call_result)
Fixes#24641
Co-Authored-By: Claude Opus 4.6 (1M context) <noreply@anthropic.com>
* fix: use _validate_format for proper MIME type to Bedrock format mapping
Address Greptile review: naive media_type.split("/")[1] produced invalid
Bedrock format names for complex MIME types (e.g. OOXML → docx, text/plain
→ txt, text/markdown → md). Now reuses BedrockImageProcessor._validate_format
which handles all MIME types correctly via mimetypes + fallback.
Also fixes test assertions to expect correct Bedrock format values and adds
text/plain and text/markdown test cases.
Co-Authored-By: Claude Opus 4.6 (1M context) <noreply@anthropic.com>
* fix: reject non-base64 document sources with a clear error
URL-type document sources (e.g. {"type": "url", "url": "..."}) would
crash with an opaque KeyError on missing 'media_type'. Guard at the top
of _process_document_message and raise a clear ValueError since Bedrock
Converse only supports base64-encoded document sources.
Co-Authored-By: Claude Opus 4.6 (1M context) <noreply@anthropic.com>
---------
Co-authored-by: Claude Opus 4.6 (1M context) <noreply@anthropic.com>
The /v2/model/info endpoint (used by the UI's Models + Endpoints page)
was not resolving access group names when filtering models by team.
When a team has models: ["Group-A"] where "Group-A" is an access group,
_filter_models_by_team_id() passed it as a literal model name to
get_model_list(), which found no deployments with that name. This caused
the UI to show all models instead of only team-accessible ones.
The request-time auth path (model_in_access_group in auth_checks.py)
correctly resolves access groups via get_model_access_groups(). This
fix applies the same resolution in _filter_models_by_team_id() for both
the in-memory router lookup and the database fallback query.
Tests added:
- test_filter_resolves_access_group_names
- test_filter_resolves_mix_of_access_groups_and_literal_names
- test_filter_excludes_models_from_other_access_group
- test_filter_db_fallback_receives_resolved_model_names
Orgs with no MCP permissions configured (the common default) previously
returned None without writing to cache, meaning every subsequent MCP
request triggered a fresh find_unique against litellm_organizationtable.
Cache a sentinel string on the negative path so the DB is queried at most
once per cache TTL per org, regardless of whether the org has an
object_permission or not.
Made-with: Cursor
Greptile P2: ``asyncio.gather`` without ``return_exceptions=True`` lets
the first failing call propagate immediately, leaving the other in-flight
inspections running until they complete on their own. Pass
``return_exceptions=True`` so every inspection finishes, then re-raise
the first exception encountered while iterating results.
Co-Authored-By: Claude Opus 4.7 (1M context) <noreply@anthropic.com>
- Cache org object_permission in user_api_key_cache to avoid a DB hit on
every MCP request (was: raw find_unique on every call).
- Expand org mcp_servers list via expand_permission_list() so name-based
entries resolve to canonical IDs, consistent with key/team/end-user path.
- Expand org mcp_tool_permissions via expand_tool_permissions() in both
_get_allowed_mcp_servers_for_org and get_allowed_tools_for_server,
closing the silent name-vs-ID mismatch that could let restricted tools
through.
- The second _get_org_object_permission call in get_allowed_tools_for_server
now hits the cache (warm from the earlier server-list check), resolving
the double DB round-trip without changing the call structure.
Made-with: Cursor
Greptile P1: Aim's ``_anonymize_request`` and Lakera v2's mask-PII path
both wrote redacted content only to ``data["messages"]``. The Responses
API backend reads ``data["input"]``, so when a request arrived via
``/v1/responses`` with a plain string ``input`` the hook would update
``messages`` (which the backend ignores) and leave ``input`` carrying
the original unredacted text. Net effect: anonymize/mask silently passed
PII through to the LLM.
Add ``apply_redacted_messages_back`` to ``_content_utils`` — it writes
the redacted messages back to ``data["messages"]`` AND, when present,
re-flattens the redacted content into ``data["input"]``. Aim and
Lakera v2 now route their mask writeback through this helper. List
``input`` (multimodal) is still handled by the upstream
block-on-multimodal guard.
Adds unit tests for the helper and regression tests asserting
``data["input"]`` is redacted for both hooks on Responses-API string
input.
Co-Authored-By: Claude Opus 4.7 (1M context) <noreply@anthropic.com>
Apply organization object_permission as a ceiling on allowed MCP servers
and tool permissions, consistent with vector store org checks.
Includes unit tests for org ceiling, intersection, and tool filtering.
Made-with: Cursor
Two more in-place rewrite paths exhibit the same regression as Lakera v2:
overwriting ``data["messages"]`` with text-only redacted versions silently
strips image/audio parts from multimodal requests.
- ``LassoGuardrail._run_lasso_guardrail``: when ``mask=True`` AND input
is multimodal/Responses-API list, fall back to the classify endpoint
(which raises on BLOCK actions but never overwrites the payload).
- ``AimGuardrail._anonymize_request``: when input is multimodal, raise
the standard 400 instead of replacing ``data["messages"]`` with the
text-only ``redacted_chat`` from Aim. The error message tells the
user to either send plain string content or rely on block-mode.
Co-Authored-By: Claude Opus 4.7 (1M context) <noreply@anthropic.com>
Mask-in-place uses the offsets that Lakera returns for the inspection
payload. ``build_inspection_messages`` flattens multimodal content into
joined text before sending to Lakera, so the offsets refer to the
flattened representation. Writing those offsets back via
``_mask_pii_in_messages`` and overwriting ``data["messages"]`` would
silently strip image/audio parts from the original request — that is a
real functional regression for Lakera + mask mode + multimodal input.
Detect multimodal input (any list-format ``content`` or non-string
``data["input"]``) up front and skip the mask-in-place branch in that
case. The hook then falls into the standard block-on-detect path so PII
is still blocked but the multimodal payload is never silently rewritten.
Per-part masking that preserves multimodal structure is the right
long-term fix; tracking that as a follow-up.
Also: add ``has_non_string_content`` to ``_content_utils`` (with tests)
and a regression test that asserts multimodal+PII raises an HTTPException
instead of returning a flattened request body.
Co-Authored-By: Claude Opus 4.7 (1M context) <noreply@anthropic.com>
Greptile P2 follow-ups on _content_utils.py:
- Drop unreachable ``_resolve_messages``. The new
``_iter_inspection_messages`` walks ``messages`` AND ``input``
independently; leaving the old fallback-only variant around invited a
future maintainer to wire it back up and silently narrow coverage.
- Rename ``iter_user_text`` → ``iter_message_text``. The helper walks
every role (user, assistant, system); the old name implied user-turn
content only. Callers and tests updated.
- Close mixed-list coverage gap. When ``data["input"]`` was a list
mixing content-part dicts and bare strings, ``iter_message_text`` and
``build_inspection_messages`` only saw the dict parts while
``walk_user_text`` already inspected both. ``_iter_text_parts_in_content``
now treats bare strings inside a content list as text fragments, so
read and write helpers agree on coverage.
Adds two regression tests for the mixed-list shape.
Co-Authored-By: Claude Opus 4.7 (1M context) <noreply@anthropic.com>
CI mypy flagged the two ``_mask_pii_in_messages(messages=new_messages, ...)``
call sites because ``build_inspection_messages`` returns
``List[Dict[str, str]]`` while ``_mask_pii_in_messages`` declares
``List[AllMessageValues]``. The runtime payload matches; widening the
return type to a Union of TypedDicts is a larger refactor than is
warranted by this hardening change, so add the localised ``# type: ignore``
to match the two we already added on the ``call_v2_guard`` calls above.
Co-Authored-By: Claude Opus 4.7 (1M context) <noreply@anthropic.com>
Several guardrail hooks short-circuit when ``message.content`` is a list
or when the request uses the Responses-API ``input`` field instead of
``messages``. Centralise the content-walking logic in a shared helper and
update the affected hooks so list-format and Responses-API payloads no
longer skip inspection.
Also: Aim's ``async_post_call_success_hook`` now inspects every choice
(via ``asyncio.gather``) instead of only ``choices[0]`` — the prior
behaviour let ``n>1`` callers hide content in subsequent completions.
Hooks updated to use the new helper:
- aim, lakera_ai_v2, lasso (post a synthesised messages list to a remote
guardrail service)
- azure_content_safety, ibm_detector, banned_keywords, openai_moderation,
google_text_moderation (iterate text fragments locally)
- secret_detection (walk-and-rewrite to redact in place)
Drive-by fix: the legacy ``data["prompt"]`` list-handling path in
secret_detection rebound the loop variable instead of mutating the list,
leaving secrets unredacted on text-completion calls; corrected to index
back into the list.
Co-Authored-By: Claude Opus 4.7 (1M context) <noreply@anthropic.com>