Without :(glob), git matches 'tests/e2e/**/*.py' with * crossing slashes, so the
pattern needs at least one directory below tests/e2e and a top-level file never
matches. PR #39209 added tests/e2e/test_junit_properties.py with three
basedpyright errors and the e2e step printed "No changed tests/e2e Python files;
skipping." The ruff format step's 'litellm/**/*.py' skipped litellm/main.py and
the other top-level modules the same way.
:(glob) makes /**/ match zero or more directories, so both gates now select
top-level and nested files alike. A regression test runs the workflow's own
pathspecs against a throwaway repo and locks that in for every diff-scoped gate.
initialize_presidio registers up to three callbacks per guardrail but the
registry only kept the first, so deleting or re-syncing the guardrail left
the post_call siblings serving the old config. The initializer now returns
every callback it registered, the registry tracks primary and siblings per
guardrail id, delete purges all of them from every callback list, and
update pushes the new params into each while siblings keep their stage.
An Anthropic buffer without a stop_reason only ran the flat text scan, so a
rewrite there was dropped while the executor trusted the translation to have
delivered it. A Responses buffer ending at response.output_item.done returned
after the tool-call scan without ever checking the text. Both now reach the
flat scan and raise UndeliverableStreamRewrite when a caller expects the
rewrite delivered, matching the existing Responses no-envelope fallback.
* fix(search): forward search-tool params through the router, complete Parallel AI v1 param mapping
SearchAPIRouter dropped every parameter configured on a search tool, forwarding
only per-request kwargs. Any tool-level setting (mode, max_results, ...) was
silently lost on the way to the adapter, for every search provider.
Also completes the Parallel AI v1 search surface: after_date, fetch_policy,
location and include_domains now nest under advanced_settings instead of being
sent as unknown top-level fields, responses preserve search_id / session_id /
warnings / raw excerpts, and search cost is derived from the request mode and
the provider's reported usage rather than a single flat rate.
* fix(parallel_ai): stop a caller from pricing its own search request
`_parallel_ai_usage` carries the provider's reported usage into cost
calculation. It was only written when the response contained a usage block, so
a caller could pass `_parallel_ai_usage=[{"name": "sku_search", "count": 0}]`
and, whenever the provider omitted usage, bill $0.00 instead of $0.005 — the
value also reached the upstream request body as an unknown field.
The key is now stripped from inbound params and written unconditionally from
the parsed response, so only the provider can populate it.
* fix(parallel_ai): price fast search mode correctly
* test(parallel_ai): fake search at HTTP boundary
* fix(parallel_ai): tolerate null search result fields
---------
Co-authored-by: khushishelat <shelatkhushi@gmail.com>
An agent's object_permission.mcp_toolsets could be persisted through the
new edit form and PATCH /v1/agents but never reached the request-time
checks: _get_allowed_mcp_servers_for_agent read only mcp_servers and
mcp_access_groups, so an agent granted a toolset alone resolved to [] and
placed no ceiling on its keys, and _get_agent_tool_permissions_for_server
ignored the tools those toolsets grant. Both helpers now resolve toolsets
through the shared _toolset_tool_permissions / _toolset_tools_for_server
helpers the key, team, and org levels already use, and a declared toolset
that resolves to nothing raises UnloadableEntitlementError so the
resolver denies instead of reading the agent as unrestricted
tests/e2e/test_junit_properties.py fed a hand-rolled FakeItem to
result_properties and attach_result_properties, both typed pytest.Item,
so uv run basedpyright tests/e2e reported 3 reportArgumentType errors on
litellm_internal_staging and every make check that scopes a litellm/ or
tests/e2e/ Python file failed.
Each test now looks up its own collected Item in request.session.items
and applies the covers marker at run time through request.applymarker,
so the coverage registry's collect-only pass never sees the test ids and
the production functions keep their pytest.Item signatures. No casts, no
ignores.
Resolves LIT-6669
Brings the stack base up to origin/litellm_internal_staging. Staging's
get_streaming_string_so_far now joins delta-only Responses parts itself, so
the handler's separate delta joiner and the flag-gated delta scan go away
and the fallback keeps failing closed on a rewrite it cannot deliver. The
two tests that pinned the old flag-gated behaviour are dropped in favour of
staging's delta-assembly tests, and _process_ended_stream takes the same
UserAPIKeyAuth | None the typed process_output_response now expects
Any scoped MCP request (/mcp/<name> path or x-mcp-servers header) that resolves to zero allowed servers now returns one generic 403, so unknown, unauthorized, and access-group names are indistinguishable and the error cannot be used to enumerate servers. The agent-attributed variant reruns the same scope resolver with the agent binding stripped and fires only when that rerun resolves, naming the vetoed server or access group and both fix paths
AgentObjectPermission now declares mcp_toolsets so PATCH /v1/agents keeps it, and the agent edit form round-trips toolsets alongside servers and access groups instead of dropping them on save. The two scope-resolver params only iterate their names, so they take Sequence[str] and the LIT001 budget ratchets down by the two annotations this widens
* fix(proxy): keep passthrough logging metadata and model_info dicts when team callbacks are wired
Passing team callback vars into Logging(kwargs=...) makes get_litellm_params materialize a full litellm_params, where metadata and model_info default to None instead of being absent. Readers that resolve them as .get(key, {}).get(...) then raise, so any passthrough request from a team with logging callbacks 500s once a pre-call guardrail is on, and the router strategy loggers log a traceback per request.
* test(proxy): annotate the closure dicts the passthrough logging tests record into
The LLM Obs callback copied litellm's OpenAI-shaped objects into the span
verbatim, so every field Datadog names differently landed somewhere it does
not read: tool calls kept their nested `function` wrapper instead of DD's
name/arguments/tool_id, tool messages carried no result linking them to their
call, the request's tools were never sent, and prompt-cache counts sat inside
meta.metadata rather than the span metrics its cache dashboards chart.
One rule governs the message mapper: add the fields Datadog declares, and never
destroy content it did not understand. Content collapses to its text only when
it has text, so a content list carrying tool or image blocks rides along
unchanged, and absent messages map to an empty input rather than a fabricated
turn. Tool calls and results are read from both dialects, the OpenAI
`tool_calls` / `role: tool` shape and the Anthropic `tool_use` / `tool_result`
content blocks, so /v1/messages sessions gain tool linking they never had.
Cache counts come from the same owners the savings dashboard uses, so every
provider spelling resolves through one place rather than a second local guess.
The three cache metrics partition the input count: litellm's normalized prompt
total includes both cache categories, as the cost calculator's pricing helper
documents, so the non-cached residual subtracts reads AND writes. Counting a
primed prefix as ordinary input had inflated non-cached usage by exactly the
cache-write count on every priming request.
Correlating a result to its call reads ids and names structurally and parses no
arguments, so a tool call's arguments are decoded once per span rather than
once per pass, and arguments past a size bound ship as the raw string instead
of paying a decode that multiplies memory on hostile compact JSON.
The flat `output_tool_calls.*` metadata copies go away with this: they were a
second representation of a fact that now has its own field on the same span.
Chat streaming write-backs now match chunks by the choice's index field
instead of its list position, delivering rewrites to the right choice on
n>1 streams; an ended-stream rewrite on a multi-choice buffer fails
closed since stream_chunk_builder collapses the choices. The Responses
fallback joins output_text.delta events when delivery is expected, so a
delta-only buffer is guardrail-checked instead of released raw.
Bedrock rejects requests carrying cachePoint blocks for models whose entry in the cost map does not declare supports_prompt_caching (403 "You invoked an unsupported model or your request did not allow prompt caching"). Clients like Claude Code attach cache_control to every request, so any such model behind the gateway failed on every call. The new bedrock_model_accepts_cache_points predicate drops cachePoint emission for map-known non-caching models at all three emission funnels, keeps emitting for unmapped ids (application inference profile ARNs), and skips the gateway injection credit when the tool_config point is not placed.
aiohttp shields its DNS resolution task; when the connector closes it cancels
that child, so the request task sees CancelledError without ever being
cancelled itself. map_aiohttp_exceptions() only caught Exception, so the
BaseException skipped transport mapping, router retries and proxy error
handling, and /v1/responses answered 500 "No response returned".
Catch CancelledError in the mapper, re-raise when the current task is really
being cancelled (Task.cancelling() > 0), and otherwise map it to
httpx.ConnectError so the usual retry, fallback and error mapping apply.
Co-authored-by: yassin <yassin@berri.ai>
Co-authored-by: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>
The Dockerfile, docker/Dockerfile.non_root and docker/Dockerfile.database uv sync stages never passed --extra bedrock-realtime, so aws-sdk-bedrock-runtime was absent from the image venv and Bedrock Nova Sonic /v1/realtime sessions failed with 'Missing aws_sdk_bedrock_runtime'. gateway/Dockerfile already had the extra (PR #34426).
Adds a static check over every uv sync in the proxy Dockerfiles and an image-level import probe that the image-scan workflow runs against the built root, non-root and gateway images.
Co-authored-by: yassin <yassin@berri.ai>
Co-authored-by: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>