* fix(proxy): return the real status code when a credential update is rejected
update_credential ended its except clause with 'return handle_exception_on_proxy(e)'. Returning the exception makes it the response body, so FastAPI answers 200 and every rejection on this route reads as a successful write to any caller that checks the status; the admin dashboard's API client is one. Patching a name that does not exist answered 200 with the real 404 buried in the body.
The sibling handlers in this file already raise. The route had no test coverage, which is why it survived.
* Update tests/test_litellm/proxy/credential_endpoints/test_endpoints.py
Co-authored-by: greptile-apps[bot] <165735046+greptile-apps[bot]@users.noreply.github.com>
---------
Co-authored-by: greptile-apps[bot] <165735046+greptile-apps[bot]@users.noreply.github.com>
Guards _declared_query_params against a regression in the get_flat_params
migration: the flatten step returns path, query, header and cookie params
together, so a dropped ParamTypes.query filter would wrongly treat path or
header names as declared query params and accept unknown ones. Removing the
filter fails these tests.
The end-to-end extras test now injects a MagicMock(spec=HTTPHandler) via
the client parameter instead of patching post on a real handler, and the
docstrings on the new regression tests are removed, addressing the
remaining Greptile review feedback.
Mirror the chat extras translation for /v1/completions, adapted to the
typed OpenAI SDK: anything completions.create() rejects (reasoning_effort,
response_format, fireworks-native extras) rides inside extra_body, which
the SDK merges server-side. Top-level reasoning_effort and response_format
are moved into extra_body (they raised TypeError before), truncate
aliases, chat_template_kwargs effort keys, and guided_* resolve into
extra_body fields, and the strip set removes the rest. Verified live:
/v1/completions rejects prompt_truncate_len, so both truncate names are
stripped on this path rather than renamed.
The Azure Sentinel logger hardcoded the commercial Entra authority and the
commercial Azure Monitor audience, so Log Analytics ingestion could not work in
Azure Government even when the ingestion endpoint pointed at a sovereign Data
Collection Endpoint.
Resolve the authority from AZURE_AUTHORITY_HOST and derive the matching Logs
Ingestion audience from it. Moving only the token URL is not enough: sovereign
Entra would then be asked for a token scoped to the commercial audience, which
the sovereign endpoint rejects.
The http handler merges extra_body after transform_request, so a
response_format nested in an explicit extra_body would silently clobber
the explicit top-level response_format. Drop the nested copy with a
debug log so the top-level value wins, closing the precedence hole in
the guided-param native-wins path.
Streaming /chat/completions and /v1/responses emit nothing, not even response headers, until the upstream yields its first chunk, so an ingress with an idle read timeout (nginx proxy-read-timeout) drops long time-to-first-token streams.
Reuses the existing Anthropic keepalive wrapper with a configurable ping payload, emitting an SSE comment on the OpenAI-shaped routes so conformant clients ignore it. Off unless litellm_settings.sse_keepalive_ping_interval_seconds is set.
Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>
test_get_deployment_credentials_with_provider_bedrock_batch_fields already
covers s3_encryption_key_id on the base branch, and the new test passes with
every production file in this branch reverted, so it guards nothing.
create_websocket_passthrough_route existed but /openai and
/openai_passthrough only registered HTTP methods, so WS upgrades were
rejected at routing. Add catch-all websocket routes mirroring the HTTP
passthrough target construction.
Fixes#36088
An intercepted web search called litellm.asearch() with only the search tool's litellm_params, so the search request carried no owner. The proxy's spend hook skips any call with no key, user or team attached, so the search's provider cost never reached SpendLogs; it was missing from the Logs page and never counted against the caller's budget. The same path never ran the rate limiter either, so an intercepted search was free of the key's RPM/TPM limits.
The search now carries the originating key's attribution metadata (key hash, alias, user, team, org, plus model_group set to the resolved search tool) and runs the caller's rate limit checks before hitting the provider, matching what a direct /v1/search request gets. SDK calls with no proxy auth context are unchanged.
Resolves LIT-5033
Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>
create_a2a_client took the raw client off a process-wide cached handler and
called headers.update() on it, then leaned on folding the header set into the
cache key (through the unrelated disable_aiohttp_transport field) to keep one
caller's credentials away from the next.
Per-caller headers now ride with each request through the a2a SDK's call
context, and the agent card fetch gets them through resolver_http_kwargs, so
the shared client is never written to and its cache key no longer varies by
header set. Since the proxy puts a fresh trace id in every request's headers,
that key previously changed on every call, giving each request its own httpx
client and flushing the 200-entry client cache that every other provider
shares. All A2A callers on one timeout now reuse a single pooled client.
Sharing that client also means sharing its httpx cookie jar, which httpx fills
from every Set-Cookie and replays on any later request to a matching domain, so
one agent's session cookie would arrive at another agent on the same host. The
pooled client now carries a cookie policy that stores and sends nothing, which
neither litellm nor the a2a SDK relies on: the SDK's auth interceptor skips
cookie-borne API keys outright.
The constraint was 1.0, so anything above that was silently clamped down.
SCX accepts [0.0, 2.0), verified live against both GLM-5.2 and Qwen3.8
Max: 1.5, 1.99 and 1.999 all return 200, while 2.0 returns 400 with
"Temperature should be in [0.0, 2.0)"
Since the clamp is an inclusive min(), 2.0 cannot be the ceiling or it
would pass through a value the endpoint rejects. 1.99 is the practical
maximum
The clamp test now pins both ends: 2.5 comes back as 1.99, and 1.7 rides
through untouched where it used to be flattened to 1.0
Replaces the five launch models with the two that SCX.ai now leads on.
Both are live on api.scx.ai and both were verified against it for tool
calling, json_object and json_schema output, reasoning, prompt caching,
and, for Qwen3.8 Max, image input
Pricing follows SCX's published USD rates. GLM-5.2 lands at $0.55/M input
and $1.9255/M output, tracking the recent GLM-5.2 market repricing;
Qwen3.8 Max at $1.815/M and $5.4461/M sits under the only other seller of
that model, and is the first Qwen3.8 Max entry in the catalog
Also corrects a metadata bug the removed entries carried: they set
max_tokens equal to max_input_tokens, conflating the context window with
the output cap. Both new entries declare a max_output_tokens of 131072,
which is what the endpoint's own validator enforces
The Add Model placeholder moves to scx-ai/GLM-5.2 now that MiniMax-M2.7
is no longer in the catalog
get_litellm_params returned metadata=None whenever only litellm_metadata was
supplied, which overwrote the fallback function_setup had already applied and
left litellm_params["metadata"] empty. On the /v1/responses
completion-transformation bridge, used by every provider without a native
Responses API config, and on /v1/messages, that discarded the caller's trace
fields a second time after the proxy had promoted them.
Resolve metadata to a copy of litellm_metadata when metadata is empty, guarding
on isinstance because the proxy leaves an unparseable litellm_metadata string in
place and a null metadata would otherwise suppress the backfill and break the
merge. update_from_kwargs copies rather than aliases for the same reason: on
these routes it is handed the caller's provider-bound dict and would otherwise
write user_api_key_auth into it.
* fix(proxy): promote caller metadata trace fields into litellm_metadata
Routes in LITELLM_METADATA_ROUTES keep the caller's metadata as a provider
passthrough field and track proxy state in litellm_metadata, which is the dict
the logging integrations read. The caller's trace_id, session_id, trace_user_id
and trace_metadata therefore never reached any callback on /v1/responses,
/v1/messages, /v1/batches or /v1/files, and mask_input / mask_output were
dropped with them so a caller asking for redaction had their prompt logged in
full.
Promote an explicit allow-list of those fields from the requester_metadata
snapshot into litellm_metadata, never overwriting a value already set so
header-derived ids keep precedence. Trace-mutation controls (existing_trace_id,
update_trace_keys) and trace_public are deliberately excluded: langfuse applies
them to an arbitrary caller-chosen trace with no ownership check. tags is
excluded because per-tag budget enforcement runs earlier, at auth time.
This covers providers with a native Responses API config. Providers reaching
/v1/responses through the chat-completions bridge need the companion change to
get_litellm_params.
* ci: retrigger workflows
* warn at startup when a proxy-wide budget is set but no DB is connected
litellm.max_budget is only enforced via DB-loaded global spend, so a DB-less proxy silently ignores it. Log a one-time startup warning.
Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>
* refactor(proxy): inject max_budget into DB-less budget warning
Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>
* test(proxy): cover DB-less budget warning startup call
Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>
* test(proxy): pin DB-less budget warning call site
Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>
* test(proxy): stabilize budget warning call-site pin
Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>
---------
Co-authored-by: tin <tin@berri.ai>
Co-authored-by: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>
Bare fireworks_ai/<slug> only resolved to accounts/fireworks/models/<slug>,
so Fireworks routers (served at accounts/fireworks/routers/<id>, e.g.
glm-latest and firerouter) could not be reached without passing the full
resource id. Add a shared resolve_fireworks_resource_name helper that maps an
explicit routers/<id> or models/<id> segment to the right resource path, keeps
the existing -fast router heuristic, and defaults bare slugs to models/ for
backward compatibility. Wire it into both the chat and text-completion
transforms, which had drifted (completion lacked router handling entirely)
Add test_abort_upstream_logs_warning_when_aclose_raises: verifies that _abort_upstream swallows and logs any exception raised by the upstream's aclose() method instead of propagating it.
Add test_enqueue_for_client_returns_false_when_already_detached: verifies that _enqueue_for_client returns False immediately without touching the queue when client_detached is already set before the call.
Add test_enqueue_for_
Strip net-new inline # blocks from streaming_iterator.py, the unit test file,
and the live-proxy regression test to comply with the no-new-comments rule.
Add test_async_sse_wrapper_aborts_upstream_when_detached_drain_cap_reached:
verifies that when the detached-drain cap is already full, the pump calls
aclose() on the upstream so the provider stops generating and billing
instead of continuing to stream while we record only the partial prefix.
Also fixes LIT001 (bare dict in AsyncIterator union) by replacing dict
with Mapping[str, object] across all three stream-type annotations, and
adds the required LIT003 reason strings to the three noqa: BLE001 directives.
The relay queue was unbounded, so a client reading a long stream more slowly
than Bedrock produced it let the pump accumulate every pending SSE chunk in
memory, and detached post-disconnect drains had no concurrency bound, so an
authenticated client could open many large streams and read slowly to pin
unbounded worker state.
Bound the relay queue and make the pump apply backpressure while the client is
connected (it blocks on a full queue, racing the disconnect signal), so a slow
reader throttles the upstream read exactly as the old direct yield did. Cap how
many detached drains run at once; over the cap a disconnected pump bills what it
collected instead of draining further. Detached-drain lifetime is otherwise
bounded by the upstream stream/read timeout. Both limits are tunable via env.
The detached pump previously caught every upstream exception (Bedrock read,
decode, provider-response, or chunk-conversion error) and terminated the
client stream normally, masking the original provider exception and its
status so downstream failure handling never ran.
Now, when the upstream fails while the client is still connected, forward the
original exception through the queue so the client-facing generator re-raises
it and the proxy's failure handling (status code, post_call_failure_hook)
runs unchanged. Only when the client has already disconnected, where there is
no one to propagate to and no failure hook will fire, fall back to salvaging
partial spend from the collected chunks.
On the /v1/messages -> bedrock/ invoke streaming path a client disconnect
raises CancelledError inside the httpx socket read, which unwinds the whole
upstream generator chain before any finally can drain it. Bedrock keeps
generating and billing the full response, so spend tracking logged only the
truncated partial the client drained (output tokens ~1-15 vs the real count)
and undercounted against AWS invocation logs.
Move the upstream read into a detached background task that fully drains the
provider stream to its terminal message_delta/message_stop and bills there.
The client-facing generator only relays chunks off a queue, so a disconnect
tears down the relay but not the pump. A client_detached event stops
enqueueing after disconnect so the queue can't grow unbounded.
Map the remaining gateway-documented effort keys: thinking as an alias
for enable_thinking (enable_thinking wins when both are present),
reasoning_budget to an integer reasoning_effort (skipped when thinking
is explicitly off), and low_effort=true to reasoning_effort=low (budget
wins when both are set). guided_json and guided_choice response_format
wrappers now include the name field (response and choice) to match the
gateway wire shape.
The native files and batches routes declare /{provider}/v1/... and their routers are mounted before the passthrough router, so /openai_passthrough/v1/files and /openai_passthrough/v1/batches matched them with provider="openai_passthrough" and 500'd on the LlmProviders lookup instead of reaching openai_proxy_route.
Move the dedicated /openai_passthrough prefix onto its own router mounted ahead of the batches and files routers. /openai/... and every other provider prefix keep their current behavior.
Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>