The resolver placed both operator-configured key headers at the top of its
precedence, but user_api_key_auth only overrides with litellm_key_header_name;
a pass_through_endpoints litellm_user_api_key is checked last. So a request that
authenticated via Authorization while also sending a pass-through header could
have the wrong value chosen, leaving the authenticated Authorization key
forwarded. Order the resolver exactly like get_api_key: override first, built-in
headers next, pass-through header last.
A migration that defines an uncalled row-rewriting routine and elsewhere
references a quoted column, table, index, or constraint sharing the routine's
name was wrongly flagged: the guard restored every double-quoted identifier
before the name search, so a like-named identifier read as a call. Restore only
quoted names that open a call, followed by "(", so a routine invoked through a
quoted identifier is still caught while a like-named non-call identifier stays
masked and no rewrite-free migration is rejected
user_api_key_auth also accepts the caller key from a pass_through_endpoints
entry's headers.litellm_user_api_key, not just litellm_key_header_name. Drop
every operator-configured caller-key header by name and treat them as
top-precedence caller-key sources, so a virtual key sent through one is never
forwarded to Google.
LoggingWorker._ensure_queue nulled self._queue on a loop change, discarding every
pending LoggingTask (each an un-awaited spend-logging coroutine) with no counter and
only a debug log. SDK callers using asyncio.run() per request and mixed sync/async
processes rebind the queue's loop and silently lose spend rows and observability events.
Drain the stale queue and move the pending tasks onto a fresh queue bound to the new
loop, warn with the carried-over count, and keep flush()/join() honest since the queue
is no longer thrown away. Adds a regression test that fills the queue before the loop
change and asserts every task survives and still executes.
The LIT-4761 streaming-classification tests passed only the bring-your-own
Google OAuth token in Authorization and mocked get_litellm_virtual_key, a shape
that cannot authenticate in production. The credential-less filter now resolves
the caller key by auth precedence, so a lone Authorization value reads as the
key and is stripped. Send the virtual key in x-litellm-api-key, matching a real
request, so Authorization is preserved and the classification assertions run.
A migration that defines a row-rewriting routine and calls it as
"backfill"() at the top level slipped past the checker, since masking
blanks double-quoted identifiers before the routine-call search runs, so
the call could not be found by name and the body read as uncalled. mask()
now returns those identifier spans and outside_definition puts them back,
so a call written through a quoted identifier reads as the call it is and
the routine's body gets scanned the same as a bare call
the x-litellm-model upload path returns ids wrapped with
encode_file_id_with_model (litellm:<raw>;model,<m> base64'd). chat
completions + /v1/responses forwarded the wrapped id straight to the
provider, breaking openai (file not found / >64 chars), gemini
(unknown mime), etc. wire up the existing get_original_file_id +
is_model_embedded_id helpers in update_messages_with_model_file_ids
and update_responses_input_with_model_file_ids — falls through after
the managed-files path so existing flows are unchanged. 3 new
regression tests + dem proof len 71 -> 26.
Adds the `gemini_family` bundled template to the auto-router tab, a
heuristic-classifier preset alongside the existing Anthropic and OpenAI
family presets.
Tiers ascend in cost across the Gemini lineup:
SIMPLE gemini-2.5-flash-lite $0.10 / $0.40
MEDIUM gemini-3.1-flash-lite $0.25 / $1.50
COMPLEX gemini-3.7-flash $0.75 / $3.75
REASONING gemini-3.1-pro-preview $2.00 / $12.00
Uses concrete model ids rather than Google's `gemini-*-latest` aliases.
Those aliases hot-swap to the newest release of their variation (stable,
preview or experimental) with only a two-week notice, while their rows in
model_prices_and_context_window.json are pinned at 2.5-generation rates,
so a swap onto a 3.x model would bill at the stale price and silently
undercount auto-router spend. A pin test asserts no tier resolves to a
`-latest` alias and that all four rungs are distinct.
A downstream disconnect mid-relay was recording the chunk whose write never
landed, so replay would hand back a byte the record run never delivered. Append
each chunk after its yield returns, and label the truncation from the generator
close, so the recording holds exactly what the proxy received.
The caching-local, proxy-extras and enterprise-package shards each budget
pytest 20m but cap the whole job at 55m. Setup can consume up to 35m, and
the runner adds 5m of overhead, so the job deadline can preempt pytest
inside its own advertised budget and the shard dies without a test report.
check_workflow_startup_safety enforces that invariant and is currently
failing on litellm_internal_staging, which reds the code-quality job for
every open PR. Raising the three caps to 60m satisfies 20 + 35 + 5.
The filter's own Bearer-only stripping missed the other schemes
user_api_key_auth accepts, so a virtual key echoed as `Authorization: Basic
<key>` alongside a higher-precedence auth header did not match the caller key
and was forwarded to Google. Reuse the auth module's _get_bearer_token so the
comparison strips exactly what authentication does (Bearer / bearer / Basic /
AWS4-HMAC-SHA256), falling back to the raw value for a bare token.
Reject shebangs even when preceded by a UTF-8 BOM or leading whitespace,
and stop misclassifying UTF-8 text that happens to start with the
ASCII-printable magics BZh (bzip2) or dex\\n (Android DEX) as archives
or executables by applying the same UTF-8 carve-out already used for MZ
The DML scan reconstructed a DO/EXECUTE'd literal with undouble, which collapses
each doubled quote to one character and shrinks the text. Every offset after a
collapsed pair then shifted, so a row-rewrite scanned out of the literal reported
an earlier file line and could miss a data-migration-ok marker placed on its real
line. Reconstruct with defuse_escapes instead, turning each doubled quote into a
quote and a space so the pair keeps its two characters and every offset holds,
while a nested -- or /* still stays inside its string.
The credential-less filter derived the caller key only from x-litellm-api-key,
Authorization, and the custom header, but the route authenticates through
Depends(user_api_key_auth), which also accepts the key from x-goog-api-key. A
virtual key sent only in x-goog-api-key therefore authenticated yet was kept as
a preserved upstream header and forwarded to Google. Resolve the caller key by
the same precedence get_api_key uses and value-strip exactly that, so a key in
x-goog-api-key is stripped while a real Google key alongside a higher-precedence
virtual key is preserved.
A reasoning model whose map entry names no effort flag now resolves to None, so
the API omits the field and the dashboard keeps its six-level fallback, and a
deployment counts as catalog-known only when the map supplied its mode, so an
operator writing model_info on an off-map deployment no longer empties the
levels its mapped siblings agree on.
Also drops the ultra level nothing asked for, forwards every level the public
literal names across the chat to Responses bridge, and removes the unreachable
supported_reasoning_efforts validator.
The author does not want fetchAvailableModelsForTeam fanning out a second request,
so it goes back to the single /models call it was before. The team path carries no
capability metadata again, which is what it did prior to this branch, and the effort
options still come from /model_group/info on the non-team path.
The ModelGroupInfo splat let a supported_reasoning_efforts value left in a deployment's
model_info seed the group, so it narrowed the group from whichever deployment was read
first and was silently ignored on every other one. The field is derived from the group's
deployments, so start it unset and let the intersection fill it in.
Also correct two docstring claims that did not match the code: the anthropic chat path
gates xhigh and max on the output_config path only, and the mode signal separates an
unknown deployment from a known non-reasoning one only while that deployment carries no
model_info of its own.
get_model_info answers supports_reasoning None both for a model absent from
the map, which the router registers under a synthesized entry, and for a
mapped model that simply is not a reasoning model. Reading both as "adds no
levels" let one custom deployment wipe every level its mapped siblings agreed
on.
The synthesized entry carries no mode, which every real map entry for a
routable model does, so an unset flag with no mode now resolves to unknown and
never narrows its group. A group that genuinely shares no level still
advertises none, and the dashboard drops the effort control for it instead of
offering levels routing would refuse.
The chat-completions gate only ever owned xhigh. Widening it to max and ultra
made gpt-5.6 answer 400 on requests litellm itself converts to /v1/responses,
where max is valid, because the gate runs before the bridge decision. No map
entry asserts either flag, so the widened gate could only ever reject.
An empty per-group intersection now falls back to the capability-blind level
list in the dashboard, matching what the picker showed before the field
existed, and ModelGroupInfo tolerates whatever shape an operator writes under
supported_reasoning_efforts instead of failing the whole /model_group/info
response.
/v1/chat/completions rejects reasoning_effort=max on every gpt-5.6 snapshot with
"Unsupported value: 'reasoning_effort' does not support 'max' with this model.
Supported values are: 'none', 'low', 'medium', 'high', and 'xhigh'", so the 16
gpt-5.6 entries that asserted supports_max_reasoning_effort had model groups
advertising a level routing would always get a 400 for.
The flag reads as chat-surface capability everywhere else in the map: before
this branch no openai or azure entry carried it at all, only anthropic-family
ones whose chat endpoint really does take max. Dropping it lines the gpt-5.6
advertisement up with what the endpoint accepts and lets the chat gate refuse
the level with our own error instead of forwarding it into a provider 400.
/v1/responses does accept max for gpt-5.6 and is unaffected: the responses
config gates none alone, so the cursor thinking-max variant keeps resolving.
AzureOpenAIGPT5Config raises UnsupportedParamsError on reasoning_effort='none'
unless the model map flags it, while OpenAI never refuses the level, so a single
opt-out polarity advertised none for 61 azure deployments that reject it. Defer
to _supports_factory when the azure flag is absent so the advertised list and the
request gate agree on every azure gpt-5 model in the map.
Also wrap the widened reasoning_effort Literal in main.py, which ruff format
flagged over the line limit.
Enumerating credential-bearing kwargs in RETRY_BREADCRUMB_EXCLUDED_KWARGS is always one
new kwarg behind: it missed top-level extra_headers and provider token fields, which
log_retry still copied into router.previous_models verbatim. Scrub the breadcrumb with
mask_credentials_in_payload instead, so credential-named values are masked at any depth
(extra_headers.authorization, api_key, aws_secret_access_key, vertex_credentials,
azure_ad_token, and future kwargs), and leave the exclusion set to the request payload and
router walk state only.
This hardens the in-memory breadcrumb; it is not a fix for a reproduced SpendLogs leak. The
SpendLogs metadata allowlist and the universal previous_models stripping already keep this
breadcrumb off every persisted surface.
Parametrize the regression test over provider_specific_header, extra_headers, and api_key,
asserting the raw credential value never survives into previous_models for any shape while
the container key still reaches the breadcrumb
The call-detection restore splices into a fixed-position list, so its
.ljust(end - start) holds that length invariant and a test guards it. The
DML-scan recursion instead hands the undoubled literal to a fresh scan_region
as its own region, whose length feeds nothing, so the pad only appends
trailing spaces that shift no keyword and change no reported line. Drop it and
the docstring clause that claimed it kept the offsets landing
Extract a shared _valid_max_results predicate that rejects bools (an int
subclass) and non-positive values, and reuse it from both the connection-mode
request count and the response-side cap so both paths honor the same contract.
The record/replay harness stored a streamed provider response as one
buffered body, so a replayed stream arrived coalesced and the
/v1/messages streaming test could not be edge-wired. Keep each SSE
transfer chunk in the bundle in the order the provider sent it (a new
streamed response shape at BUNDLE_FORMAT_VERSION 4) so replay reproduces
the provider's split points, the recorded usage chunk keeps its
position, and a mid-stream upstream error replays as the same
mid-stream error rather than a clean body.
Resolves LIT-5742
Flatten dict-backed multipart bodies so a scalar list becomes one field with a
tuple value, which httpx emits as a repeated part per element, instead of
collapsing to the last element under dict.update. Nested objects still flatten
to key[subkey] like the OpenAI SDK, and the file-tuple video path is untouched.
The hand-rolled drop set missed Ocp-Apim-Subscription-Key, so a caller
Azure APIM secret in that header was forwarded to Google on the
credential-less branch. Derive the name-drop set from the canonical
SpecialHeaders.litellm_credential_header_names(), minus Authorization and
x-goog-api-key which double as real Google credentials and are value-stripped
instead. New credential headers added there are now dropped automatically.
Uploaded files reaching the RAG ingest path were trusted by client
filename and content-type, so archives and executable scripts were
ingested and malicious content was never screened. Enforce controls at
the upload boundary before the file leaves the proxy:
- classify content by magic bytes and a strict UTF-8 decode, never by
the client filename or content-type
- allowlist PDF and UTF-8 text; reject archives and executables/scripts
- cap upload size (512MB) via a bounded read
- run every accepted upload through a dependency-injected malware
scanner, failing closed on scan error; the default scanner flags the
EICAR test file so the hook is validated end to end
- give accepted uploads a server-generated filename so the client name
never reaches storage
- set Content-Disposition attachment and X-Content-Type-Options nosniff
on vector-store file downloads
The router hop _ageneric_api_call_with_fallbacks canonicalises the passthrough
call type onto litellm_metadata, and the cost callback reads spend attribution
from that bucket while only backfilling user_api_key* keys from metadata. The
helper was building on metadata, so agent_id and user_api_end_user_max_budget
were silently dropped before the callback ever saw them. Build and pass the
attribution under litellm_metadata so every field survives.
log_retry copied every kwarg into the previous_models breadcrumb, so a client's
forwarded Authorization (provider_specific_header) and the deployment api_key /
headers rode along in an in-memory structure whose comment says it reaches spend
logs and logging callbacks. Those values have no diagnostic use in a breadcrumb.
Add provider_specific_header, headers, and api_key to RETRY_BREADCRUMB_EXCLUDED_KWARGS
so the credential is never placed there in the first place. This is defense in depth:
no persisted leak exists today, since the SpendLogs metadata allowlist and every
logging integration already drop previous_models before serialization. Removing the
credential at the source means a future logging path cannot expose it either
user_api_key_auth also authenticates a caller from the operator-configured
general_settings.litellm_key_header_name, reading that header straight off
the request, so a virtual key sent there survived the credential-less Vertex
forwarding filter and reached Google alongside a real bring-your-own
credential. Value-strip every header whose value matches the caller's key
from any accepted source, including that custom header.
Address Greptile review comments and the strict lint budgets:
- read supports_adaptive_thinking from the model cost map instead of
substring-matching the model name, so aliases and newly onboarded
adaptive-only models need no code change
- add tencent/minimax-m3 to the pricing JSON (and backup), which also
fixes cost tracking for the model
- type the thinking/extra_body payloads with ReadOnly TypedDicts
- build the merged extra_body without rebinding or in-place mutation
- send a caller api_key via the Azure api-key header instead of Authorization: Bearer
- cap web_search results to the requested max_results (the tool has no count knob)
- surface a Foundry failed/incomplete response status as a 502 error
- zero the per-query cost in web_search mode; keep the map price for connection mode
- trim the example config to terse env-var pointers