/health/test_connection returns "Missing credentials. Please pass one of
api_key, azure_ad_token, ..." for a perfectly healthy configured model, as
soon as the request also carries any custom-pricing field.
_config_base_for_health_check strips the configured connection fields when
the request sets one of _BANNED_REQUEST_BODY_PARAMS, on the reasoning that
such a request describes a connection of its own. That tuple ends with
every CustomPricingLiteLLMParams field, which is correct for the request-body
check it was built for — those fields are banned because they poison the
shared model-cost registry — but they are not connection parameters. A test
that names a configured model and its price therefore has the configuration
emptied out from under it, litellm_credential_name included, and the probe
fails with no credential at all.
It lands on the Admin UI's Add Model wizard, where "Test Connection" sits
next to the pricing fields that a model missing from the cost map has to
have, so the two are filled in together and the button reports a working
deployment as broken.
Splits the connection-relevant subset out as
_CONNECTION_OVERRIDE_REQUEST_PARAMS — the same tuple minus the pricing
fields — and gates the health-check merge on that. The request-body check is
unchanged, so pricing fields stay banned there.
Measured against a configured azure deployment: {"model": ...} succeeds,
{"model": ..., "input_cost_per_token": 1e-9} failed and now succeeds, and
{"api_base": ..., "input_cost_per_token": 1e-9} still drops the configured
credentials.
Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
PR #41343 and PR #41112 both added cache_read_input_token_cost to the six amazon.nova-{micro,lite,pro}-v1:0 and us.amazon.nova-* entries, one at the top of each entry and one at the bottom. The merge kept both, so every PR now fails test_price_map_has_no_duplicate_keys. Both copies carried the same value, so this only removes the trailing duplicate in both price files
A budget bypass that ships off by default stays open for every deployment
that does not know to look for the flag, so `enforce_fallback_budget` now
defaults to true and `general_settings.enforce_fallback_budget: false` is
the opt-out for anyone who wants the old unguarded behaviour back.
BREAKING CHANGE: a paid fallback target is now refused for callers who are
over their key or user `max_budget`. Deployments relying on fallbacks to
keep serving over-budget callers must set enforce_fallback_budget: false.
Delete litellm/ocr/input.py and the native _ocr_file_document, _ocr_upload_document
and _ocr_mime_type helpers. File documents now project to a typed OcrDocumentInput
and the core lifecycle reads local paths, encodes bytes and asks the host to read
file-like objects through a ReadDocument operation before the provider request
Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>
The request path used get_end_user_object, which falls back to a database
lookup on a cache miss. Read the LiteLLM_EndUserTable row auth already cached
instead, with the default budget already attached, and leave misses to the
periodic refresh
Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>
Without a master key the auth layer echoes whatever key the caller presented as the authenticated key, so the passthrough's strip-by-value matched the caller's own Anthropic key and dropped it: a bring-your-own-key request that returned 200 on main answered 401 telling the caller to send the key they had just sent. Only the auth module's own no-auth dev-mode definition, shared through is_no_auth_dev_mode, decides that nothing was authenticated, and only when no custom auth is installed; JWTs, OAuth2 tokens, and custom-auth credentials are still stripped there. The sk- prefix heuristic goes with it.
The Vertex credential-less test now sets a master key, since a virtual key can only authenticate under one: the auth layer returns before any key lookup when the master key is unset.
The metric attribute filter now removes every attribute whose value is
None, and the content and inference-details events pass their attributes
through drop_none before emitting, so a call with no provider label or no
model name never hands the OTLP exporter a NoneType attribute. This closes
the gen_ai.request.model report on #36759 the same way the gen_ai.system
one was closed, and the regression tests cover both keys.
A stream counted before its usage is known increments TPM by zero, so the
worker that served it never refreshed its local TPM value from Redis and
the first byte headers reported the token count another worker had already
consumed. Both pipeline operations now always run, matching the pre-change
callback, so the returned values refresh both worker local keys
Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>