The merge with litellm_internal_staging resolved the test file by taking this
branch's copy whole, which discarded the two tests staging had appended to the
same end-of-file region:
test_delete_evicts_cache_after_row_is_gone
test_update_evicts_old_and_new_cache_keys_after_write
Both sides only appended, so the conflict was additive and nothing had to be
chosen between them. The eviction code those tests cover did survive in
jwt_key_mapping_endpoints.py, so the branch was shipping it untested.
Appending staging's block restores them alongside the nine resolver tests here.
The file is now a superset of both sides, verified line by line, and the test
quality count stays at the baseline of 28 because staging's block carries its
own test-quality-ok suppressions.
A JWT key mapping can now name its virtual key by the SHA-256 hash the proxy
already stores, instead of only by the plaintext key.
litellm_key makes its generated key write-only so raw keys stay out of Terraform
state, and write-only attributes cannot be referenced at all, so the natural
wiring fails while planning, in every apply ordering:
Error: Missing required argument
with litellm_jwt_key_mapping.example
key = litellm_key.example.key
The argument "key" is required, but no definition was found.
The only way out today is supplying the plaintext from a variable or a secret
manager, which means the mapped key cannot be one the proxy generated and the
configuration has to carry a credential. The value the mapping stores is
hash_token(key), which is the same hash litellm_key already exports as
token_id, and a hash is not a credential, so accepting it closes the gap:
resource "litellm_jwt_key_mapping" "service" {
jwt_claim_name = "client_id"
jwt_claim_value = "reporting-service"
token_id = litellm_key.service.token_id
}
CreateJWTKeyMappingRequest and UpdateJWTKeyMappingRequest gain an optional
token. Create requires exactly one of key or token, update accepts at most one,
and omitting both still leaves the mapped key alone. A supplied token must be 64
lowercase hex characters, because hash_token() hashes unconditionally and a
plaintext key sent as token would be stored as a hash of a hash, then silently
match nothing at auth time. Both rejections are 400s raised before the row is
written.
On the provider side, key becomes Optional with ExactlyOneOf{key, token_id} and
token_id is added next to it. token_id is not marked sensitive since a hash is
not a credential, both fields are omitempty on the wire so the proxy receives
only the one that was configured, and a failed update reverts token_id for the
same reason it already reverts key.
key keeps working unchanged and existing state is untouched. The only change to
it is Required to Optional, which no existing configuration can violate.
get_llm_provider() and Router._add_deployment() only knew the built-in
provider_list and JSON providers, so a provider registered through
litellm.custom_provider_map was rejected until custom_llm_setup() had
run inside the first completion() call. Both now check the map directly.
Resolves LIT-1742
Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>
The Vertex passthrough route resolved a router deployment only to rewrite the
upstream URL and dropped its model_info, so the standard logging payload and
the Prometheus litellm_deployment_success_responses_total counter carried
model_id="". Carry the deployment's model_info through request.state into the
passthrough logging metadata, where it overrides any client-supplied model_info.
Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>
_handle_clientside_credential registered the per-request Deployment it built for
a client-supplied api_key/api_base via upsert_deployment, which added it to
self.model_list under the shared model_name. That made a request-scoped
credential a permanent, load-balanced deployment that any later caller of the
same model group could be routed onto, reaching the provider with someone
else's forwarded credential.
The per-request Deployment still gets its own stable id for cooldown and
logging identity; it is just never registered with the router.
Resolves LIT-7811
The failure logger skips fallback hops (has_logged_async_failure is already set), so
model_call_details.end_time still belongs to the previous hop and predates this hop's
api_call_start_time. The fallback cooldown guard measured a negative elapsed time and
cooled down deployments for caller-set timeouts.
Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>
client_side_timeout records that the caller configured a timeout, not that
the timeout fired. A 408 the provider returns before that deadline is a
deployment failure and must still count toward cooldown.
Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>
A 408 produced by a timeout the caller set (a timeout body field or an
x-litellm-timeout header, which the proxy marks as client_side_timeout)
says nothing about the deployment's health, yet the router's primary
failure callback counted it toward allowed_fails and cooled the
deployment down. The fallback path already skipped it.
The marker never reached that callback because get_litellm_params drops
kwargs outside OPTIONAL_KWARGS_KEYS, so it is listed there now, and
deployment_callback_on_failure returns before the failure counter when
is_caller_timeout_408 holds. A 408 from a timeout the deployment or the
provider set still counts and still cools the deployment down.
budget_duration and allowed_models ride on /team/member_add. tpm_limit and rpm_limit are sent through /team/member_update, the only endpoint that accepts them. Removing any of the four from config sends an explicit clear (null, or an empty list for allowed_models) since member_update is a merge-patch. The resource ID is set before the post-add limits call so a failure there taints the resource instead of orphaning the memberships
Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>