Profiling the sidecar-enabled gateway at 700 rps showed ~2.4% of all samples
in get_model_group_info called per request from budget reservation, plus
get_deployment_model_info for tiered pricing tables. Both are read-only lookups
over the model list, so serve them from the Router's lru caches and clear the
deployment cache alongside the group cache when the model list changes.
The deployment-info cache is a per-router lru_cache built in __init__ rather
than a class-level decorated method, so it does not pin Router instances in a
process-wide cache and is dropped with the router.
A price data reload replaces litellm.model_cost without touching model_list, so
the reload replay also clears both caches; otherwise reservation would keep
pricing against the old catalog until an unrelated model-list change.
Co-authored-by: yassin <yassin@berri.ai>
* feat(mock): report admission-time input token count in mock_response usage
Mock completions always reported prompt_tokens=10, so spend tracking, TPM metrics, budgets and the tokens-per-minute autoscaling signal saw 10 tokens for a 100k-token request. Budget reservation now carries the admission-time input token count in the reservation record, and mock_completion reads it back so mock traffic exercises the same spend and TPM paths as real traffic without any extra tokenizer work.
* fix(mock): keep a zero admission input token count instead of falling back to 10
---------
Co-authored-by: yassin <yassin@berri.ai>
Every Bedrock and SageMaker call site hand-copied the same nine aws_* kwargs
into BaseAWSLLM.get_credentials, so each new auth param has to be threaded
into a dozen places and any site that misses one silently assumes the role
with the wrong parameters.
Introduce AwsAuthParams, a frozen pydantic model whose fields are exactly the
credential-shaped params get_credentials accepts, plus resolve_credentials on
BaseAWSLLM and pop_aws_auth_params for the call sites that must strip the keys
out of optional_params. Deriving AWS_AUTH_PARAM_KEYS from the model's fields
means the mirror list in common_utils can no longer drift from the struct.
Behavior is unchanged: the same values reach STS from the same call sites.
Dropping any one field from the resolver fails one of the new tests.
Claude-Session: https://claude.ai/code/session_01E6zsK1DBcXfbetkgX86fw2
A deployment with base_model set was priced from the region's own row once
completion_cost forwarded the response region into cost_per_token, which now
strips the provider prefix and finds bedrock/<region>/<base_model>. Explicit
pricing (base_model or custom pricing) suppresses the region for cost_per_token
the same way _select_model_name_for_cost_calc already does
Strengthens the serverless test to assert ResourceType is signed rather
than merely present, ties _uses_tls to the redis-py kwarg coercion so the
two cannot drift, locks the stripped-kwarg name tuple to the test's
expectations, and adds "off" and "True" sentinel flag values.
Renames the provider builder's parameter to redis_settings.
The ElastiCache IAM gate read `aws_iam_auth` and `ssl` with a helper that
only accepted the literal string "true", while the kwarg coercion that runs
later accepts "true", "1" and "yes". Type coercion happens after the gate, so
`REDIS_AWS_IAM_AUTH=1` silently skipped IAM auth and `REDIS_SSL=1` made the
"requires TLS" check fail closed on a connection that was in fact TLS. Both
helpers now share `_str_to_bool`.
AWS signs serverless cache tokens with an extra `ResourceType=ServerlessCache`
query parameter, so tokens minted for a serverless cache were rejected. Adds an
`aws_iam_serverless` setting (`REDIS_AWS_IAM_SERVERLESS`) that puts the
parameter into the signed URL, and lowercases the cache name because
ElastiCache lowercases it at creation time.
get and post already take follow_redirects. put built the request and sent it
with the client default, so a caller uploading to a URL it did not choose had
no way to refuse a redirect. Same plumbing as the other two methods.