Profiling the sidecar-enabled gateway at 700 rps showed ~2.4% of all samples
in get_model_group_info called per request from budget reservation, plus
get_deployment_model_info for tiered pricing tables. Both are read-only lookups
over the model list, so serve them from the Router's lru caches and clear the
deployment cache alongside the group cache when the model list changes.
The deployment-info cache is a per-router lru_cache built in __init__ rather
than a class-level decorated method, so it does not pin Router instances in a
process-wide cache and is dropped with the router.
A price data reload replaces litellm.model_cost without touching model_list, so
the reload replay also clears both caches; otherwise reservation would keep
pricing against the old catalog until an unrelated model-list change.
Co-authored-by: yassin <yassin@berri.ai>
* feat(mock): report admission-time input token count in mock_response usage
Mock completions always reported prompt_tokens=10, so spend tracking, TPM metrics, budgets and the tokens-per-minute autoscaling signal saw 10 tokens for a 100k-token request. Budget reservation now carries the admission-time input token count in the reservation record, and mock_completion reads it back so mock traffic exercises the same spend and TPM paths as real traffic without any extra tokenizer work.
* fix(mock): keep a zero admission input token count instead of falling back to 10
---------
Co-authored-by: yassin <yassin@berri.ai>
The squash of #40512 onto a base that already carried #40514 left the Partial call on the metadata pre-read error path only, so a rejected /key/update again persisted the planned values into state and TestResourceKeyUpdateFailureKeepsPriorState fails on the default branch.
Co-authored-by: yassin <yassin@berri.ai>
Co-authored-by: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>
A deployment with base_model set was priced from the region's own row once
completion_cost forwarded the response region into cost_per_token, which now
strips the provider prefix and finds bedrock/<region>/<base_model>. Explicit
pricing (base_model or custom pricing) suppresses the region for cost_per_token
the same way _select_model_name_for_cost_calc already does
Strengthens the serverless test to assert ResourceType is signed rather
than merely present, ties _uses_tls to the redis-py kwarg coercion so the
two cannot drift, locks the stripped-kwarg name tuple to the test's
expectations, and adds "off" and "True" sentinel flag values.
Renames the provider builder's parameter to redis_settings.
The ElastiCache IAM gate read `aws_iam_auth` and `ssl` with a helper that
only accepted the literal string "true", while the kwarg coercion that runs
later accepts "true", "1" and "yes". Type coercion happens after the gate, so
`REDIS_AWS_IAM_AUTH=1` silently skipped IAM auth and `REDIS_SSL=1` made the
"requires TLS" check fail closed on a connection that was in fact TLS. Both
helpers now share `_str_to_bool`.
AWS signs serverless cache tokens with an extra `ResourceType=ServerlessCache`
query parameter, so tokens minted for a serverless cache were rejected. Adds an
`aws_iam_serverless` setting (`REDIS_AWS_IAM_SERVERLESS`) that puts the
parameter into the signed URL, and lowercases the cache name because
ElastiCache lowercases it at creation time.
The comment documents that the raw Azure client id, tenant id and secret
are deliberately kept off the connect function, so this branch should
never have dropped it
Deferred annotation evaluation keeps the type-checking-only botocore
import off the runtime path, so the alias only reintroduced typing.Any,
which the strict ruff budget now bans
get and post already take follow_redirects. put built the request and sent it
with the client default, so a caller uploading to a URL it did not choose had
no way to refuse a redirect. Same plumbing as the other two methods.
* feat(prometheus): bucket latency by input sequence length
* style: format startup resolver call
* fix(prometheus): handle unknown input lengths
* fix(prometheus): preserve disabled custom input length labels
* test(prometheus): seed startup snapshot in mocked logger test
* fix(proxy): preserve database setting types during startup
* fix(prometheus): distinguish missing usage and preserve config persistence
Keep quoted database-storage config values intact for legacy persistence readers while using a local boolean for early callback discovery. Distinguish absent provider usage from an explicitly reported zero when labeling latency metrics.
* fix(proxy): normalize input length flag from secret managers
* fix(prometheus): isolate input buckets and preserve missing usage
Keep built-in buckets on latency histograms, preserve unrelated custom labels, and classify raw incomplete usage and upstream total-only headers as unknown. Cover count conservation, failure callbacks, explicit zero, startup snapshots, and direct caller compatibility. Drop earlier branch budget changes.
* fix(proxy): defer Prometheus alerting until stored settings load
Reuse successful startup storage resolution and preserve callback deduplication across alerting reloads.
* test(prometheus): restore input length flag between tests
Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>
* Revert "test(prometheus): restore input length flag between tests"
This reverts commit f302e0b7bd.
* refactor(prometheus): make input length flag config/env only
Drop the Admin UI General Settings row, the safe DB override entry, the startup reorder that loaded DB litellm_settings before Prometheus callbacks, and the alerting-only Prometheus path. The flag now behaves like prometheus_emit_stream_label: litellm_settings in config.yaml or an os.environ reference, applied on restart.
Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>
---------
Co-authored-by: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>