* fix(responses): keep a flagged hosted deployment's prompt cache breakpoint on the chat bridge
Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>
* fix(responses): read the serving provider's row for prompt cache breakpoints on the chat bridge and base_model
* fix(anthropic_cache_control_hook): rename the module-level provider resolver so the recursion check passes
Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>
* test(ui): follow the credential dialog's new name field and auth method id in the federation e2e spec
Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>
---------
Co-authored-by: mateo <mateo@berri.ai>
Co-authored-by: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>
Co-authored-by: mateo-berri <277851410+mateo-berri@users.noreply.github.com>
* test: wait on conditions instead of wall-clock in guardrail parallelism and scope-option tests
Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>
* test: annotate the guardrail barrier locals as Final
Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>
---------
Co-authored-by: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>
* refactor(types): replace Any with proven types in 5 files
Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>
* refactor(types): drop Azure AI Search and OpenRouter image edit seams that reject previously accepted payloads
Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>
---------
Co-authored-by: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>
* fix(mcp): never forward the caller's LiteLLM key to MCP servers
Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>
* refactor(mcp): tidy caller-key scrub
Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>
* test(mcp): avoid master-key scanner false positive
Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>
* fix(mcp): match admission when scrubbing the caller key
Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>
* refactor(mcp): drop header copy churn
Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>
* fix(mcp): normalize bearer variants in caller-key match
Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>
* fix(mcp): inject admission header name and cover provider keys
Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>
* test(mcp): exercise real REST header extraction
Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>
* test(mcp): cover caller-key scrub on the real REST path
Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>
* fix(mcp): exclude gateway admission header from upstream forwarding
* fix(mcp): skip empty admission headers when scrubbing keys
* fix(mcp): use scrubbed server credentials for auth probes
* fix(mcp): match probe credential precedence to upstream calls
* test(mcp): simplify probe header assertion
---------
Co-authored-by: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>
Co-authored-by: Joshua Valluru <326636767+joshua-berri@users.noreply.github.com>
* feat(ui): configure OpenAI workload identity federation from the LLM Credentials and Add Model forms
* fix(ui): let a federated OpenAI credential omit the service account when the proxy env provides it
* fix(ui): clear the federation API Base error when the admin switches the credential to an API key
* fix(proxy): read the UTC clock once per operation in gateway tracking, PTU rollup and Mavvrik export
* refactor(proxy): default the injected clocks to get_utc_datetime instead of three private copies
---------
Co-authored-by: yuneng <yuneng@berri.ai>
* fix(ui): let ModelSelect deselect selections that are no longer offered
A selected model that is no longer served (removed from config, deleted,
or outside the org ceiling) never appeared in the dropdown, so once it sat
past the fifth chip nothing could remove it. List such selections in an
Unavailable group ahead of the offered options so they can be found and
unchecked.
* fix(ui): skip the Unavailable group when no live models were loaded
A failed or empty model list made every selection look unavailable. Only
flag selections as unavailable when there is a live model list to compare
against, and cover ordering with several unavailable selections.
* fix(ui): only suppress the Unavailable group when the model list failed to load
Keying the guard on the context-filtered list hid the group when an
organization ceiling excluded every selection, which brought the dead end
back for team forms. Key it on the proxy model list having loaded.
* fix(ui): keep other selections when removing one beside a special option
Removing an unavailable model while a special option stayed selected
collapsed the whole selection to the special option, dropping the other
saved values. Collapse only when a special option is newly picked. Also
skip the Unavailable group while an org team's model ceiling is unknown,
since the offered list is empty for that reason alone.
* feat(credentials): add credential_alias and make credential_name immutable on PATCH
Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>
* test(ui): query credential rows by role to stay under the no-node-access budget
Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>
* test(proxy): expect credential_alias in load_credential_list dump
Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>
* test(ui): drive the credential SearchSelect by placeholder and option roles
Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>
* fix(proxy): pass the decrypted CredentialItem to update_db_credential during master key rotation
Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>
* fix(ui): trim credential_alias in CredentialModal so whitespace-only input clears the alias
Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>
* feat(credentials): replace credential_alias with display_name and reject edits to config credentials
credential_name stays the immutable reference key. display_name is a nullable, trimmed, max 255
character label set on POST or PATCH (omit keeps it, null clears it, blank is a 400). Reads return
display_name plus an in-memory source tag (db or config), and PATCH or DELETE on a config-defined
credential answers 400 instead of a misleading 404
* feat(ui): show credential display names everywhere and lock config credentials
The credentials table shows the display name with the credential name beneath it, badges config
credentials and disables their edit and delete actions. The edit modal adds an editable display
name next to the read-only credential name and only sends it when it changed. Every credential
picker and reference (add model, model info, models table, vector store form and info) shows the
label while still submitting credential_name
* feat(cli): set display names on credentials and show them in the list
* refactor(credentials): keep the new credential types inside the lint gate budgets
* refactor(cli): print credential command JSON through one helper
* fix(credentials): treat an empty credential_name on PATCH as omitted and pin the 404 for vanished rows
A blank credential_name never renamed anything, so PATCH accepts it again instead of answering 400. New tests pin a trimmed display_name on PATCH, the repository carrying display_name, and a 404 (not the config-owned 400) for a DB credential another worker already deleted. The hydration helper no longer copies display_name, since none of its callers read it, and the CLI update passes display_name straight through because the exactly-one check already makes it None when clearing.
* fix(ui): keep a model's credential when its picker text is emptied and skip the admin-only list for other roles
Clearing the search text in the model edit form's credential picker used to submit null, silently detaching the model's credential on save. None is now the only way to clear it. The models table fetches /credentials only for proxy admins, since everyone else got a 403 on each page load and falls back to the raw name anyway. Also fixes a type error in the re-use credential dialog and adds tests for the display name surviving a provider switch, the 255 character limit, the request payload, and the vector store None choice.
* fix(client): percent-encode the credential name when updating its display name
A name with ? or # was cut short in the URL, so the PATCH landed on a different credential or 404'd.
* fix(credentials): answer 405 with Allow: GET for PATCH and DELETE on config credentials
A config-defined credential exists (GET returns it) but is read-only through the API, which is what 405 Method Not Allowed means. 400 described a malformed request, and 404 would claim the credential does not exist.
* fix(proxy): ignore a display_name set on config credential_list entries so non-string values cannot fail boot
---------
Co-authored-by: yuneng <yuneng@berri.ai>
Co-authored-by: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>
Restrict 26 build_and_test jobs to main with job-level branch filters, delete litellm_router_unit_testing in favor of a router-unit-tests GHA shard, and narrow guardrails_testing to the license-dependent test while the rest of tests/guardrails_tests runs in a guardrails-tests GHA shard
Co-authored-by: yuneng <yuneng@berri.ai>
* fix(anthropic): count a leading system run through count_tokens' system parameter
Anthropic's count_tokens rejects role "system" at the head of messages, so
/v1/responses/input_tokens with instructions, and /v1/messages/count_tokens
with a system-role message, fell back to the local tokenizer. The shared
Anthropic count_tokens transformation now lifts the leading run of system
messages into the top-level system parameter, the way the chat path sends
it, after any system the caller set. Anthropic direct, Azure AI Anthropic,
and Bedrock Mantle share that transformation.
* test(integration): cover the count_tokens leading-system lift across Anthropic, Azure AI and Bedrock Mantle
---------
Co-authored-by: mateo-berri <277851410+mateo-berri@users.noreply.github.com>
* ci: restructure lint, unit and smoke workflows
* ci: preserve source formatting in tier workflows
* ci: retire per-shard coverage flags and tighten the tier workflow guards
* test(ci): freeze the workflow startup safety models
* ci: keep the current required check names running until the ruleset moves to the tier collectors
---------
Co-authored-by: yuneng <yuneng@berri.ai>
* fix(token_counter): price a base64 PDF document per page instead of as one image
A `document` or `file` block carrying inline PDF bytes was priced like a single
image (85 tokens), so when a provider's count-tokens endpoint rejected the
model (Bedrock Opus) the local fallback answered 116 for a 12-page PDF the
provider then billed at 35941 input tokens. The counter now reads the PDF with
pypdf and prices each page as its extracted text plus the image Anthropic
renders it to (1568 px long edge, 1.15 MP, 750 pixels per token), falling back
to the old image pricing when pypdf is missing or the bytes are not a readable PDF.
* refactor(token_counter): count PDF pages as they are read
Sum each page's text and rendered-image tokens straight from the pypdf
reader instead of materializing a page list first, keep the fallback
to image pricing atomic when a page cannot be read, and annotate the
new tests' locals as Final
* test(integration): audit cells for page-priced PDF documents in count_tokens, pre-call checks and spend
* test(integration): release the held peer when the concurrent budget wait times out
---------
Co-authored-by: mateo-berri <277851410+mateo-berri@users.noreply.github.com>
* fix(router): match deployment pricing ids against the cost map only within the deployment's provider
A deployment model_info.id that equals another provider's catalog key
(e.g. baseten/zai-org/glm-5.2 on an openai-compatible api_base) was merged
into that provider's built-in row by register_model, keeping
litellm_provider=baseten on the entry. _check_provider_match then rejected
the row at request time and the deployment billed $0. register_model now
takes a keyword-only custom_llm_provider used to scope
_get_builtin_model_info_for_registration and
_resolve_builtin_model_cost_entry, and the router passes the deployment's
provider through. The stored entry stays provider-less for
non-colliding ids.
Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>
* refactor(utils): pass custom_llm_provider straight to the registration lookup
Drop the per-entry lookup-provider expression in register_model per review:
the kwarg alone scopes _get_builtin_model_info_for_registration, and
_resolve_builtin_model_cost_entry keeps its main signature and caller.
_register_custom_pricing_for_request passes the provider through so
router-originated per-request registrations get the same scoped lookup.
Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>
* test(utils): drop docstring and simplify set restore in per-request collision test
Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>
* test(router): type the colliding-id deployment fixture
Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>
---------
Co-authored-by: kerry <kerry@berri.ai>
Co-authored-by: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>
* fix(cli): pin pi compat flags for gateway models
Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>
* fix(cli): annotate pi compatibility field
Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>
* fix(cli): clarify pi compatibility suppression
Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>
* fix(cli): drop the stale mutable-ok suppression on the pi compat block
---------
Co-authored-by: jesus <jesus@berri.ai>
Co-authored-by: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>
Co-authored-by: mateo-berri <277851410+mateo-berri@users.noreply.github.com>
* fix(router): drop the encrypted reasoning a fallback hop's target cannot decrypt
An order-based or configured fallback hop replayed the failed provider's
encrypted reasoning items to the next deployment, which answered 400
(Bedrock Mantle: invalid encrypted reasoning; OpenAI:
invalid_encrypted_content), so every multi-turn Responses fallback for
Codex-style clients failed. The hop now drops the encrypted reasoning
its target cannot decrypt and keeps each item's readable summary. With
encrypted_content_affinity on, the pin narrows to the hop's target
order instead of emptying it, so the hop reaches the next order
instead of failing with no deployments available.
* fix(router): keep encrypted reasoning a same-boundary fallback hop can decrypt
Unmarked encrypted reasoning on a hop is attributed to the deployment that just
failed, read from the retry breadcrumb, so a hop to a deployment on the same
api_base and api_key keeps it and a cross-provider hop still drops it. The hop
tests script the upstream at the httpx boundary instead of doubling the handler,
and the router coverage script lists the two hop helpers with their tests
* fix(router): read the hop's failed deployment from its own metadata bucket and carry it into the Responses mid-stream snapshot
* test(router): use a real Router without the origin deployment in the hop strip test
* test(router): cover the hop strip when no failed deployment is known
* fix(proxy): drop the router's fallback hop state keys from the client body
* fix(proxy): keep a request's max_fallbacks cap, drop only the hop state keys
* test(integration): audit cells for the fallback hop encrypted reasoning strip
Forty-five checked-in cells under tests/integration/routing prove the hop strips the previous deployment's encrypted reasoning on /v1/responses, /v1/chat/completions and /v1/messages (httpx, OpenAI and Anthropic SDKs, sync and async, streaming and not), that the affinity pin yields to the hop, that a client-sent fallback_depth, _target_order and attempted_targets never move or strip a request, and that a concurrent burst, an order-1 outage and a killed worker keep every request stripped and logged once. Every call goes through a lane pinned to one worker that already lists the deployments it needs, because the peer worker learns a /model/new row through the config-sync resync up to sixteen seconds later
---------
Co-authored-by: mateo-berri <277851410+mateo-berri@users.noreply.github.com>
* feat(guardrails): support logging_only mode for the Akto guardrail
* test(guardrails): wait for the spend row before reading Akto calls and cover mixed modes
* test(guardrails): cover Akto logging_only on MCP tool calls
* test(guardrails): type the Akto logging_only unit tests and inject the HTTP handler
* test(guardrails): cover failing Akto replies, provider failures and a mixed outage burst under logging_only
* test(guardrails): cover an unreachable Akto with fail_open under logging_only
* ci(unit): add unit passed collector job and fold proxy-db shards into test-unit.yml
* test(ci): cover the unit passed gate's success, failure, cancelled and skipped results
* test(ci): run the unit passed gate test from the code quality workflow
---------
Co-authored-by: yuneng <yuneng@berri.ai>
* refactor(python-bridge): ship a signature base and read the resolved call
NativeCall carries base (positionals by name plus signature defaults) instead
of the fully bound dict. resolved lays kwargs over base, which is what bound
held, so every pre-hook read and every route host keeps seeing the same values.
Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
* fix(bedrock): build the transcription NativeCall with an empty base
Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
* test(dispatch): read the resolved call instead of bound
Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
* refactor(host-python): rename effective to effective_py_args and note the shallow copy
Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>
---------
Co-authored-by: Claude Opus 5.5 <noreply@anthropic.com>
Co-authored-by: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>
#44800 re-raises the provider's own error from the chat adapter, so a bridged /v1/messages stream's error frame now reads litellm.RateLimitError without the litellm.MidStreamFallbackError: prefix that #44989's tests pinned shortly before #44800 merged. The three tests now expect the provider error text once and no sentinel, matching the chat_limited case and every other assertion in the two files
Co-authored-by: mateo-berri <277851410+mateo-berri@users.noreply.github.com>
* fix(bedrock): split the <reasoning> tag for gpt-oss only on native Chat Completions
The native Chat Completions route moved a leading <reasoning>...</reasoning>
block into reasoning_content for every model, while only gpt-oss writes its
reasoning inline. A GPT 5.6 or Grok answer that starts with a literal
<reasoning> tag lost that text, streaming and non-streaming alike. Both paths
now split only when the model id is gpt-oss.
* fix(bedrock): drive the inline reasoning split from a cost-map flag
---------
Co-authored-by: mateo-berri <277851410+mateo-berri@users.noreply.github.com>
* test: replace 61 live logging and otel tests with offline unit and integration coverage
* test: restore datadog formatting, deliver datadog logs and redis failures over the wire, assert full router hook payloads, tighten stream usage and otel checks
Restores the pre-existing test_datadog.py lines the branch had reflowed. Datadog success, failure and redis-failure replacements now assert the gzip body posted to the intake, with redis failing through a real cache on a closed local port. Router hook sequence recorder checks the legacy field types and asserts concrete payloads, exact streaming and fallback sequences. Stream usage asserts the default include_usage request body and the redacted messages value. Otel asserts response id and token counts.
* test: assert budget envelope figures, guardrail inspection and exact prometheus samples per request
* test: isolate the prometheus latency test on its own deployment so counts do not depend on order
* test: cover timed slack delivery, keep the redis failure test off the network, and make new payload types read-only
* test: drain the redis test's logging and scope otel span checks to the test's own trace
* test: drive the periodic slack flush without a wall-clock interval
* test: scope router hook events to the test, script the db clock, split a nested comprehension
---------
Co-authored-by: yuneng <yuneng@berri.ai>
* test(ui-e2e): update usage page selectors after the #45221 redesign
* test(ui-e2e): drop the top keys locator comment
---------
Co-authored-by: yuneng <yuneng@berri.ai>
* fix(otel): emit OpenInference tool calls and metadata on Arize OTel v2 spans
Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>
* fix(otel): shed OpenInference output tool calls individually under the span attribute budget
Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>
* test(otel): avoid mutation in Arize OTel v2 integration helpers
Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>
* test(otel): audit Arize OTel v2 OpenInference spans across endpoints, modes and chaos
Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>
* test(otel): wait for each exported span before the next request in Arize OTel v2 audit tests
* test(otel): make Arize OTel v2 audit absence and outage checks deterministic
Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>
* test(otel): collect Arize OTel v2 outage spans through an in-order sentinel
Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>
* test(otel): add skipped BUG cells for pre-existing Arize OTel v2 gaps
Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>
* fix(otel): repair Arize OTel v2 regressions from #43698 (linear fit, metadata slot, repr tool args) (#44488)
* fix(otel): keep Arize OTel v2 regressions in check — O(n) fit, metadata slot, repr tool args
Three regressions from #43698's OpenInference tool-call/metadata emission:
1. Metadata evicted indexed message attributes: the new `metadata` key
competed for the 128-attribute span budget, and the fit sheds whole
message groups BEFORE `span.set_attribute`, so the SDK's dropped
counter stayed 0 — an invisible eviction (live A/B: input-message
attributes 86 -> 84). Two-part fix: the fit pins `metadata` behind
every message group (it sheds only once all indexed messages are
gone), and the span budget no longer charges pre-set attributes the
mappers overwrite in place — a boundary-opened LLM span already
carries keys like `gen_ai.request.model`, so the old accounting
reserved slots the fit could never spend. Live: 86 input-message
attributes with the metadata attribute riding alongside.
2. Quadratic shed on long prompts: `_message_shed_groups` rescanned the
full group map once per message (measured on a real acompletion:
0.032/0.128/0.478/1.910s at 1000/2000/4000/8000 messages vs
0.007/0.010/0.022/0.029s at base). Index the tool-call groups once by
(family, message index): the fit is linear again (0.008/0.007/0.014/
0.031s, same rig).
3. Malformed Python tool arguments lost the whole span: provider
adapters and `model_construct` responses hand over raw objects, and
`json.dumps` raises on tuple-keyed dicts (TypeError) and cycles
(ValueError) before the span is exported. Serialize with a repr
fallback; both cases now export with a readable arguments attribute.
The attribute budget change affects every boundary-opened LLM-call span
(strictly more attributes retained, never fewer); the mapper changes
only touch the OpenInference vocabulary.
* refactor(otel): build the tool-call group index in one shot
Review follow-up: the dict.setdefault/append seeding in
_tool_call_groups_by_message violated the no-mutation coding convention
(AGENTS.md: build values in one shot with comprehensions or generators
wrapped in tuple()/MappingProxyType()). Rebuild it as a sorted groupby
comprehension; randomized parity harness confirms the shed order is
byte-identical to the seeded version (400 trials).
Also pin the overflow corner Greptile asked about: a pre-set
indexed-message key the fit sheds keeps its earlier value in place, so
the span total can never exceed the SDK limit (new emitter test).
* fix(otel): key the groupby with an explicit tuple to keep basedpyright at budget
The slice-keyed groupby (group[:2]) widened the key to tuple[str | int],
adding one reportGeneralTypeIssues over the codebase ceiling. Key by the
explicit (family, message index) pair instead; shed order unchanged
(300-trial randomized parity harness).
* fix(otel): read pre-set span keys through a helper typed for both runtime shapes
The SDK annotates ReadableSpan.attributes as a Mapping, but an ended span
hands back a tuple of pairs, so the inline isinstance branch narrowed to
Never and pushed reportGeneralTypeIssues one over the codebase ceiling.
Extract _carried_keys with the runtime union declared on the parameter;
behavior unchanged.
* test(otel): drop the wall-clock bound from the long-prompt attribute fit test
Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>
* fix(otel): annotate to_openai_dict as Mapping after the rebase onto main
Main's type-discipline budget tightened since the branch point; the plain
dict return annotation is the one violation the rebased branch adds.
Callers only serialize the result, so the read-only view is accurate.
* chore(otel): drop the restating docstring from to_openai_dict
---------
Co-authored-by: yucheng <yucheng@berri.ai>
Co-authored-by: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>
* fix(proxy): accept team-scoped models by their public name on POST /fallback
create_fallback validated the primary and fallback models against the
router's stored model names only, so a model created through POST
/model/new with model_info.team_id was rejected with a 404 unless the
caller used the generated model_name_<team_id>_<uuid> name. The
endpoint now also accepts the team public model names, which request
time fallback matching already keys on, and lists both kinds of names
in the 404's available_models.
* fix(proxy): read fallback rules fresh and clear the config cache after a fallback write
A second POST or DELETE /fallback within the 60 s config cache TTL started from
a cached copy of router_settings and dropped every rule stored since that copy
was taken, by any instance. Both endpoints now evict the cached row before the
read and invalidate it after the upsert
* test(proxy): type the fallback endpoint tests and prove a team request fails over by public name
* fix(proxy): keep fallback writes working when the Redis config cache is down and type the stored settings read
* fix(proxy): keep the non-standard fallback shapes the router accepts on writes and resolve the config cache at call time
* test(proxy): mark the config cache outage test's result as Final
* fix(proxy): replace a same-key fallback rule in place and read a null rule list as empty
* test(proxy): let the stored router settings fixture carry a null rule list
* test(proxy): audit cells for fallback rules by team public name
* test(proxy): delete the fallback rules the audit cells save
---------
Co-authored-by: mateo-berri <277851410+mateo-berri@users.noreply.github.com>
* fix(ui): link model access group chips to the access group filter
Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>
* style(ui): format access group chip link changes
Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>
* test(ui): mock access group hook in affected suites
Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>
* fix(ui): keep access group chips unlinked until the group lookup resolves
Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>
* fix(ui): keep cached access group names when a refetch fails
Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>
---------
Co-authored-by: nate <nate@berri.ai>
Co-authored-by: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>
* fix(ui): show the Add Model picker once the model catalog loads after a provider is picked
* fix(ui): show the Add Model picker when a name typed before the catalog loaded was cleared
* feat(spend_logs): configure which metadata fields are stored in LiteLLM_SpendLogs
Adds general_settings.spend_logs_metadata_fields with mutually exclusive include and exclude lists. The filter runs on a copy of the row right before it is queued for Postgres, so daily spend rollups, budgets and callbacks still see every key.
Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>
* fix(spend_logs): keep excluded auto-router savings keys out of published spend log metadata
Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>
* test(spend_logs): cover metadata retention across endpoints, failures, cache hits, batches, runtime updates and auto-router publication
Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>
* chore(ui): regenerate schema.d.ts for spend_logs_metadata_fields
Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>
* fix(spend_logs): filter metadata at DB write so guardrail usage sees full rows
Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>
* test(spend_logs): poll guardrail daily metrics instead of reading once
Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>
* test(spend_logs): drop timeout comment
Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>
* fix(spend_logs): read spend_logs_metadata_fields through typed general settings
Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>
---------
Co-authored-by: mrinal <mrinal@berri.ai>
Co-authored-by: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>