* fix(health): probe test_connection with the credential the request names
/health/test_connection matches the request's model string against the
configured deployments and merges the match's litellm_params underneath the
request. A request that named a stored credential but no key of its own
still satisfied the "request sets no connection fields" test, so it inherited
the matched deployment's api_key and api_base, and load_credentials_from_list
then skipped the named credential because api_key was already set.
A wildcard route covering the model is enough to match, so the Add Model
page's Test Connect probed with an unrelated deployment's key while echoing
back the credential that was selected.
Naming a credential the configuration does not name now withholds the
configuration's credential fields, the same set already withheld from a
request that supplies its own endpoint. Naming no credential still inherits
them, as documented.
* test(health): drop test docstrings that restate their own names
* test(health): assert the credential probe on the wire, not on the call args
The connection-test regressions patched litellm.ahealth_check and read the
params handed to it. Driving the endpoint through the app with respx faking
the upstream instead lets the real credential resolution run, so the tests
assert the key and host that actually go out, which is what the bug was about.
It also drops three of the five patched proxy internals; the two that are left
are proxy-global wiring with no injection seam, the same ones the image_edit
connection test already has to reach for.
* chore(ui): regenerate schema.d.ts for the test_connection docs change
An open classifier circuit routed through the ordinary heuristic or
classifier_fallback path, and both causes are pin-worthy, so a session
whose turn landed on the cooldown fallback held that model for the whole
session_affinity TTL and never reclassified after the breaker closed.
The circuit-open signal now blocks the pin, and _classifier_failure_outcome
tags its outcomes through one helper instead of reassigning a Final.
* feat(team): report per-user spend within a team for JWT traffic
Add GET /team/spend/by_user, which groups raw spend logs by (team_id, user)
so JWT/SSO requests with no virtual key are attributed to the user inside
each selected team. Team admins see every member, plain members see only
their own row. The Team Usage page gets a Spend Per User Within Team card
with CSV export backed by the same endpoint.
Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>
* test(team): cover /team/spend/by_user in behavior suite, tf audit allowlist and EntityUsage unit test
Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>
* refactor(team): drop explanatory docstrings from /team/spend/by_user and regen schema.d.ts
Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>
---------
Co-authored-by: yassin <yassin@berri.ai>
Co-authored-by: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>
test_event_loop_stall_timeout_burst_keeps_breaker_closed built its timeout
burst by wrapping a healthy fake call in asyncio.wait_for. Before 3.12,
wait_for returns the inner result when the inner future also completed while
the loop was blocked, so no call timed out, the burst never materialised, and
the test's own liveness guard failed with 0 >= 3.
The fake now checks its own client deadline against the clock, the way a client
library does, so the stall produces a real redis TimeoutError burst on every
interpreter. The breaker itself is unchanged: its duration gate is plain
time.time() bookkeeping and never depended on the version.
* feat(ui): page the public model hub table off /public/v1/model_hub
The public Model Hub page loaded every published model group in one call and
did all of its searching, sorting and filtering in the browser, so a proxy with
a few thousand groups sent megabytes to render one screen.
The models table now asks /public/v1/model_hub for one page at a time. Paging,
sorting, search and the provider and mode filters are query parameters on that
route, and the pagination footer counts from the response envelope's
total_count rather than the rows on screen. Column sortability is derived from
the fields the route declares sortable, so a header can no longer ask it for a
sort it answers with a 400.
The feature filter is dropped: supports_* are booleans and the route has no
boolean filter, so it could only ever have filtered the page in view.
* fix(ui): offer the model modes litellm actually prices in the hub filter
The mode filter listed 'moderations', which no model group's mode is ever set
to, so picking it could only ever return nothing; 'anthropic_messages' was
dead the same way. Four real modes the catalogue does use, search, ocr,
guardrail and vector_store, were missing entirely.
The list is now the mode vocabulary in model_prices_and_context_window.json,
and a test reads that file so an option that matches nothing, or a mode with
no option, fails instead of silently filtering to an empty table. A failed
page fetch logs the route's error detail again, as it did before the table
moved to the paginated route.
* test(ui): keep the model hub health rows out of the inline-object budget
frontend-lint's local/no-large-inline-object-arg budget went 555 to 556: the
health check rows became arguments to the row helper. They are plain literals
spreading a shared default again, which is what they were before, and the gate
reports 554 against a max of 555.
* feat: keep every model hub filter when the table pages
Moving the table onto /public/v1/model_hub cost it two controls the route
could not serve: the provider filter fell back to one substring because
providers only declared contains, and the feature filter went away entirely
because supports_* are booleans with no filter at all. Options for the
dropdowns went with them, since a page of rows only knows the values on that
page.
The route now declares providers in, a features field whose value is the
capability names a row has, and providers, rpm and tpm as sortable. Features
is one repeated field rather than a boolean per flag so selecting two of them
matches either, which is what the multi-select has always meant. Three facet
routes serve the distinct providers, modes and features across the published
groups, carrying the parent's filters, per section 12 of the list design.
All of it is additive: the route rejects unknown parameters, so no request
that worked before changes, and the design's stability policy calls new
filters and parameters safe within a version.
Health status stays unsortable. Health is read for the rows on the page, and
ordering the match set by it would mean reading it for every published group,
which is the cost the paging exists to avoid.
* fix(ui): put the model hub facet types where the generator emits them
The generated file lists paths in sorted order and operations in path order.
Both new blocks were spliced in one entry too late, after
/queue/chat/completions rather than before it, so the schema.d.ts sync check
regenerated the file and found them misplaced. Same blocks, byte for byte,
moved to the position the generator gives them.
* fix(proxy): type a facet payload as the sequence the framework hands it
The lint job's basedpyright gate flagged one new reportArgumentType: handle_facet
passes a tuple, and FacetListResponse declared data as list[str]. A list would
have traded that error for an LIT002 mutable construction, and both budgets are
already at their ceiling on the base.
Sequence[str] is what the framework actually produces and what the model always
accepted: pydantic emits the same array schema either way, verified against
model_json_schema, so the OpenAPI spec and schema.d.ts are unchanged, and the
existing list-passing caller in spend_logs still type checks.
* test(proxy): pin the facet route's rejection contract
handle_facet answers six ways before it ever reaches the executor, and
none of them was covered: a denied scope, a filter operator the spec does
not offer, a repeated parameter, a non-positive page or page_size, and the
where clause those last two feed. Every one is a 400 or 403 an
unauthenticated caller can reach, so each gets a test that fails when the
branch stops firing.
Semantic cache keys omit the prompt, so every end user behind one virtual key
shares a bucket and can be served another user's semantically similar response.
Add an opt-in cache_params.semantic_cache_scope (key | end_user) that appends the
authenticated end-user id to the tenant scope, read from metadata and
litellm_metadata so /v1/chat/completions, /v1/responses and /v1/messages are all
covered, falling back to the key scope when no end-user id is present. Expose the
setting in the cache settings API and the Admin UI cache settings form
Co-authored-by: yassin <yassin@berri.ai>
Co-authored-by: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>
Team and service keys often have no user_email after a successful token
join. Treating empty email as a miss hashed every verification token on
routine CloudZero and Focus exports.
Co-authored-by: Mateo Wang <mateo-berri@users.noreply.github.com>
Export rows already join user_email from DailyUserSpend.user_id. Recovered
key-owner email must not replace that when only the alias join missed.
Co-authored-by: Mateo Wang <mateo-berri@users.noreply.github.com>
* fix: remove wildcard routes from /v1/models response
Wildcard routes like bedrock/* were leaking into the /v1/models response
because _get_wildcard_models only removed them from unique_models in the
fallback branches (no router or no deployment), but not when the router
had a matching deployment. Now wildcards are always removed from the base
list; they are only re-added to the result when return_wildcard_routes=True
is explicitly passed.
Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>
* refactor(model_checks): collapse wildcard expansion branches and tighten regression tests
Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>
---------
Co-authored-by: Krrish Dholakia <krrishdholakia@berri.ai>
Co-authored-by: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>
Co-authored-by: yassin <yassin@berri.ai>
* fix(azure): restrict the storage credential chain to deployment identities
The keyless Azure Storage path walks the full DefaultAzureCredential chain, so a
proxy with no storage service principal authenticates as whichever identity the
host happens to carry: an operator's az login on a workstation, or the
AZURE_CLIENT_ID/AZURE_CLIENT_SECRET service principal set for Azure OpenAI.
Neither is the identity granted Storage Blob Data Contributor.
Narrow the chain to workload identity and managed identity, the two credentials
a deployment legitimately holds. Azure OpenAI, Postgres IAM auth and the other
callers of get_azure_ad_token_provider keep the full chain.
* test(azure): read the credential chain off the mock instead of an accumulator
* chore: drop a stray launch traceback committed at the repo root
* fix(azure): let the storage chain reach a system assigned managed identity
DefaultAzureCredential keeps one managed identity link and pins it to
AZURE_CLIENT_ID, so a host that sets that variable for Azure OpenAI and runs as
a system assigned identity never got asked for a storage token. Build the chain
from the three credentials a deployment can carry instead of subtracting the
ones it cannot.
* fix(snowflake): normalize Cortex Claude request shapes
Co-authored-by: Kamron Javaherpour <kamron@kargo.com>
Co-authored-by: Oleksandr Kononov <oleks.konov@kargo.com>
* style(snowflake): format Cortex request transformations
* fix(snowflake): annotate Cortex wire payloads
* fix(snowflake): route Cortex content through the shared Anthropic converters
* fix(snowflake): surface Cortex prompt-cache usage and thinking blocks
Parse Cortex's Anthropic-dialect responses and SSE with Anthropic's own parser so cache_creation/cache_read counts, thinking blocks and signatures reach the caller. Restore thinking for every Claude model: Cortex documents extended thinking broadly and only adaptive thinking is 4.6-gated.
* fix(snowflake): echo signed thinking blocks on every assistant turn
The reference converter extends signed thinking blocks on each assistant turn, not just tool-call turns, so a replayed thinking-plus-text response keeps its signed block. Content-less thinking turns send no empty text block.
* fix(snowflake): preserve thinking list content
---------
Co-authored-by: Oleksandr Kononov <oleks.konov@kargo.com>
* fix(router): evict stale global pattern_router entries on upsert/delete
upsert_deployment and delete_deployment cleaned team_pattern_routers but left
the outgoing deployment in the global pattern_router, so wildcard requests kept
round-robining onto the stale entry after a PATCH /model/{id}/update.
Fixes#29064
Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>
* test(router): dedupe test_get_configured_mode_reads_deployment_model_info name
Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>
* fix(router): restore global pattern_router eviction dropped by previous commit
Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>
---------
Co-authored-by: yassin <yassin@berri.ai>
Co-authored-by: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>
* perf(mcp): cache SSO identity assertion reads on the ID-JAG path
Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>
* fix(mcp): guard sso assertion cache against stale relogin reads
Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>
* fix(mcp): keep sso assertion cache entries and generation markers in separate namespaces
Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>
* test(router): rename duplicate get_configured_mode test so ruff F811 passes
Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>
* fix(mcp): use a process-wide epoch for sso assertion cache invalidation
Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>
---------
Co-authored-by: yassin <yassin@berri.ai>
Co-authored-by: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>