* feat(ui): page the public model hub table off /public/v1/model_hub
The public Model Hub page loaded every published model group in one call and
did all of its searching, sorting and filtering in the browser, so a proxy with
a few thousand groups sent megabytes to render one screen.
The models table now asks /public/v1/model_hub for one page at a time. Paging,
sorting, search and the provider and mode filters are query parameters on that
route, and the pagination footer counts from the response envelope's
total_count rather than the rows on screen. Column sortability is derived from
the fields the route declares sortable, so a header can no longer ask it for a
sort it answers with a 400.
The feature filter is dropped: supports_* are booleans and the route has no
boolean filter, so it could only ever have filtered the page in view.
* fix(ui): offer the model modes litellm actually prices in the hub filter
The mode filter listed 'moderations', which no model group's mode is ever set
to, so picking it could only ever return nothing; 'anthropic_messages' was
dead the same way. Four real modes the catalogue does use, search, ocr,
guardrail and vector_store, were missing entirely.
The list is now the mode vocabulary in model_prices_and_context_window.json,
and a test reads that file so an option that matches nothing, or a mode with
no option, fails instead of silently filtering to an empty table. A failed
page fetch logs the route's error detail again, as it did before the table
moved to the paginated route.
* test(ui): keep the model hub health rows out of the inline-object budget
frontend-lint's local/no-large-inline-object-arg budget went 555 to 556: the
health check rows became arguments to the row helper. They are plain literals
spreading a shared default again, which is what they were before, and the gate
reports 554 against a max of 555.
* feat: keep every model hub filter when the table pages
Moving the table onto /public/v1/model_hub cost it two controls the route
could not serve: the provider filter fell back to one substring because
providers only declared contains, and the feature filter went away entirely
because supports_* are booleans with no filter at all. Options for the
dropdowns went with them, since a page of rows only knows the values on that
page.
The route now declares providers in, a features field whose value is the
capability names a row has, and providers, rpm and tpm as sortable. Features
is one repeated field rather than a boolean per flag so selecting two of them
matches either, which is what the multi-select has always meant. Three facet
routes serve the distinct providers, modes and features across the published
groups, carrying the parent's filters, per section 12 of the list design.
All of it is additive: the route rejects unknown parameters, so no request
that worked before changes, and the design's stability policy calls new
filters and parameters safe within a version.
Health status stays unsortable. Health is read for the rows on the page, and
ordering the match set by it would mean reading it for every published group,
which is the cost the paging exists to avoid.
* fix(ui): put the model hub facet types where the generator emits them
The generated file lists paths in sorted order and operations in path order.
Both new blocks were spliced in one entry too late, after
/queue/chat/completions rather than before it, so the schema.d.ts sync check
regenerated the file and found them misplaced. Same blocks, byte for byte,
moved to the position the generator gives them.
* fix(proxy): type a facet payload as the sequence the framework hands it
The lint job's basedpyright gate flagged one new reportArgumentType: handle_facet
passes a tuple, and FacetListResponse declared data as list[str]. A list would
have traded that error for an LIT002 mutable construction, and both budgets are
already at their ceiling on the base.
Sequence[str] is what the framework actually produces and what the model always
accepted: pydantic emits the same array schema either way, verified against
model_json_schema, so the OpenAPI spec and schema.d.ts are unchanged, and the
existing list-passing caller in spend_logs still type checks.
* test(proxy): pin the facet route's rejection contract
handle_facet answers six ways before it ever reaches the executor, and
none of them was covered: a denied scope, a filter operator the spec does
not offer, a repeated parameter, a non-positive page or page_size, and the
where clause those last two feed. Every one is a 400 or 403 an
unauthenticated caller can reach, so each gets a test that fails when the
branch stops firing.
Team and service keys often have no user_email after a successful token
join. Treating empty email as a miss hashed every verification token on
routine CloudZero and Focus exports.
Co-authored-by: Mateo Wang <mateo-berri@users.noreply.github.com>
Export rows already join user_email from DailyUserSpend.user_id. Recovered
key-owner email must not replace that when only the alias join missed.
Co-authored-by: Mateo Wang <mateo-berri@users.noreply.github.com>
* fix: remove wildcard routes from /v1/models response
Wildcard routes like bedrock/* were leaking into the /v1/models response
because _get_wildcard_models only removed them from unique_models in the
fallback branches (no router or no deployment), but not when the router
had a matching deployment. Now wildcards are always removed from the base
list; they are only re-added to the result when return_wildcard_routes=True
is explicitly passed.
Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>
* refactor(model_checks): collapse wildcard expansion branches and tighten regression tests
Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>
---------
Co-authored-by: Krrish Dholakia <krrishdholakia@berri.ai>
Co-authored-by: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>
Co-authored-by: yassin <yassin@berri.ai>
* perf(mcp): cache SSO identity assertion reads on the ID-JAG path
Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>
* fix(mcp): guard sso assertion cache against stale relogin reads
Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>
* fix(mcp): keep sso assertion cache entries and generation markers in separate namespaces
Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>
* test(router): rename duplicate get_configured_mode test so ruff F811 passes
Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>
* fix(mcp): use a process-wide epoch for sso assertion cache invalidation
Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>
---------
Co-authored-by: yassin <yassin@berri.ai>
Co-authored-by: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>
Okta sends profile updates as full PUTs with no groups or groups: [], since SCIM User.groups is readOnly and membership is synced through /Groups. The PUT handler diffed that empty list against the stored teams, removed the user from every team (which also deletes their team keys) and recomputed the role from an empty group list. Treat an empty groups list on PUT as unspecified: keep the stored teams and leave the role alone. Explicit non-empty groups still replace memberships as before
Claude-Session: https://claude.ai/code/session_01CqwUV4Ywnu5aUjXx1UhJrM
Calling the endpoint without going through FastAPI leaves the new search param set to its Query default object, which is not None, so the grouped-session and request_id lookup tests started taking the search branch
Claude-Session: https://claude.ai/code/session_01Q5sbiogJzPcCRmYSbaHxZf
The search= param on /key/list, /audit, and /spend/logs/ui, plus key_hash= on /key/list, now compare the pasted value verbatim. Only a copied key ID (the hash) matches, so a raw virtual key never needs to travel in a GET query string
Claude-Session: https://claude.ai/code/session_01Q5sbiogJzPcCRmYSbaHxZf
Routes without /projects/<project>/locations/<location>/ built the upstream host from the URL's
still-empty location and 500ed even with default_vertex_config set. Build the base URL once after
the configured project and location are applied, drop the hook that re-derived it afterwards, and
answer 400 with a fix-it message when no location is available at all.
Resolves LIT-6905
GET /key/list?search= matches the key hash (a raw sk- key is hashed
first) or a case-insensitive alias substring, and key_hash= now hashes a
raw sk- value too. GET /v1/memory?search= matches a key prefix or an
exact memory_id. GET /audit?search= matches id, object_id, changed_by,
or changed_by_api_key. GET /spend/logs/ui?search= matches request_id
across all time and api_key, team_id, user, end_user, session_id, or
model_id inside the date window; session grouping is skipped while a
search is active.
Claude-Session: https://claude.ai/code/session_01Q5sbiogJzPcCRmYSbaHxZf
* fix(mcp): resolve OAuth broker endpoints by server_id with IP access checks
Resolve named OAuth lookups through server IDs while retaining client IP checks\n\nCo-authored-by: KK291860 <krishnakumar.kocherykumaran@sephora.com>
Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>
* ci: retrigger e2e pipeline
Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>
---------
Co-authored-by: yassin <yassin@berri.ai>
Co-authored-by: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>
Listing agents (GET /v1/agents and MCP agent_search) treated the absence of any
agent grant on the key or team as permission to see every agent. Non-admin keys
now list only the union of explicit grants, and dashboard sessions resolve that
union through the user's real teams and user row instead of the shared
dashboard team. Proxy admins still see everything and direct access to a named
agent is unchanged.
Resolves LIT-6862
Co-authored-by: yassin <yassin@berri.ai>
Co-authored-by: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>
Without the auto_router feature in the signed enterprise license a proxy may hold
one complexity router with classifier_type heuristic_v2 across config.yaml and the
DB; with it the limit is lifted. The ceiling is derived once from LicenseCheck and
handed to the Router, which refuses the extra router at registration. config.yaml
over the limit refuses to start, and /model/new, /model/update and
PATCH /model/{id}/update refuse the write with a 403 before touching the DB.
Expiry follows the existing max_users/max_teams pattern: judged when the
license is verified, not on every call, and a verify that rejects the license
(expired or unreadable) leaves no signed payload behind. The rollback after a
failed upsert re-admits state that was already serving, so it is exempt from the
ceiling: an edit that fails, including one refused by a ceiling that has since
tightened, leaves the router serving its previous configuration.
A write that leaves a row on heuristic_v2 under a limited license runs in one
transaction that takes a Postgres advisory lock before counting the DB rows plus
this proxy's config.yaml routers, so concurrent writes on any pod cannot both
claim the sole slot and no surplus row is ever persisted.
Only the row insert runs under that lock: the team model bookkeeping, which
needs a second pool connection, runs after the transaction has committed.
PATCH /model/{id}/update follows the same order as create: the row is written
through the slot first and the team's model list is updated only afterwards, so
a refused write leaves the team as it was.
The slot transaction bypasses the repository's publish-on-write, so it
publishes the config change once after commit, as delete_team_models does.
* fix(vector_stores): only list vector stores the caller was granted
/vector_store/list returned every managed vector store with no team_id to any
key, and let a dashboard session see stores created from the dashboard because
every session shares the litellm-dashboard team id. Non-admin listings now show
a store only when the key or one of the caller's real teams is allowlisted for
it via object_permission.vector_stores, or the team owns it
Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>
* fix(vector_stores): keep a dashboard session key's own grants when the user has no teams
Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>
---------
Co-authored-by: yassin <yassin@berri.ai>
Co-authored-by: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>