* feat(mcp): hand listed-tool metadata to pre-call hooks with per-caller catalog identity
Track the tools each MCP server listed per caller identity so pre_mcp_call and during_mcp_call hooks receive the tool description and input schema the client saw. Servers with no caller-dependent inputs share one slot; user identity, forwarded headers, stdio env, relayed bearers, and server-specific auth get their own. Local registry and OpenAPI paths pass the registered metadata and admin description overrides. The Agent 365 guardrail reads the new fields into its evaluate payload.
Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>
* refactor(mcp): drop the listed-tools empty sentinel and routine test docstrings
Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>
* fix(mcp): mark the listed-tools cache digest as a non-security hash
Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>
* fix(mcp): key the listed-tools cache by the OBO subject token
token_exchange servers list upstream with the caller's own Entra bearer, so two callers on one
LiteLLM key with different subjects were sharing a catalog slot
Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>
* fix(mcp): resolve the BYOK credential before keying the listed-tools slot
Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>
* fix(mcp): drop the OAuth discovery cache when a server definition changes
Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>
* refactor(mcp): drop a diff-narrating comment
Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>
* fix(mcp): never validate a supplied header on the tools/list BYOK path
The pre-listing resolver ran the tool-call byok_auth_required check even
when the caller already supplied x-mcp-auth, and it ran outside the
per-server error boundary, so a single deprecated-header caller dropped
the server from the aggregate list. Listing now returns a supplied header
unchanged and falls back to the stored credential without raising
Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>
* test(mcp): assert the BYOK listing lands in the caller's listed-tool slot
Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>
* test(mcp): cover the deprecated string x-mcp-auth header on a BYOK tools/list
Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>
* fix(mcp): key the per-caller listed-tool slot by the hashed token
Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>
* fix(mcp): key discovery cache by the hashed token
Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>
* fix(mcp): key discovery caches per caller correctly and drop stale caches on server updates
Discovery-list cache identity now uses the hashed token instead of the raw
api_key and treats MCPJWTSigner-signed servers as per caller. Server
definition changes also drop the cached upstream OAuth metadata. OpenAPI
listings look tools up under the normalized registry prefix with the
separator, so an overlapping sibling prefix no longer leaks into the list.
Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>
* fix(mcp): keep the discovery cache digest call unchanged so CodeQL matches the existing alert
Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>
* refactor(mcp): derive the listed-tool caller identity from the discovery cache key
Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>
* fix(mcp): guard OAuth metadata cache writes with a per-server generation and drop unproven per-caller discovery keys
An upstream metadata fetch that started before a server edit could store its stale reply after
invalidate_oauth_metadata_cache ran. Invalidation now bumps a per-server generation and the fetch
only stores when the generation it captured before I/O is unchanged.
The MCPJWTSigner-based per-caller discovery classification and the api_key to token key change had no
reproduction (the signer only injects on tools/list, and UserAPIKeyAuth hashes api_key in place), so
both go back to the merge-base behavior.
Integration coverage under tests/integration/mcp: overlapping OpenAPI aliases, a config-declared
server name with a space, OAuth metadata refetch after a save, and the in-flight stale-write race
Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>
* fix(mcp): keep OAuth metadata generations only while a fetch is in flight
Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>
* fix(mcp): count queued OAuth metadata fetchers so invalidation survives lock handoff
Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>
* fix(mcp): keep a held OAuth metadata lock registered even when no fetcher slot claims it
Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>
* test(mcp): prove a peer worker drops stale upstream OAuth metadata after a save elsewhere
Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>
* test(mcp): return one masked text per scanned string in the selected-guardrail REST test
Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>
* fix(mcp): fold the signed caller into the discovery digest instead of a second key hash
Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>
* fix(mcp): satisfy type discipline gate on listed-tool identity
Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>
* fix(mcp): hand tools/call hooks the exact catalog entry tools/list served
get_listed_tool re-applied the admin description override on top of the cached listing, so a
guardrail-masked description was restored to its original wording at call time, and the OpenAPI /
local-registry call path built its metadata from the registry instead of the guarded caller catalog.
Both paths now return the cached entry as served, falling back to the registry only when no listing
was recorded
Adds tests/integration/mcp/test_mcp_listed_tool_metadata.py (red on the prior head for the two
regressions, red on the merge base for the feature, green on this head)
Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>
* fix(mcp): key OpenAPI listed-tool entries per caller so tools/call reads its own guarded listing
Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>
* test(mcp): align listed-tool slot tests with per-caller keying
Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>
* fix(mcp): keep oauth2 listing on the minted or signed credential, not the stored BYOK secret
The listing helper that keys the per-caller catalog by the stored BYOK credential also handed that
credential to the upstream client, which on an oauth2 server short-circuited the client_credentials
mint and the MCPJWTSigner gate. Split the two: the catalog identity keeps the stored credential so
tools/call finds the caller's slot, while an oauth2 server's tools/list sends only the per-request
header, letting the M2M mint or signed JWT proceed as on main
Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>
* test(mcp): oauth2 BYOK listing sends the minted token, not the stored secret, through the real proxy
Integration cell for the listing fix: a client_credentials BYOK server with a stored user credential,
one tools/list as that user, the peer must see a live minted bearer and one /token mint. Red at the
pre-fix tip (zero mints, stored secret upstream), green at the fixed head and at the merge base
Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>
* fix(mcp): keep the stored BYOK credential for catalog identity only on tools/list
Listing used the resolved stored credential both to key the caller's catalog slot and as the
upstream transport header, so REST api_key and bearer_token listings sent the user's secret
instead of the server's static token and the MCPJWTSigner gate went quiet. The upstream client
and the signer gate now read the caller-supplied mcp_auth_header for every auth type, exactly
as before the catalog existed, and the stored credential only names the slot tools/call reads
Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>
* fix(mcp): type the listed-tool metadata read from pre-call kwargs for the basedpyright gate
Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>
* fix(mcp): hand never-listed tools/call hooks name and arguments only
The local-registry call path fell back to the registry entry with the admin description override
when no tools/list had been recorded for the caller, so a pre_mcp_call guardrail scanned a
description the caller was never served and blocked OpenAPI calls that passed before, and base's
own selected-guardrail REST test failed on the two-text redaction. _registered_tool_metadata now
returns the listed entry or None, so a tools/call with no prior listing sends name and arguments
only as promised, and that REST test double goes back to its base shape
* fix(mcp): keep during_mcp_call hooks on name and arguments only
call_tool handed the caller's listed entry to the during-hook task as well, so during_mcp_call
guardrails scanned the description line and schema leaves of any listed tool after the upstream
call had already run, blocking calls that passed before whenever the policy matched the
description, returned a fixed-length texts list, or hit the depth guard on a deep schema. The
listed entry is only disclosed for pre_mcp_call, so the during task no longer receives it and its
request object carries no description or schema, as before
* fix(mcp): key the BYOK catalog slot by the client's header, not the stored credential
tools/list resolved the stored BYOK credential to pick the caller's catalog slot, which read the
credential store before the classified try block. With Postgres down and a cold per-worker cache
that made every REST tools/list on an is_byok server fail with tools=[] and no upstream call, and
the read seeded the per-worker cache (including a negative entry), so a tools/call on another
worker after a store, rotate or revoke on this one kept using the stale value.
The slot is now keyed by what the client supplied plus the caller's hashed key, on both sides.
_get_tools_from_server and call_tool take a keyword-only catalog_auth_header that defaults to
mcp_auth_header as received (the default is the builtin Ellipsis so it survives a module reload).
The /mcp fan-out and execute_mcp_tool, which swap the resolved credential into mcp_auth_header,
pass the client's value explicitly. What goes upstream is unchanged. _byok_catalog_auth_header is
gone.
* fix(mcp): drop a listed catalog recorded across a server save
_record_listed_tools ran after the awaited upstream fetch, so a PUT /v1/mcp/server that landed
mid-fetch had its invalidation undone when the fetch completed: hooks then saw the pre-save
description next to the post-save definition until the next listing, instead of name and
arguments only.
The manager now keeps a per-server listed-tools generation, bumped by
_invalidate_server_definition_caches. _get_tools_from_server reads it before the fetch and
_record_listed_tools skips the write when it moved; the next listing records normally.
* fix(mcp): drop the catalog again once a saved OpenAPI server's registry is rebuilt
add_server and update_server publish the saved definition before the OpenAPI registry entries are
rebuilt from the spec, so a listing recorded during that fetch held the pre-save entries under the
new generation. The generation is bumped a second time after the registry refresh.
The during-hook task no longer accepts a listed entry, the one-line wrapper over get_listed_tool is
inlined at its two call sites, and the per-server generation map is a plain dict.
* fix(mcp): keep discovery and OAuth metadata caches across an OpenAPI spec re-read
add_server and update_server ran the full server-definition invalidation a
second time after the awaited OpenAPI spec fetch, which also dropped the
prompts/resources/templates discovery entries and the OAuth protected-resource
metadata filled under the already-published definition, so the next request
went upstream again. Only the listed-tool catalog recorded during the fetch
holds pre-save entries, so the post-fetch pass now drops just that catalog
and bumps its generation via the new _drop_listed_tools helper, which the
full invalidation also calls.
* fix(mcp): look a called tool up in the listed catalog by its bare name only
get_listed_tool stripped the server prefix a second time when the exact
name was absent from the caller's listing, so a never-listed upstream tool
whose bare name starts with the server prefix resolved to the listed
sibling and that sibling's description and input schema reached the
pre-call hooks for a call to a different tool. Every caller already passes
the once-stripped bare name, so the lookup is now exact.
Tests that looked the catalog up by a prefixed name now use the bare name
the callers pass; two new tests pin the never-listed sibling case at the
manager and at the tools/call path.
* fix(mcp): record a listed-tool catalog only for a listing the caller is served
_get_tools_from_server now records the catalog into the caller's
listed-tools slot only when asked (record_listing=True), which the
served listings pass: the /mcp and Responses API tools/list handlers via
_get_tools_from_mcp_servers, MCPServerManager.list_tools, and the REST
listing via _list_server_tools. Four internal listings stop recording,
so a later tools/call hands pre_mcp_call hooks name and arguments only,
as on main:
- _list_tools_before_first_call, the implicit listing inside tools/call
when this worker does not yet expose the tool
- fetch_pinnable_tool_catalog, the admin pin snapshot listed without the
catalog guard and without description overrides
- _initialize_tool_name_to_mcp_server_name_mapping, the startup fill
- get_tools_for_server, used by the semantic tool filter
_create_prefixed_tools returns to its tool-name mapping job only; the
record follows it in _get_tools_from_server.
* fix(mcp): opt every listing out of catalog recording unless it is served
The aggregate listing and _list_mcp_tools now default to record_listing=False,
so a catalog fetched inside a tools/call no longer fills the caller's
listed-tools slot. The /mcp/proxy meta-tools (call_tool, search_tools,
get_tool_schema) and the tool-search virtual tool stop recording: /mcp/proxy
serves only the meta-tools and the search serves only its hits, so a later
pre_mcp_call hook was reading a description the caller never listed.
The tools/list handler, the Responses MCP handler and the /v1/mcp/tools
management listing opt in with record_listing=True, since each serves the
catalog to the caller.
* fix(mcp): key the listed-tool slot by the caller's admission identity and forwarded bearer
The slot a tools/list records for a later tools/call was keyed by (user_id, api_key)
only, so every team-only JWT caller shared one slot and one JWT user acting in two
teams shared a slot; a tools/call then handed pre_mcp_call hooks a description another
caller was served. The slot is now keyed by the hashed key, user, team and organization,
plus the admission credential of a caller admitted with neither a key nor a user.
The caller bearer split the slot only on client-forwarded-token and token-exchange
servers; a legacy delegated oauth2 server (delegate_auth_to_upstream without client
credentials) also forwards it upstream and served a different catalog per bearer into
one slot. The bearer now splits the slot on every server whose egress forwards it
(_consumes_caller_authorization) or exchanges it.
* fix(mcp): record only tools served by the bridge
Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>
* fix(mcp): keep bridge tool metadata request-local
* refactor(mcp): centralize listed catalog recording guard
* fix(mcp): preserve base TPM reservations for listed tool calls
* fix(mcp): preserve project token reservations for listed calls
* fix(mcp): record served catalogs and preserve call message bytes
* test(mcp): align listing expectation with deferred recording
* test(mcp): audit listed metadata across callers and bridge lifecycles
* test(mcp): preserve guardrail fixture worker affinity
* refactor(mcp): expose listed catalog recording API
* feat(mcp): pass served_tools through the anthropic messages bridge
/v1/messages auto-execution now hands this request's resolved tool
definitions to _execute_tool_calls, matching the Responses and chat
completions bridges: the pre_mcp_call hook receives the description
and input schema the model was shown for that call. Request-local
only; the shared listed-tools catalog is untouched.
* test(mcp): pin served_tools handoff on the anthropic messages bridge
Mirrors the credentials-forwarding test: the request's resolved tool
definitions must reach _execute_tool_calls under served_tools so
pre_mcp_call hooks judge the call on the description and input schema
the model was shown. Fails without the previous commit's one-liner.
* style(mcp): sort the local import block ruff flagged
* fix(mcp): keep the admin include_disabled_tools view off the listed-tools catalog
GET /mcp-rest/tools/list?include_disabled_tools=true is the admin-only
configuration view: apply_tool_filters is False, so it serves the full
server catalog. Recording that response into the caller's listed-tools
slot warmed tools/call metadata no runtime listing ever served,
breaking the only-a-served-listing-records invariant (Bugbot).
The record is now gated on apply_tool_filters; disabled tools stay
unreachable (the call-time allowlist 403 fires before hooks), so the
observable fix is the slot no longer warming from a settings view.
Verified live: the new test fails on the unfixed head and passes here,
and the rest of the listed-tool-metadata suite is unchanged.
---------
Co-authored-by: yucheng <yucheng@berri.ai>
Co-authored-by: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>
* test(ci): fix four CircleCI regressions on main
Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>
* test(ci): stop reloading auth_checks in unit tests
Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>
* test(constants): cover CLI JWT expiry env parsing
Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>
* test(gateway): restore proxy lifespan after importing gateway.main in launch tests
Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>
---------
Co-authored-by: mateo <mateo@berri.ai>
Co-authored-by: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>
* fix(lens): pin worker dependencies and support approved image digests
* fix(lens): include locked dependencies and release identity in build context
* fix(lens): select the dev worker package for SHA-tagged charts
* feat(anthropic): workload identity federation and pluggable identity sources
Backend half of #38818 (internal copy of the fork PR #38013), rebuilt as one
commit on top of litellm_internal_staging without the dashboard changes.
Deployments on anthropic/ without a static api_key can exchange an OIDC
workload assertion for a short-lived sk-ant-oat01 token through a shared
RFC 7523 JWT-bearer engine. The assertion comes from a mounted token file,
an env token, a LiteLLM-signed issuer, or Keycloak, chosen per deployment,
per named credential, or through ANTHROPIC_IDENTITY_SOURCE. The federation
fields are server-owned: refused inline in request bodies and on
POST /model/new, proxy-admin only on credentials, and the token exchange
is pinned to api.anthropic.com unless LITELLM_ANTHROPIC_WIF_ALLOWED_HOSTS
adds a host. GET /credentials/{name}/jwks exports the public key set of a
LiteLLM-signed credential for the Claude Console.
The OpenAI federation trio from #39613 rides along on the backend side with
the same server-owned handling.
Fixes#28607
Resolves LIT-6107
Co-authored-by: derhornspieler <15236687+derhornspieler@users.noreply.github.com>
* fix(anthropic): let batch-result downloads mint from deployment params and accept host:port allowlist entries
The files handler enabled workload identity on batch-result downloads but never received the
deployment's litellm_params, so a deployment authenticating through a named credential could only
mint from process-wide env vars. It now threads litellm_params through to the auth header the way
the batch retrieve path already does.
LITELLM_ANTHROPIC_WIF_ALLOWED_HOSTS entries written as host:port were read by urlsplit as a scheme,
so the allowlist kept the raw entry while the exchange compared bare hostnames and refused the
gateway. Entries are now parsed as network locations whether or not they carry a scheme.
* fix(types): move the WIF kwargs key sets to a leaf module so the kwargs funnel imports without a cycle
* test(anthropic): pin case-insensitive matching of WIF exchange-host allowlist entries
* fix(anthropic): end workload identity federation errors without a period so the router suffix reads cleanly
* fix(proxy): decrypt stored litellm_params before the WIF write gate
* fix(proxy): hide WIF secret references from /health output
* fix(proxy): keep the proxy error shape on credential endpoint refusals
* fix(proxy): hide identity token file paths from /health output
* fix(anthropic): rename the federation workspace param so Bedrock's anthropic_workspace_id keeps working
The Bedrock Claude Platform route already reads anthropic_workspace_id from
optional_params, so banning that spelling as a server-owned federation
parameter broke a pre-existing client capability. The federation field is now
anthropic_federation_workspace_id (env ANTHROPIC_FEDERATION_WORKSPACE_ID),
which restores the base branch's behavior for Bedrock callers, drops the
Bedrock-specific hint from the refusal message, and deletes the unconditional
ban constant that no longer had a reader
* fix(auth): share one exchanged token across workers reading the same assertion
Anthropic accepts each identity assertion exactly once, so two uvicorn
workers reading the same token file both minting from it means the second
exchange is denied with jti_reused. Minted tokens now land in a per-user
0700 cache directory guarded by a file lock, so workers on the same host
reuse one exchange until the token expires or the assertion rotates. A 401
is only retried when the re-read assertion actually differs, and the denial
hint explains jti_reused. LITELLM_TOKEN_EXCHANGE_CACHE_DIR moves the cache
and an empty value disables it
* fix: keep anthropic federation from being shadowed or leaked
An empty or whitespace-only ANTHROPIC_API_KEY counted as set, so a federated
deployment sent an empty x-api-key on every call instead of minting a token.
Blank values now read as unset, and a real static key on a federated deployment
logs once that it outranks federation and nothing is being federated.
The exchange-host allowlist matched hostnames only, so a second process on
another port of an allowed host was trusted with the workload's identity token.
An entry that names a port now trusts that port alone, while a bare host still
trusts every port.
The shared token store exists so the workers reading one projected token file do
not each spend its single-use jti. A source that mints its own assertion per
exchange shares nothing with another worker, so it no longer writes a live token
to disk for a lookup that can never hit.
* fix: unlink a staged token file a failed write leaves behind
The 401 denial hint now also says federation ignores ANTHROPIC_WORKSPACE_ID, which the Bedrock Claude platform provider already reads.
* refactor: move anthropic jwks derivation behind a provider-owned tagged union
* fix: unlink the staged token file when its write fails at close
A buffered write only reaches the disk when the handle closes, so a full disk surfaces at close and left the staging file behind holding a usable token.
* fix(anthropic): close the staging descriptor before writing the shared token file
* fix(wif): judge federation writes by what they set, not what is stored
The admin gate read the stored deployment, so a team admin lost edit, delete
and Test Connection on any deployment carrying federation params. It now
returns early unless the submitted fields touch the federation surface, and a
Test Connection probe that points the deployment at its own api_base is still
refused, with the 403 no longer wrapped into a 500
The rest of the same review pass: POST /model/new refuses only a blocking
value of `blocked`, so a client that always sends `blocked: false` is not
turned away; a request body can no longer pick which federated identity to
mint as by naming a stored credential; an advisory refresh the executor
refuses disarms the entry instead of wedging the identity until the follower
timeout; the static-key shadow warning resolves its env fallback inside the
cache instead of once per request; credential writes drop nulls before
storing them; the token exchange validates the endpoint URL before reading an
assertion and keeps refusing redirects across a client heal; /health hides
every server-owned federation field from non-admins; and the async create_file
and create_batch paths say which setting is missing when the provider resolves
no URL
* fix(proxy): let a deployment write name a federated credential
reject_federated_credential_reference runs from is_request_body_safe, which
pre_db_read_auth_checks calls on every route, so it also fired on POST
/model/new, /model/update, /model/{id}/update and /health/test_connection. A
proxy admin could no longer attach a federated credential to a deployment over
the API or the Admin UI, leaving a static config.yaml entry as the only way to
configure the feature the rejection told the caller to go configure, and
_reject_non_admin_wif_write never got to make the call it exists to make.
is_request_body_safe now takes the route and skips only the credential-reference
check on the routes that reach can_user_make_model_call. Federation fields typed
inline into a body stay refused everywhere, and a call naming a federated
credential still cannot pick the identity it mints as.
* refactor(proxy): derive health display policy from the federation key sets
The health check module hand-copied the five workload identity fields whose
value is a credential, so a shared proxy surface named provider-specific
parameters and a newly added secret-bearing field would have gone on being
displayed until someone remembered both places
WIF_SECRET_BEARING_KEYS now sits beside the key sets it splits out of,
types/utils derives secret_bearing_wif_litellm_params from it, and the health
layer splats that tuple the same way it already splats the admin-only one
* fix(anthropic_wif): treat blank identity-source fields as unset
* test(proxy): classify the federation params in the credential slot registry
main's registry test (#43298) now fails the build for any credential-named
deployment param without a classification. The five federation fields that
carry a token, a token file path, or a signing or client secret reference are
Unplanted, matching WIF_SECRET_BEARING_KEYS; the four remaining Keycloak
settings name a URL, a client id, an auth method, or a scope and are NotSecret
* fix(anthropic_wif): declare federation params as owned connection leaves and chart their metrics
Register the 18 Anthropic and 3 OpenAI federation params as frozen
ConnectionSettings leaves so the owned-kwarg registry, the kwargs funnel
and the request-body ban list read one declaration. Pass the deployment
api_base through to the count-tokens handler instead of a pre-suffixed
URL, which doubled the /count_tokens path on main's prompt-cache
predictor. Add the five litellm_anthropic_wif_* families to the
all-metrics Grafana dashboard.
* fix(credentials): gate PATCH on WIF fields resolved from model_id
The credential PATCH handler checked server-owned workload identity
federation fields only on the values the caller sent, while a body that
named a deployment through model_id had its credential values resolved
after that check. A non-admin could therefore copy a federated
deployment's WIF fields onto an ordinary credential. Resolve the incoming
values first and run the non-admin gate on them, matching the POST path
* fix(anthropic): count tokens with ANTHROPIC_AUTH_TOKEN through the shared auth header
Count-tokens walked its own credential ladder: a static key, else skip minting when
ANTHROPIC_AUTH_TOKEN is set, else mint a federated token. With only the auth token set it
forwarded nothing and the proxy silently fell back to its local tokenizer while chat on the
same deployment authenticated with that token. The handler now takes the auth header that
AnthropicModelInfo.aget_auth_header resolves, the same ladder chat, files, batches and skills
use, and merges the oauth beta a minted or consumer token carries with the token-counting beta
---------
Co-authored-by: mateo-berri <277851410+mateo-berri@users.noreply.github.com>
Co-authored-by: derhornspieler <15236687+derhornspieler@users.noreply.github.com>
Co-authored-by: mateo-berri <happymvw@gmail.com>
* fix(health): probe Bedrock Mantle Claude deployments over the Anthropic Messages API
Bedrock Mantle serves Claude ids only on /anthropic/v1/messages, but health
checks probed every chat-mode deployment over /v1/chat/completions, so a
bedrock_mantle Claude deployment showed unhealthy while real /v1/messages
traffic to it succeeded
Add an anthropic_messages health check mode and make it the default for
bedrock_mantle Claude models. An explicit model_info.mode still wins, and
/health/test_connection and the Add Model form accept the new mode
* fix(health): resolve the test connection mode from the deployment when the request omits it
The Admin UI model page sent the mode /model/info had filled in from the cost
map back as the probe mode, so Test Connection on a Bedrock Mantle Claude
deployment still went over chat completions. The page now forwards only the
row's id, and /health/test_connection resolves a missing mode the way /health
does: the stored model_info.mode, then the mode the provider requires, then the
cost map.
* fix(health): resolve an omitted ahealth_check mode the way the proxy does
* fix(health): test connection honors a stored mode only for the stored model and rejects a non-string mode
A request that selects a stored deployment and sends a different litellm_params.model now resolves the probe mode from that model instead of the stored model_info.mode. A litellm_params.mode that is not a string answers 400 instead of 500. The Bedrock Mantle rule that Claude models are probed over the Messages API moves into the provider package.
* fix(health): shape test connection probe params for the model the request probes
A request that selects a stored deployment by id and overrides the model
resolved its probe mode from the overridden model but still injected
max_tokens from the stored mode, so an embedding override of an
anthropic_messages deployment failed with a Mistral 422 extra_forbidden
* fix(health): report an early ahealth_check failure as itself, not as a missing mode
With the mode resolved automatically when the caller omits it, a failure
before that resolution (no model, a non-string model, a provider that does
not resolve) was wrapped as "Missing mode", a hint that pointed at the wrong
fix and dropped raw_request_typed_dict from the result. Every failure now
returns the same shape.
---------
Co-authored-by: mateo-berri <277851410+mateo-berri@users.noreply.github.com>
* fix(auth): resolve hidden model_group_alias entries in the zero-cost budget check
Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>
* fix(auth): key the zero-cost cache by the resolved model group
Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>
* fix(auth): keep the zero-cost verdict per requested name and include hidden aliases
* test(integration): audit the zero-cost bypass through hidden model_group_alias names
* test(router): cover the extracted routing strategy switch
* test(integration): record a pre-flip burst before the alias flip
---------
Co-authored-by: jesus <jesus@berri.ai>
Co-authored-by: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>
Co-authored-by: mateo-berri <277851410+mateo-berri@users.noreply.github.com>
* feat(ui): prototype observed engineering ROI dashboard
* feat(roi): replace effort estimates with measured repository metrics
* fix(roi): finish connection recovery and generated API contracts
* fix(roi): show merged changes before accounts are linked
* fix(roi): preserve selected report tab across refreshes
* fix(roi): recover app authorization and keep detail values readable
* fix(roi): reuse the shared OAuth HTTP client
* feat(roi): combine providers and compare equal reporting periods
* docs: explain ROI metrics for first-time readers
* fix(roi): preserve connections and scheduled reports during setup
* ci(roi): assign database contracts to the active Postgres shard
* fix(roi): preserve issue counts and normalized connections
* fix(ui): compact ROI dashboard header and metrics
* fix(ui): show ROI repository count with expandable list
* fix(ui): wrap ROI controls within narrow panels
* fix(roi): restore sample report preview and simplify setup
* feat(proxy): serve a Codex-native model catalog with per-model service_tiers from /v1/models
GET /v1/models and /models answer Codex CLI's catalog fetch (the request
carrying its client_version query parameter) with Codex's own
{"models": [...]} shape: a model Codex knows keeps the metadata of its
bundled 0.159.3 catalog (vendored), any other model gets Codex's fallback
entry, and model_info.service_tiers becomes each entry's service tiers so
Codex offers them as slash commands that send service_tier upstream.
Without the parameter the OpenAI list shape is unchanged. The CLI's
litellm agents codex catalog shares the same builder.
* fix(proxy): offer a Codex service tier only when every deployment of the model lists it
* fix(codex-catalog): an invalid service_tiers value offers no tier for the model
* fix(codex-catalog): read service tiers off the deployments the key's team can route to
A tier is offered to Codex only when every deployment of the model name a
request from the key's team can route to lists it, so another team's
deployment of the name and a deployment an admin paused via model_info.blocked
no longer withhold or add tiers for requests that never reach them
The catalog's always-null fields are annotated NoneType so the module imports
under pydantic 2.12.0 on Python 3.14, the lowest pin the MCP resolve job
installs, which rejects a None annotation with a None default
* test(codex-catalog): drop the redundant module docstring and sort the imports
* test(integration): add the Codex catalog audit cells and the multi-worker convergence note
* test(integration): clean up every catalog test model and answer the refresh GET
* fix(proxy): keep tiered models under Codex's catalog cut and resolve alias tiers
Under Codex's 1 MiB catalog limit the entries offering a service tier are kept
ahead of those offering none, each group in model_list order, with every kept
entry at its listing position, so the model an operator configured tiers for
survives a wide key's long listing. A model_group_alias row reads its target's
deployments, so it carries the target's tiers and stock metadata under the
alias name.
* fix(proxy): pick Codex catalog metadata per team and skip entries too large for the cut
The upstream model that selects Codex's stock entry was read off the first deployment of a name
without checking the key's team, so a team whose requests route to a different deployment could be
handed another team's prompt, reasoning levels, and tiers. The upstream model and the tiers now come
from the same team-aware selection routing uses, and a caller with no team reads the deployments no
team owns
The byte cut kept a prefix of the tier-first order, so one entry larger than the whole limit emptied
the catalog. An entry too large for the bytes left is now passed over and the smaller ones after it
are still kept
---------
Co-authored-by: mateo-berri <277851410+mateo-berri@users.noreply.github.com>
* feat(lens): record run steps, trigger and exact run windows on investigations
* feat(lens): scan only traces since the last run and keep a capped step log
* feat(lens): log each analysis model call with its model, tokens and cost
* feat(lens): accept agent and time window on run now and add turn-all-on
* test(lens): cover new-traces-only windows, manual runs and the step cap
* test(lens): cover run now overrides and turning paused investigations on
* chore(ui): regenerate api types for lens run steps and run now options
* feat(lens): group open findings into one row per problem with agent filters
* test(lens): cover the findings table grouping, filters and schedule labels
* feat(lens): add a findings table across all investigations
* feat(lens): show a live step feed with the model behind each call
* feat(lens): offer turning all paused investigations on
* feat(lens): let run now pick an agent and time window
* test(lens): cover run now request building
* feat(lens): show the step feed and run now dialog on an investigation
* feat(lens): open run now choices instead of running immediately
* feat(lens): open findings first and peek a finding without leaving the table
* feat(lens): fold investigation actions into the findings toolbar
* feat(lens): show each investigation's schedule and open findings
* feat(lens): name the agent on a finding
* feat(lens): keep new investigations watching every 15 minutes by default
* feat(lens): show the watch schedule outside advanced options
* feat(lens): send run now options and turn-all-on from the dashboard
* test(lens): give demo runs steps and a trigger
* feat(lens): let the findings table fill the screen
* test(lens): add steps and trigger to progress fixtures
* test(lens): add steps and trigger to status fixtures
* test(lens): cover the default watch schedule in setup
* test(lens): cover run now choices from an investigation
* test(lens): open saved investigations from the manage view
* feat(lens): use one tab bar for traces, findings and investigations
* feat(lens): place page actions on the lens tab row
* feat(lens): drop the nested tabs and edit investigations in place
* feat(lens): show investigations as a table with run now and edit
* feat(lens): name each findings row for screen readers
* test(lens): open saved investigation links on findings
* test(lens): reach findings and investigations from the top tabs
* fix(lens): mark run now jobs manual and keep them from moving the scheduled scan
* fix(lens): keep run now since-last-run windows even with an agent override
* test(lens): cover that manual runs never skip scheduled traces
* test(lens): cover run now windows with agent and lookback overrides
* fix(lens): group findings without Map.groupBy and expose sampled runs
* fix(lens): open older findings and review every merged copy from one row
* fix(lens): hide edit and run now from read-only viewers
* chore(lens): drop restating comments from the findings table
* chore(lens): drop restating comments from the step feed
* chore(lens): drop restating comments from the paused banner
* chore(lens): drop restating comments from header actions
* chore(lens): drop restating comments from run now
* test(lens): cover merged findings and read-only investigation rows
* fix(lens): record a model step even when the response has no usage
* test(lens): cover model steps with and without reported usage
* fix(lens): keep a merged finding open when one of its updates fails
* refactor(lens): accept update results from the investigations view
* refactor(lens): accept update results in investigation actions
* test(lens): cover retrying a merged finding after a failed update
* feat(decisions): add unified /v1/decisions endpoint for Jev-compatible providers
Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>
* fix(decisions): register typesafe as a provider so Jev deployments load
Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>
* refactor(decisions): move provider endpoints under llms and validate proxy bodies
Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>
* feat(decisions): add Cloudflare Clef and Strands Decider backends
Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>
* fix(decisions): register decisions routes for managed agents and gateway
Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>
* test(decisions): use raw regex for cloudflare missing account match
Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>
* fix(decisions): avoid cast in Cloudflare response unwrapping
Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>
* fix(decisions): default model, evaluation health probe, short Cloudflare names
The proxy validates only state and questions, so a request without a
model falls through to the configured default model like every other
route. Health checks probe evaluation-mode deployments through the
Decisions API instead of failing with an unsupported mode, and
cloudflare/clef and cloudflare/clef-flash get cost-map rows so the short
names resolve a mode and a price. The registry no longer claims typed
decisions for a provider with no backend.
* fix(decisions): let health_check_params override the evaluation probe
Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>
* test(integration): audit the decisions endpoint across providers, limits, health and chaos
Adds the /v1/decisions audit cells: one wire contract per provider (path, key, body and cost-map billing), the gateway-only fields and tags, the sad paths (invalid bodies, unknown model, key checks, api_base in the body, upstream 401/429/500, a 200 without answers, an unreachable upstream), the two evaluation-mode health probes, and three chaos cells (a mixed-failure burst over both routes, a worker SIGKILL mid-burst, an upstream outage and restart on the same port).
The PR's cost case read the upstream observations through the gateway, which answers 404 for that path; it now reads them from the upstream URL. The owned proxy harness takes extra CLI arguments, and its graceful stop waits as long as a worker boot may take, since a worker still starting honors SIGTERM only once it is up and the 30 second wait forced a cleanup under load.
* fix(decisions): send env API keys to a configured api_base
Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>
* feat(decisions): add zero-cost evaluation cost-map entry for Strands Decider
Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>
* fix(decisions): register the routes through the lazy feature registry
The Decisions router was included at import, ahead of the config and DB
pass-through endpoints, so a pass-through configured at /v1/decisions
was skipped and answered 400 as an unknown Decisions provider. The
routes now register through LAZY_FEATURES, which splices them in after
every eager route, so a pass-through at /v1/decisions keeps its route
while /decisions still serves natively. The lazy OpenAPI snapshot carries
the two paths so the schema shows them before the first call.
The audit cells add the env-key egress to a configured api_base, the
client api_base opt-in shared with chat, the pass-through precedence on
an owned proxy, and the Strands evaluation health check resolved from
the cost map. The integration config exports the Perplexity env key the
first cell needs.
* fix(decisions): keep the Cloudflare api_base message in its transformation and read the audit upstream once per cell
* fix(proxy): let a config pass-through beat a lazily registered route in eager mode
With LITELLM_DISABLE_LAZY_ROUTES set the decisions routes are registered at
startup, so SafeRouteAdder treated a config pass-through at exactly
/v1/decisions as already registered and dropped it. In lazy mode a pass-through
created through the API after the first native call was skipped the same way.
Routes a lazy feature owns no longer count as registered, and a route added at
one of their paths is placed ahead of them, the precedence lazy mode gives a
config pass-through when the feature has not loaded yet.
---------
Co-authored-by: mateo <mateo@berri.ai>
Co-authored-by: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>
Co-authored-by: mateo-berri <277851410+mateo-berri@users.noreply.github.com>
is_prompt_caching_valid_prompt ran the full Python token_counter over every message to compare against the deployment's prompt cache minimum, 500 to 1000 ms at 440k to 740k tokens on every request that reaches the prompt_caching pre-call check, Rust on or off. messages_reach_token_count does the same arithmetic as token_counter(...) >= threshold and stops at the first message that reaches the threshold. Groups with one healthy deployment skip the prefix hash and pin lookup, which cannot change the result for them
Four fixed name span events make the pre-LLM phases measurable with OTel v2: litellm.request.body_received (with body_bytes) once per body read before parsing, on the JSON, binary and form branches, body_parsed, pre_call_completed, and deployment_selected emitted once per pick inside Router.async_get_available_deployment and get_available_deployment with attempt, reason and model group, so every router surface, retry and fallback is covered. Measured locally on /v1/chat/completions, /v1/messages and /v1/responses at 440k tokens with Rust on and off against a fake upstream
Co-authored-by: yassin <yassin@berri.ai>
Co-authored-by: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>
Response cache reads and writes open cache.get llm_response and cache.set llm_response phase spans with their Redis spans nested underneath, on the Python path and on the native Rust path, and deployment selection runs inside a route {model_group} phase so the cooldown, usage and model-id reads the router issues nest under it before chat {model}. The autorouter classifier call nests under that route phase as well and carries its typed internal origin on litellm.request.purpose, so it is told apart from the provider attempt. Service spans are named {service}.{verb} {target} from a low-cardinality key family the producer declares (llm_response, auth_objects, spend_counters, router_cooldowns, claude_code_session_router_binding, rate_limits, pod_lock, budget_reset, ...) instead of the raw method or a per-request pipeline length; a pipeline flush is targeted by the one family its ops share or by mixed with the sorted families on litellm.redis.families, a batch op keeps the family it was declared under whichever pipeline or standalone read settles it, and the ambient family labels Redis spans only, never the DB write-back a task spawned inside that context performs later. The raw method stays on litellm.service.call_type and on the Prometheus and Datadog labels. Caller attribution is carried across asyncio task boundaries on a ContextVar so forwarder-only chains no longer surface, the raw cache key is dropped from Redis span metadata, pipeline op counts land as an integer attribute, every call_type the Redis cache layer emits maps to a verb, and a scan over litellm/ and enterprise/ fails when a Redis producer, batch reservation included, declares no key family.
A V2 logger built for a key or team logging entry while the operator's V2 logger is already registered keeps only the exporters its own preset contributed, whether or not the operator holds credentials for that backend, so every chat span no longer reaches the operator's collector twice. A span the success callback has to open itself, with no pre-call carrier, starts at the provider handoff (api_call_start_time) instead of the logging object's creation.
Co-authored-by: yassin <yassin@berri.ai>
Co-authored-by: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>
log_db_metrics wrapped whole cache-first auth helpers and always emitted a ServiceTypes.DB success event, so in-memory cache hits showed up as postgres <fn> spans and DB service metrics. The decorator now installs a ContextVar witness that _TrackedPrismaEngine marks on every Prisma query and transaction call, and the DB event is emitted only when the witness was marked. Real reads keep their existing call_type names, the failure path and the PROXY batch-write branch are unchanged, and Redis instrumentation is untouched.
A decorated helper that reaches Prisma only through another decorated helper (get_key_object -> get_object_permission, get_team_object_by_alias -> get_object_permission, get_tag_object -> get_tag_objects_batch) used to emit two events for one query. The inner wrapper now marks its witness as reported when it emits a success or DB failure event, and only unreported activity is handed up to the enclosing witness, so the inner event is the one that survives. An outer helper that also queries Prisma directly or through undecorated callees still gets its own event.
Tests: get_user_object and get_org_object cache hits emit no DB event; a get_user_object miss through the generated Prisma client emits exactly one postgres get_user_object event; decorator-level tests cover nested calls emitting only the inner event, outer calls with their own query, inner non-DB failures, bounded lookups and sibling-request isolation.
Co-authored-by: yassin <yassin@berri.ai>
Co-authored-by: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>
* feat(dev): seed linked tracing and spend fixtures
* chore(dev): use OpenAI model in tracing config
* chore(dev): align tracing credentials with UI E2E
* fix(dev): update fixture seeder query scope
* feat(dev): seed linked tracing and spend fixtures
* chore(dev): use OpenAI model in tracing config
* chore(dev): align tracing credentials with UI E2E
* fix(dev): update fixture seeder query scope
* wip
* wip
* wip
* chore(trace): checkpoint ongoing Rust migration
* refactor(trace): group Python bridge under trace package
* refactor(traces): read span conventions through a Convention trait
Each span format (Claude Code, LangSmith, OpenInference, gen_ai) now lives under
normalize/convention/ as a unit struct implementing Convention, owning both its
detection and its extraction. Precedence is one ordered registry instead of an
if-chain in mod.rs that reached into each module differently.
The modules now share one way to read attributes: present() for the first
non-empty key and Payload for a text that also reports the key it consumed,
replacing three different idioms and the &mut Vec threaded through payload
readers. Instrumentation::adjust returns a new Extraction instead of mutating
one, with each SDK rule as its own function, and the LangChain middleware
suffix list exists once.
* feat(trace): export Rust-owned wire schemas and enforce contract bounds
* fix(trace): bound quoted counts in ClickHouse wire schemas
* feat(trace): generate Python wire contracts with datamodel-code-generator
* test(trace): validate migrated callers and generated contracts at the native boundary
* refactor(traces): rename normalization convention to format
* fix(traces): reconcile spend evidence and preserve unknown costs
* feat(traces): normalize additional telemetry formats
* test(traces): cover captured normalization fixtures
* refactor(traces): isolate SDK normalization rules
* feat(tracing): seed all trace exports for local dashboard
* fix(clickhouse): preserve custom LiteLLM request metadata
* docs(traces): define normalization module boundaries
* docs(traces): define resolution and OTLP boundaries
* fix(ui): normalize nullable trace message names
* refactor(traces): split resolver modules and cover resolution behavior
* test(traces): replace normalization snapshots with behavior assertions
* fix(ui): align dashboard API contracts with generated types
* refactor(traces): type normalization and storage boundaries
* fix(traces): seed captured SDK spend and preserve provider identities
* wip
* test(traces): verify guide discovery and content ordering
---------
Co-authored-by: Yujong Lee <yujong@berri.ai>
Co-authored-by: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>
* feat(lens): simplify trace inspection in the gateway drawer
* feat(lens): add a conversation view for full traces
* test(lens): keep normalized message fixtures type safe
* refactor(lens): make conversation view read like a chat
* refactor(lens): use a quiet trace view menu
* refactor(lens): make trace view tabs explicit
* refactor(lens): restore compact trace view switch
* fix(lens): preserve complete conversation history and tool types
* feat(lens): add trace full-screen and close controls
* fix(lens): show forwarded answers and agent errors once
* fix(lens): reset full screen when closing a trace
* fix(roi): estimate linked authors and clarify model selection
* fix: preserve trace errors and ROI results across partial failures
* fix(roi): correct pagination variable typing
* fix(ui): place loaded root failures in conversation order
* fix(roi): read estimator recommendations from model catalog
* revert: remove catalog-driven ROI recommendations
* fix(proxy): stamp the client alias on a copy of each streamed chunk so pricing sees the deployment model
Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>
* test(spend): streamed alias matching a capability rule bills the deployment price
Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>
* test(spend): assert every streamed chunk carries the client alias
Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>
* fix(logging): log the client alias on the priced streamed response, the same as non-streamed
Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>
---------
Co-authored-by: kerry <kerry@berri.ai>
Co-authored-by: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>
* feat(roi): support GitLab and tagged branch costs
* fix(roi): count tagged branches independently of estimation status
* test(roi): capture live GitHub and GitLab report validation
* fix(roi): open estimate details at the start
* fix(roi): clarify cost views and unify report layout
* feat(roi): showcase per-PR costs in the sample report
* fix(roi): separate report tabs and preserve branch cost attribution
* fix(roi): preserve demo previews and align progress spacing
* fix(roi): isolate demo loading and parallelize fork lookups
Preserve active sync status when source changes finish saving, keep live reports available when demo requests fail, and cover each review regression
* fix(roi): separate demo and live loading states
Clear the demo URL on fallback, wait for live requests on exit, and retain request errors until the corresponding operation recovers
* fix(roi): ignore refreshes from a previous source
* fix: trust gateway context for ROI estimator exclusion
* fix: preserve historical ROI estimator exclusion
* fix(search_tools): encrypt search tool litellm_params at rest
Encrypt every string value of a search tool's litellm_params on create and
update and decrypt on every DB read, so legacy plaintext rows load unchanged.
Include the table in master key rotation, LITELLM_MIGRATE_FROM_MASTER_KEY and
the migrate-encryption scan.
* fix(search_tools): keep edits made while the master key rotates
Write each rotated search tool row only if it still holds the litellm_params
that were read, and re-read and rotate it again if it was edited in between,
so a PUT that lands during /key/regenerate is not overwritten.
* fix(search_tools): retry rotation writes until the row stops changing
Rotate a search tool row again for as long as it keeps being edited instead of
giving up after five attempts, and stop with a warning only when the conditional
write fails on an unchanged row. Build the decrypted read result without
mutating it in place.
* refactor(search_tools): rotate edited rows in a loop, drop the step comment
Retry the conditional rotation write in a loop instead of recursion so sustained
edits cannot deepen the call stack, drop the step comment on the rotation call,
and stop mutating local state in the rotation tests.
* test(search_tools): drop the rotation test docstring
* Store search tool params as written when no encryption key is configured
* Rotate search tools under the salt key, keep non-ciphertext values and loaded tools that do not decrypt
* Treat a search tool as undecryptable only when its provider is ciphertext-length
* Drop suppressions the type discipline gate on main now reports as unused
* Show the loaded search tool in the admin list and info views when its DB params do not decrypt
* Keep the DB row's other fields when the admin views substitute loaded params
* fix(guardrails): encrypt guardrail litellm_params secrets at rest
* fix(guardrails): keep salt-key encryption on master key rotation and retry rows edited mid-rotation
- rotate guardrail params under LITELLM_SALT_KEY when set, matching the key reads decrypt with
- re-read and retry a row whose updated_at moved during rotation, up to GUARDRAIL_ROTATION_ATTEMPTS
- build decrypted Guardrail rows and the rotation count without mutating locals
* refactor(guardrails): retry guardrail rotation by bounded recursion instead of a rebound cursor
- each attempt re-reads the row and recurses with attempts_left - 1, so no loop variable is rebound
- cover the give-up path after GUARDRAIL_ROTATION_ATTEMPTS writes
* test(guardrails): drive the real guardrail rotator from the master key rotation test
- inject an encrypted guardrail row through the prisma client instead of replacing the GuardrailRegistry method
- assert the written params decrypt under the new master key
* Annotate guardrail param encryption collections for type-discipline gate
* Type guardrail param recursion through validated JSON containers
* Type guardrail registry test helpers and drop section comment
* Reject client-supplied encrypted values in guardrail litellm_params
* Allow depth-bounded contains_encrypted_marker in the recursion detector
* Keep a loaded guardrail when its DB params do not decrypt with the current key
* Apply other DB edits while keeping loaded values that do not decrypt, including PATCH models
* Keep the loaded guardrail when an undecryptable param has no loaded value
* Drop suppressions the type discipline gate on main now reports as unused
* Assert what the reinitialized guardrail holds after an edit to an undecryptable one
* Drive the rotation sync tests through a registered guardrail instead of patching reinitialize
* Type the rotation test helpers and drop the new test docstrings
* fix(guardrails): refuse to approve a submission whose params do not decrypt
Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>
---------
Co-authored-by: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>
* fix(proxy): return 400 instead of 500 for missing required params and invalid pagination
Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>
* ci: run search_endpoints tests in proxy-endpoints shard
Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>
* fix(proxy): return 4xx for missing required params across all LLM routes and propagate provider status on lookups
Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>
* fix(llm_http_handler): keep provider error text when re-raising mapped errors
Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>
* fix(proxy): allow promptless image edits and default search models
Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>
* fix(proxy): default missing image edit image to None
Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>
* refactor(proxy): build image edit defaults without mutating request data
Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>
* fix(proxy): keep image edit defaults within type-discipline budget and give request mocks a scope
Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>
* test(proxy): inject a fake router for the search default model test
Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>
* test(llms): cover provider error status on vector store and file lookup handlers
Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>
* test(llms): keep the lookup handler raise block to a single statement
Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>
* test(llms): cover provider error status on eval, eval run, skill and vector store file content lookups
Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>
* test(integration): cover missing required body params and provider lookup status codes
Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>
* ci: run tests/unit/proxy/search_endpoints in the proxy-endpoints shard
Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>
* test(integration): bind spend-row request id with partial to satisfy B023
Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>
* fix(proxy): only reject non-positive page_size on vector store list
Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>
* fix(proxy): remove unreachable fine-tuning body validation
Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>
* test(integration): cover streaming anthropic messages reaching the upstream
Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>
* test(integration): count only provider calls when asserting missing params never reach the upstream
Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>
* fix(proxy): preserve merge-base request compatibility
Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>
* fix(proxy): preserve interaction completion model defaults
Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>
* fix(proxy): retry model read-through before rejecting params a DB-only deployment may default
Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>
---------
Co-authored-by: shivam <shivam@berri.ai>
Co-authored-by: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>
Co-authored-by: yucheng <yucheng@berri.ai>
* fix(vector_stores): return managed file ids from vector store file list
Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>
* test(vector_stores): cover managed file list route
Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>
* fix(vector_stores): only map round-trippable managed ids and index flat file ids
Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>
* fix(proxy-extras): build managed file gin index concurrently
Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>
* chore(proxy-extras): move the managed file gin index migration after main's newest
* fix(vector_stores): satisfy lint gate
Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>
* chore(proxy): drop the stale no-index note on the raw file-id guard
* test(vector_stores): cover managed file ids on the vector store file list end to end
Integration cells for GET /v1/vector_stores/{vs}/files mapping provider file ids back to
the caller's owner-scoped managed ids and decoding managed after and before cursors: raw
httpx, the OpenAI SDK sync and async pagers, the three credential routing modes, the owner
filter branches, raw and unmappable cursors, provider errors, duplicate and non-string ids,
a provider outage mid-burst, a worker SIGKILL mid-burst, and the GIN index migration applied
by the migration entrypoint and by db push
---------
Co-authored-by: shivam <shivam@berri.ai>
Co-authored-by: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>
Co-authored-by: mateo-berri <277851410+mateo-berri@users.noreply.github.com>
* feat(interactions): durable cross-pod settlement for background interaction billing
Background interaction billing lived only in the creating replica's memory, so a
DELETE routed to another replica, or a restart of the creating one, never billed
the completed provider work and the budget reservation was refunded at the poll
timeout. The create now registers the billing context in a settlement store
before returning, the proxy installs a Prisma-backed store at boot
(LiteLLM_BackgroundInteractionSettlement, schema-only migration), any replica
claims the row once through a conditional update before billing or releasing,
startup resumes every unclaimed row with its remaining timeout, and a give-up
records an unsettled outcome instead of silently reconciling to zero. The SDK
keeps an in-memory store and behaves as before.
* fix(interactions): survive a settlement install failure at boot and stop carrying request headers
* fix(interactions): drop the stored request context once a settlement row is settled
* fix(interactions): bill the completed response a poll already saw when its claim only answers at the deadline
* fix(interactions): carry a missing model through the settlement context for agent-only background creates
An interaction created with an agent and no model reaches the poll with no model name, exactly as on main. The settlement context now stores that None instead of rejecting the create, which answered the client with a 500 after the provider had already accepted it.
* fix(interactions): leave an unfetchable background interaction to its creating poll when a delete lands elsewhere
The remote pre-delete path fetches with only the delete's credentials, so a fetch it cannot make says nothing about the interaction. It used to claim the settlement row and release the reservation anyway, which stopped the creating replica's poll and lost the bill when the delete then failed the same way. It now returns without claiming; the in-process path keeps releasing on an unfetchable state, since its context carries the create's own credentials.
* fix(interactions): fail a cross-replica delete when its pre-delete fetch fails so the creating poll keeps the bill
* fix(interactions): keep the stored settlement gate when registration raises after landing, and fail resumed-poll deletes closed
A registration that raised after its row committed moved the poll to a private in-memory gate, so the creating worker billed while the stored row stayed unclaimed for another replica's delete or the next boot to bill again. The row is now read back once and, when it landed, the poll claims through it like every other settler.
A worker that resumed the poll after a restart is not the creator, so its delete on a failed pre-delete fetch now fails with the fetch's error instead of releasing and deleting. After a fleet restart every worker holds resumed polls, which left the fail-closed path applying nowhere.
* fix(interactions): settle an unverified registration through the durable claim
A create whose settlement-store write raised no longer bills through a
private in-memory gate that a later boot's resume cannot see. The claim
asks the durable store first and falls back to the local gate only when
the store answers that no row exists, and a missing settlement table reads
as no rows so a replica without the migration still settles in process.
* test(proxy): keep the settlement test where the proxy-infra shard collects it
The merge of main moved test_background_interaction_settlement.py under
tests/unit/proxy/spend_tracking, but the proxy-db shards claim tests/unit/proxy
files one by one in .circleci/scripts/unit_selection.sh, so no CI shard ran it
and codecov/patch dropped. tests/test_litellm/proxy/spend_tracking is collected
whole by the proxy-infra shard, which is where the test ran before the merge.
* fix(interactions): raise on a non-2xx Gemini interaction fetch
AsyncHTTPHandler.get never raises for status and the Gemini GET transform
only raised when the body was not JSON, so a 500 or 404 carrying Gemini's
JSON error body parsed as an interaction with no status. A delete on a
replica other than the creator then claimed the settlement as released and
forwarded the delete instead of failing closed, and the bill was lost. The
transform now raises GeminiError with the vendor's status, as the delete
transform already does; the in-process poll already retries a fetch that
raises
* test(integration): audit durable background interaction settlement across replicas
Twenty-six deterministic cells drive a one-worker creator and a two-worker
settler against an owned scripted Gemini upstream: cross-replica deletes
bill once, failed and cancelled interactions release, a later replica
resumes unclaimed rows, custom deployment pricing bills at the deployment
rate, a fetch the settler cannot make fails the delete closed, odd ids are
refused, a missing settlement table keeps in-process billing, polling
disabled registers nothing, the budget reservation is released by the
settler, an upstream outage mid-burst fails closed and recovers, killed
workers hand their polls to the respawned ones, and concurrent deletes on a
slow upstream settle exactly once. The support upstream gains a scripted
interaction store with per-id GET status and delay, and the process helper
gains an owned upstream a test can stop and restart
* test(integration): refuse a repeated delete in the scripted upstream and pin the settlement budget below one estimate
* chore(ui): regenerate dashboard API types after merging main
* test(integration): accept the 422 budget refusal and a respawned worker's resume
The budget cell pinned a 400 that the proxy stopped answering when budget refusals moved to 422, so it now asserts the status and the budget_exceeded error type the sibling budget tests pin. The later-booting replica cell accepts a claimer that is any worker started after the creates, since uvicorn's supervisor can respawn the creator's worker under load and the respawned worker's boot resume claims the rows by design; the single spend row check is unchanged
* test(integration): delete the pinned key's interaction with a second key
A key whose budget is filled by its own reservation is refused on every route, the DELETE included, so the cell now asserts that 422 and sends the delete with a second key, which is what the reservation release on another replica needs in order to be observable at all
---------
Co-authored-by: mateo-berri <277851410+mateo-berri@users.noreply.github.com>
* fix(proxy): use OpenAI workload identity federation tokens on /openai_passthrough
The OpenAI passthrough routes (HTTP and websocket) only looked up a static
OpenAI API key, so proxies authenticating to OpenAI through workload identity
federation failed with "Required 'OPENAI_API_KEY'". When no non-empty static
key is configured, exchange the workload identity subject token for a bearer
token, scoped to the passthrough's own OPENAI_API_BASE so the token is never
sent to a non-OpenAI host
* fix(proxy): close the OpenAI websocket passthrough cleanly when the workload identity exchange fails
A rejected, unreachable or unreadable workload identity exchange raised out of
the websocket route before the handshake was accepted, so clients saw a bare
handshake failure. Log the cause and close with 1011 and a fixed reason instead
* refactor(openai): resolve workload identity bearer tokens for an api base in the OpenAI provider module
The passthrough route only owns the static key lookup now and asks the OpenAI
provider module for a workload identity bearer token scoped to its api base
---------
Co-authored-by: Krrish Dholakia <krrish+github@berri.ai>
Per-router auto-router savings were missing on proxies with
disable_spend_logs and for requests without a session id under
missing_session_id: omit. The router-day row is written by the
auto-router turn, whose enqueue sat behind the spend-logs flag and whose
builder dropped sessionless requests, while LiteLLM_DailyUserSpend has
neither gate. That money then showed only as unattributed savings.
Enqueue the turn whether or not spend logs are kept, and keep a
sessionless turn with an empty session id that writes the router-day row
while the session upserts skip it. The day and session rows still commit
in one statement, so late baseline corrections keep their ordering.
Without spend logs, write only the router-day aggregate: the turn drops
its session id, so no per-session row is stored, and baseline capture is
skipped, since a baseline observation can only publish once its
request's spend log exists. The flag keeps its meaning of no per-request
or per-session data, while the daily per-router money matches
LiteLLM_DailyUserSpend.
Co-authored-by: Claude Opus 5.5 <noreply@anthropic.com>
* refactor(lens): move analysis prompts into markdown files
* feat(lens): ask the investigator for a scoped agent fix brief with two options
* test(lens): cover the agent fix brief through investigation and merges
* chore(ui): regenerate api types for the lens fix brief
* feat(ui): build copyable lens fix prompts
* feat(ui): show the lens fix brief with copy buttons for claude code and codex
* test(ui): cover copying a lens fix option
* refactor(lens): replace the fix options with a plain issue brief
* feat(lens): ask for problem, user goal, outcome and test cases without prescribing code changes
* test(lens): cover the issue brief through investigation and merges
* chore(ui): regenerate api types for the lens issue brief
* refactor(ui): drop the lens fix prompt builders
* feat(ui): add a lens issue brief panel
* feat(ui): show the lens issue brief in the finding drawer
* test(ui): cover the lens issue brief and the legacy fallback
* feat(ui): render a lens issue brief as a markdown document
* test(ui): pin the lens issue brief markdown layout
* feat(ui): show the issue brief as a copyable file with claude code and codex buttons
* feat(ui): pass the finding title into the issue brief
* test(ui): cover copying the issue brief for claude code and codex
* feat(ui): bold the input and expected labels in lens test cases
* test(ui): pin the bold test case labels in the issue brief
* feat(ui): render the issue brief as formatted markdown
* test(ui): cover the rendered issue brief sections and raw markdown copy
* feat: add Bespoke Nimble gateway and OSS classifier support
* feat: accept Ollama's nimble model name for the Bespoke provider
* test: exempt the POST-only bespoke decisions route from the all-methods check
test_pass_through_routes_support_all_methods requires every built-in
pass-through route to accept every HTTP method unless it is listed in
PROTOCOL_CONSTRAINED_PASS_THROUGH_ROUTES. /bespoke/v1/systemone is
POST-only like /laya/v1/systemone, so the test failed at this branch
and passed at the merge base. List it alongside Laya.
Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com>
---------
Co-authored-by: mateo-berri <277851410+mateo-berri@users.noreply.github.com>
Co-authored-by: Claude Fable 5.1 <noreply@anthropic.com>
* refactor(proxy): extract shared spend log read policy
* test(proxy): use named bindings for spend scope regression
* test(proxy): reuse existing spend log query harness
* test(proxy): cover spend log permission lookup adoption
* chore(proxy): relocate existing spend query baseline
* refactor(proxy): make scope query returns explicit
* refactor(proxy): inject deferred log permission lookup
* test(proxy): cover teamless management compatibility lookup
* refactor(proxy): compose user and team log grants
* refactor(proxy): share generic authorization composition
* refactor(proxy): compose trace read permissions
* refactor(proxy): centralize spend and trace authorization
* refactor(proxy): strengthen spend and trace scope types
* refactor(proxy): flatten log read scope into owned logs
Replace the AnyOf grant tree with a flat OwnedLogs(user_id, team_ids) scope,
and OwnedTraces(logs, api_key_hash) for traces, since every consumer flattened
the tree back into that shape.
A caller with no user id now gets an empty scope instead of matching ownerless
rows through Prisma's IS NULL. The dead request_id guard in ui_view_spend_logs
is removed, and the management facets inject the log team lookup and reuse
read_scope_sql instead of the list shim.
Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
* test(proxy): run spend scope tests through one SQLite emulator
Replace the string-matching payload emulator and the hand-rolled Prisma where
interpreter with one SQLite helper that runs the real scope SQL. Session scope
tests now go through the endpoint, including the no-user caller that must not
match ownerless rows. Drop duplicated lookup-failure and trace mapping cases.
load_permitted_log_team_ids returns no teams without a database instead of
relying on the resolver's broad except.
Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
* fix(proxy): unify log and trace ownership permissions
* test(tracing): align fixtures with ownership read scopes
* refactor(tracing): align query scopes with row ownership
* refactor(spend): make ownership SQL predicates explicit
* test(spend): validate ownership SQL against PostgreSQL
* docs(traces): drop key-row visibility from query help guide
Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>
* test(spend): reach the empty-memberships branch in team lookup test
Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>
* chore(ui): regenerate dashboard API types
Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>
---------
Co-authored-by: Claude Opus 5.5 <noreply@anthropic.com>
Co-authored-by: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>
* fix(proxy): stop queued registry read-throughs spending the resync budget
RegistryReadThrough.attempt serializes misses behind one lock, but every
request that was queued behind the first one still spent a unit of the
20-per-5s resync budget and re-ran the DB resync, even though the first
request had already loaded the object. A burst of more than 20 requests
for a model created on another worker therefore exhausted the budget and
the rest got 400 "Invalid model name".
attempt now checks whether the key is already loaded once it holds the
lock and returns early without touching the budget. Models check the
router's model names and deployment ids, guardrails and agents reuse
their existing registry lookups.
* test(proxy): gate the queued read-through test on events and cover each registry's loaded check
The queued-requests test now holds the first resync on an asyncio.Event instead of
a timed sleep and records calls in a recorder with tuple and frozenset state. New
tests show the model, guardrail and agent read-throughs each answer an object that
is already loaded without reading the database, so rewiring any registry's loaded
check now fails a test
* test(proxy): keep the queued read-through recorder inside its test and type the agent registry fixture
* refactor(traces): type the ClickHouse query help response
Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>
* test(traces): include agent names and frameworks in named contract round trips
Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>
* test(traces): cover native query help validation in the storage adapter
Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>
---------
Co-authored-by: Yujong Lee <yujong@berri.ai>
Co-authored-by: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>
* test(integration): answer the model-info refresh GET in the mixed MCP responses wire peer
The proxy's periodic model-info refresh sends GET /v1/models to the deployment api_base, which tripped the peer's /responses-only assertion when a tick landed mid-test
* test(e2e): wait for the api-keys URL after clicking Virtual Keys in onboarding
/ui already renders the Virtual Keys heading, so the helper returned before navigation finished. The late route change moved focus and closed the account menu popover in hideLiteAdmin
* test(proxy): stop two unit modules leaking app.openapi_schema and a session-wide Router
test_custom_openapi cached a stripped schema on app.openapi_schema and never cleared it, breaking later openapi route tests. test_proxy_reject_logging built a module-level Router that stayed in the live router registry all session and re-added cost-map keys during a reload. Reset the schema via monkeypatch and make the Router a function-scoped fixture
* perf(proxy): aggregate daily model usage per flush instead of upserting per request
* perf(proxy): drain queued model usage in the spend log flush job
* perf(proxy): queue model usage at request time instead of writing to the db
* test(proxy): cover batched daily model usage aggregation and retries
* test(proxy): read back model insights written by the batched flush
* fix(proxy): drain the whole model usage queue each flush so it cannot grow unbounded
* test(proxy): give the mock prisma client a model usage queue
* test(proxy): cover draining a model usage queue larger than one spend log batch
* refactor(mcp): extract upstream preparation and support modern clients
* fix(mcp): reject incompatible upstream transport before saving
* fix(mcp): serialize protocol validation with server updates
---------
Co-authored-by: Joshua Valluru <326636767+joshua-berri@users.noreply.github.com>
* wip
* wip
* test(traces): separate root status from diagnostic error counts
* test(traces): cover normalization precedence and fallbacks
* chore(cache): remove stray comments from trace PR
* test(traces): name lens test for shared query path
* fix(traces): place query implementation before test module
* test(traces): use unified read scope in migration tests
* ci(rust): allow feature checks to finish
* ci(mcp): allow dependency resolution to finish
* fix(traces): preserve key visibility and safe spend attribution