* fix(lens): pin worker dependencies and support approved image digests
* fix(lens): include locked dependencies and release identity in build context
* fix(lens): select the dev worker package for SHA-tagged charts
* feat(anthropic): workload identity federation and pluggable identity sources
Backend half of #38818 (internal copy of the fork PR #38013), rebuilt as one
commit on top of litellm_internal_staging without the dashboard changes.
Deployments on anthropic/ without a static api_key can exchange an OIDC
workload assertion for a short-lived sk-ant-oat01 token through a shared
RFC 7523 JWT-bearer engine. The assertion comes from a mounted token file,
an env token, a LiteLLM-signed issuer, or Keycloak, chosen per deployment,
per named credential, or through ANTHROPIC_IDENTITY_SOURCE. The federation
fields are server-owned: refused inline in request bodies and on
POST /model/new, proxy-admin only on credentials, and the token exchange
is pinned to api.anthropic.com unless LITELLM_ANTHROPIC_WIF_ALLOWED_HOSTS
adds a host. GET /credentials/{name}/jwks exports the public key set of a
LiteLLM-signed credential for the Claude Console.
The OpenAI federation trio from #39613 rides along on the backend side with
the same server-owned handling.
Fixes#28607
Resolves LIT-6107
Co-authored-by: derhornspieler <15236687+derhornspieler@users.noreply.github.com>
* fix(anthropic): let batch-result downloads mint from deployment params and accept host:port allowlist entries
The files handler enabled workload identity on batch-result downloads but never received the
deployment's litellm_params, so a deployment authenticating through a named credential could only
mint from process-wide env vars. It now threads litellm_params through to the auth header the way
the batch retrieve path already does.
LITELLM_ANTHROPIC_WIF_ALLOWED_HOSTS entries written as host:port were read by urlsplit as a scheme,
so the allowlist kept the raw entry while the exchange compared bare hostnames and refused the
gateway. Entries are now parsed as network locations whether or not they carry a scheme.
* fix(types): move the WIF kwargs key sets to a leaf module so the kwargs funnel imports without a cycle
* test(anthropic): pin case-insensitive matching of WIF exchange-host allowlist entries
* fix(anthropic): end workload identity federation errors without a period so the router suffix reads cleanly
* fix(proxy): decrypt stored litellm_params before the WIF write gate
* fix(proxy): hide WIF secret references from /health output
* fix(proxy): keep the proxy error shape on credential endpoint refusals
* fix(proxy): hide identity token file paths from /health output
* fix(anthropic): rename the federation workspace param so Bedrock's anthropic_workspace_id keeps working
The Bedrock Claude Platform route already reads anthropic_workspace_id from
optional_params, so banning that spelling as a server-owned federation
parameter broke a pre-existing client capability. The federation field is now
anthropic_federation_workspace_id (env ANTHROPIC_FEDERATION_WORKSPACE_ID),
which restores the base branch's behavior for Bedrock callers, drops the
Bedrock-specific hint from the refusal message, and deletes the unconditional
ban constant that no longer had a reader
* fix(auth): share one exchanged token across workers reading the same assertion
Anthropic accepts each identity assertion exactly once, so two uvicorn
workers reading the same token file both minting from it means the second
exchange is denied with jti_reused. Minted tokens now land in a per-user
0700 cache directory guarded by a file lock, so workers on the same host
reuse one exchange until the token expires or the assertion rotates. A 401
is only retried when the re-read assertion actually differs, and the denial
hint explains jti_reused. LITELLM_TOKEN_EXCHANGE_CACHE_DIR moves the cache
and an empty value disables it
* fix: keep anthropic federation from being shadowed or leaked
An empty or whitespace-only ANTHROPIC_API_KEY counted as set, so a federated
deployment sent an empty x-api-key on every call instead of minting a token.
Blank values now read as unset, and a real static key on a federated deployment
logs once that it outranks federation and nothing is being federated.
The exchange-host allowlist matched hostnames only, so a second process on
another port of an allowed host was trusted with the workload's identity token.
An entry that names a port now trusts that port alone, while a bare host still
trusts every port.
The shared token store exists so the workers reading one projected token file do
not each spend its single-use jti. A source that mints its own assertion per
exchange shares nothing with another worker, so it no longer writes a live token
to disk for a lookup that can never hit.
* fix: unlink a staged token file a failed write leaves behind
The 401 denial hint now also says federation ignores ANTHROPIC_WORKSPACE_ID, which the Bedrock Claude platform provider already reads.
* refactor: move anthropic jwks derivation behind a provider-owned tagged union
* fix: unlink the staged token file when its write fails at close
A buffered write only reaches the disk when the handle closes, so a full disk surfaces at close and left the staging file behind holding a usable token.
* fix(anthropic): close the staging descriptor before writing the shared token file
* fix(wif): judge federation writes by what they set, not what is stored
The admin gate read the stored deployment, so a team admin lost edit, delete
and Test Connection on any deployment carrying federation params. It now
returns early unless the submitted fields touch the federation surface, and a
Test Connection probe that points the deployment at its own api_base is still
refused, with the 403 no longer wrapped into a 500
The rest of the same review pass: POST /model/new refuses only a blocking
value of `blocked`, so a client that always sends `blocked: false` is not
turned away; a request body can no longer pick which federated identity to
mint as by naming a stored credential; an advisory refresh the executor
refuses disarms the entry instead of wedging the identity until the follower
timeout; the static-key shadow warning resolves its env fallback inside the
cache instead of once per request; credential writes drop nulls before
storing them; the token exchange validates the endpoint URL before reading an
assertion and keeps refusing redirects across a client heal; /health hides
every server-owned federation field from non-admins; and the async create_file
and create_batch paths say which setting is missing when the provider resolves
no URL
* fix(proxy): let a deployment write name a federated credential
reject_federated_credential_reference runs from is_request_body_safe, which
pre_db_read_auth_checks calls on every route, so it also fired on POST
/model/new, /model/update, /model/{id}/update and /health/test_connection. A
proxy admin could no longer attach a federated credential to a deployment over
the API or the Admin UI, leaving a static config.yaml entry as the only way to
configure the feature the rejection told the caller to go configure, and
_reject_non_admin_wif_write never got to make the call it exists to make.
is_request_body_safe now takes the route and skips only the credential-reference
check on the routes that reach can_user_make_model_call. Federation fields typed
inline into a body stay refused everywhere, and a call naming a federated
credential still cannot pick the identity it mints as.
* refactor(proxy): derive health display policy from the federation key sets
The health check module hand-copied the five workload identity fields whose
value is a credential, so a shared proxy surface named provider-specific
parameters and a newly added secret-bearing field would have gone on being
displayed until someone remembered both places
WIF_SECRET_BEARING_KEYS now sits beside the key sets it splits out of,
types/utils derives secret_bearing_wif_litellm_params from it, and the health
layer splats that tuple the same way it already splats the admin-only one
* fix(anthropic_wif): treat blank identity-source fields as unset
* test(proxy): classify the federation params in the credential slot registry
main's registry test (#43298) now fails the build for any credential-named
deployment param without a classification. The five federation fields that
carry a token, a token file path, or a signing or client secret reference are
Unplanted, matching WIF_SECRET_BEARING_KEYS; the four remaining Keycloak
settings name a URL, a client id, an auth method, or a scope and are NotSecret
* fix(anthropic_wif): declare federation params as owned connection leaves and chart their metrics
Register the 18 Anthropic and 3 OpenAI federation params as frozen
ConnectionSettings leaves so the owned-kwarg registry, the kwargs funnel
and the request-body ban list read one declaration. Pass the deployment
api_base through to the count-tokens handler instead of a pre-suffixed
URL, which doubled the /count_tokens path on main's prompt-cache
predictor. Add the five litellm_anthropic_wif_* families to the
all-metrics Grafana dashboard.
* fix(credentials): gate PATCH on WIF fields resolved from model_id
The credential PATCH handler checked server-owned workload identity
federation fields only on the values the caller sent, while a body that
named a deployment through model_id had its credential values resolved
after that check. A non-admin could therefore copy a federated
deployment's WIF fields onto an ordinary credential. Resolve the incoming
values first and run the non-admin gate on them, matching the POST path
* fix(anthropic): count tokens with ANTHROPIC_AUTH_TOKEN through the shared auth header
Count-tokens walked its own credential ladder: a static key, else skip minting when
ANTHROPIC_AUTH_TOKEN is set, else mint a federated token. With only the auth token set it
forwarded nothing and the proxy silently fell back to its local tokenizer while chat on the
same deployment authenticated with that token. The handler now takes the auth header that
AnthropicModelInfo.aget_auth_header resolves, the same ladder chat, files, batches and skills
use, and merges the oauth beta a minted or consumer token carries with the token-counting beta
---------
Co-authored-by: mateo-berri <277851410+mateo-berri@users.noreply.github.com>
Co-authored-by: derhornspieler <15236687+derhornspieler@users.noreply.github.com>
Co-authored-by: mateo-berri <happymvw@gmail.com>
* fix(health): probe Bedrock Mantle Claude deployments over the Anthropic Messages API
Bedrock Mantle serves Claude ids only on /anthropic/v1/messages, but health
checks probed every chat-mode deployment over /v1/chat/completions, so a
bedrock_mantle Claude deployment showed unhealthy while real /v1/messages
traffic to it succeeded
Add an anthropic_messages health check mode and make it the default for
bedrock_mantle Claude models. An explicit model_info.mode still wins, and
/health/test_connection and the Add Model form accept the new mode
* fix(health): resolve the test connection mode from the deployment when the request omits it
The Admin UI model page sent the mode /model/info had filled in from the cost
map back as the probe mode, so Test Connection on a Bedrock Mantle Claude
deployment still went over chat completions. The page now forwards only the
row's id, and /health/test_connection resolves a missing mode the way /health
does: the stored model_info.mode, then the mode the provider requires, then the
cost map.
* fix(health): resolve an omitted ahealth_check mode the way the proxy does
* fix(health): test connection honors a stored mode only for the stored model and rejects a non-string mode
A request that selects a stored deployment and sends a different litellm_params.model now resolves the probe mode from that model instead of the stored model_info.mode. A litellm_params.mode that is not a string answers 400 instead of 500. The Bedrock Mantle rule that Claude models are probed over the Messages API moves into the provider package.
* fix(health): shape test connection probe params for the model the request probes
A request that selects a stored deployment by id and overrides the model
resolved its probe mode from the overridden model but still injected
max_tokens from the stored mode, so an embedding override of an
anthropic_messages deployment failed with a Mistral 422 extra_forbidden
* fix(health): report an early ahealth_check failure as itself, not as a missing mode
With the mode resolved automatically when the caller omits it, a failure
before that resolution (no model, a non-string model, a provider that does
not resolve) was wrapped as "Missing mode", a hint that pointed at the wrong
fix and dropped raw_request_typed_dict from the result. Every failure now
returns the same shape.
---------
Co-authored-by: mateo-berri <277851410+mateo-berri@users.noreply.github.com>
* fix(auth): resolve hidden model_group_alias entries in the zero-cost budget check
Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>
* fix(auth): key the zero-cost cache by the resolved model group
Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>
* fix(auth): keep the zero-cost verdict per requested name and include hidden aliases
* test(integration): audit the zero-cost bypass through hidden model_group_alias names
* test(router): cover the extracted routing strategy switch
* test(integration): record a pre-flip burst before the alias flip
---------
Co-authored-by: jesus <jesus@berri.ai>
Co-authored-by: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>
Co-authored-by: mateo-berri <277851410+mateo-berri@users.noreply.github.com>
* feat(ui): prototype observed engineering ROI dashboard
* feat(roi): replace effort estimates with measured repository metrics
* fix(roi): finish connection recovery and generated API contracts
* fix(roi): show merged changes before accounts are linked
* fix(roi): preserve selected report tab across refreshes
* fix(roi): recover app authorization and keep detail values readable
* fix(roi): reuse the shared OAuth HTTP client
* feat(roi): combine providers and compare equal reporting periods
* docs: explain ROI metrics for first-time readers
* fix(roi): preserve connections and scheduled reports during setup
* ci(roi): assign database contracts to the active Postgres shard
* fix(roi): preserve issue counts and normalized connections
* fix(ui): compact ROI dashboard header and metrics
* fix(ui): show ROI repository count with expandable list
* fix(ui): wrap ROI controls within narrow panels
* fix(roi): restore sample report preview and simplify setup
* feat(proxy): serve a Codex-native model catalog with per-model service_tiers from /v1/models
GET /v1/models and /models answer Codex CLI's catalog fetch (the request
carrying its client_version query parameter) with Codex's own
{"models": [...]} shape: a model Codex knows keeps the metadata of its
bundled 0.159.3 catalog (vendored), any other model gets Codex's fallback
entry, and model_info.service_tiers becomes each entry's service tiers so
Codex offers them as slash commands that send service_tier upstream.
Without the parameter the OpenAI list shape is unchanged. The CLI's
litellm agents codex catalog shares the same builder.
* fix(proxy): offer a Codex service tier only when every deployment of the model lists it
* fix(codex-catalog): an invalid service_tiers value offers no tier for the model
* fix(codex-catalog): read service tiers off the deployments the key's team can route to
A tier is offered to Codex only when every deployment of the model name a
request from the key's team can route to lists it, so another team's
deployment of the name and a deployment an admin paused via model_info.blocked
no longer withhold or add tiers for requests that never reach them
The catalog's always-null fields are annotated NoneType so the module imports
under pydantic 2.12.0 on Python 3.14, the lowest pin the MCP resolve job
installs, which rejects a None annotation with a None default
* test(codex-catalog): drop the redundant module docstring and sort the imports
* test(integration): add the Codex catalog audit cells and the multi-worker convergence note
* test(integration): clean up every catalog test model and answer the refresh GET
* fix(proxy): keep tiered models under Codex's catalog cut and resolve alias tiers
Under Codex's 1 MiB catalog limit the entries offering a service tier are kept
ahead of those offering none, each group in model_list order, with every kept
entry at its listing position, so the model an operator configured tiers for
survives a wide key's long listing. A model_group_alias row reads its target's
deployments, so it carries the target's tiers and stock metadata under the
alias name.
* fix(proxy): pick Codex catalog metadata per team and skip entries too large for the cut
The upstream model that selects Codex's stock entry was read off the first deployment of a name
without checking the key's team, so a team whose requests route to a different deployment could be
handed another team's prompt, reasoning levels, and tiers. The upstream model and the tiers now come
from the same team-aware selection routing uses, and a caller with no team reads the deployments no
team owns
The byte cut kept a prefix of the tier-first order, so one entry larger than the whole limit emptied
the catalog. An entry too large for the bytes left is now passed over and the smaller ones after it
are still kept
---------
Co-authored-by: mateo-berri <277851410+mateo-berri@users.noreply.github.com>
* feat(lens): record run steps, trigger and exact run windows on investigations
* feat(lens): scan only traces since the last run and keep a capped step log
* feat(lens): log each analysis model call with its model, tokens and cost
* feat(lens): accept agent and time window on run now and add turn-all-on
* test(lens): cover new-traces-only windows, manual runs and the step cap
* test(lens): cover run now overrides and turning paused investigations on
* chore(ui): regenerate api types for lens run steps and run now options
* feat(lens): group open findings into one row per problem with agent filters
* test(lens): cover the findings table grouping, filters and schedule labels
* feat(lens): add a findings table across all investigations
* feat(lens): show a live step feed with the model behind each call
* feat(lens): offer turning all paused investigations on
* feat(lens): let run now pick an agent and time window
* test(lens): cover run now request building
* feat(lens): show the step feed and run now dialog on an investigation
* feat(lens): open run now choices instead of running immediately
* feat(lens): open findings first and peek a finding without leaving the table
* feat(lens): fold investigation actions into the findings toolbar
* feat(lens): show each investigation's schedule and open findings
* feat(lens): name the agent on a finding
* feat(lens): keep new investigations watching every 15 minutes by default
* feat(lens): show the watch schedule outside advanced options
* feat(lens): send run now options and turn-all-on from the dashboard
* test(lens): give demo runs steps and a trigger
* feat(lens): let the findings table fill the screen
* test(lens): add steps and trigger to progress fixtures
* test(lens): add steps and trigger to status fixtures
* test(lens): cover the default watch schedule in setup
* test(lens): cover run now choices from an investigation
* test(lens): open saved investigations from the manage view
* feat(lens): use one tab bar for traces, findings and investigations
* feat(lens): place page actions on the lens tab row
* feat(lens): drop the nested tabs and edit investigations in place
* feat(lens): show investigations as a table with run now and edit
* feat(lens): name each findings row for screen readers
* test(lens): open saved investigation links on findings
* test(lens): reach findings and investigations from the top tabs
* fix(lens): mark run now jobs manual and keep them from moving the scheduled scan
* fix(lens): keep run now since-last-run windows even with an agent override
* test(lens): cover that manual runs never skip scheduled traces
* test(lens): cover run now windows with agent and lookback overrides
* fix(lens): group findings without Map.groupBy and expose sampled runs
* fix(lens): open older findings and review every merged copy from one row
* fix(lens): hide edit and run now from read-only viewers
* chore(lens): drop restating comments from the findings table
* chore(lens): drop restating comments from the step feed
* chore(lens): drop restating comments from the paused banner
* chore(lens): drop restating comments from header actions
* chore(lens): drop restating comments from run now
* test(lens): cover merged findings and read-only investigation rows
* fix(lens): record a model step even when the response has no usage
* test(lens): cover model steps with and without reported usage
* fix(lens): keep a merged finding open when one of its updates fails
* refactor(lens): accept update results from the investigations view
* refactor(lens): accept update results in investigation actions
* test(lens): cover retrying a merged finding after a failed update
* feat(decisions): add unified /v1/decisions endpoint for Jev-compatible providers
Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>
* fix(decisions): register typesafe as a provider so Jev deployments load
Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>
* refactor(decisions): move provider endpoints under llms and validate proxy bodies
Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>
* feat(decisions): add Cloudflare Clef and Strands Decider backends
Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>
* fix(decisions): register decisions routes for managed agents and gateway
Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>
* test(decisions): use raw regex for cloudflare missing account match
Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>
* fix(decisions): avoid cast in Cloudflare response unwrapping
Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>
* fix(decisions): default model, evaluation health probe, short Cloudflare names
The proxy validates only state and questions, so a request without a
model falls through to the configured default model like every other
route. Health checks probe evaluation-mode deployments through the
Decisions API instead of failing with an unsupported mode, and
cloudflare/clef and cloudflare/clef-flash get cost-map rows so the short
names resolve a mode and a price. The registry no longer claims typed
decisions for a provider with no backend.
* fix(decisions): let health_check_params override the evaluation probe
Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>
* test(integration): audit the decisions endpoint across providers, limits, health and chaos
Adds the /v1/decisions audit cells: one wire contract per provider (path, key, body and cost-map billing), the gateway-only fields and tags, the sad paths (invalid bodies, unknown model, key checks, api_base in the body, upstream 401/429/500, a 200 without answers, an unreachable upstream), the two evaluation-mode health probes, and three chaos cells (a mixed-failure burst over both routes, a worker SIGKILL mid-burst, an upstream outage and restart on the same port).
The PR's cost case read the upstream observations through the gateway, which answers 404 for that path; it now reads them from the upstream URL. The owned proxy harness takes extra CLI arguments, and its graceful stop waits as long as a worker boot may take, since a worker still starting honors SIGTERM only once it is up and the 30 second wait forced a cleanup under load.
* fix(decisions): send env API keys to a configured api_base
Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>
* feat(decisions): add zero-cost evaluation cost-map entry for Strands Decider
Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>
* fix(decisions): register the routes through the lazy feature registry
The Decisions router was included at import, ahead of the config and DB
pass-through endpoints, so a pass-through configured at /v1/decisions
was skipped and answered 400 as an unknown Decisions provider. The
routes now register through LAZY_FEATURES, which splices them in after
every eager route, so a pass-through at /v1/decisions keeps its route
while /decisions still serves natively. The lazy OpenAPI snapshot carries
the two paths so the schema shows them before the first call.
The audit cells add the env-key egress to a configured api_base, the
client api_base opt-in shared with chat, the pass-through precedence on
an owned proxy, and the Strands evaluation health check resolved from
the cost map. The integration config exports the Perplexity env key the
first cell needs.
* fix(decisions): keep the Cloudflare api_base message in its transformation and read the audit upstream once per cell
* fix(proxy): let a config pass-through beat a lazily registered route in eager mode
With LITELLM_DISABLE_LAZY_ROUTES set the decisions routes are registered at
startup, so SafeRouteAdder treated a config pass-through at exactly
/v1/decisions as already registered and dropped it. In lazy mode a pass-through
created through the API after the first native call was skipped the same way.
Routes a lazy feature owns no longer count as registered, and a route added at
one of their paths is placed ahead of them, the precedence lazy mode gives a
config pass-through when the feature has not loaded yet.
---------
Co-authored-by: mateo <mateo@berri.ai>
Co-authored-by: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>
Co-authored-by: mateo-berri <277851410+mateo-berri@users.noreply.github.com>
is_prompt_caching_valid_prompt ran the full Python token_counter over every message to compare against the deployment's prompt cache minimum, 500 to 1000 ms at 440k to 740k tokens on every request that reaches the prompt_caching pre-call check, Rust on or off. messages_reach_token_count does the same arithmetic as token_counter(...) >= threshold and stops at the first message that reaches the threshold. Groups with one healthy deployment skip the prefix hash and pin lookup, which cannot change the result for them
Four fixed name span events make the pre-LLM phases measurable with OTel v2: litellm.request.body_received (with body_bytes) once per body read before parsing, on the JSON, binary and form branches, body_parsed, pre_call_completed, and deployment_selected emitted once per pick inside Router.async_get_available_deployment and get_available_deployment with attempt, reason and model group, so every router surface, retry and fallback is covered. Measured locally on /v1/chat/completions, /v1/messages and /v1/responses at 440k tokens with Rust on and off against a fake upstream
Co-authored-by: yassin <yassin@berri.ai>
Co-authored-by: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>
Response cache reads and writes open cache.get llm_response and cache.set llm_response phase spans with their Redis spans nested underneath, on the Python path and on the native Rust path, and deployment selection runs inside a route {model_group} phase so the cooldown, usage and model-id reads the router issues nest under it before chat {model}. The autorouter classifier call nests under that route phase as well and carries its typed internal origin on litellm.request.purpose, so it is told apart from the provider attempt. Service spans are named {service}.{verb} {target} from a low-cardinality key family the producer declares (llm_response, auth_objects, spend_counters, router_cooldowns, claude_code_session_router_binding, rate_limits, pod_lock, budget_reset, ...) instead of the raw method or a per-request pipeline length; a pipeline flush is targeted by the one family its ops share or by mixed with the sorted families on litellm.redis.families, a batch op keeps the family it was declared under whichever pipeline or standalone read settles it, and the ambient family labels Redis spans only, never the DB write-back a task spawned inside that context performs later. The raw method stays on litellm.service.call_type and on the Prometheus and Datadog labels. Caller attribution is carried across asyncio task boundaries on a ContextVar so forwarder-only chains no longer surface, the raw cache key is dropped from Redis span metadata, pipeline op counts land as an integer attribute, every call_type the Redis cache layer emits maps to a verb, and a scan over litellm/ and enterprise/ fails when a Redis producer, batch reservation included, declares no key family.
A V2 logger built for a key or team logging entry while the operator's V2 logger is already registered keeps only the exporters its own preset contributed, whether or not the operator holds credentials for that backend, so every chat span no longer reaches the operator's collector twice. A span the success callback has to open itself, with no pre-call carrier, starts at the provider handoff (api_call_start_time) instead of the logging object's creation.
Co-authored-by: yassin <yassin@berri.ai>
Co-authored-by: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>
log_db_metrics wrapped whole cache-first auth helpers and always emitted a ServiceTypes.DB success event, so in-memory cache hits showed up as postgres <fn> spans and DB service metrics. The decorator now installs a ContextVar witness that _TrackedPrismaEngine marks on every Prisma query and transaction call, and the DB event is emitted only when the witness was marked. Real reads keep their existing call_type names, the failure path and the PROXY batch-write branch are unchanged, and Redis instrumentation is untouched.
A decorated helper that reaches Prisma only through another decorated helper (get_key_object -> get_object_permission, get_team_object_by_alias -> get_object_permission, get_tag_object -> get_tag_objects_batch) used to emit two events for one query. The inner wrapper now marks its witness as reported when it emits a success or DB failure event, and only unreported activity is handed up to the enclosing witness, so the inner event is the one that survives. An outer helper that also queries Prisma directly or through undecorated callees still gets its own event.
Tests: get_user_object and get_org_object cache hits emit no DB event; a get_user_object miss through the generated Prisma client emits exactly one postgres get_user_object event; decorator-level tests cover nested calls emitting only the inner event, outer calls with their own query, inner non-DB failures, bounded lookups and sibling-request isolation.
Co-authored-by: yassin <yassin@berri.ai>
Co-authored-by: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>
* feat(dev): seed linked tracing and spend fixtures
* chore(dev): use OpenAI model in tracing config
* chore(dev): align tracing credentials with UI E2E
* fix(dev): update fixture seeder query scope
* feat(dev): seed linked tracing and spend fixtures
* chore(dev): use OpenAI model in tracing config
* chore(dev): align tracing credentials with UI E2E
* fix(dev): update fixture seeder query scope
* wip
* wip
* wip
* chore(trace): checkpoint ongoing Rust migration
* refactor(trace): group Python bridge under trace package
* refactor(traces): read span conventions through a Convention trait
Each span format (Claude Code, LangSmith, OpenInference, gen_ai) now lives under
normalize/convention/ as a unit struct implementing Convention, owning both its
detection and its extraction. Precedence is one ordered registry instead of an
if-chain in mod.rs that reached into each module differently.
The modules now share one way to read attributes: present() for the first
non-empty key and Payload for a text that also reports the key it consumed,
replacing three different idioms and the &mut Vec threaded through payload
readers. Instrumentation::adjust returns a new Extraction instead of mutating
one, with each SDK rule as its own function, and the LangChain middleware
suffix list exists once.
* feat(trace): export Rust-owned wire schemas and enforce contract bounds
* fix(trace): bound quoted counts in ClickHouse wire schemas
* feat(trace): generate Python wire contracts with datamodel-code-generator
* test(trace): validate migrated callers and generated contracts at the native boundary
* refactor(traces): rename normalization convention to format
* fix(traces): reconcile spend evidence and preserve unknown costs
* feat(traces): normalize additional telemetry formats
* test(traces): cover captured normalization fixtures
* refactor(traces): isolate SDK normalization rules
* feat(tracing): seed all trace exports for local dashboard
* fix(clickhouse): preserve custom LiteLLM request metadata
* docs(traces): define normalization module boundaries
* docs(traces): define resolution and OTLP boundaries
* fix(ui): normalize nullable trace message names
* refactor(traces): split resolver modules and cover resolution behavior
* test(traces): replace normalization snapshots with behavior assertions
* fix(ui): align dashboard API contracts with generated types
* refactor(traces): type normalization and storage boundaries
* fix(traces): seed captured SDK spend and preserve provider identities
* wip
* test(traces): verify guide discovery and content ordering
---------
Co-authored-by: Yujong Lee <yujong@berri.ai>
Co-authored-by: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>
* feat(lens): simplify trace inspection in the gateway drawer
* feat(lens): add a conversation view for full traces
* test(lens): keep normalized message fixtures type safe
* refactor(lens): make conversation view read like a chat
* refactor(lens): use a quiet trace view menu
* refactor(lens): make trace view tabs explicit
* refactor(lens): restore compact trace view switch
* fix(lens): preserve complete conversation history and tool types
* feat(lens): add trace full-screen and close controls
* fix(lens): show forwarded answers and agent errors once
* fix(lens): reset full screen when closing a trace
* fix(roi): estimate linked authors and clarify model selection
* fix: preserve trace errors and ROI results across partial failures
* fix(roi): correct pagination variable typing
* fix(ui): place loaded root failures in conversation order
* fix(roi): read estimator recommendations from model catalog
* revert: remove catalog-driven ROI recommendations
* fix(proxy): stamp the client alias on a copy of each streamed chunk so pricing sees the deployment model
Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>
* test(spend): streamed alias matching a capability rule bills the deployment price
Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>
* test(spend): assert every streamed chunk carries the client alias
Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>
* fix(logging): log the client alias on the priced streamed response, the same as non-streamed
Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>
---------
Co-authored-by: kerry <kerry@berri.ai>
Co-authored-by: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>
* feat(roi): support GitLab and tagged branch costs
* fix(roi): count tagged branches independently of estimation status
* test(roi): capture live GitHub and GitLab report validation
* fix(roi): open estimate details at the start
* fix(roi): clarify cost views and unify report layout
* feat(roi): showcase per-PR costs in the sample report
* fix(roi): separate report tabs and preserve branch cost attribution
* fix(roi): preserve demo previews and align progress spacing
* fix(roi): isolate demo loading and parallelize fork lookups
Preserve active sync status when source changes finish saving, keep live reports available when demo requests fail, and cover each review regression
* fix(roi): separate demo and live loading states
Clear the demo URL on fallback, wait for live requests on exit, and retain request errors until the corresponding operation recovers
* fix(roi): ignore refreshes from a previous source
* fix: trust gateway context for ROI estimator exclusion
* fix: preserve historical ROI estimator exclusion
* fix(search_tools): encrypt search tool litellm_params at rest
Encrypt every string value of a search tool's litellm_params on create and
update and decrypt on every DB read, so legacy plaintext rows load unchanged.
Include the table in master key rotation, LITELLM_MIGRATE_FROM_MASTER_KEY and
the migrate-encryption scan.
* fix(search_tools): keep edits made while the master key rotates
Write each rotated search tool row only if it still holds the litellm_params
that were read, and re-read and rotate it again if it was edited in between,
so a PUT that lands during /key/regenerate is not overwritten.
* fix(search_tools): retry rotation writes until the row stops changing
Rotate a search tool row again for as long as it keeps being edited instead of
giving up after five attempts, and stop with a warning only when the conditional
write fails on an unchanged row. Build the decrypted read result without
mutating it in place.
* refactor(search_tools): rotate edited rows in a loop, drop the step comment
Retry the conditional rotation write in a loop instead of recursion so sustained
edits cannot deepen the call stack, drop the step comment on the rotation call,
and stop mutating local state in the rotation tests.
* test(search_tools): drop the rotation test docstring
* Store search tool params as written when no encryption key is configured
* Rotate search tools under the salt key, keep non-ciphertext values and loaded tools that do not decrypt
* Treat a search tool as undecryptable only when its provider is ciphertext-length
* Drop suppressions the type discipline gate on main now reports as unused
* Show the loaded search tool in the admin list and info views when its DB params do not decrypt
* Keep the DB row's other fields when the admin views substitute loaded params
* fix(guardrails): encrypt guardrail litellm_params secrets at rest
* fix(guardrails): keep salt-key encryption on master key rotation and retry rows edited mid-rotation
- rotate guardrail params under LITELLM_SALT_KEY when set, matching the key reads decrypt with
- re-read and retry a row whose updated_at moved during rotation, up to GUARDRAIL_ROTATION_ATTEMPTS
- build decrypted Guardrail rows and the rotation count without mutating locals
* refactor(guardrails): retry guardrail rotation by bounded recursion instead of a rebound cursor
- each attempt re-reads the row and recurses with attempts_left - 1, so no loop variable is rebound
- cover the give-up path after GUARDRAIL_ROTATION_ATTEMPTS writes
* test(guardrails): drive the real guardrail rotator from the master key rotation test
- inject an encrypted guardrail row through the prisma client instead of replacing the GuardrailRegistry method
- assert the written params decrypt under the new master key
* Annotate guardrail param encryption collections for type-discipline gate
* Type guardrail param recursion through validated JSON containers
* Type guardrail registry test helpers and drop section comment
* Reject client-supplied encrypted values in guardrail litellm_params
* Allow depth-bounded contains_encrypted_marker in the recursion detector
* Keep a loaded guardrail when its DB params do not decrypt with the current key
* Apply other DB edits while keeping loaded values that do not decrypt, including PATCH models
* Keep the loaded guardrail when an undecryptable param has no loaded value
* Drop suppressions the type discipline gate on main now reports as unused
* Assert what the reinitialized guardrail holds after an edit to an undecryptable one
* Drive the rotation sync tests through a registered guardrail instead of patching reinitialize
* Type the rotation test helpers and drop the new test docstrings
* fix(guardrails): refuse to approve a submission whose params do not decrypt
Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>
---------
Co-authored-by: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>
* fix(proxy): return 400 instead of 500 for missing required params and invalid pagination
Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>
* ci: run search_endpoints tests in proxy-endpoints shard
Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>
* fix(proxy): return 4xx for missing required params across all LLM routes and propagate provider status on lookups
Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>
* fix(llm_http_handler): keep provider error text when re-raising mapped errors
Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>
* fix(proxy): allow promptless image edits and default search models
Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>
* fix(proxy): default missing image edit image to None
Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>
* refactor(proxy): build image edit defaults without mutating request data
Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>
* fix(proxy): keep image edit defaults within type-discipline budget and give request mocks a scope
Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>
* test(proxy): inject a fake router for the search default model test
Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>
* test(llms): cover provider error status on vector store and file lookup handlers
Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>
* test(llms): keep the lookup handler raise block to a single statement
Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>
* test(llms): cover provider error status on eval, eval run, skill and vector store file content lookups
Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>
* test(integration): cover missing required body params and provider lookup status codes
Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>
* ci: run tests/unit/proxy/search_endpoints in the proxy-endpoints shard
Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>
* test(integration): bind spend-row request id with partial to satisfy B023
Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>
* fix(proxy): only reject non-positive page_size on vector store list
Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>
* fix(proxy): remove unreachable fine-tuning body validation
Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>
* test(integration): cover streaming anthropic messages reaching the upstream
Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>
* test(integration): count only provider calls when asserting missing params never reach the upstream
Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>
* fix(proxy): preserve merge-base request compatibility
Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>
* fix(proxy): preserve interaction completion model defaults
Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>
* fix(proxy): retry model read-through before rejecting params a DB-only deployment may default
Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>
---------
Co-authored-by: shivam <shivam@berri.ai>
Co-authored-by: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>
Co-authored-by: yucheng <yucheng@berri.ai>
* fix(vector_stores): return managed file ids from vector store file list
Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>
* test(vector_stores): cover managed file list route
Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>
* fix(vector_stores): only map round-trippable managed ids and index flat file ids
Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>
* fix(proxy-extras): build managed file gin index concurrently
Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>
* chore(proxy-extras): move the managed file gin index migration after main's newest
* fix(vector_stores): satisfy lint gate
Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>
* chore(proxy): drop the stale no-index note on the raw file-id guard
* test(vector_stores): cover managed file ids on the vector store file list end to end
Integration cells for GET /v1/vector_stores/{vs}/files mapping provider file ids back to
the caller's owner-scoped managed ids and decoding managed after and before cursors: raw
httpx, the OpenAI SDK sync and async pagers, the three credential routing modes, the owner
filter branches, raw and unmappable cursors, provider errors, duplicate and non-string ids,
a provider outage mid-burst, a worker SIGKILL mid-burst, and the GIN index migration applied
by the migration entrypoint and by db push
---------
Co-authored-by: shivam <shivam@berri.ai>
Co-authored-by: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>
Co-authored-by: mateo-berri <277851410+mateo-berri@users.noreply.github.com>
* feat(interactions): durable cross-pod settlement for background interaction billing
Background interaction billing lived only in the creating replica's memory, so a
DELETE routed to another replica, or a restart of the creating one, never billed
the completed provider work and the budget reservation was refunded at the poll
timeout. The create now registers the billing context in a settlement store
before returning, the proxy installs a Prisma-backed store at boot
(LiteLLM_BackgroundInteractionSettlement, schema-only migration), any replica
claims the row once through a conditional update before billing or releasing,
startup resumes every unclaimed row with its remaining timeout, and a give-up
records an unsettled outcome instead of silently reconciling to zero. The SDK
keeps an in-memory store and behaves as before.
* fix(interactions): survive a settlement install failure at boot and stop carrying request headers
* fix(interactions): drop the stored request context once a settlement row is settled
* fix(interactions): bill the completed response a poll already saw when its claim only answers at the deadline
* fix(interactions): carry a missing model through the settlement context for agent-only background creates
An interaction created with an agent and no model reaches the poll with no model name, exactly as on main. The settlement context now stores that None instead of rejecting the create, which answered the client with a 500 after the provider had already accepted it.
* fix(interactions): leave an unfetchable background interaction to its creating poll when a delete lands elsewhere
The remote pre-delete path fetches with only the delete's credentials, so a fetch it cannot make says nothing about the interaction. It used to claim the settlement row and release the reservation anyway, which stopped the creating replica's poll and lost the bill when the delete then failed the same way. It now returns without claiming; the in-process path keeps releasing on an unfetchable state, since its context carries the create's own credentials.
* fix(interactions): fail a cross-replica delete when its pre-delete fetch fails so the creating poll keeps the bill
* fix(interactions): keep the stored settlement gate when registration raises after landing, and fail resumed-poll deletes closed
A registration that raised after its row committed moved the poll to a private in-memory gate, so the creating worker billed while the stored row stayed unclaimed for another replica's delete or the next boot to bill again. The row is now read back once and, when it landed, the poll claims through it like every other settler.
A worker that resumed the poll after a restart is not the creator, so its delete on a failed pre-delete fetch now fails with the fetch's error instead of releasing and deleting. After a fleet restart every worker holds resumed polls, which left the fail-closed path applying nowhere.
* fix(interactions): settle an unverified registration through the durable claim
A create whose settlement-store write raised no longer bills through a
private in-memory gate that a later boot's resume cannot see. The claim
asks the durable store first and falls back to the local gate only when
the store answers that no row exists, and a missing settlement table reads
as no rows so a replica without the migration still settles in process.
* test(proxy): keep the settlement test where the proxy-infra shard collects it
The merge of main moved test_background_interaction_settlement.py under
tests/unit/proxy/spend_tracking, but the proxy-db shards claim tests/unit/proxy
files one by one in .circleci/scripts/unit_selection.sh, so no CI shard ran it
and codecov/patch dropped. tests/test_litellm/proxy/spend_tracking is collected
whole by the proxy-infra shard, which is where the test ran before the merge.
* fix(interactions): raise on a non-2xx Gemini interaction fetch
AsyncHTTPHandler.get never raises for status and the Gemini GET transform
only raised when the body was not JSON, so a 500 or 404 carrying Gemini's
JSON error body parsed as an interaction with no status. A delete on a
replica other than the creator then claimed the settlement as released and
forwarded the delete instead of failing closed, and the bill was lost. The
transform now raises GeminiError with the vendor's status, as the delete
transform already does; the in-process poll already retries a fetch that
raises
* test(integration): audit durable background interaction settlement across replicas
Twenty-six deterministic cells drive a one-worker creator and a two-worker
settler against an owned scripted Gemini upstream: cross-replica deletes
bill once, failed and cancelled interactions release, a later replica
resumes unclaimed rows, custom deployment pricing bills at the deployment
rate, a fetch the settler cannot make fails the delete closed, odd ids are
refused, a missing settlement table keeps in-process billing, polling
disabled registers nothing, the budget reservation is released by the
settler, an upstream outage mid-burst fails closed and recovers, killed
workers hand their polls to the respawned ones, and concurrent deletes on a
slow upstream settle exactly once. The support upstream gains a scripted
interaction store with per-id GET status and delay, and the process helper
gains an owned upstream a test can stop and restart
* test(integration): refuse a repeated delete in the scripted upstream and pin the settlement budget below one estimate
* chore(ui): regenerate dashboard API types after merging main
* test(integration): accept the 422 budget refusal and a respawned worker's resume
The budget cell pinned a 400 that the proxy stopped answering when budget refusals moved to 422, so it now asserts the status and the budget_exceeded error type the sibling budget tests pin. The later-booting replica cell accepts a claimer that is any worker started after the creates, since uvicorn's supervisor can respawn the creator's worker under load and the respawned worker's boot resume claims the rows by design; the single spend row check is unchanged
* test(integration): delete the pinned key's interaction with a second key
A key whose budget is filled by its own reservation is refused on every route, the DELETE included, so the cell now asserts that 422 and sends the delete with a second key, which is what the reservation release on another replica needs in order to be observable at all
---------
Co-authored-by: mateo-berri <277851410+mateo-berri@users.noreply.github.com>
* fix(proxy): use OpenAI workload identity federation tokens on /openai_passthrough
The OpenAI passthrough routes (HTTP and websocket) only looked up a static
OpenAI API key, so proxies authenticating to OpenAI through workload identity
federation failed with "Required 'OPENAI_API_KEY'". When no non-empty static
key is configured, exchange the workload identity subject token for a bearer
token, scoped to the passthrough's own OPENAI_API_BASE so the token is never
sent to a non-OpenAI host
* fix(proxy): close the OpenAI websocket passthrough cleanly when the workload identity exchange fails
A rejected, unreachable or unreadable workload identity exchange raised out of
the websocket route before the handshake was accepted, so clients saw a bare
handshake failure. Log the cause and close with 1011 and a fixed reason instead
* refactor(openai): resolve workload identity bearer tokens for an api base in the OpenAI provider module
The passthrough route only owns the static key lookup now and asks the OpenAI
provider module for a workload identity bearer token scoped to its api base
---------
Co-authored-by: Krrish Dholakia <krrish+github@berri.ai>
Per-router auto-router savings were missing on proxies with
disable_spend_logs and for requests without a session id under
missing_session_id: omit. The router-day row is written by the
auto-router turn, whose enqueue sat behind the spend-logs flag and whose
builder dropped sessionless requests, while LiteLLM_DailyUserSpend has
neither gate. That money then showed only as unattributed savings.
Enqueue the turn whether or not spend logs are kept, and keep a
sessionless turn with an empty session id that writes the router-day row
while the session upserts skip it. The day and session rows still commit
in one statement, so late baseline corrections keep their ordering.
Without spend logs, write only the router-day aggregate: the turn drops
its session id, so no per-session row is stored, and baseline capture is
skipped, since a baseline observation can only publish once its
request's spend log exists. The flag keeps its meaning of no per-request
or per-session data, while the daily per-router money matches
LiteLLM_DailyUserSpend.
Co-authored-by: Claude Opus 5.5 <noreply@anthropic.com>
* refactor(lens): move analysis prompts into markdown files
* feat(lens): ask the investigator for a scoped agent fix brief with two options
* test(lens): cover the agent fix brief through investigation and merges
* chore(ui): regenerate api types for the lens fix brief
* feat(ui): build copyable lens fix prompts
* feat(ui): show the lens fix brief with copy buttons for claude code and codex
* test(ui): cover copying a lens fix option
* refactor(lens): replace the fix options with a plain issue brief
* feat(lens): ask for problem, user goal, outcome and test cases without prescribing code changes
* test(lens): cover the issue brief through investigation and merges
* chore(ui): regenerate api types for the lens issue brief
* refactor(ui): drop the lens fix prompt builders
* feat(ui): add a lens issue brief panel
* feat(ui): show the lens issue brief in the finding drawer
* test(ui): cover the lens issue brief and the legacy fallback
* feat(ui): render a lens issue brief as a markdown document
* test(ui): pin the lens issue brief markdown layout
* feat(ui): show the issue brief as a copyable file with claude code and codex buttons
* feat(ui): pass the finding title into the issue brief
* test(ui): cover copying the issue brief for claude code and codex
* feat(ui): bold the input and expected labels in lens test cases
* test(ui): pin the bold test case labels in the issue brief
* feat(ui): render the issue brief as formatted markdown
* test(ui): cover the rendered issue brief sections and raw markdown copy
* feat: add Bespoke Nimble gateway and OSS classifier support
* feat: accept Ollama's nimble model name for the Bespoke provider
* test: exempt the POST-only bespoke decisions route from the all-methods check
test_pass_through_routes_support_all_methods requires every built-in
pass-through route to accept every HTTP method unless it is listed in
PROTOCOL_CONSTRAINED_PASS_THROUGH_ROUTES. /bespoke/v1/systemone is
POST-only like /laya/v1/systemone, so the test failed at this branch
and passed at the merge base. List it alongside Laya.
Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com>
---------
Co-authored-by: mateo-berri <277851410+mateo-berri@users.noreply.github.com>
Co-authored-by: Claude Fable 5.1 <noreply@anthropic.com>
* refactor(proxy): extract shared spend log read policy
* test(proxy): use named bindings for spend scope regression
* test(proxy): reuse existing spend log query harness
* test(proxy): cover spend log permission lookup adoption
* chore(proxy): relocate existing spend query baseline
* refactor(proxy): make scope query returns explicit
* refactor(proxy): inject deferred log permission lookup
* test(proxy): cover teamless management compatibility lookup
* refactor(proxy): compose user and team log grants
* refactor(proxy): share generic authorization composition
* refactor(proxy): compose trace read permissions
* refactor(proxy): centralize spend and trace authorization
* refactor(proxy): strengthen spend and trace scope types
* refactor(proxy): flatten log read scope into owned logs
Replace the AnyOf grant tree with a flat OwnedLogs(user_id, team_ids) scope,
and OwnedTraces(logs, api_key_hash) for traces, since every consumer flattened
the tree back into that shape.
A caller with no user id now gets an empty scope instead of matching ownerless
rows through Prisma's IS NULL. The dead request_id guard in ui_view_spend_logs
is removed, and the management facets inject the log team lookup and reuse
read_scope_sql instead of the list shim.
Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
* test(proxy): run spend scope tests through one SQLite emulator
Replace the string-matching payload emulator and the hand-rolled Prisma where
interpreter with one SQLite helper that runs the real scope SQL. Session scope
tests now go through the endpoint, including the no-user caller that must not
match ownerless rows. Drop duplicated lookup-failure and trace mapping cases.
load_permitted_log_team_ids returns no teams without a database instead of
relying on the resolver's broad except.
Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
* fix(proxy): unify log and trace ownership permissions
* test(tracing): align fixtures with ownership read scopes
* refactor(tracing): align query scopes with row ownership
* refactor(spend): make ownership SQL predicates explicit
* test(spend): validate ownership SQL against PostgreSQL
* docs(traces): drop key-row visibility from query help guide
Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>
* test(spend): reach the empty-memberships branch in team lookup test
Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>
* chore(ui): regenerate dashboard API types
Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>
---------
Co-authored-by: Claude Opus 5.5 <noreply@anthropic.com>
Co-authored-by: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>
* fix(proxy): stop queued registry read-throughs spending the resync budget
RegistryReadThrough.attempt serializes misses behind one lock, but every
request that was queued behind the first one still spent a unit of the
20-per-5s resync budget and re-ran the DB resync, even though the first
request had already loaded the object. A burst of more than 20 requests
for a model created on another worker therefore exhausted the budget and
the rest got 400 "Invalid model name".
attempt now checks whether the key is already loaded once it holds the
lock and returns early without touching the budget. Models check the
router's model names and deployment ids, guardrails and agents reuse
their existing registry lookups.
* test(proxy): gate the queued read-through test on events and cover each registry's loaded check
The queued-requests test now holds the first resync on an asyncio.Event instead of
a timed sleep and records calls in a recorder with tuple and frozenset state. New
tests show the model, guardrail and agent read-throughs each answer an object that
is already loaded without reading the database, so rewiring any registry's loaded
check now fails a test
* test(proxy): keep the queued read-through recorder inside its test and type the agent registry fixture
* refactor(traces): type the ClickHouse query help response
Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>
* test(traces): include agent names and frameworks in named contract round trips
Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>
* test(traces): cover native query help validation in the storage adapter
Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>
---------
Co-authored-by: Yujong Lee <yujong@berri.ai>
Co-authored-by: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>
* test(integration): answer the model-info refresh GET in the mixed MCP responses wire peer
The proxy's periodic model-info refresh sends GET /v1/models to the deployment api_base, which tripped the peer's /responses-only assertion when a tick landed mid-test
* test(e2e): wait for the api-keys URL after clicking Virtual Keys in onboarding
/ui already renders the Virtual Keys heading, so the helper returned before navigation finished. The late route change moved focus and closed the account menu popover in hideLiteAdmin
* test(proxy): stop two unit modules leaking app.openapi_schema and a session-wide Router
test_custom_openapi cached a stripped schema on app.openapi_schema and never cleared it, breaking later openapi route tests. test_proxy_reject_logging built a module-level Router that stayed in the live router registry all session and re-added cost-map keys during a reload. Reset the schema via monkeypatch and make the Router a function-scoped fixture
* perf(proxy): aggregate daily model usage per flush instead of upserting per request
* perf(proxy): drain queued model usage in the spend log flush job
* perf(proxy): queue model usage at request time instead of writing to the db
* test(proxy): cover batched daily model usage aggregation and retries
* test(proxy): read back model insights written by the batched flush
* fix(proxy): drain the whole model usage queue each flush so it cannot grow unbounded
* test(proxy): give the mock prisma client a model usage queue
* test(proxy): cover draining a model usage queue larger than one spend log batch
* refactor(mcp): extract upstream preparation and support modern clients
* fix(mcp): reject incompatible upstream transport before saving
* fix(mcp): serialize protocol validation with server updates
---------
Co-authored-by: Joshua Valluru <326636767+joshua-berri@users.noreply.github.com>
* wip
* wip
* test(traces): separate root status from diagnostic error counts
* test(traces): cover normalization precedence and fallbacks
* chore(cache): remove stray comments from trace PR
* test(traces): name lens test for shared query path
* fix(traces): place query implementation before test module
* test(traces): use unified read scope in migration tests
* ci(rust): allow feature checks to finish
* ci(mcp): allow dependency resolution to finish
* fix(traces): preserve key visibility and safe spend attribution
* test(proxy): stop the proxy_server app fixture leaking LITELLM_LOG
The session app fixture set LITELLM_LOG=ERROR with os.environ.setdefault and never removed it, so later tests on the same xdist worker inherited it. test_drop_params_env_var spawns a subprocess with os.environ and lost the warning it asserts on. Scope the variable to the import with a MonkeyPatch context
* test(secret-detection): give the hand-built redaction request an ASGI path
Since #43975 _read_request_body checks the route path via request.scope, and a scope without path raised KeyError that was swallowed into an empty body, so chat_completion failed with a missing messages parameter. Real ASGI scopes always carry path
* test(integration): isolate litellm callback lists per sdk test
usage-based-routing-v2 Routers register their selector in litellm.callbacks and nothing removes it, not even Router.reset(). The counter TTL and Redis service metrics tests left their selectors behind, and the next usage routing test ran their pre-call checks against its own rpm=1 deployments, raising "Deployment over defined rpm limit". An autouse fixture now gives each sdk test copies of the callback lists and restores the originals afterwards
* test(integration): keep the owner-lookup fault proxy off the shared read replica
The owned proxy points DATABASE_URL at a scratch database but inherited
DATABASE_URL_READ_REPLICA from the replica job, so auth read the shared
database and rejected the freshly created key with token_not_found_in_db.
Drop the replica variable like the other scratch-database owned proxies
* test(integration): request every seeded key in the team owner breakdown
The aggregated team activity endpoint now caps breakdown.api_keys at the top
100 keys by default (#43398), so the 300 seeded keys came back as 100 rows.
The test guarantees each key is reported with its own owner, so ask for an
api_key_limit that covers all seeded keys
* test(integration): give every owned Redis its own port in the redis-cache container
On CircleCI every owned Redis ran on the fixed port 16379 inside the shared
redis-cache container. When an earlier server still held that port, the new
one failed to bind, readiness pinged the old server, the pidfile read failed
and cleanup then reported "Owned Redis still serves after shutdown"
Reserve an ephemeral port for the docker-exec path the same way the local
binary path already does, and refuse to start when something already serves
the chosen port so the failure names the real cause
* test(e2e): skip the Vertex Mistral partner case the e2e project cannot reach
The e2e Vertex project gets a 404 publisher model not found for vertex_ai/mistral-small-2503, so the case can only fail
* test(e2e): skip the Vertex gpt-oss partner case the e2e project never serves
vertex_ai/openai/gpt-oss-120b-maas has hit a 60s read timeout with no response headers on every run in the e2e Vertex project since the case was ported, and no other Vertex partner chat model passes there to switch to
* test(e2e): check only stored message content for a leaked card number
The Presidio spend-log check ran the card-number pattern over the whole serialized response, so a Luhn-valid usage.cost float (0.0003466000000000001) failed the streaming /v1/messages case although the stored content was <CREDIT_CARD>. The check now reads the content and text strings of the stored response, which is where a raw card would land, and still requires the placeholder there
* test(e2e): assert the proxy decodes token-array embeddings for titan
The port in #44120 carried over a legacy SDK-direct test that expected Bedrock to reject token ids with a 400. Through the proxy, /embeddings decodes token arrays to text for providers that cannot embed tokens, so titan answers 200. The test now sends a token array and its decoded sentence and requires the two vectors to match, which fails if the proxy stops decoding or decodes with the wrong tokenizer
* test(e2e): run the Bedrock extended-thinking round trip on a model that honors enabled thinking
us.anthropic.claude-sonnet-5-5 is adaptive-only, so litellm sends thinking.type=enabled with a 1024 budget as adaptive with low effort, and Bedrock returned no reasoning blocks on 5 of 5 identical Converse calls (boto3 direct agreed). us.anthropic.claude-sonnet-4-6 accepts the legacy shape verbatim and returned reasoning on 5 of 5. The non-thinking Bedrock case stays on sonnet-5-5
* test(proxy): stop unit modules forcing DEBUG logging into the event-loop lag tests
Five tests/unit modules set verbose_proxy_logger to DEBUG at import, so every xdist worker that collected them logged the 2.4MB pass-through response from a worker thread, and secret redaction of that line held the GIL for ~0.8s+ inside the timed window. The lag tests now pin the LiteLLM loggers to WARNING and freeze gc while timing, and the module-level DEBUG overrides are removed
* test(e2e): cite the tokenizer and date behind the titan token-array fixture
* test(e2e): let migration seed replicas finish their request-log indexes before cloning
Since #43948 a serving proxy builds the two LiteLLM_SpendLogs indexes on a background thread after it reports ready. The seed fixtures stopped the replica at readiness, so every cloned legacy database lacked an index no real deployment would be missing, and the v2 baseline diff refused it. Seeds now wait until both indexes exist and are valid in the database's schema
* test(passthrough): give the pass-through MockRequest an httpx URL and ASGI scope
#43626 made get_request_route read request.scope during pass-through kwarg setup; the MockRequest in tests/unit/passthrough had neither a scope nor a URL object, so both stream-param tests raised before reaching the code they check. Mirrors the repair #43626 made to the tests/pass_through_unit_tests fake
* test(integration): ignore foreign allow_all_keys MCP servers in the access matrix tool list
test_toolset_gateway_url_serves_a_team_granted_toolset_to_a_key_without_its_own_grant (#43908) registers an allow_all_keys server on the shared gateway, and allow_all_keys servers are listed to every key by design, so a matrix case running on another xdist worker at the same time saw its tools. The matrix now drops tools of allow_all_keys servers it did not create, read from LiteLLM_MCPServerTable before and after listing, and still compares everything else exactly
* fix(lens): run investigations with configured wildcard models
* fix(lens): validate worker model access and pricing before analysis
* fix(lens): bound worker validation and preserve unrelated edits
* fix(scim): apply path-less group PATCH ops instead of storing them under an empty metadata key
A path-less add/replace op (RFC 7644 3.5.2, what Okta Push Groups sends on a
rename) carries a partial Group resource. Each of its attributes now applies as
if sent with that path, so displayName updates the team alias and externalId
and members get their usual handling, and the pushed attributes merge into the
scim_data snapshot the PUT path already writes. A path-less remove or a
path-less op without an object value is rejected with a 400. Any group PATCH
drops an empty metadata key an earlier push left behind, and the Admin UI
metadata form skips an empty key so an affected team can save its settings.
* fix(scim): let a later path op win over an earlier path-less value in the group snapshot
* fix(scim): type the stored team metadata before the JSON object check
* test(scim): run the real group transformation in the path-less replace test
* test(scim): assert the renamed group comes back from the path-less replace
* test(scim): audit the path-less group PATCH on the live proxy
---------
Co-authored-by: mateo-berri <277851410+mateo-berri@users.noreply.github.com>
* fix(proxy): carry key, team and project tags into pass-through spend logs
Pass-through endpoints built their request metadata without the key, team and project controls that native routes apply, so spend rows for configured routes and provider pass-throughs like /anthropic dropped the key, team and project tags and the key and team spend_logs_metadata. The native team and project controls now live in a shared helper that both paths call, key spend_logs_metadata is copied instead of aliased from the cached key, client metadata cannot overwrite user_api_key_ fields, and header tags dedupe with the same merge used on native routes
Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>
* test(integration): drop covers markers from pass-through tag tests
Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>
* refactor(proxy): type the shared team and project control helper
Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>
* test(integration): validate pass-through endpoint list instead of suppressing pyright
Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>
* test(integration): cover body, streaming, hostile, forged and native cells for pass-through tags
Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>
* test: mock pass-through request url as httpx.URL after rebase on main
* refactor(proxy): merge spend_logs_metadata sources without a stacked comprehension
* refactor(proxy): name the team and request spend_logs_metadata merge
---------
Co-authored-by: ryan <ryan@berri.ai>
Co-authored-by: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>