* fix(proxy): delete large teams without per-member transaction fan-out
Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>
* fix(proxy): evict email-only member caches and reset team members metric on delete
Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>
* fix(proxy): keep new delete-team literals within the LIT002 ceiling
Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>
* fix(proxy): resolve deleted-team member ids before the locked delete
Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>
* fix(proxy): resolve email-only deleted-team members with one case-insensitive lookup
`_deleted_team_member_user_ids` looked each email-only roster entry up with its own
`find_users_by_email` call inside an unbounded `asyncio.gather`: one exact-match query
per email, so a large roster fanned out against the pool again and a roster email that
differed in case from its user row was missed. Add `UserRepository.find_by_emails`, a
single case-insensitive `in` query, and call it once before the locked delete.
`management_helpers/utils.py` goes back to its main-branch shape since the single-email
helper no longer needs exporting.
* fix(repositories): slice find_by_emails into bounded IN statements
The unbounded-IN lint flagged the case-insensitive email lookup added for
/team/delete cache eviction. chunked_in.find_many_in cannot carry Prisma's
insensitive mode, so the repository slices the deduplicated list into
IN_LIST_CHUNK_SIZE statements itself and concatenates the pages. Empty input
still returns () without a query.
* fix(proxy): delete a team once when /team/delete repeats its id
The audit sent {"team_ids": [T, T]}: main answered 400 "User not found in
team" after deleting the keys and memberships and writing two tombstones,
leaving the team row behind; this branch answered 200 but still wrote the
tombstone, audit row and eviction twice. DeleteTeamRequest now collapses
repeated ids in order, so every later step sees each team once and the
response lists each deleted team once.
* test(integration): audit cells for /team/delete on large, legacy and concurrent teams
Thirty-eight deterministic cells in tests/integration/management/ (the CircleCI
integration-management group) covering the /team/delete happy, sad, edge and chaos rows:
250 members against a pool limit of five on two workers, the advisory-lock wait, email-only
legacy roster entries in every casing, member and team cache eviction on both proxies for
every client and endpoint, the Prometheus gauge, audit rows, malformed and duplicate input,
the route gate, and a worker kill, a Redis outage and a proxy restart mid-burst.
Every cell runs against the real proxy, Postgres and Redis with the scripted upstream; no
component is mocked. On the merge base the rows this fix changes are red (P2028 on the
250-member team, two lock waiters, case-mismatched email lookups, duplicate ids, orphaned
LiteLLM_UserTable.teams references under a concurrent burst); on the tip every cell is green
twice with identical selections.
Two pre-existing behaviours are pinned as observed rather than fixed here: a roster entry with
neither user_id nor user_email answers 500, and the LiteLLM_DeletedTeamTable row is committed
before the locked transaction, so a delete that dies in between leaves a tombstone for a live
team and the retry adds a second.
* test(integration): pin each chaos outage to a live /team/delete
The three chaos cells applied the outage once three deletes had answered, which on a fast
run let the whole burst finish before the worker kill, Redis stop or SIGTERM landed, so the
cells passed without exercising the failure. Each cell now holds the first team's advisory
lock from a test-owned transaction, waits until that team's delete is queued behind it in
Postgres with its request unanswered, applies the outage, and only then releases the lock,
so an in-flight delete meets the failure on every run and both legs. The pinned team's
outcome and the number of deletes answered before the outage are recorded as junit
properties (pinned_delete, answered_before_outage).
---------
Co-authored-by: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>
Co-authored-by: ryan-crabbe-berri <ryan@berri.ai>
* fix(hosted_vllm): keep reasoning_content on assistant messages in _transform_messages
vLLM accepts reasoning_content (200 on the wire) and qwen/deepseek/glm
chat templates consume it, so popping it made reasoning models lose
earlier reasoning across tool loops. thinking_blocks is still removed
for vLLM compatibility.
Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>
* fix(hosted_vllm): forward replayed reasoning_content only when it is a string
* test(integration): cover hosted_vllm reasoning_content replay across endpoints
* test(integration): require the surviving worker to serve its held requests in the sigkill chaos cell
---------
Co-authored-by: mateo <mateo@berri.ai>
Co-authored-by: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>
Co-authored-by: mateo-berri <277851410+mateo-berri@users.noreply.github.com>
Co-authored-by: yassin <yassin@berri.ai>
* feat(otel): excluded_services opt-out for datastore spans on tenant destinations
Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>
* test(integration): keep upstream support unchanged
Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>
* fix(otel): excluded_services resolves from the otel callback config only
Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>
* fix(otel): name the otel callback logger so excluded_services owner lookup matches
Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>
* test(integration): assert no aux datastore traces reach the tenant sink
Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>
* test(integration): read bogus-start proxy log from the results dir
Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>
* test(integration): read only this invocation's bogus-start proxy log
Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>
* test(integration): assert operator kept db spans over the whole recorded window
Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>
* test(integration): split operator db-span asserts by trace scope
Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>
* fix(otel): build the otel logger after preset callbacks and validate the exclusion env at boot
Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>
* fix(otel): tolerate a bogus exclusion env when callback config wins
Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>
* fix(otel): keep bogus exclusion env fatal when a preset parses it
Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>
* fix(otel): hoist the preset check out of the callback loop
Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>
* test(otel): use a rule-scoped pyright suppression
Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>
* fix(otel): log and drop unknown excluded_services instead of failing boot
Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>
* test(integration): wait for the operator spend-writer span before checking the tenant for postgres
Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>
* fix(otel v2): leave callback init and boot untouched when excluded_services is unset
Read callback_settings.otel.excluded_services directly instead of making the otel callback build its own logger, and drop the new boot-time parse of callback_settings.otel, so a proxy without the setting behaves exactly as on main
Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>
* fix(otel v2): normalize callback_settings excluded_services without rereading OTel env vars
Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>
* fix(otel v2): log and ignore malformed excluded_services instead of failing startup
Lowercase and trim names, drop non-string items, and add an integration matrix over endpoints, clients, cache hits, destination outages and setting shapes
Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>
* test(otel v2): cover failed upstream calls in the excluded_services matrix
Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>
* test(otel v2): pin operator Langfuse credentials in preset-only excluded_services tests
Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>
---------
Co-authored-by: yucheng <yucheng@berri.ai>
Co-authored-by: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>
Co-authored-by: mrinal <mrinal@berri.ai>
* wip
* feat(traces): establish shared Rust storage foundation
* fix(traces): escape ClickHouse text parameters
* test(traces): exercise response cap with bounded strings
* fix(traces): remove unnecessary lint expectation
* fix(traces): encode ClickHouse timestamp units in Rust
* test(traces): mark exception match as a regex
* refactor(traces): execute schema setup in Rust
* refactor(traces): use shared logging execution wrapper
* docs(traces): replace foundation README with boundary rules
* fix(traces): use current bridge execution facade
* fix(traces): account for protocol cast in lint budget
Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>
---------
Co-authored-by: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>
* fix(responses): scan and mask top-level instructions with guardrails
The Responses guardrail translation handler put a non-empty top-level instructions field into structured_messages as a system row but never into the flat texts list, so guardrails that scan texts skipped it, flat-text masking could not rewrite it, and PANW latest-only selection failed its alignment guard whenever instructions were present.
Seed texts with the instructions row, carry that offset into the flat-text write-back so a rewritten row lands on data["instructions"], and account for the leading row in the PANW Responses alignment.
Resolves LIT-8931
Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>
* fix(responses): reject empty guardrail rewrites instead of forwarding raw input
An explicit texts=[] answer from a guardrail now fails the count check and
raises UnappliableRequestRewrite like any other misaligned rewrite; only a
missing texts key means no rewrite. Types the out-param as dict[str, object]
and adds integration coverage for instructions blocking, masking, empty
instructions, tool loops, latest-only and concurrent workers
Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>
* test(responses): type the texts-replacing guardrail helper explicitly
Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>
* fix(responses): honor skip_system_message_in_guardrail for instructions and system input items
Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>
* test(responses): cover skip_system_message_in_guardrail on the live proxy
Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>
* fix(responses): keep skipped rows through full-coverage rewrites and align latest-only with skip_system
Trust a guardrail's structured_messages_cover_full_request claim only when it
returns as many rows as the full normalized request, otherwise merge the scoped
rows back so skipped instructions and system items survive the write-back.
Make PANW's Responses reasoning alignment skip-aware so latest-only still picks
the latest user turn when system content is excluded from texts.
Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>
* test(responses): annotate new guardrail tests with return types
Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>
* fix(responses): treat an empty guardrail texts answer as no rewrite like chat completions
Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>
* test(responses): type the guardrail test doubles explicitly
Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>
---------
Co-authored-by: yucheng <yucheng@berri.ai>
Co-authored-by: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>
* fix(proxy): strip caller credentials from websocket passthrough
Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>
* test(proxy): cover configured x-api-key in websocket passthrough credential test
Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>
---------
Co-authored-by: oliver <oliver@berri.ai>
Co-authored-by: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>
* feat(agents): authoritative permissions
* fix: enforce authoritative managed agent permissions
* fix(agents): only consult the identity store for managed targets
is_agent_allowed entered the identity-store path whenever a prisma client
was configured, so an ordinary agent paired with an internal user returned
503 instead of 200. Classify the target from the registry first and fall
back to the store only when the registry has no entry, so an unmanaged
target never depends on the store being reachable.
* fix(agents): gate the managed path on an admitted policy object
Ten call sites branched on `managed_agent_policy is not None`, which any
MagicMock attribute satisfies, so the managed path fired on unmanaged
subjects and died in Pydantic validation as a 503. Route every check
through a shared helper that requires a real AgentResponse.
* test(mcp): stub the writer replica the fresh-policy reads use
reload_admitted_user now passes check_db_only through to get_user_object,
so the user row is read from writer_db. Point the mocks at the replica the
code actually reads and give each parametrized case its own user id.
* fix(agents): cap a managed agent at the invoking team's agents
resolve_agent_access returned the managed policy's grants before the
agent_caller ceiling was applied, so a managed agent acting on behalf of a
user reached agents that user's team was never granted. Intersect with the
caller ceiling the unmanaged path already honours.
* fix(agents): restore token narrowing and scope the private-access suppressions
The managed-model check lost its valid_token narrowing when it moved to the
shared helper. Make the caller-access resolver public rather than reaching
into it from module scope, and give each remaining private access a reason.
* docs(agents): drop the comment claiming admins skip the A2A permission check
The check has never had an admin bypass on this path, so the comment
described behaviour the code does not implement.
* test(proxy): stub the writer reads and restore the MCP manager singleton
Fresh-policy user lookups read writer_db, so the team and rest-endpoint
mocks stubbed a replica the code no longer reads, and the dashboard
session fake still had the pre-kwarg signature. The manager reload also
rebound global_mcp_server_manager in every MCP module without restoring
it, leaking an empty manager into later files.
* style: sort imports under the litellm package ruff config
* fix(mcp): cap a managed agent's servers and tools at the invoking caller
managed_agent_servers and managed_agent_tools returned the agent's own
grants without the agent_caller ceiling the unmanaged resolvers apply, so
a managed agent reached MCP servers and tools the echoed caller could not.
Call the existing ceiling helpers on both axes.
* refactor(mcp): return the caller-capped tools without an interim list
The ceiling helper already returns a sequence, so materializing it into a
list added a mutable collection for nothing. Sort at the return sites
instead, which also makes the tool order stable across both branches.
* fix(agents): preserve actor ceilings during managed target checks
* fix(agents): keep managed permission ceilings authoritative
* fix(mcp): fail closed on authoritative caller team outages
---------
Co-authored-by: Joshua Valluru <326636767+joshua-berri@users.noreply.github.com>
* fix(proxy): keep request-body aws credentials out of stored spend-log requests
* fix(proxy): redact every credential-named request-body field in stored spend-log requests
Replace the hard-coded AWS key check in the spend-log request-body sanitizer with
SensitiveDataMasker's key classification, so Azure, Vertex, watsonx, OCI, GigaChat,
Gemini and header credentials are redacted too. Proxy-stamped key identity metadata
is kept.
* fix(proxy): keep request identifiers named like keys in stored spend-log requests
* refactor(proxy): drop the AWS-only snapshot exclusion now that spend-log redaction is name-based
* refactor(proxy): use SensitiveDataMasker's key classification without an exclusion list
* refactor(proxy): always redact credential-named fields in stored spend-log payloads
* fix(proxy): attribute completed batch cost rows to /batches in daily activity
Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>
* test(spend): wait for priced batch tokens before asserting team endpoint activity
Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>
---------
Co-authored-by: kerry <kerry@berri.ai>
Co-authored-by: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>
* test(router): settle the shared logging worker before recording shadow callbacks
Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>
* test(router): always stop the shared logging worker after settling it
Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>
---------
Co-authored-by: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>
* refactor: clean up fresh tech debt from 2026-09-29
Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>
* test(routing): pin usage-based routing Redis reads through the proxy and SDK
Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>
---------
Co-authored-by: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>
* fix(ui): keep MCP permissions visible after key, team and MCP server saves
Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>
* fix(proxy): type the object_permission include as a prisma TypedDict
Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>
* fix(ui): do not block key save confirmation on cache refetch
Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>
---------
Co-authored-by: ryan <ryan@berri.ai>
Co-authored-by: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>
* fix(proxy): look up hashed key names with two spend log rows per key
The spend-log fallback for keys missing from the key table read every row per key to check that all named rows agreed, which passed the 5s statement timeout on busy keys even with the (api_key, startTime) index. Probe only the oldest and newest named row per key, so the lookup stays two index reads per key however much the key logged.
* fix(proxy): cap each spend log name probe at 100 rows per key
* fix(proxy): bound the newest-row probe at where the oldest probe stopped
The newest-row probe now starts at the row where the oldest-row probe gave up, so a key with under 200 rows in the window is read once instead of twice, and the lookup transaction turns bitmap scans off so the planner walks the (api_key, startTime) index instead of every row of a busy key when statistics or the visibility map are stale.
* test(integration): add spend log alias probe cells for the daily activity routes
---------
Co-authored-by: mateo-berri <277851410+mateo-berri@users.noreply.github.com>
* fix(ui): surface x-litellm-call-id in Logs search, table and drawer
Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>
* chore(ui): regenerate api types for spend logs search description
Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>
* test(e2e): drop redundant comments from the call id logs helpers
Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>
* style(e2e): format logs call id helper and spec
Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>
* fix(ui): keep one id per Logs row, move x-litellm-call-id to hover and drawer
The Request ID cell shows only request_id again. When the row's litellm_call_id
differs, the cell tooltip lists it as x-litellm-call-id with its own copy button,
and the drawer header labels the second line x-litellm-call-id: instead of the
call id caption. Stacking two ids in every row made the column noisy for the
common case where the viewer only needs the row they searched for.
* test(e2e): cover the Request ID tooltip hover and copy path
Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>
* test: poll the clipboard after the tooltip copy and drop a jsdom aside
The e2e read navigator.clipboard right after the click, so a slow async write
could fail the check even though copy works. The unit test's fireEvent choice
(jsdom has no layout, so a real pointer move off the trigger closes the tooltip
before the click lands) is documented here instead of inline.
---------
Co-authored-by: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>
Co-authored-by: ryan-crabbe-berri <ryan@berri.ai>
* refactor(auth): bind UI/CLI session tokens to their own AES-GCM context
UI and CLI session tokens are now always encrypted with AES-256-GCM and a
fixed session associated-data value, and the session-token check only accepts
AES-GCM values carrying that same value. Stored secrets keep their current
encryption and decrypt unchanged, so nothing needs migrating.
encrypt_value_helper and decrypt_value_helper take an optional aad. XSalsa20
cannot bind associated data, so an AAD-bound value is always written as
AES-256-GCM, and an AAD-bound decrypt refuses the legacy format.
Session tokens issued before the upgrade stop validating, so UI and CLI users
sign in once more after upgrading.
* test(e2e): cover real SSO login through the dashboard and the lite CLI
Adds two specs under tests/e2e/ui/oidc, run by playwright.oidc.config.ts
against a live Keycloak stack. The dashboard spec checks that the SSO
session authorizes the Virtual Keys and Models data requests. The CLI
spec runs a real lite login in an isolated HOME with the keyring
disabled, then lists models and sends one chat completion with the
stored session. The main Playwright config now ignores oidc/.
* fix(auth): encode UI/CLI session tokens as unpadded base64url
Session tokens carried the v2:gcm: storage prefix and base64 padding. Basic-auth parsers split on the first colon and browsers reject ':' and '=' in WebSocket subprotocols, so Langfuse pass-through and the realtime playground could not use them
Tokens are now plain unpadded base64url, the same header-safe shape as any bearer token
* fix(auth): prefix UI/CLI session tokens with litellm_login_
A prefix-less token starts with sk- about once in 262,144 logins and is then routed as a virtual key, so that login gets a 401. The prefix also makes session tokens easy to spot in logs
The prefix doubles as the token's AES-GCM associated data, so the visible kind and the encrypted kind cannot disagree
---------
Co-authored-by: ryan-crabbe-berri <ryan@berri.ai>
Identity objects (key, end user) load through the request MGET and their write-backs, the registry
reads and the management-object SETs ride the request pipeline. A team refresh invalidates its alias
with a pipelined DEL instead of a synchronous DEL plus a duplicate async one, and an MGET miss is
remembered so no per-key GET follows it in the same request.
Resolves LIT-9012
Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>
Co-authored-by: yassin <yassin@berri.ai>
Post-call owners declare into one request-scoped RedisBatch per Redis backend: spend counter
increments and reservation reconciliation, rate-limit token Lua updates and refunds, parallel-slot
release (freed locally at once), deployment TPM, and compatible async response-cache SETs. The batch
is sent once the success and failure callbacks have run, or on a deadline, and pending batches are
drained at shutdown before Redis disconnects. nx writes, non-Redis caches and calls outside a request
stay direct; numeric string TTLs keep the direct-path coercion.
Resolves LIT-8883
Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>
Co-authored-by: yassin <yassin@berri.ai>
* fix(mcp): key discovery caches per caller correctly and drop stale caches on server updates
Discovery-list cache identity now uses the hashed token instead of the raw
api_key and treats MCPJWTSigner-signed servers as per caller. Server
definition changes also drop the cached upstream OAuth metadata. OpenAPI
listings look tools up under the normalized registry prefix with the
separator, so an overlapping sibling prefix no longer leaks into the list.
Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>
* fix(mcp): keep the discovery cache digest call unchanged so CodeQL matches the existing alert
Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>
* fix(mcp): guard OAuth metadata cache writes with a per-server generation and drop unproven per-caller discovery keys
An upstream metadata fetch that started before a server edit could store its stale reply after
invalidate_oauth_metadata_cache ran. Invalidation now bumps a per-server generation and the fetch
only stores when the generation it captured before I/O is unchanged.
The MCPJWTSigner-based per-caller discovery classification and the api_key to token key change had no
reproduction (the signer only injects on tools/list, and UserAPIKeyAuth hashes api_key in place), so
both go back to the merge-base behavior.
Integration coverage under tests/integration/mcp: overlapping OpenAPI aliases, a config-declared
server name with a space, OAuth metadata refetch after a save, and the in-flight stale-write race
Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>
* fix(mcp): keep OAuth metadata generations only while a fetch is in flight
Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>
* fix(mcp): count queued OAuth metadata fetchers so invalidation survives lock handoff
Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>
* fix(mcp): keep a held OAuth metadata lock registered even when no fetcher slot claims it
Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>
* test(mcp): prove a peer worker drops stale upstream OAuth metadata after a save elsewhere
Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>
---------
Co-authored-by: yucheng <yucheng@berri.ai>
Co-authored-by: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>
RedisBatch: one pipeline per Redis backend for independently declared operations (MGET, GET, Lua
scripts, INCRBYFLOAT, SET, DEL), a future per operation so each owner keeps its own fallback, Redis
Cluster hash-slot fallback. A request-scoped batch middleware shares that pipeline across the auth
identity reads and write-back, the spend counter MGET, the rate limiter Lua groups and the routing
read. A rate-limit denial stands when another pipelined group fails; every pipelined group is refunded
on rejection; local cooldowns win over the prefetch.
The routing prefetch failure log line strips request line breaks (CodeQL py/log-injection)
Resolves LIT-8882
Co-authored-by: yassin <yassin@berri.ai>
Co-authored-by: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>
* fix(proxy): recover daily spend key owners
Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>
* fix(proxy): simplify daily spend owner recovery
Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>
* style(proxy): format daily activity metadata
Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>
* test(proxy): cover recovered owner metadata merge
Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>
* fix(proxy): bound the daily spend owner lookup with the statement timeout
* test(integration): audit the daily activity key owner fallback on every usage route
Thirty five integration cells under tests/integration/spend cover the daily
spend owner fallback on all nine daily activity routes and /usage/ai/chat:
the happy path per route, the unanimity rules (two users, blank and null
rows, an owner the user table lacks, live and deleted keys with and without
their own user, a spend log alias), a non admin reader, an invalid key, a 5 KB
key, a locked LiteLLM_DailyUserSpend, 300 keys of one team, repeated reads, a
second user landing between reads, a concurrent burst across the unified
endpoints, a killed worker, and a proxy restart
The traffic cells ignore the GET /v1/models call the proxy's five minute token
limit refresh makes to every registered OpenAI compatible deployment, since it
lands on a test's provider wire whenever the refresh instant falls inside the
test
---------
Co-authored-by: jesus <jesus@berri.ai>
Co-authored-by: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>
Co-authored-by: mateo-berri <277851410+mateo-berri@users.noreply.github.com>
Claude Code background sessions (claude --bg) stamp x-app: cli-bg on every
request, including main-loop turns. The session router binding only
accepted x-app: cli, so a background session never bound and its
subagents' concrete-model calls bypassed the router.
The binding write already requires the requested model to resolve to a
pre-routing strategy, so background side calls naming plain models still
never bind. Since f6eff1bde0 removed the clear path, the x-app check
guarded nothing else.
Co-authored-by: Claude Opus 5.5 <noreply@anthropic.com>
* perf(responses): run aresponses through the async wrapper so the cache is read once
aresponses() sets kwargs["aresponses"] = True and runs the decorated sync
responses() on an executor, but _is_async_request() did not recognise that
flag, so the sync wrapper did a second cache lookup on the executor thread
with a differently ordered cache-key input. Every /v1/responses request paid
two cache GETs against two different keys. Recognising aresponses in
_is_async_request() leaves the async wrapper as the only cache reader and
writer for the async path, one GET per request, same key on read and write
Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>
* fix(caching): let a responses cache entry cover aresponses so responses-only configs keep caching
Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>
---------
Co-authored-by: yassin <yassin@berri.ai>
Co-authored-by: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>
Auth's spend counter MGET scope spans common checks, model budget check and reservation;
reservation increments go out as one pipeline; post-call reconcile adjustments ride the ordinary
increment pipeline and update_cache uses one batched read. Over-budget reservation counters are
charged one at a time so a rejection never touches the counters after it; post-call counter keys are
derived from ids without validating a UserAPIKeyAuth.
Resolves LIT-8881
Co-authored-by: yassin <yassin@berri.ai>
Co-authored-by: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>
* perf(router): fetch cooldown state and usage counters in one Redis round trip
The cooldown filter (CooldownCache) and usage-based-routing-v2 selection
(LowestTPMLoggingHandler_v2) each issued their own MGET on every request
because they live in different objects. RoutingReadBatch fetches both key
sets through DualCache.async_batch_get_cache_shared while the healthy
deployments are resolved and hands the usage slice to the strategy, so
selection does not read again. Each cache keeps its own memory tier,
throttling, reservation rollback and circuit-breaker handling, and the
strategy falls back to its own read when the prefetch does not cover its
keys. simple-shuffle keeps reading only cooldowns.
aresponses no longer issues a second, blocking response-cache read from
the worker thread that runs the sync wrapper.
Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>
* fix(caching): keep per-cache tier failures inside the shared batch read
Wrap the memory-tier prepare and backfill steps of DualCache.async_batch_get_cache_shared
so a failing tier degrades that cache's read to None the way async_batch_get_cache does,
instead of escaping into routing. Drop the aresponses sync-cache guard: for native
Responses models the worker-thread read is the one whose key matches the write, so
skipping it broke cached /v1/responses replays.
Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>
* refactor(router): rename usage key builder so the async cache-call check reads it as a key helper
Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>
* fix(alerting): narrow daily-report cache values before numeric comparison
Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>
* refactor(caching): type the shared batch-read helpers and merge Redis results without mutation
Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>
* style: fix import sort in test_dual_cache
Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>
* refactor(caching): flatten shared batch read keys without a stacked comprehension
Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>
---------
Co-authored-by: yassin <yassin@berri.ai>
Co-authored-by: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>
* fix(streaming): keep the provider's served service_tier on streamed chunks and spend rows
Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>
* fix(streaming): satisfy type-discipline and strict ruff budgets
Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>
* fix(streaming): stamp the served service_tier on every Responses bridge chunk
Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>
* fix(anthropic-adapter): expose streamed chunks so disconnects bill partial spend
Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>
* test(service-tier): cover anthropic and responses served-tier billing paths
Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>
* fix(anthropic-adapter): return a chunks-exposing stream so disconnects bill partial spend
Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>
* test(service-tier): bill disconnects through the router's anthropic stream wrapper
Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>
* style: apply ruff format to the anthropic stream changes
Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>
* test(coverage): ignore delegating properties the ast scan cannot see
Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>
* style: keep the cast-ok reasons on the cast call line
Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>
* test(integration): cover served service_tier billing for streamed chat and messages, complete and disconnected
Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>
* fix(anthropic-cache): delegate chunks/messages/model through the messages stream cache writer
Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>
* fix(streaming): keep service_tier on OpenAI-compatible parsed chunks
Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>
* refactor(streaming): parameterize delegated chunks and messages types
Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>
* fix(tests): follow the anthropic pass_through rename after merging main
Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>
* test(anthropic): drain the logging worker between response cache tests
Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>
* test(spend): cover azure, databricks, responses bridge and gemini served tiers in the stream billing integration test
Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>
* fix(databricks): keep the served service_tier on streamed chunks and bill it
Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>
* fix(databricks): type the served service_tier chunk without a loose kwargs dict
Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>
* fix(cost): bill the served service_tier over the requested one
Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>
* refactor(cost): drop explanatory comment from the tier resolution
Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>
---------
Co-authored-by: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>
Co-authored-by: kerry <kerry@berri.ai>
* fix(model-prices): correct azure/eu/gpt-6-astra to Data Zone rates
Co-authored-by: rain <1504569896@qq.com>
Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>
* fix(model-prices): align groq, gemini and openai entries with official docs
Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>
* fix(model-prices): roll in verified Vertex, Gemini, OpenRouter and Azure AI registry fixes
Absorbs the fields from #43609, #43666, #43671 and #43644 that match the provider's own docs or price API today, and adds a cost test for the azure/eu/gpt-6-astra Data Zone tiers
Co-authored-by: bunnysayzz <stfuazzo@gmail.com>
Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>
* fix(model-prices): add Copilot, Bedrock Kimi K3, Gemini Robotics and OpenRouter values from official sources
Co-authored-by: Michal Formanek <michal.formanek@generaliceska.cz>
Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>
---------
Co-authored-by: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>
Co-authored-by: rain <1504569896@qq.com>
Co-authored-by: bunnysayzz <stfuazzo@gmail.com>
Co-authored-by: Michal Formanek <michal.formanek@generaliceska.cz>
* feat(cost_calculator): add cost_per_second for chat per-second pricing
Keep legacy input_cost_per_second and output_cost_per_second as aliases for chat, completion, embedding and responses. When both legacy fields are set, input_cost_per_second wins
Move Bedrock commitment rows to cost_per_second so they bill once
Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>
* refactor(cost_calculator): drop legacy per-second fields from chat paths
Keep Azure chat token pricing generic and update inert Voxtral rates and SageMaker examples
Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>
* fix(cost_calculator): recognize output-only per-second rates
Include output_cost_per_second when checking whether a deployment cost entry has pricing so output-only legacy aliases remain attached to the deployment during cost selection
Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>
* test(pricing): cover cost_per_second and legacy per-second aliases through the proxy
Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>
* refactor(cost_calculator): drop output_cost_per_second as a chat per-second alias
Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>
* feat(cost_calculator): restore output_cost_per_second as a chat per-second fallback
Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>
* fix(cost-map): keep input_cost_per_second on bedrock commitment rows for older clients
Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>
---------
Co-authored-by: kerry <kerry@berri.ai>
Co-authored-by: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>
* test(proxy): classify every credential-bearing param for the canary suite
* test(proxy): classify gcs_path_service_account as secret, run registry in auth-checks shard, check slot ids at import
* test(proxy): use one generic slot id for callback and request-body credential params
* test(proxy): name a canary slot only for params an integration test plants
* test(proxy): classify the SigNoz callback params
* test(proxy): move the slot sync note into the module docstring
* test(proxy): model unplanted credential params as their own classification
* test(security): classify request-body api_key as unplanted until D1 exists; check registry slots against the harness
* test(security): classify request-body api_key under slot D1
* test(security): classify Langfuse and Datadog callback secrets under slots C1 and C3
* test(integration): credential canary suite harness
Adds tests/integration/security with canary generation and search, sweeps over the database, GET routes, client responses, sink doubles and Redis, an owned proxy rig, a sweep sensitivity self-test and the config deployment api_key slot. Registers the security group in run.py, the manifest and the CircleCI integration matrix.
* test(integration): widen canary route sweep and harden the rig
Enumerate lazily registered feature routers, call parameterized routes with placeholder ids, fail on routes that return no response, skip provider pass-through routes, add an explicit admin-only route allowance, let the sink double use a configurable token, inflate gzip members anywhere in a blob, sweep Redis before the route walk, and trap outbound connections from the owned proxy.
* test(integration): descend into any decoded value that can still hold an encoded canary
* test(integration): bound canary decoding by depth and decoded bytes
* test(integration): scope log-table and spend-log reads to the scenario window
* test(integration): sweep spend-log rows in the scenario date window
* test(integration): keep spend-log date window summarized
* test(integration): sweep proxy logs, metrics, a gzip Datadog intake and the Logs drawer for credential canaries
* test(e2e): treat an unset prompt-storage setting as unset and restore it
* test(integration): name the Datadog sink slot G1d
* test(e2e): search the Logs page for base64 forms of the deployment key
* test(integration): resolve deployment ids, scope paginated log lists, key allowances by slot
* test(integration): pass the resolved deployment id to the Datadog route sweep
* test(integration): expect 404 from the caller-scoped team membership route
* test(integration): use the rig's own master key and expect 404 from submission lookups
* test(integration): check the overridden rig key without assuming the default key is unknown
* test(integration): credential canary suite harness
Adds tests/integration/security with canary generation and search, sweeps over the database, GET routes, client responses, sink doubles and Redis, an owned proxy rig, a sweep sensitivity self-test and the config deployment api_key slot. Registers the security group in run.py, the manifest and the CircleCI integration matrix.
* test(integration): widen canary route sweep and harden the rig
Enumerate lazily registered feature routers, call parameterized routes with placeholder ids, fail on routes that return no response, skip provider pass-through routes, add an explicit admin-only route allowance, let the sink double use a configurable token, inflate gzip members anywhere in a blob, sweep Redis before the route walk, and trap outbound connections from the owned proxy.
* test(integration): descend into any decoded value that can still hold an encoded canary
* test(integration): bound canary decoding by depth and decoded bytes
* test(integration): scope log-table and spend-log reads to the scenario window
* test(integration): sweep spend-log rows in the scenario date window
* test(integration): keep spend-log date window summarized
* test(integration): request-path credential canary slots D1-D4
* test(integration): read the Logs drawer and spend-log filter for failed request rows
* test(integration): check the marker in each failed row's spend-log filter; run header slots on chat-family routes
* test(integration): run D2 on embeddings again; only the client-header slot runs on chat-family routes
* test(integration): resolve deployment ids, scope paginated log lists, key allowances by slot
* test(integration): pass the slot deployment's model_info id to the route sweep
* test(integration): expect 404 from the caller-scoped team membership route
* test(integration): use the rig's own master key and expect 404 from submission lookups
* test(integration): check the overridden rig key without assuming the default key is unknown
* test(integration): credential canary suite harness
Adds tests/integration/security with canary generation and search, sweeps over the database, GET routes, client responses, sink doubles and Redis, an owned proxy rig, a sweep sensitivity self-test and the config deployment api_key slot. Registers the security group in run.py, the manifest and the CircleCI integration matrix.
* test(integration): widen canary route sweep and harden the rig
Enumerate lazily registered feature routers, call parameterized routes with placeholder ids, fail on routes that return no response, skip provider pass-through routes, add an explicit admin-only route allowance, let the sink double use a configurable token, inflate gzip members anywhere in a blob, sweep Redis before the route walk, and trap outbound connections from the owned proxy.
* test(integration): descend into any decoded value that can still hold an encoded canary
* test(integration): bound canary decoding by depth and decoded bytes
* test(integration): scope log-table and spend-log reads to the scenario window
* test(integration): sweep spend-log rows in the scenario date window
* test(integration): keep spend-log date window summarized
* test(integration): resolve deployment ids, scope paginated log lists, key allowances by slot
* test(integration): expect 404 from the caller-scoped team membership route
* test(integration): use the rig's own master key and expect 404 from submission lookups
* test(integration): check the overridden rig key without assuming the default key is unknown
* test(integration): callback credential canary slots C1-C3 and D5
Team callback, team callback_settings, config default_team_settings and key metadata.logging Langfuse secrets, a team Datadog dd_api_key, and request-body Langfuse keys (allow_client_side_credentials) must reach only their sink. Each scenario checks its sink received the canary as auth and that the marker is visible at the stored body, the Logs drawer route and the sink. Adds a unit test that the stored request body snapshot carries no callback parameter.
* test(integration): give the callback sink waits a wider bound
* test(integration): sweep provider requests for callback credentials
* feat(guardrails): send a configured gateway_name from noma_v2 to Noma
The noma_v2 guardrail accepts a gateway_name param, falling back to the
NOMA_GATEWAY_NAME env var. The value is stripped, and when it is non-empty
it goes out as a top-level gateway_name field on /litellm/guardrail. The
param works for both guardrail: noma_v2 and guardrail: noma with use_v2,
and it is appended after the existing constructor params so positional
callers keep their meaning
* chore(ui): regenerate OpenAPI snapshot and dashboard types for gateway_name
The new noma_v2 gateway_name param shows up in the proxy OpenAPI spec, so
the lazy snapshot and the generated dashboard types need regenerating
* Update litellm/proxy/guardrails/guardrail_hooks/noma/noma_v2.py
Co-authored-by: greptile-apps[bot] <165735046+greptile-apps[bot]@users.noreply.github.com>
---------
Co-authored-by: Claude <noreply@anthropic.com>
Co-authored-by: greptile-apps[bot] <165735046+greptile-apps[bot]@users.noreply.github.com>
* fix(router): stream anthropic messages lifecycle frames live when no fallback can take over
The /v1/messages streaming wrapper buffered message_start and
content_block_start until the first content_block_delta and dropped
pings behind buffered frames unconditionally, even for requests no
fallback could ever recover. With adaptive thinking on Bedrock or
Vertex the client saw no bytes for the whole thinking pass and hit
read timeouts.
Buffering now applies only while a fallback can still take over
(generic or refusal chain resolving), and a ping is always forwarded
live since it carries no lifecycle and keeps the connection alive.
Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>
* fix(router): mirror every dispatcher fallback path in the anthropic stream gate
Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>
* fix(router): skip already-tried order levels in the anthropic stream gate
Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>
* fix(router): keep a transport-split ping behind buffered lifecycle frames instead of forwarding its head live
Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>
* chore(router): credit the #39566 branch this fix supersedes
Co-authored-by: Radu Swigler <radu.porumba@gmail.com>
Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>
---------
Co-authored-by: mateo <mateo@berri.ai>
Co-authored-by: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>
Co-authored-by: yassin <yassin@berri.ai>
Co-authored-by: Radu Swigler <radu.porumba@gmail.com>