Anonymous protected-resource discovery now runs the same should_run_guardrail probe as a keyed
connect, so a default_on guardrail whose Mode only has tags and no default advertises the gateway
authorization server as the merge base did instead of the Entra issuer and scope
Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>
The caller sign-in subject no longer depends on how the caller was admitted: a non-virtual bearer in
Authorization is the subject unless it repeats x-litellm-api-key, so built-in OAuth2 and JWT admissions
forward the caller's token to Agent 365 as the merge base did. The connect gate pre-flights that subject
only while connecting; on an open session the tool-call hook runs the single exchange and answers a
rejection inside the JSON-RPC envelope with its guardrail Logs row, again as the merge base did
Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>
The connect gate only needs the body to tell initialize from other methods, so the peek now runs on demand through a callback and the consumed ASGI messages replay to the handler. A gated route answers 401 before the body arrives again, as the merge base did
Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>
The regression test holds the POST body back forever and expects the 403 for a session owned by someone else within two seconds, red on 7603412660 where the connect-time peek waited for the body first
Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>
A POST that names an existing mcp-session-id skips the connect-time body peek, so another caller's request is refused with 403 before any body byte is awaited. Session-bearing bodies are read after the owner check and replayed to the session manager unchanged
Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>
The connect-time sign-in preflight raised its fail-closed 503 on every method of a single-server route, so a post-connect tools/call whose Entra exchange hit a gateway fault got a bare 503 with no JSON-RPC envelope, no hook run and no guardrail Logs row. The gate now reads the body before the preflight and only raises the 503 while the request is an initialize (or the SSE GET); any other request reaches the tool-call hook, which answers 200 isError with the verdict and writes the failure row as the merge base did
The guardrail posts to the Entra token endpoint through its own handler again, bounded by request_timeout, and classifies the answer itself through the shared OAuth error reader, so the gateway-fault reason keeps the OAuth code (invalid_client, invalid_scope, ...), a malformed caller assertion (AADSTS 50027xx) stays a 401 caller fault reading rejected (invalid_client), and transport, HTTP 500, non-JSON and missing access_token answers read as the merge base's distinct reasons instead of one collapsed sentence. build_token_exchanger takes the HTTP post as a dependency in place of request_timeout
Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>
The unselected aggregate connect test asserted only on a patched has_user_oauth_token
call count, which the test-quality gate flags as mock-echo (TQ002). It now calls the
preflight unpatched and asserts it returns without a sign-in challenge, so the mutant
that probes every registered server is still killed by the 401 it would raise.
Also applies ruff format to test_caller_sign_in.py, a file new in this PR
Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>
_scoped_server returned a denial whenever the registry placed a scoped name on a server the caller
does not hold, so a key granted only access group docs lost the group's servers when an ungranted
server was also named docs. It now returns None there, as it does for a name hidden from the caller's
IP, and the caller retries the name as an access group it holds, which is what the merge base did.
Regression tests pin the moved oauth_delegate and gateway oauth2 route shapes (alias case, server id,
aggregate x-mcp-servers, authorization-server document, authorize relay) to the exact-name route, and
the unselected aggregate connect to no challenge
Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>
A malformed caller assertion that Entra reports as invalid_client with an AADSTS50027xx error_codes entry is now the caller's 401 sign-in challenge in the shared token exchange provider, never a 503 gateway-credential fault or a fail-open pass, so both the Agent 365 guardrail and plain OBO servers classify it the same way
The Agent 365 sign-in provider advertises a server's configured scopes for every server and falls back to api://<client_id>/access_as_user only when none are configured
Connect-time challenges from both the Agent 365 provider and plain OBO carry resource_metadata as an absolute URL built from the request origin (trusted forwarded headers honored) and naming the route the client actually used, so /<server>/mcp and /mcp/<server> each point at their own protected-resource document
Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>
The connect preflight runs the scoped router's grant-first lookup for every single-server connect, so a granted plain server named like an ungranted OBO server's alias is served instead of intercepted by that server's 401, and obo_without_subject is read off the server the lookup picks
Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>
The connect preflight passes connected_as only for caller sign-in servers, so an oauth2_token_exchange server challenges with the merge-base resource_metadata path (its alias) on /mcp/<server_name>, the aggregate route and the legacy route as well as on /mcp/<alias>. The case-variant integration assertion that expected 403 is corrected to the 200 both base and head return
Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>
A key granted only the server named `docs` was refused with 403 on
/mcp/docs when an ungranted server held `docs` as its alias, and a key
granted only `github` was refused on /mcp/GITHUB when an ungranted
`GitHub` existed: the scoped router took the registry-wide pick and
answered "denied" whenever that pick was not among the caller's servers.
The router now runs the registry's own pass order (exact alias,
server_name, name, server_id; the same case-insensitively; prefix forms)
over the caller's granted servers first, through
get_mcp_server_answering_to(among=...), with IP hiding applied at the
pass that found the server. A name hidden from the client IP stays denied
before any grant lookup, and a name the registry places only on an
ungranted server stays denied rather than being retried as an access
group.
The connect preflight reuses the router's selection for a single scoped
name as the server it challenges, signs in and exchanges for, so a 401
names the granted server; an ungranted caller keeps the registry pick and
the downstream 403, and the no-key path is unchanged. Unauthenticated
RFC 9728 discovery has no grant list and stays on the registry pick.
Tests: the `only_b` assertion in
test_scoped_router_selects_the_server_the_connect_preflight_resolves and
the no-IP `internal` assertion in
test_scoped_router_hides_a_private_server_from_an_external_ip_like_the_connect_preflight
now expect the granted server, which is what the base branch returned for
both shapes; the unentitled-key exchange test's fixture becomes the empty
selection the router returns for such a key; listing tests that stub the
manager now point its lookup at a real empty manager so the `among` pass
runs. New tests cover the grant-first router, the manager `among` pass
order and IP hiding, and the connect preflight under an alias collision.
The slot a tools/list records for a later tools/call was keyed by (user_id, api_key)
only, so every team-only JWT caller shared one slot and one JWT user acting in two
teams shared a slot; a tools/call then handed pre_mcp_call hooks a description another
caller was served. The slot is now keyed by the hashed key, user, team and organization,
plus the admission credential of a caller admitted with neither a key nor a user.
The caller bearer split the slot only on client-forwarded-token and token-exchange
servers; a legacy delegated oauth2 server (delegate_auth_to_upstream without client
credentials) also forwards it upstream and served a different catalog per bearer into
one slot. The bearer now splits the slot on every server whose egress forwards it
(_consumes_caller_authorization) or exchanges it.
The timeout reliability tests rely on a 1ms deadline the real backend always
misses. With E2E_PROVIDER_CACHE on, the deployment pointed at the cache edge,
and its healthy sibling in the same test had already recorded a response for
the same canonical request, so the edge answered from Redis inside the 1ms read
window. Build 342 of litellm-e2e saw test_timeout_trips_cooldown_then_recovers
get a 200 from the timing-out deployment itself, with a recording made about
12 hours earlier. Both timeout helpers now register on the live provider path,
which PROVIDER_CACHE.md reserves for tests that need real provider timing
* feat(proxy): add LITELLM_DISABLE_LAZY_ROUTES to register optional routers at startup
Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>
* fix(proxy): register eager lazy routes at startup so late eager routes keep precedence
Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>
* refactor(proxy): share one optional-feature install path between lazy and eager registration
Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>
* docs(proxy): say eager lazy routes register at worker startup
Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>
* refactor(proxy): name the import callable passed to _install
Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>
* test(integration): prove a startup hook can drop eager lazy routes for good
Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>
* test(integration): strip the lazy-routes flag from the lazy-mode control proxies
Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>
* test(proxy): cover the lazy warmup route registering a feature
Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>
* ci(unit): run tests/unit/proxy/test__lazy_features.py in the proxy-server-core shard
Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>
* test(integration): audit cells for LITELLM_DISABLE_LAZY_ROUTES
Flag spellings, /openapi.json at boot, the warm-up route in both modes, /mcp/proxy ahead of the /mcp mount,
route table and OpenAPI parity with a fully warmed lazy proxy, a broken optional import in both modes, every
client SDK against the completion endpoints, a boot burst with a killed worker, and a restart
* fix(proxy): keep config pass-through routes ahead of eagerly registered features
With LITELLM_DISABLE_LAZY_ROUTES set, features registered before the proxy lifespan
added config pass-through endpoints, so a pass-through overlapping a feature path
(e.g. a self-hosted /langfuse) lost to the built-in route. Restore lazy mode's
registry order once startup finishes, without bringing back routes a startup hook removed
---------
Co-authored-by: ryan <ryan@berri.ai>
Co-authored-by: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>
The aggregate listing and _list_mcp_tools now default to record_listing=False,
so a catalog fetched inside a tools/call no longer fills the caller's
listed-tools slot. The /mcp/proxy meta-tools (call_tool, search_tools,
get_tool_schema) and the tool-search virtual tool stop recording: /mcp/proxy
serves only the meta-tools and the search serves only its hits, so a later
pre_mcp_call hook was reading a description the caller never listed.
The tools/list handler, the Responses MCP handler and the /v1/mcp/tools
management listing opt in with record_listing=True, since each serves the
catalog to the caller.
* feat(ui): add a call that posts an OTLP export to the proxy
* feat(ui): build a sample agent run as an OTLP export
* feat(ui): add a pulsing dot for active tracing
* feat(ui): render a crisp sample run preview from real trace components
* chore(ui): remove the pixelated agent traces preview image
* feat(ui): add test trace, tracing key, otel endpoints and more frameworks to tracing setup
* test(ui): cover tracing setup test trace, key masking, endpoints and frameworks
* feat(ui): open the received test trace and mark tracing as active
* test(ui): mock the new tracing setup network calls
* fix(ui): let tracing setup use the full page width
* fix(ui): send OTLP/JSON sample trace ids as hex per the spec
* test(ui): cover the sample trace export shape and hex ids
* fix(bedrock): accept Converse messages with no content key
A user or tool message whose content key is missing (or null, which the
message cleanup strips) made every Bedrock Converse request fail with
APIConnectionError 'content' before reaching Bedrock. The Converse
transform now reads content with .get for those messages, as it already
did for assistant messages: a content-less user message adds no block
and a content-less tool message becomes a toolResult with empty content.
The str branch also sends the continue message text instead of the
original whitespace-only text.
* fix(bedrock): send the continue message for a content-less user turn
* test(bedrock): type the content-less Converse message test parameters
* fix(bedrock): accept a Converse system message with no content key
* refactor(bedrock): read the system message content with get
* test(integration): audit Bedrock Converse messages without content
Adds the /audit cells for a chat message whose content key is missing or
null on a Converse-routed Bedrock deployment: happy, sad, edge, and chaos
rows through the OpenAI SDK, the Anthropic SDK, and raw httpx against the
scripted upstream, asserting the caller's response, the body the peer
received, and the spend row. The owned-proxy readiness deadline in the
integration harness is now INTEGRATION_PROXY_READY_SECONDS (default 70).
* test(integration): bound stray spend rows in the mid-burst restart cell
---------
Co-authored-by: mateo-berri <277851410+mateo-berri@users.noreply.github.com>
_get_tools_from_server now records the catalog into the caller's
listed-tools slot only when asked (record_listing=True), which the
served listings pass: the /mcp and Responses API tools/list handlers via
_get_tools_from_mcp_servers, MCPServerManager.list_tools, and the REST
listing via _list_server_tools. Four internal listings stop recording,
so a later tools/call hands pre_mcp_call hooks name and arguments only,
as on main:
- _list_tools_before_first_call, the implicit listing inside tools/call
when this worker does not yet expose the tool
- fetch_pinnable_tool_catalog, the admin pin snapshot listed without the
catalog guard and without description overrides
- _initialize_tool_name_to_mcp_server_name_mapping, the startup fill
- get_tools_for_server, used by the semantic tool filter
_create_prefixed_tools returns to its tool-name mapping job only; the
record follows it in _get_tools_from_server.
* fix(mcp): expand team and dashboard grants when listing and serving toolsets
Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>
* fix(mcp): honor team toolset grants on the responses gateway path and expose a public team permission lookup
Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>
* chore(mcp): drop the toolset route docstring tweak so the OpenAPI snapshot stays unchanged
Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>
* fix(mcp): scope inherited toolset grants by the key's own MCP ceiling
A key that declares any MCP grant of its own keeps only its own toolsets, one
that declares none inherits its team's, and require_key_mcp_access_defined
stops a virtual key inheriting while dashboard sessions and admitted users
still do. Adds the direct, no-grant, admin and key-ceiling integration cases
and makes the LLM gateway toolset case discriminate a scoped toolset from the
aggregate grant
Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>
* fix(mcp): resolve toolset grants per admitted source and enforce the live team roster
Dashboard sessions and gateway-admitted users now expand into the admitted
subject's per-team sources when resolving toolset grants, so a team-granted
toolset is not capped by the user's own MCP row and is reachable on the
namespaced route. A cached team id no longer grants a toolset unless the live
roster still lists the user, a team lookup fault only drops that team's
inherited grant, and /team/member_add evicts the cached team object so the new
member is authoritative immediately
Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>
* fix(mcp): read a key's named object permission before letting it inherit team toolsets
Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>
* refactor(mcp): inject the toolset grant resolver into scope helpers so tests stop patching MCPRequestHandler
Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>
* fix(mcp): drop the duplicate admitted_subject_sources wrapper after merging main and follow its renamed resolvers
Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>
* fix(mcp): honour fresh policy and the session resource scope on pinned toolsets
A pinned toolset scope now reads the toolset through the writer when the admitted session
requires fresh policy, so a tool revoked from the toolset is gone on the next request. A
gateway bearer scoped to one server can only open a toolset that names that server
Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>
* fix(mcp): keep operator-open servers out of toolset gateway urls
Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>
* test(mcp): audit cells for team-granted toolsets across every surface
Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>
* test(mcp): bound toolset edit convergence by both cache layers
Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>
* test(mcp): drop toolset integration cells that test behavior this PR does not change
Repeat-read byte identity, 20 concurrent calls, a stopped peer and a killed worker are covered
generically by test_mcp_resilience.py and test_mcp_user_env_vars.py. The toolset edit cell asserts
pre-existing cache propagation and flaked locally with connection resets while polling
* chore(mcp): drop mutable-ok suppressions that main's LIT013 now flags as unused
* chore(mcp): keep the require_key_mcp_access_defined read from adding an unknown-argument type error
---------
Co-authored-by: ryan <ryan@berri.ai>
Co-authored-by: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>
The two SpendLogs index migrations shipped in v1.103.0 each break one table shape: the plain CREATE INDEX holds a SHARE lock on a large unpartitioned table and the CONCURRENTLY one fails with 0A000 on a partitioned parent. Both files are now inert and the indexes are built by a table-driven, shape-aware step after migrate deploy: CONCURRENTLY on a plain table, ON ONLY the parent plus per-partition CONCURRENTLY and ATTACH PARTITION on a partitioned one. The migration job builds them synchronously and exits non-zero on failure; a serving proxy that ran migrate deploy itself builds them in the background off the readiness path. A valid index of the same definition under another name is renamed and reused, an invalid one is rebuilt, and extra copies are reported with their DROP INDEX statement instead of being dropped. The migration checker rejects any CREATE INDEX on LiteLLM_SpendLogs or LiteLLM_ErrorLogs in future migrations
Co-authored-by: yassin <yassin@berri.ai>
Co-authored-by: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>
The Lens rename check added in #44034 refuses prisma db push when it cannot
reach the database. This test pointed DATABASE_URL at a dead port to keep the
real CI database out of the run, so it now trips that check before the push
timeout it exists to cover. Unsetting DATABASE_URL skips every database probe
and keeps the test focused on the timeout hint.
* add test case for /rag/query and stronger auth check
* style(proxy): ruff format auth_checks.py
Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>
* fix(proxy): build rag query vector store ids immutably and test the no-registry path
Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>
* test(proxy): integration coverage for /v1/rag/query vector store allowlist
Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>
* test(proxy): audit cells for /v1/rag/query vector store allowlist
Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>
* test(proxy): consolidate vector store allowlist audit coverage
Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>
* fix(proxy): type RAG vector store request body
Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>
---------
Co-authored-by: Mrinal Chanshetty <mrinal@berri.ai>
Co-authored-by: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>
The usage card hid actual and baseline spend unless every older session
could be rebuilt from SpendLogs within two seconds, which on a real
gateway it never was. Each complexity router's actual spend is now its
rollup spend and its baseline is spend plus recorded savings, for old
and new requests alike, so the benchmarks and session endpoints never
scan SpendLogs. Adaptive and quality routers record no savings baseline
and stay out of the compared totals; savings_estimated_classifier_cost
is kept and covers the same compared requests
* feat(proxy): gzip buffered responses for clients that accept it
Large JSON reads like /user/daily/activity/aggregated shipped tens of MB
uncompressed. Compress single-message bodies of 500B or more when the
client's Accept-Encoding allows gzip (q-values and the wildcard honored).
Streamed and etagged responses pass through untouched, every negotiable
response carries Vary: Accept-Encoding, and bodies of 1MB or more are
compressed in a worker thread
* fix(proxy): skip partial and no-transform responses in gzip and always release the held start
The gzip gate now also skips 206 Partial Content and Cache-Control: no-transform,
since compressing either breaks byte ranges or ignores an explicit ban on transforms.
A response start without a headers key no longer raises, and a start the app never
follows with a body message is forwarded when the app returns instead of being dropped.
Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
---------
Co-authored-by: Claude Opus 5.5 <noreply@anthropic.com>
* feat(tool-policies): show the user who owns the key that discovered a tool
GET /v1/tool/list and GET /v1/tool/{tool_name} resolve the discovering key's owner from the verification token and user tables at response time and return it as a nullable user field. The Tool Policies page adds a User column that shows alias, then email, then ID, with the same cell the Virtual Keys page uses. Keys without an owner, deleted owners, and rows without a key hash show no user, and a database failure in the owner lookup keeps the tools listed with user null
Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>
* fix(tool-policies): bound the owner lookup with chunked membership queries
The key-by-token and user-by-id lookups behind the tool rows' user field
put every distinct key hash into one IN list. BaseRepository gains
find_many_in, which runs the repository's chunked membership query and
converts the rows like find_many does, and the owner lookup uses it
Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>
* test(tool-policies): cover the owner column across the tool routes and the dashboard
Integration cells for the direct, detail and filtered tool routes, owners without alias or email, deleted owners and keys, keyless and unknown-key historical rows, more keys than one membership chunk, repeated reads, two-worker reads during discovery and a failed owner lookup. A Playwright cell drives the bundled Tool Policies page against the live proxy and follows the owner link
Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>
---------
Co-authored-by: ryan <ryan@berri.ai>
Co-authored-by: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>