Commit graph

18941 commits

Author SHA1 Message Date
oliver
560e5da320 fix(proxy): archive the old key only after regenerate validation passes
Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>
2026-09-11 08:02:18 +00:00
yucheng
dfbca29e5e Merge remote-tracking branch 'origin/litellm_internal_staging' into litellm_agent365_mcp_guardrail 2026-09-11 07:55:30 +00:00
yucheng
df739bdd9e fix(mcp): hand pre-call hooks the metadata of the registered tool that actually runs
Building the guardrail's tool metadata from the local registry entry that
dispatch resolved, instead of re-deriving it from the tool name, keeps an
OpenAPI operation whose name starts with its own server prefix from being
reported with the shorter operation's description and schema. The registry
branch in get_listed_tool is gone with it, and the test doubles for the
local registry now carry a string description and dict schema like the real
entries do

Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>
2026-09-11 07:54:36 +00:00
yucheng
82902e83c2 fix(proxy): write failed-login counters and their expiry in one Redis call
Use RedisCache.async_increment_with_floor (a single Lua INCRBY + EXPIRE) for the
shared login counters instead of the two-step INCRBYFLOAT then EXPIRE, so a
counter can never be committed to Redis without its expiry. The repair in
_remaining_window now only covers expiries stripped out of band (PERSIST, a
restore) and uses the same atomic call

Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>
2026-09-11 07:54:19 +00:00
oliver
e205be80df fix(proxy): enforce custom_key_update policy on /key/regenerate
Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>
2026-09-11 07:41:13 +00:00
joshua-berri
9a715df212
Merge pull request #40665 from BerriAI/litellm_fix_openapi_mcp_health_4896
fix(mcp): check OpenAPI specifications without native MCP handshakes
2026-09-10 22:02:45 -07:00
Kerry Lu
33ec56ed75 test(e2e): rewrite the Redis timeout test as a locust chaos load test
The sequential version sent one request at a time, so a Redis outage never
reached the concurrency where the failed-tracking alert body actually grows.
This drives the proxy with locust against one model group of three mock
deployments, two failing at order 1 and one serving at order 2, so every
request spends its retries on the failing pair and lands on the serving
deployment through the order-based fallback. Two phases, a healthy baseline
and a CLIENT PAUSE WRITE window, and every request must succeed in both.

Latency, RSS and CPU are reported as p50/p90/p99 per phase rather than
asserted on: RSS and CPU come from psutil on the proxy's process tree, since
a multi-worker proxy serves /metrics from the prometheus multiprocess
collector and that drops the process collector's series. Thresholds stay open
until weekly runs give real baselines.

Co-Authored-By: Claude Code <noreply@anthropic.com>
2026-09-10 21:55:23 -07:00
yucheng
6c865d1413 fix(mcp): resolve OpenAPI tool metadata from the local registry before any tools/list
Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>
2026-09-11 04:34:56 +00:00
Joshua Valluru
f04fb748c5 fix(mcp): explain refused OAuth registration and bound discovery retries 2026-09-10 21:02:46 -07:00
yucheng
0f7105f25e fix(mcp): hand listed tool metadata to pre-call hooks on the local registry path
OpenAPI-generated and legacy local-registry tools dispatch through execute_mcp_tool's
local branch, which called pre_call_tool_check without the cached MCPTool. Agent 365
therefore received bare {"name"} payloads for those tools while managed-server tools
carried description and inputSchema. Both local call sites now pass get_listed_tool

Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>
2026-09-11 03:43:33 +00:00
Joshua Valluru
ca03c889c9 fix(mcp): avoid caching cancelled OpenAPI health probes 2026-09-10 20:26:53 -07:00
mateo-berri
fea4970159 Merge remote-tracking branch 'origin/litellm_internal_staging' into litellm_redis_pool_timeout_counts_as_timeout 2026-09-10 20:23:25 -07:00
joshua-berri
acb9086f29
Merge pull request #40664 from BerriAI/litellm_fix_mcp_vscode_dcr_7449
fix(mcp): accept VS Code OAuth registration callbacks
2026-09-10 20:17:46 -07:00
yucheng
886220375f fix(guardrails): agent 365 sign-in for scopeless servers and gateway credential errors
Scopeless Agent 365 gated servers now advertise api://<client_id>/access_as_user instead of
staying silent, so a client can still sign in. Entra rejecting the gateway's own credentials
(invalid_client, unauthorized_client, invalid_scope, invalid_resource) follows unreachable_fallback
rather than telling the caller to sign in again with a 401

Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>
2026-09-11 03:15:59 +00:00
mateo-berri
54e247998e fix(redis): count pool wait timeouts as breaker timeouts
redis-py's blocking pool reports a saturated pool as ConnectionError chained
from asyncio.TimeoutError. The circuit breaker classified that as a hard
connectivity failure and opened at once while Redis was healthy. Follow the
explicit cause chain so it counts as a timeout and stays behind the
timeout_min_duration gate
2026-09-10 20:14:50 -07:00
Joshua Valluru
da1dfcdb24 refactor(mcp): reuse the shared HTTP handler for bounded probes 2026-09-10 19:59:07 -07:00
yucheng
803f8a69f0 test(mcp): assert the no-challenge outcome and register openapi tools through the registry api
Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>
2026-09-11 02:56:34 +00:00
yucheng
02233a2df3 refactor(mcp): build agent 365 protected resource metadata immutably to satisfy the type discipline gate
Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>
2026-09-11 02:49:13 +00:00
mateo-berri
77a0053a20 test(e2e): require the same workers at both memory checkpoints
Each checkpoint now samples until no new worker has answered for the settle
window, and the growth assertion refuses a worker set that changed between the
warm and after checkpoints instead of comparing only the intersection, so a
leaking worker reached by one checkpoint alone cannot drop out of the gate
2026-09-10 19:47:11 -07:00
Joshua Valluru
576c1bc5d6 fix(mcp): bound and coalesce OpenAPI health probes 2026-09-10 19:45:44 -07:00
mateo-berri
fdd423128e Merge remote-tracking branch 'origin/litellm_internal_staging' into litellm_redis_breaker_open_silent_miss 2026-09-10 19:40:52 -07:00
mateo-berri
f384acb840 Merge remote-tracking branch 'origin/litellm_internal_staging' into litellm_e2e_memory_regression_failing_requests 2026-09-10 19:40:11 -07:00
Joshua Valluru
5735587133 chore(ui): sync API descriptions with the current default branch 2026-09-10 19:30:10 -07:00
Joshua Valluru
fc95d22367 fix(mcp): accept VS Code OAuth registration callbacks 2026-09-10 19:28:40 -07:00
yucheng
d5b8effa99 fix(mcp): skip prefix lookup when a server has no listed tools
Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>
2026-09-11 02:27:19 +00:00
mateo-berri
2fc520329f fix(router): keep the budget push off the request callback path
The provider budget push runs inside the request success callback, so
awaiting the Redis pipeline there made every request wait for the round
trip. Hand it back to a task whose failure is logged through the breaker
aware logger, so an open breaker stays a debug line and a real Redis error
is one error line instead of an unretrieved task traceback
2026-09-10 19:26:28 -07:00
mateo-berri
eca7bb11ce fix(proxy): read RSS from /proc when psutil is missing so the release image reports memory
The release image installs only the proxy extras, and psutil is a locust and mirakuru dev
dependency, so /debug/memory/summary answered with an error and no ram_usage_mb on the e2e
gate. Fall back to /proc/self/statm and /proc/meminfo on Linux when psutil cannot be imported
2026-09-10 19:25:55 -07:00
Hayden Moulds
5b9244105c
test(proxy): use deployment listing metadata in alias coverage 2026-09-11 12:23:48 +10:00
Hayden Moulds
e18d766f53
test(proxy): consolidate team alias metadata coverage 2026-09-11 12:22:07 +10:00
Hayden Moulds
e664500003
test(proxy): cover team alias retrieve metadata 2026-09-11 12:21:57 +10:00
Hayden Moulds
49809a814b
fix(proxy): preserve metadata for public team aliases 2026-09-11 12:21:57 +10:00
mateo-berri
0c84f656b5 Merge remote-tracking branch 'origin/litellm_internal_staging' into litellm_ocr_custom_pricing
Some checks failed
LiteLLM Rust / rust-lint (push) Has been cancelled
LiteLLM Rust / rust-test (push) Has been cancelled
Terraform Modules / fmt, validate, test (aws) (push) Has been cancelled
Terraform Modules / fmt, validate, test (gcp) (push) Has been cancelled
Terraform Provider / gofmt, vet, build, test (push) Has been cancelled
Terraform Provider / Provider endpoints vs proxy OpenAPI schema (push) Has been cancelled
2026-09-10 19:18:43 -07:00
tin-berri
7419a536ad
fix(auto-router): omit Claude Code system text from classifier (#40655)
Co-authored-by: Claude Code <noreply@anthropic.com>
2026-09-10 19:18:25 -07:00
Joshua Valluru
8c82c325ac fix(mcp): check OpenAPI specifications without native MCP handshakes 2026-09-10 19:18:19 -07:00
mateo-berri
0ffe6512de fix(router): keep the routing and budget sync loops quiet while the Redis breaker is open 2026-09-10 19:15:40 -07:00
yucheng
e9c654869f Merge remote-tracking branch 'origin/litellm_internal_staging' into litellm_agent365_mcp_guardrail 2026-09-11 02:14:35 +00:00
yucheng
faa430c6c1 feat(mcp): challenge Agent 365 gated MCP servers with the Entra RFC 9728 metadata
When an Agent 365 guardrail applies to an MCP server that advertises scopes and no bearer arrives,
reuse the MCP OBO raise_token_exchange_challenge so the 401 and WWW-Authenticate header leave at the
transport layer. The protected-resource metadata for that server names the guardrail's Entra v2
issuer and the server's scopes, so Claude Code and other MCP clients run browser SSO and attach the
bearer themselves instead of the user pasting a token into the client config.

Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>
2026-09-11 02:08:09 +00:00
yucheng
35a1d017dc feat(guardrails): send MCP tool metadata to Agent 365 and treat unevaluated Defender verdicts as unavailable
Carry the listed tool's description and inputSchema from MCPServerManager through the pre-call and
during-call hook request objects into the Agent 365 evaluate payload, omitting them when the tool was
never listed. An allowed verdict whose defender.status is not Evaluated (Skipped, FailedOpen, missing)
now follows the unreachable_fallback policy instead of counting as a scanned allow. The conversationId
prefers the proxy-owned litellm_call_id over caller-controlled mcp-session-id headers.

Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>
2026-09-11 02:08:01 +00:00
mateo-berri
5fdb0860ec fix(helm): route /debug/memory/summary to the gateway so the memory gate reads the serving workers
On the release gate the e2e tests only see the nginx router, and the chart's
ingress sent /debug/memory/summary to the backend catch-all, so the RSS check
measured the backend pod instead of the gateway workers that serve the failing
requests. Render it as an Exact gateway path next to /test, name the host in the
summary response so workers behind one origin never collide on pid alone, and
key the harness readings by (origin, hostname, pid)
2026-09-10 19:02:37 -07:00
mateo-berri
01c6b50564 fix(caching): let a Redis breaker success count only for the state that admitted the call 2026-09-10 18:57:24 -07:00
Jon Walton
f358ebbc08
Merge remote-tracking branch 'origin/litellm_internal_staging' into litellm_user_budget_webhook_alerts 2026-09-11 09:49:16 +08:00
ryan
a9bd86b371 feat(ui): search, sort and role filter for the team member table
Rebuild the shared member table on DataTable so admins can search members by name, email or user id, sort by name, email, role, budget and spend, and filter by role. /team/info now returns each member's user_alias so the table can show a human-readable name

Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>
2026-09-11 01:37:48 +00:00
devin-ai-integration[bot]
880ccc76a5
fix(streaming): keep admitted mock streams alive with empty stream_options and honor zero prompt counts (#40650)
* fix(streaming): keep usage-only chunks from crashing streams with empty stream_options

The usage-only chunk branch in CustomStreamWrapper.chunk_creator indexed stream_options["include_usage"] directly, so a caller passing stream_options={} hit a KeyError that surfaced as MidStreamFallbackError. Streaming mock_response with an admission input_tokens count (#40637) now always emits such a chunk, which made the crash reachable. Reuse the send_stream_usage policy computed at init instead. Also annotate the #40637 test bindings with Final.

Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>

* fix(streaming): report admitted zero prompt tokens instead of recounting in mock streams

Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>

---------

Co-authored-by: yassin <yassin@berri.ai>
Co-authored-by: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>
2026-09-10 18:37:39 -07:00
mateo-berri
dcdd884352 fix(caching): let only the recovery probe close a half-open Redis breaker
A call admitted before the breaker opened could finish while the breaker was
HALF_OPEN and close it before the designated probe reported, so Redis traffic
resumed on a stale answer. The admission now records whether the call is the
probe and only the probe's success closes a half-open breaker.

The cron job lock manager also logged an error every cycle the open breaker
refused its Redis call, one line per job per pod. That refusal is now a debug
line like every other guarded call, while real Redis errors still log at error
2026-09-10 18:36:19 -07:00
ryan-crabbe-berri
ed90ff4a39 fix(auth): refresh lite login session token grants from the live user and team rows
A lite login token carried a snapshot of the team's models, aliases and the
user's role taken at login, so team or role changes never reached that CLI
until the user logged in again.

Pull the user, membership and team row loading that the JWT path did inline in
JWTAuthManager.get_objects into a GrantResolver under auth/resolvers, and have
the session token branch of the auth builder resolve the same rows on every
request. A user removed from the team now gets 403, a deleted user 401, and a
demoted admin no longer takes the admin early return.
2026-09-10 18:33:49 -07:00
ryan
a0b55fe68d feat(ui): search Key Activity by key alias, key hash, user id, or email
Team Usage and the main Usage page render every key in the selected
scope with no way to narrow the list. Add a client-side search box
above Key Activity that filters the loaded keys by alias, hash, user
id, or user email, and expose user_id on the daily activity key
metadata so the id is searchable

Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>
2026-09-11 01:30:00 +00:00
yuneng-jiang
41b0deb627
Merge pull request #40643 from BerriAI/litellm_fix_optional_logging_assertions
test: respect optional logging payload fields
2026-09-10 18:27:51 -07:00
mateo-berri
7169ddaef6 fix(router): treat a breaker-refused Redis read as a miss in the health state cache
The sync Redis read now raises while the circuit breaker is open, and the
health state merge caught that as a generic error, skipping the local write
and logging an error on every background health check cycle. Read the shared
snapshot through a helper that treats the refused read as a miss so the merge
falls back to the pod-local copy the way a swallowed connection error already did
2026-09-10 18:16:59 -07:00
Mateo Wang
4fbe2276a1
fix(logging): finish response metadata before the sync logging thread reads it (#39869)
* fix(logging): finish response metadata before the sync logging thread reads it

The async and sync client wrappers handed the response to the threaded success handler before computing its cost, call id, and api_base, so that thread inserted into the same metadata dict the request coroutine was still iterating and a finished chat completion turned into a 500 (dictionary changed size during iteration). Metadata is now finalized first, and the merge and header copies snapshot their dicts before iterating.

* fix(logging): snapshot metadata with a dict copy and drop redundant comment

Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>

* fix(logging): copy metadata via dict.copy and dedupe Final import

Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>

---------

Co-authored-by: yassin <yassin@berri.ai>
Co-authored-by: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>
2026-09-10 18:15:24 -07:00
mateo-berri
e4f958d7cf fix(gateway): serve /debug/memory/summary on the data plane so the memory regression test can read each gateway's RSS 2026-09-10 18:12:40 -07:00