Commit graph

42397 commits

Author SHA1 Message Date
shivam
d8c7e25296 fix(batches): paginate managed batch list by unified_object_id cursor
GET /batches served from the managed-objects table paged with a
where id > after filter, but the after cursor clients send back is a
batch's unified_object_id (the value returned as .id and last_id), and
id is the table's random-uuid primary key. Comparing the two unrelated
fields, while ordering by created_at desc but filtering with gt, made
pages repeat the same last_id (pagination loops) and silently drop
batches. Switch to Prisma cursor pagination on the unique
unified_object_id column so listing walks every batch exactly once in
reverse-chronological order, matching OpenAI

Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>
2026-07-21 23:25:26 +00:00
yuneng-jiang
28e93e42e5
test(ui): run vitest unit tests in GitHub Actions and fix stale key-info tests (#34175)
* test(ui): run vitest unit tests in GitHub Actions and fix stale key-info tests

The dashboard's vitest suite only ran on CircleCI; GitHub Actions covered the
UI build, lint and api-types sync but never the unit tests. Add a UI Unit Tests
workflow that runs the suite, sharded across a matrix so the wall-clock is not
bound by a single 4-core runner.

Porting it surfaced 17 pre-existing failures. Adding the block/unblock key
action moved Delete Key and Reset Spend into a "More key actions" dropdown and
introduced a React Query hook; KeyInfoHeader's own test was updated but the two
KeyInfoView test files were not. Reach those actions through the dropdown and
stub the new hook the way the neighbouring hook is already stubbed.

The same refactor had quietly hollowed out assertions that still passed:
"should not show Reset Spend button for regular key owner" queried for a button
role that no longer exists, so it held green regardless of the permission
check. Those now open the menu and assert on the menu item, which fails when
canResetSpend is forced true.

Also add the missing cost-optimization page description; page_utils guards that
every navigable page carries one.

* ci(ui): scope PR runs to changed tests, run the full suite on staging

Running the whole vitest suite on every pull request costs about five minutes,
and none of it is recoverable through parallelism: vitest schedules by file and
create_mcp_server.test.tsx alone accounts for 252s of the 255s total, so shards
and extra cores cannot get under that floor. Measured on this branch, css:false,
pool=threads and isolate=false all landed within noise of the baseline.

Scope pull requests to tests reachable from the diff instead, which takes 11s
here, and keep a full run on pushes to litellm_internal_staging so nothing rots
behind a gap in the module graph. Backend-only pull requests match no test files
and exit zero; --passWithNoTests states that rather than leaning on it being the
current default. The checkout needs full history for --changed to resolve the
base commit.
2026-07-21 16:23:43 -07:00
Yassin Kortam
bc374fcd9f
test(e2e): add Azure AI Foundry and Anthropic /v1/messages coverage for the Rust bridge (#34021)
* test(e2e): add Azure AI Foundry and Anthropic /v1/messages coverage

* test(e2e): add Azure AI Foundry + Anthropic messages coverage for the Rust bridge

* test(e2e): guard against an empty SSE stream in the Azure Foundry tool-use streaming case
2026-07-21 16:22:47 -07:00
mateo-berri
48295df0e9 test(e2e): raise warm cache read floor to 0.65 from measured healthy and regression baselines 2026-07-21 16:21:45 -07:00
Yassin Kortam
8a56899e1e
test(e2e): cover config and misc management routes for Management/UI coverage (#34120) 2026-07-21 16:20:50 -07:00
ryan-crabbe-berri
e17f3b6e1a
fix(proxy): populate user_email on UserAPIKeyAuth for JWT auth (#34174)
Some checks are pending
CodSpeed Benchmarks / benchmarks (push) Waiting to run
GitHub Actions Security Analysis / zizmor (push) Waiting to run
JWT auth built UserAPIKeyAuth without user_email even though the resolved
user row and the JWT email claim were both available, so the user_email
label on Prometheus metrics and user_api_key_user_email in
StandardLogging/SpendLogs metadata were always None for JWT traffic.

Plumb user_email through JWTAuthBuilderResult: auth_builder returns the
user row email when set, falling back to the user_email_jwt_field claim
(covers the scope-based proxy-admin path where no user row is loaded).
The JWT branch now stamps it on the proxy-admin return, the standard
valid_token, and the auto-registered virtual key object.

Resolves LIT-4238
2026-07-21 16:11:47 -07:00
ryan-crabbe-berri
02746eb122
fix(ui): harden provider logo map typing and bundled asset guard (#34163)
Follow-up to the static logo import PR. Types providerLogoMap as
Partial<Record<Providers, string>> so raw string keys and lookups are
compile errors, tightens the resolveLogoSrc passthrough from /_next/ to
/_next/static/ so lookalike backend paths still get root-prefixed, adds
an enum coverage test that locks the exact set of logoless providers,
and makes Logo props a discriminated union so provider and src modes
cannot be mixed and src mode requires a label.
2026-07-21 16:11:38 -07:00
ryan-crabbe-berri
42f269ccf2
fix(ui): stop dashboard key-edit form 403ing on non-budget saves (#34112)
The key-edit form sent budget_limits on every save (the stored windows, or []
when a key has none). The backend treats any budget_limits in a /key/update
request as an admin-only budget change, so a non-admin key owner editing a
non-budget field (models, MCP servers, alias) always hit 403 with "Only proxy
admins, team admins, or org admins can call /key/update".

Only include budget_limits when the user actually changed the budget windows,
mirroring how the same handler already drops an unchanged allowed_routes. The
comparison is on (duration, cap) ignoring the server-owned reset_at and window
order; [] is still sent when the user deletes the last window so clearing keeps
working. No backend or API behavior changes.
2026-07-21 16:11:31 -07:00
mateo-berri
1692170264 fix(e2e): count aborted-session turns as failures and require a spend stability window 2026-07-21 15:34:56 -07:00
yassin
51f0f40c2f test(e2e): invoke a real published a2a agent and assert it replies
Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>
2026-07-21 22:33:07 +00:00
ryan-crabbe-berri
2b2ae4ca49
refactor(ui): migrate MCP, callback, guardrail, SSO, and search tool logos to the shared Logo component (#34169)
Some checks failed
LiteLLM Rust / rustfmt, clippy, test (push) Has been cancelled
* refactor(ui): migrate MCP, callback, guardrail, SSO, and search tool logos to the shared Logo component

Third step of the logo consolidation. Every remaining rogue logo
pattern now renders through Logo: MCP well-known grid and backend
mcp_info.logo_url sites (which previously skipped resolveLogoSrc and
broke under non-root mounts), callback maps in callback_info_helpers
plus the backend-provided variant in settings.tsx, the guardrail map
with the garden dataset now deriving logos from guardrailLogoMap
instead of duplicating them, the SSO map deduped from two verbatim
copies into SSOSettings/constants.ts, search tools' filename guessing
replaced with an explicit static-import map, and the two straggler
sites in EntityUsage and model_info_view.

Static-map path strings become bundled static imports throughout;
backend-provided URLs stay runtime strings resolved via Logo src mode.
MCPLogoSelector still stores stable /ui/assets/logos paths so existing
DB rows keep matching. okta's logo remains an external hotlink pending
a vendored local asset. promptguard.svg drops a mismatched intrinsic
dimension attribute for the Turbopack import parser.

* fix(ui): make resolveLogoSrc idempotent for values already carrying the server root path

Stored mcp_info.logo_url values from sub-path deployments could bake in
the deployment root because the old bare img sites did no resolution.
Prefixing those again produced /litellm/litellm/... and a fallback
avatar. Skip prefixing when the value already starts with the current
normalized root segment; paths whose first segment merely begins with
the root text still get prefixed.
2026-07-21 22:22:54 +00:00
Mateo Wang
9d714b0b29
Merge pull request #34172 from BerriAI/litellm_scrub_customer_name
chore(tests): replace a customer name and domain with neutral placeholders
2026-07-21 15:22:35 -07:00
yuneng-jiang
72d458e416
feat(ui): add react-hook-form + zod form infrastructure (#34170)
* feat(ui): add react-hook-form + zod form infrastructure

Introduce the shared form layer the dashboard's antd forms will migrate onto,
with no user-visible change yet.

- pin react-hook-form, @hookform/resolvers, and zod (kept on 3.25.76 and
  imported via the zod/v4 entrypoint so openai's optional zod ^3 peer still
  resolves and npm ci stays clean)
- vendor the base-vega Field family into components/shared/form as forwardRef
  components on the repo's cva.config, since base-vega ships no form primitive
  and its field source imports class-variance-authority and is React 19 style
- add a FormField bridge that binds a react-hook-form Controller to the Field
  layer and wires label, description, and error ids into aria attributes
- add pickDirty, which narrows a submitted body to the top-level keys the user
  actually touched so a partial update stops re-sending untouched fields

pickDirty reads dirtiness at the top level because react-hook-form tracks it
per leaf, so an edited array arrives as [true, false] and a cleared list as an
empty array that still carries its default-length dirty markers; the falsy
clear tokens (null, [], {}, 0, false) all survive.

Tests cover the Field primitives, the FormField aria wiring against a live
zod resolver, and pickDirty both as a unit and driven through a real
react-hook-form instance.

* test(ui): lock pickDirty behavior on a pure field-array reorder

react-hook-form compares each array element to its default positionally by
value, so useFieldArray move/swap and a reordered scalar array all mark the
moved indices dirty and pickDirty sends the whole array; a swap of two equal
elements is a value-level no-op and is correctly omitted. Covers the reorder
case a review flagged as untested.
2026-07-21 15:17:58 -07:00
yuneng-jiang
58ff0e32ba
chore(deps): bump gitpython to 3.1.52 in uv.lock (#34168) 2026-07-21 22:13:07 +00:00
tin-berri
869ef0cbfb
Merge pull request #34072 from BerriAI/litellm_mcp_ema_assertion_store
feat(mcp): store the enterprise IdP identity assertion at SSO login for EMA egress
2026-07-21 15:10:35 -07:00
mateo-berri
fa025fc474 chore(tests): replace a customer name and domain with neutral placeholders 2026-07-21 15:09:15 -07:00
yuneng-jiang
bc9534085b
feat(ui): add controlled row selection to the shared DataTable (#34167) 2026-07-21 22:08:08 +00:00
Yassin Kortam
c2ae52709c
feat(proxy): make DB config-reload interval configurable via config.yaml and UI (#34130)
The add_deployment and get_credentials background jobs that keep a multi-pod
deployment in sync with config-in-DB objects (models, credentials, guardrails,
general settings, etc.) polled the database on a hardcoded 30s interval, with
no way to trade convergence latency against DB load.

Expose it as the general_setting proxy_config_reload_interval_seconds (env
PROXY_CONFIG_RELOAD_INTERVAL_SECONDS parsed via get_env_int, default 30),
threaded like the existing proxy_batch_polling_interval knob, and surface it on
the admin general-settings page so it is reachable from the dashboard and
persists to the DB for all pods. Non-positive values are rejected at the UI
(gt=0) and fall back to 30s with a warning on the env/config/DB paths.
2026-07-21 22:03:22 +00:00
Yassin Kortam
090c4b8dd8
test(e2e): cover model, tag and access group routes for Management/UI coverage (#34118) 2026-07-21 21:59:34 +00:00
Tin Chi Lo
ddae6eac6b fix(mcp): judge EMA retention only against the config and DB authorities, never the registry snapshot 2026-07-21 14:54:37 -07:00
Tin Chi Lo
cfa074d970 fix(mcp): make the EMA retention gate authoritative across pods and tighten the assertion store 2026-07-21 14:54:37 -07:00
Tin Chi Lo
83e147b553 feat(mcp): store the enterprise IdP identity assertion at SSO login for EMA egress 2026-07-21 14:54:37 -07:00
mateo-berri
b572cb80d1 test(e2e): add weekly session-anomaly load test against real providers 2026-07-21 14:54:02 -07:00
mateo-berri
43402b04bd test(e2e): register the verbatim a2a-sdk 0.3.x card and guard malformed protocolVersion 2026-07-21 14:43:09 -07:00
Yassin Kortam
48fdaaa5cd
fix(router): honor request-level num_retries over global litellm_settings.num_retries (#34124)
The async @client wrapper stamped the global litellm.num_retries onto the raised
exception via setattr(e, "num_retries", ...), even on router calls where the
request-level num_retries had already been popped and resolved. async_function_with_retries
then adopted that stamped global value, overwriting the request-level num_retries it had
correctly resolved. So a per-request num_retries (request body or x-litellm-num-retries
header) was silently ignored whenever litellm_settings.num_retries was set.

Only stamp num_retries on the exception when the call itself carried one (an explicit
request value or a deployment's litellm_params.num_retries), never the global fallback.
The router already resolves the global via self.num_retries, so leaving the exception
unset preserves the request-level value and lets the per-deployment path set it when present.

Resolves LIT-4516
2026-07-21 21:32:34 +00:00
Yassin Kortam
7f258264c7
test(e2e): cover budget, customer, user and org routes for Management/UI coverage (#34117) 2026-07-21 21:29:53 +00:00
Yassin Kortam
b3415f426c
test(e2e): cover key management routes for Management/UI coverage (#34114) 2026-07-21 21:29:21 +00:00
mateo-berri
c8e0253cd5 Merge remote-tracking branch 'origin/litellm_internal_staging' into litellm_a2a_e2e_tests 2026-07-21 14:29:15 -07:00
Yassin Kortam
6eb3cf2c50
test(e2e): cover team management routes for Management/UI coverage (#34115) 2026-07-21 21:29:13 +00:00
ryan-crabbe-berri
e1afc64baa
refactor(ui): migrate inline provider logo lookups to the shared Logo component (#34141)
* fix(ui): bundle provider logos as static imports and unify fallback in Logo component

providerLogoMap values are now content-hashed bundle URLs emitted by
static imports instead of /ui/assets/logos/ path strings, so any
deployment that serves the app JS also serves the logos: dev server,
proxy /ui mount, server_root_path sub-paths, and the split-chart nginx
image where the old route 404d in production. A missing file is now a
build error instead of a silent runtime 404.

resolveLogoSrc passes /_next/ URLs through untouched so bundled values
never get double-prefixed with the server root path. The new Logo
molecule owns resolution and the letter-avatar fallback and warns with
the failing URL on load error; ProviderLogo delegates to it. The three
bare img sites in the agents wizard render through Logo, fixing their
broken-image bug.

Dashscope now uses qwen.png, RunwayML the on-disk runway.png, and the
GradientAI entry is removed (no plausible asset exists). soniox.svg and
ai21.svg drop a single mismatched intrinsic dimension attribute that
Turbopack's import-time image parser rejects. Dead logoSrc lookup in
AddModelForm deleted. Vitest resolves image imports to Next's
StaticImageData shape via a config plugin so tests exercise the same
/_next/ URLs as production.

* fix(ui): retry logo load when src changes after an error

Track which src errored instead of a boolean so a Logo instance whose
source changes in place (agents modal title) attempts the new URL
rather than staying on the letter-avatar until remount.

* refactor(ui): migrate inline provider logo lookups to the shared Logo component

Patterns B, C, and E from the logo consolidation: every inline
providerLogoMap lookup feeding a bare img with a hand-rolled DOM
fallback now renders through Logo (credential modal, vector store
create/info views, cost tracking margin and discount forms and tables).

getProviderDisplayInfo, handleImageError, and ProviderDisplayInfo are
deleted; getProviderLogoAndName is a strict superset of the exact-match
helper. The vector store logo map no longer duplicates provider logo
paths: shared entries reference providerLogoMap and the three
vector-store-only logos become static imports. The map itself stays
because milvus and s3_vectors have no Providers enum equivalent.

Sites that rendered nothing for an unmapped provider now render the
letter avatar. Representative tests per pattern assert the rendered img
src against providerLogoMap so a wrong provider-to-enum mapping fails,
plus letter-avatar fallbacks for unmapped providers.

* fix(ui): resolve vector store slugs through the vector store logo map

The vector store info provider badge fed backend slugs like pg_vector,
milvus, and s3_vectors to getProviderLogoAndName, which only knows LLM
providers, so those stores showed a letter avatar and a raw slug. The
pre-existing inline lookup had the same wrong-domain bug via
provider_map. Reinstate getVectorStoreProviderLogoAndName resolving
through vectorStoreProviderMap first with a fallback to the LLM
resolver, so vector-store-only providers get their own logo and display
name for the first time.
2026-07-21 14:28:05 -07:00
ryan-crabbe-berri
2e8403a073
fix(ui): bundle provider logos as static imports and unify fallback in Logo component (#34125)
* fix(ui): bundle provider logos as static imports and unify fallback in Logo component

providerLogoMap values are now content-hashed bundle URLs emitted by
static imports instead of /ui/assets/logos/ path strings, so any
deployment that serves the app JS also serves the logos: dev server,
proxy /ui mount, server_root_path sub-paths, and the split-chart nginx
image where the old route 404d in production. A missing file is now a
build error instead of a silent runtime 404.

resolveLogoSrc passes /_next/ URLs through untouched so bundled values
never get double-prefixed with the server root path. The new Logo
molecule owns resolution and the letter-avatar fallback and warns with
the failing URL on load error; ProviderLogo delegates to it. The three
bare img sites in the agents wizard render through Logo, fixing their
broken-image bug.

Dashscope now uses qwen.png, RunwayML the on-disk runway.png, and the
GradientAI entry is removed (no plausible asset exists). soniox.svg and
ai21.svg drop a single mismatched intrinsic dimension attribute that
Turbopack's import-time image parser rejects. Dead logoSrc lookup in
AddModelForm deleted. Vitest resolves image imports to Next's
StaticImageData shape via a config plugin so tests exercise the same
/_next/ URLs as production.

* fix(ui): retry logo load when src changes after an error

Track which src errored instead of a boolean so a Logo instance whose
source changes in place (agents modal title) attempts the new URL
rather than staying on the letter-avatar until remount.
2026-07-21 14:27:36 -07:00
ryan-crabbe-berri
1cc70f84c0
fix(ui): distinguish response cache from provider prompt caching (#34138)
* fix(ui): distinguish response cache from provider prompt caching

The log detail drawer labeled LiteLLM's response cache result as
"Cache Hit" and rendered a red "false" tag next to provider prompt
cache token counts, which read as prompt caching being broken. The
row is now labeled "Response Cache" with an explanatory tooltip,
shows a neutral "Miss" tag instead of a red one, and the prompt
cache token rows are prefixed with "Prompt Cache" and get their own
tooltips. Cost breakdown line items get the same prefix.

The Caching dashboard only reports response cache analytics but never
said so; it is renamed to "Response Cache" in the sidebar, gains a
scope description pointing to the Usage page and Logs for prompt
caching, and the ambiguous "Cached Tokens" stat card is renamed to
"Cached Completion Tokens".

* test(ui): cover renamed Response Cache sidebar item in e2e

The sidebar navigation spec now clicks the renamed "Response Cache"
item and asserts it routes to /ui/caching, and the menu label fixture
maps the new label while keeping "Caching" as a legacy alias.
Verified by running sidebar.spec.ts through run_e2e.sh (full harness:
built UI served by the proxy, seeded postgres); both tests pass.

* feat(ui): link cache tooltips and dashboard description to docs

The Response Cache tooltip links to the proxy caching docs and the
two prompt cache token tooltips link to the prompt caching docs, so
users can jump straight to the explanation of whichever mechanism
they are looking at. The Response Cache dashboard description links
both docs pages the same way.
2026-07-21 14:26:26 -07:00
ryan-crabbe-berri
b47fe730a4
fix(ui): stop cloning body-carrying requests into stream uploads in fetchClient middleware (#34122)
* fix(ui): stop cloning body-carrying requests into stream uploads in fetchClient middleware

The openapi-fetch middleware rebuilt every outgoing request with
new Request(url, request), which converts a string JSON body into a
ReadableStream with duplex=half. Chromium only allows streaming uploads
over HTTP/2 or HTTP/3, so against any HTTP/1.1 hop (uvicorn serves
HTTP/1.1 only) the fetch dies at the network layer with
net::ERR_ALPN_NEGOTIATION_FAILED, surfaced as "Failed to fetch".

GET callers were unaffected (null body); the first body-carrying caller
arrived with the MCP BYOK credential modal, breaking that flow on plain
http deployments in the v1.94.0 RCs.

The middleware now mutates headers on the original request when no
runtime base is registered, and when rebasing onto a runtime base it
rebuilds the request with the body materialized as bytes via
arrayBuffer(), which fetch sends with Content-Length instead of a
streaming upload

* Update ui/litellm-dashboard/src/lib/http/api.ts

Co-authored-by: greptile-apps[bot] <165735046+greptile-apps[bot]@users.noreply.github.com>

---------

Co-authored-by: greptile-apps[bot] <165735046+greptile-apps[bot]@users.noreply.github.com>
2026-07-21 14:26:02 -07:00
Yassin Kortam
e9ac84dc8b
fix(proxy): stop save_config from snapshotting environment_variables into the DB (#34119)
save_config wrote the entire merged config to the DB config table on every call. Because get_config() resolves os.environ/ placeholders to plaintext and merges the environment_variables section, any endpoint that does get_config() then mutates one section then calls save_config() (/add/allowed_ip, delete_callback, model and cost-tracking settings, and others) accidentally persisted an environment_variables row holding all YAML/OS-sourced env vars. Once that row existed the DB overlay shadowed YAML and container env on every subsequent startup, so config/env changes were silently ignored

save_config now pops environment_variables from the DB write unless the caller passes include_env_vars=True. The dedicated /config/update path already writes env vars per-section via _upsert_section, so no current caller needs to opt in

Resolves LIT-2009
2026-07-21 21:23:28 +00:00
ryan-crabbe-berri
f9300d4312
fix(e2e): pin harness Python and surface proxy boot crash output (#34157)
run_e2e.sh let uv resolve any system Python; on machines where that is
3.14 the locked uvloop 0.21.0 fails to import (BaseDefaultEventLoopPolicy
was removed from asyncio.events) and the proxy dies at boot, which the
harness reported only as a misleading 180s health timeout. Pin the
interpreter to 3.13 (overridable via UV_PYTHON) and redirect proxy output
to a log file whose tail is printed when the proxy exits early or never
becomes healthy.
2026-07-21 21:17:43 +00:00
Yassin Kortam
fcaf673418
fix(router): stop per-deployment num_retries from double-counting as provider max_retries (#34129)
* fix(router): stop per-deployment num_retries from double-counting as provider max_retries

A model group with one deployment and num_retries set in the deployment's
litellm_params sent (1 + num_retries) ** 2 requests upstream instead of
1 + num_retries. The deployment's num_retries reached litellm.completion, which
copied it onto max_retries and set it on the provider client, so the provider SDK
retried num_retries times inside each of the Router's 1 + num_retries attempts.

The Router is the sole retry owner for routed calls, so completion() now forces the
provider-SDK max_retries to 0 whenever the call originates from the Router/proxy
(detected via model_group in the request metadata) and only keeps the num_retries
to max_retries alias for direct, non-routed litellm calls (the instructor use case).
This also stops a request- or deployment-level max_retries from nesting on top of
the Router's retries.

Resolves LIT-4385

* fix(router): make router-origin check robust and close test clients

Address review: detect the router marker in both metadata and litellm_metadata
independently (a non-empty metadata without model_group no longer hides a
model_group in litellm_metadata), and close the injected async clients in the
test fixture.
2026-07-21 21:13:32 +00:00
tin-berri
914079a68c
Merge pull request #34065 from BerriAI/litellm_lit4629_issuer_discovery
fix(mcp): let an admin-pinned issuer drive OAuth discovery for url-less servers
2026-07-21 14:06:55 -07:00
mubashir1osmani
ac5b51253a
test(e2e): add Other suite and Guardrails coverage incl. an MCP tool-call guardrail (#34149)
* test(e2e): add other suite covering master-key auth and health lifecycle

Covers the other.* holding-pen cells that were uncovered: master-key
valid_allows/invalid_denied on the admin /user/list gate, and the
lifecycle probes liveness.ping, readiness.public_probe,
readiness.reports_db_status, and readiness_details.authenticated_diagnostics.

New tests/e2e/other/ suite on the shared ProxyClient; the health probes
send no auth header to prove the public routes need no credential, and the
details route is asserted to reject an anonymous caller while exposing
version/db diagnostics to the master key.

* test(e2e): cover block_code_execution and openai_moderation guardrails

Extends the guardrails suite with two built-in guardrails registered per
request (default_on=False, opted in via the chat body's guardrails selector)
so neither intercepts unrelated traffic on the shared proxy.

block_code_execution.pre_call.blocks: a python code block plus a run-this
request is intercepted with the canned content-blocked message and the model
never runs, while the same code block asked about with don't-run-it reaches
the model. Verified live.

openai_moderations.pre_call.blocks: a flagged prompt is rejected 400 naming
the moderation policy while a benign prompt passes. The guardrail calls
OpenAI's moderation API; verifying it needs an OpenAI key with moderation
quota (this account currently 429s the moderation endpoint).

Adds a shared create_backend_model helper and a generic register() plus
per-request guardrails/max_tokens on the client so more built-ins can reuse
the same path.

* test(e2e): cover presidio PII masking (pre_call + post_call)

Registers a presidio guardrail per request (default_on=False) with the
analyzer/anonymizer bases supplied in the registration params, so the test
controls its own dependency and needs no proxy restart.

presidio.pre_call.masks: a repeat-verbatim request comes back with the
<EMAIL_ADDRESS> placeholder and never the raw email, proving the prompt was
anonymized before the model saw it.

presidio.post_call.masks: with apply_to_output the model's own emitted email
is masked on the way out, so the caller never receives the raw value.

Both verified live against real presidio analyzer + anonymizer containers.
logging_only is intentionally not covered: /spend/logs exposes no prompt
messages to read back the masked log, and a logging_only run also masked the
response, contradicting its contract; noted in the module docstring for a
follow-up.

* test(e2e): cover presidio logging_only masking via OTEL read-back

Adds the third presidio cell, guardrail.presidio.logging_only.masks. The
logging_only contract (mask what is logged, do not block) is verified by
reading the request's gen-AI span back from the real OTEL destination: the
span's gen_ai.input.messages attribute carries the <EMAIL_ADDRESS> placeholder,
never the raw email, and the call itself is not blocked.

Reads the trace via the shared OtelReader, promoted from logging/ to the suite
root so both suites use it. The masked prompt is polled to a deadline because
logging_only masks the payload asynchronously and the span can briefly export
before the mask lands. Drops the throwaway chat_send in favor of the existing
transport.send for the call-id capture.

* fix(e2e): tolerate cross-pod guardrail sync delay in team-opt-out test

Stage runs multiple gateway pods behind the shared key. POST /guardrails
registers a new default-on guardrail in-process immediately only on the
pod that served the create call; every other pod picks it up on its next
periodic DB sync (proxy_server.py, every 30s), so the very next chat call
can race a pod that has not synced yet. Poll to a 40s deadline instead of
asserting on the first response, matching the existing pattern in
test_budget_reset_advances_e2e.py.

* test(e2e): cover a guardrail on the MCP tool-call path (content_filter pre_mcp_call)

Adds guardrail.litellm_content_filter.pre_mcp_call.blocks: against the real
Datadog MCP server, a content_filter guardrail configured mode=pre_mcp_call
blocks a banned keyword in an MCP tool call's arguments with HTTP 400 attributed
to the pre_mcp_call hook, and lets a clean argument reach the upstream server.

The guardrail attaches with default_on because per-key/request guardrail
selection is dropped from the synthetic MCP request the hook sees; the banned
keyword is unique per run so default_on only intercepts this test's own call.
mode must be pre_mcp_call - a pre_call config silently no-ops on tools/call
because the event type is rewritten for call_mcp_tool.

Drives the tool directly via /mcp-rest/tools/call for a deterministic check of
the same pre_mcp_call enforcement the OpenAI-SDK chat path hits when a model
invokes an MCP tool.

* fix(e2e): mid-conversation messages test uses client.proxy not client.gateway

EndpointsClient exposes .proxy after the Gateway->ProxyClient rename; the
mid-conversation system test still referenced .gateway, which fails the e2e
basedpyright gate. Aligns it with the rest of the harness.

* test(e2e): address review on the guardrail coverage

MCP tool-call guardrail: poll the banned call until the guardrail is enforced
instead of asserting on the first call, so the control-plane -> data-plane
guardrail sync cannot race the check into a false pass-through; add a repeat
banned call after enforcement to guard against a partial-propagation state.

OpenAI moderation: distinguish a moderation-endpoint 429 (rate limit / no
moderation quota) from a guardrail failure, so an account-capability gap reads
as such rather than as "did not block". Runs green with a moderation-capable key.

* test(e2e): close partial-propagation false-pass in MCP guardrail block test

The single post-block repeat call could be load-balanced back to the same
already-synced data-plane pod, so the test could pass while another pod still
lacked the guardrail and let the banned MCP call reach Datadog. Anchor a wait to
the guardrail create time (every pod is guaranteed to have DB-synced only after a
full ~30s sync interval), then require the banned call to stay blocked across
several attempts; a pass-through after that window is a real leak, not a race.

* test(e2e): drop xfail-style rate-limit branch from openai_moderation test

OpenAI's /v1/moderations is free and returns 200 with the env key (verified
directly), so the RateLimitedError branch mislabeled the failure: a 429 there is
insufficient_quota (no account billing), not throttling. The branch also only
printed a softer message before failing anyway, an xfail-in-disguise the e2e rules
forbid. A 429 now falls through and fails loudly with the full result.
2026-07-21 14:06:29 -07:00
ryan-crabbe-berri
76c9eca25d
refactor(auth): derive temp budget increase without mutating the token (#34121)
* refactor(auth): derive temp budget bump without mutation, tz-aware auth datetimes

_update_key_budget_with_temp_budget_increase mutated max_budget in place, so correctness depended on every resolution path handing it a fresh copy of the cached token; one future re-cache of a live token would compound the bump per request. Return a model_copy instead so no caller can leak an increased budget into shared state.

Also fixes the three remaining DTZ005 naive datetime.now() calls in user_api_key_auth.py (auth span start, builder start_time, service-log end_time; all consumers convert to epoch or subtract same-pair datetimes) and ratchets the DTZ005 strict budget 244 -> 241.

* test: pin non-mutation of the temp budget helper input

Adversarial mutation-testing showed reverting the helper to in-place mutation still passed every test: the cache's copy-on-read layer masks the mutation in the integration test and the direct unit test only inspected the return value. Assert the input object is left untouched and the result is a distinct object so the purity guarantee itself is load-bearing.
2026-07-21 21:02:40 +00:00
Yassin Kortam
30ed840ff5
fix(e2e): use EndpointsClient.proxy after Gateway to ProxyClient rename (#34127)
The mid-conversation-system messages test still referenced the removed
client.gateway attribute, so the tests/e2e basedpyright gate reported 9
errors and went red on every e2e PR. EndpointsClient exposes .proxy, so
point the /v1/messages post helper at client.proxy.transport.
2026-07-21 20:53:40 +00:00
Tin Chi Lo
6959d9de69 test(mcp): pin the token and register wall messages for url-less servers 2026-07-21 13:53:14 -07:00
Tin Chi Lo
930676a9bf fix(mcp): let an admin-pinned issuer drive OAuth discovery for url-less servers 2026-07-21 13:53:14 -07:00
tin-berri
1d36290551
Merge pull request #34063 from BerriAI/litellm_lit4629_openapi_oauth_egress
fix(mcp): attach resolved OAuth credentials to OpenAPI spec_path tool calls
2026-07-21 13:51:17 -07:00
ryan-crabbe-berri
d2819baf0a
feat(ui): add block/unblock key action to key info page (#34116)
Adds a Block Key / Unblock Key action to the key info page, wired to the
existing /key/block and /key/unblock endpoints which previously had no UI.
The Reset Spend and Delete Key buttons move together with it into a new
overflow dropdown next to Regenerate Key, and a red Blocked tag shows next
to the key alias while the key is blocked.
2026-07-21 13:41:10 -07:00
Mateo Wang
a7e4c83013
Merge pull request #34154 from BerriAI/litellm_a2a_protocol_version_semver
fix(a2a): accept semver protocolVersion values like 0.3.0 in agent cards
2026-07-21 13:37:05 -07:00
ryan-crabbe-berri
ee0028a841
feat(ui): surface key budget_reset_at in key info and keys table (#34113)
* feat(budgets): add configurable budget_reset_time of day

Budgets reset at midnight in the configured timezone with no way to control
the time of day, so a drained daily budget surfaces as an overnight incident.
Add a litellm_settings.budget_reset_time option (e.g. "12:00") that shifts
day/week/month resets to a configurable wall-clock time in the existing
timezone, so the end of the budget window lands during business hours.

The reset time is parsed once into an immutable BudgetResetSettings and
injected into the reset job (constructor) and computation, rather than read
from a module-level global at call time. A malformed value fails fast at
startup. Sub-day durations ignore the offset. Unset preserves midnight resets.

* feat(ui): surface key budget_reset_at in key info and keys table
2026-07-21 13:35:14 -07:00
ryan-crabbe-berri
efa997dfe0
feat(budgets): add configurable budget_reset_time of day (#31007)
Budgets reset at midnight in the configured timezone with no way to control
the time of day, so a drained daily budget surfaces as an overnight incident.
Add a litellm_settings.budget_reset_time option (e.g. "12:00") that shifts
day/week/month resets to a configurable wall-clock time in the existing
timezone, so the end of the budget window lands during business hours.

The reset time is parsed once into an immutable BudgetResetSettings and
injected into the reset job (constructor) and computation, rather than read
from a module-level global at call time. A malformed value fails fast at
startup. Sub-day durations ignore the offset. Unset preserves midnight resets.
2026-07-21 13:35:01 -07:00
mateo-berri
2310211531 fix(a2a): reject malformed protocolVersion suffixes while keeping semver prereleases 2026-07-21 13:24:58 -07:00
yassin
1315ebd1f9 test(e2e): guard 0.3.0-style semver protocolVersion registration
Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>
2026-07-21 20:14:21 +00:00
mateo-berri
062e58fb1d fix(a2a): accept semver protocolVersion values like 0.3.0 in agent cards 2026-07-21 13:03:47 -07:00