CI code-quality check flags recursive functions. Split into two
non-recursive helpers: _strip_null_bytes delegates to
_strip_null_bytes_from_value for each value, handling two levels
of nesting (sufficient for spend log payloads).
Co-Authored-By: Ishaan Jaffer <155045088+ishaan-berri@users.noreply.github.com>
PostgreSQL rejects text containing \x00 with error 22P05. When a request
input contained null bytes (e.g. Perl stack traces), the spend log write
failed on every scheduler tick, causing sustained high CPU on RDS.
Add _strip_null_bytes() that recursively removes \x00 from all string
values in a spend log payload before it reaches Prisma create_many.
Co-Authored-By: Ishaan Jaffer <155045088+ishaan-berri@users.noreply.github.com>
When redis_startup_nodes is set the async cluster client was built with no health check and no TCP keepalive, so a connection silently dropped by a cluster restart (e.g. ElastiCache Serverless maintenance) stayed in the pool and got reused while dead; the first command after the restart stalled in re-initialization until the LoggingWorker timeout cancelled it, surfacing as CancelledError then TimeoutError on the spend-counter path
Build the async cluster client with a 25s health_check_interval and socket_keepalive so an idle connection is PING-validated and reconnected before reuse, and expose both through the cluster kwarg allow-list so an explicit value from config still wins
Resolves LIT-4083
update_spend_logs flushes the queue with a single create_many per batch, so one
row carrying bytes Postgres refuses (a residual NUL byte is the canonical case)
fails the entire insert and drops every good spend log alongside it. PR #29515
strips NUL bytes from the JSON columns, but the scalar string columns (end_user,
model, session_id, ...) still flow through unsanitized, so a poisoned row can
still reach the write and take a batch of up to 1000 good rows down with it.
On a genuine data-layer rejection the batch is now bisected so the good rows
still persist and only the offending row is dropped and logged with its
request_id. The classification lives in PrismaDBExceptionHandler.is_prisma_data_error
(matched by exact type so systemic subclasses like a missing table are not
mistaken for a single poison row), which keeps prisma an in-function import and
litellm.proxy.utils importable without the proxy extra. Transport failures,
including the "can't reach database server" outage that prisma mislabels as a
DataError, are re-raised unchanged so the existing connection-retry path still
runs and a transient outage never turns into silent per-row data loss.
The bisection carries a per-batch isolation budget so an authenticated caller
flooding poisoned rows cannot amplify one failed bulk insert into ~2N failed
inserts and N log lines; once the budget is spent the still-failing remainder
is dropped wholesale under a single log line.
Resolves LIT-4103
* feat(messages): passthrough /v1/messages to native endpoints via supported_endpoints
The unified /v1/messages proxy endpoint always translated inbound Anthropic
requests down to /v1/chat/completions (or the Responses API for openai) when the
deployment's provider lacked a native Anthropic-messages config, dropping
Anthropic-only features like cache_control and thinking. Some customers run
OpenAI-compatible servers (self-hosted vLLM, DeepSeek's Anthropic endpoint, etc.)
that also natively expose /v1/messages and want the raw Anthropic payload
forwarded untranslated, while keeping provider openai so /v1/chat/completions to
the same deployment stays native.
Opt in per deployment via model_info.supported_endpoints containing
/v1/messages. When present, the gate routes to a generic, provider-agnostic
OpenAILikeAnthropicMessagesConfig that POSTs the Anthropic payload to
{api_base}/v1/messages with Bearer auth, instead of translating. Default
behavior is unchanged. Generalizes and supersedes the hosted_vllm-only,
env-var-toggled PR #28745.
* fix(messages): preserve standard-cased caller headers in native passthrough
The OpenAI-like Anthropic passthrough config only checked for lowercase header
names before injecting Bearer auth, anthropic-version, and content-type
defaults. A caller sending standard-cased Authorization, Anthropic-Version, or
Content-Type was treated as missing those headers, so LiteLLM added duplicate
lowercase variants and overwrote the caller's credential/version at the HTTP
layer. Header presence is now checked case-insensitively and the merge no longer
mutates the caller dict.
Also moves the feature docs out of the main repo (docs live in litellm-docs).
* fix(openai_like/messages): delegate to parent transform and inject anthropic-beta headers
The passthrough config bypassed the parent transform and skipped header beta injection. Both gaps cause native /v1/messages features (context management, advisor tool, fast mode, structured outputs, reasoning_effort, advisor stripping) to silently degrade on opted-in deployments. Reuse the parent's pipeline and call _update_headers_with_anthropic_beta after merging defaults
* fix: normalize anthropic-beta header key case before beta injection
* style: collapse anthropic-beta header normalization to single line
ruff format --check requires the comprehension on one line (it fits within
the 120 char limit); fixes the lint job failure on the bugbot autofix commit
* fix(messages): forward anthropic-beta to native passthrough upstream
The shared anthropic_messages HTTP handler ran update_headers_with_filtered_beta
with the deployment's custom_llm_provider after validate. For the native
/v1/messages passthrough that provider is openai, which has no beta-header
mapping, so every anthropic-beta value (caller-supplied or feature-derived for
speed/context_management/etc.) was stripped to empty before the upstream
request, breaking beta passthrough to the Anthropic-compatible endpoint.
Beta filtering only makes sense on cross-provider translation paths where the
upstream cannot understand Anthropic betas. Gate it on a new
should_filter_anthropic_beta_headers() that defaults to True (bedrock, vertex_ai,
native anthropic unchanged) and is overridden to False by
OpenAILikeAnthropicMessagesConfig, whose upstream is a native Anthropic endpoint,
so betas pass through verbatim.
* chore: remove accidentally committed local QA logs and config
---------
Co-authored-by: Cursor Agent <cursoragent@cursor.com>
* fix(mcp): stop one unauthenticated server from emptying the aggregate tools/list
On the aggregate MCP route (/mcp), the gateway fans out to every server the caller can access and
flattens their tools. _fetch_and_filter_server_tools re-raises MCPUpstreamAuthError unconditionally
(added with the OAuth passthrough feature in #28356) so it surfaces a 401 on single-server routes,
but on the aggregate route that exception propagates through the asyncio.gather fan-out and the
outer handler turns it into an empty list. The result: a single delegate/passthrough OAuth server
the user has not authenticated (e.g. a delegate-auth server) zeroes the tools of every other server,
including the ones that resolve fine, so the client connects and sees no tools.
Surface the upstream auth error only when a single server was explicitly targeted (so that route
still drives the upstream OAuth flow); across the aggregate, absorb it to [] for that one server so
the rest still list their tools. This restores the graceful per-server degradation that predated
#28356.
Adds regression tests: the aggregate keeps a healthy server's tools when a sibling raises
MCPUpstreamAuthError, and a single-server listing still surfaces it.
* fix(mcp): decide aggregate vs single-server listing by route scope, not server count
Addresses review: keying the surface-vs-absorb decision off the server count (len(allowed_mcp_servers),
and even len(mcp_servers)) misclassifies an aggregate /mcp request from a key that can access exactly
one server as a targeted single-server listing, so that one server's MCPUpstreamAuthError re-raises and
empties the aggregate again for one-server permission sets.
Use the path-derived single-server scope instead: _mcp_gateway_server_name, set by
_gateway_initialize_instructions_request_scope only when the request path names exactly one upstream
server (/<server>/mcp) and never from client headers, is None on the aggregate route (/mcp) regardless
of how many servers the key can access. Single-server routes still surface the upstream-auth challenge;
the aggregate absorbs it per server.
Adds a regression test that an aggregate request with a single accessible server still absorbs, plus
renames the single-server test to drive the route scope explicitly. The new test fails on the
count-based logic.
* fixing aggregation error
* style(mcp): collapse single-line debug log to satisfy ruff format
Adds `!` prefix negation to tag-based routing so callers can exclude deployments by exact tag value without enumerating every allowed alternative. `!provider:anthropic` removes all deployments tagged exactly `provider:anthropic` before routing, and positive and negation tags compose. Matching is exact literal membership (frozenset intersection), so there is no regex or ReDoS surface for client-supplied tags. Ban-only requests that carry only negation tags stay within the default pool, mirroring untagged-request semantics so callers can't use negation to escape it. Fallback chains keep working because get_deployments_for_tag runs on each routing hop
Copy of #31680; implementation credit to @deepanshululla
Co-authored-by: deepanshululla <15312873+deepanshululla@users.noreply.github.com>
* feat(guardrails): expose streaming knobs on generic_guardrail_api
Wire streaming_end_of_stream_only and streaming_sampling_rate through
optional params, initialize_guardrail, and get_config_model so the
generic guardrail API participates in UnifiedLLMGuardrails streaming
checks with configurable cadence and end-of-stream-only mode.
* fix(guardrails): use builtin type[] in get_config_model return
Avoids a new UP006 violation that tripped the ruff strict-rule budget
gate on the PR lint job.
* fix(guardrails): default optional streaming knobs to None
Non-None Pydantic defaults on GenericGuardrailAPIOptionalParams made
_get_config_value treat unset nested fields as explicit values, which
shadowed top-level litellm_params streaming flags whenever any other
optional_params key was present. Real defaults stay in the constructor.
* fix(guardrails): address review nits on generic_guardrail_api streaming
Validate streaming_sampling_rate >= 1 in the constructor and Pydantic
optional_params (ge=1), and add /v1/responses streaming coverage through
the unified post-call hook so Responses API usage is exercised alongside
chat completions.
* fix(guardrails): read nested streaming config from dict optional_params
Guardrail API/UI delivers optional_params as a plain dict, so getattr was
silently ignoring streaming_sampling_rate and streaming_end_of_stream_only.
Handle both dict and model shapes in _get_config_value with regression tests.
* fix(guardrails): clear ruff findings in generic_guardrail_api tests/types
* style(guardrails): ruff format generic_guardrail_api modules
---------
Co-authored-by: Marton Schneider <marton@schneider.co.nl>
* fix(ui): allow any git host on the skills add form (LIT-4053)
The skills add form only accepted GitHub URLs: its URL parser bailed on
any host that did not start with github.com, so GitLab, Bitbucket, and
self-hosted repos (and any repo subfolder on them) were rejected before a
request was ever sent. The backend already accepts arbitrary git hosts
via its url and git-subdir sources, with no host allowlist, so this was a
client-side restriction only.
Generalize the parser into an exported, host-agnostic parseSkillSource:
GitHub URLs keep their github / git-subdir shorthand, every other host is
treated as a raw repo url, and an optional Subfolder path field turns any
repo into a git-subdir source (url + path). When a pasted GitHub
tree/blob URL already encodes a subfolder, the field is cleared and
disabled so a contradictory source can never be submitted.
The parser is hardened to match the backend contract: query strings and
fragments are stripped, the host match is case-insensitive and drops a
leading www., the extracted and field-entered subfolder paths are both
validated against the same regex the server uses, a real file-extension
allowlist (not "any dot") decides whether a trailing blob segment is a
file, a branch-only tree URL falls back to the repo, non-GitHub URLs
require at least an org/repo, and the suggested skill name is kebab-cased
so it satisfies the name field's own rule.
The git-subdir source is now handled in the display helpers
(getSourceDisplayText, getSourceLink, formatInstallCommand), which
previously showed it as "Unknown source" with no link. The submit path
is fully typed (RegisterPluginRequest plus an AddPluginFormValues
interface), removing the two prior any usages; as a result an
author with an email but no name is dropped rather than sent, since the
backend requires the author name.
No backend changes. Tests cover the full host/subfolder matrix at the
parser level plus form-submit assertions on the exact source payload.
* refactor(ui): sync skill register types to the generated OpenAPI schema, surface backend errors
Replace the hand-maintained, already-drifted API types for the skills add
flow with the generated ones from schema.d.ts: PluginAuthor now aliases
components["schemas"]["PluginAuthor"], the registration payload is a new
SkillRegisterRequest (the generated RegisterPluginRequest envelope with
source narrowed to our PluginSource union, since the backend types source
as a loose string map, and version kept optional since the backend
defaults it), and the dead, mismatched RegisterPluginResponse is deleted.
registerClaudeCodePlugin's inline payload type (which was missing the
git-subdir path field entirely) is replaced with SkillRegisterRequest, so
the networking layer and the form can no longer drift from the backend.
Error handling: the add-skill form swallowed the real failure and always
showed "Failed to register skill". registerClaudeCodePlugin already
derives the backend message and throws it, so the form now surfaces it
("Failed to register skill: <reason>"), and the networking helper falls
back to the raw body / status when the error response is not JSON instead
of throwing a JSON parse error. A regression test asserts the backend
message reaches the user.
* fix(ui): reject credentialed git URLs on the skills form
A repo URL with embedded user-info (user:token@host) passed the raw-host
parser and was stored verbatim as the skill source, which is served on
the unauthenticated /public/skill_hub and marketplace.json feeds, leaking
the credentials. Reject any host segment containing '@'.
* fix(ui): validate skill repo URLs through one WHATWG URL gate
Replace the ad-hoc string parsing (stripScheme / splitHost / manual
scheme, @, ?# checks) with a single parseRepoUrl gate built on the URL
parser, so every malformed/unsafe class is handled in one place and the
URL stored on the public skill feeds is always canonical. It enforces
https (rejecting http/ssh/git/file/javascript/data and protocol-relative
//host), rejects embedded credentials (user:token@host, including
userinfo-confusion like github.com@evil.com), rejects IP-literal hosts
(loopback/private/metadata and obfuscated/IPv6 forms), and rebuilds the
stored url from origin+pathname so query strings, fragments, and trailing
slashes can never be published. The GitHub org/repo shorthand is now
charset-validated like the other paths, so junk can't reach the stored
repo. Closes both Veria findings (credentialed and http sources) plus the
adversarial-review follow-ups, with regression tests for each class.
Guard the per-request CPU cost of the chat completion, MCP tool and A2A
message transforms against regressions on every commit. All benchmarks are
pure in-process work with no network I/O so they stay deterministic under
CodSpeed's simulation mode, and they import under the base dependency set the
benchmark job installs.
Inference covers the full SDK overhead via mock_response (simple, multi-turn,
tools, streaming) plus convert_to_model_response_object as a deterministic
anchor. MCP covers the client-side tool translation and the proxy server-side
tool-name prefix round-trip. A2A covers the client request/response transforms
and the proxy server-ingress message conversion.
Adds the mcp and a2a-sdk packages to the benchmark run since those transform
modules need them, and broadens the workflow triggers to litellm_internal_staging
so the internal branch flow is benchmarked too.
* feat(otel): emit a tools/list CLIENT span for MCP discovery under otel_v2
Under otel_v2 an MCP tools/call already produced a dedicated CLIENT span, but tools/list produced none. The discovery call surfaced only as the bare POST /{mcp_server_name}/mcp server span with no MCP attributes, indistinguishable from initialize and impossible to query by method
The list success event already reaches the v2 logger with call_type list_mcp_tools, but _emit_mcp_tool_call only matched call_mcp_tool, so listing fell through to the LLM-call path and emitted nothing. This adds a dedicated MCP_LIST_TOOLS span role with its own MCPListToolsSpanData, emitted from a sibling _emit_mcp_list_tools branch that mirrors the tools/call path
Per the OTel GenAI MCP semantic conventions the span is named tools/list (the method name alone, since there is no low-cardinality target), is a CLIENT span parented to the request span, and carries mcp.method.name plus the call id. It deliberately omits gen_ai.operation.name and gen_ai.tool.name, which the convention reserves for tool executions, since listing runs no tool
* fix(otel): anchor MCP spans to params._meta trace context, not the transport span
MCP streamable-HTTP multiplexes many JSON-RPC messages over one session, so the request-root anchor captured on initialize persisted and every later message's span (tools/call, tools/list) nested under it. A tools/list run 44s after the initialize rendered 44s to the right of its parent with a clock-skew warning, because the MCP message and the HTTP transport are independent lifecycles
Following the OTel GenAI MCP semantic conventions, an MCP span now parents to the W3C trace context the client propagated in the request's params._meta (a remote parent, per SEP-414), records the transport/session span as a span link rather than the parent, and starts its own root trace when nothing was propagated. The MCP gateway captures traceparent/tracestate/baggage from each message's params._meta into a per-message contextvar that the otel_v2 emitter reads; opentelemetry stays an optional dependency via guarded lazy imports
This applies to tools/call as well as the new tools/list span, since both shared the same transport-anchoring bug
* fix(otel): drop client baggage from MCP params._meta to prevent identity spoofing
The MCP trace propagation added a W3CBaggagePropagator, so resolve_mcp_span_context
extracted the client's W3C Baggage from params._meta into the span's parent context.
The LiteLLMBaggageSpanProcessor then stamps allowlisted baggage keys onto the span,
and the list-tools/tool-call mappers don't set those identity keys, so nothing
overwrites them. A malicious MCP client could send
params._meta.baggage: litellm.team.id=...,litellm.metadata.user_api_key_user_id=...
and have those identity attributes attributed to its spans.
Extract trace context only (traceparent/tracestate) in the propagator, and stop
collecting the baggage key at the source in _mcp_meta_trace_carrier. Parenting to the
client's trace context, the actual goal, needs only trace context; remote baggage had
no legitimate consumer here. Regression tests at both layers assert a spoofed
params._meta.baggage never lands as a span identity attribute.
* style(mcp): clear ruff strict-budget breach in otel trace-carrier helpers
The otel MCP trace-carrier helpers added in this branch pushed the BLE001 and
UP006 strict-rule totals past their ceilings. Use PEP 585 `dict[str, str]` instead
of `Dict`, and narrow the optional-import guards to `except ImportError` (the only
failure these can hit, matching the "when otel_v2 is unavailable" intent) instead of
a blind `except Exception`.
* fix(otel): stamp authenticated identity baggage onto MCP spans
Parenting MCP spans to the client's params._meta trace context over an empty
Context() meant the tool-call and tools/list spans carried no team/key/metadata
identity at all, so they couldn't be attributed or filtered by team in a traces
backend. The LLM-call span already re-seeds identity from the parsed, authenticated
StandardLoggingPayload rather than trusting ambient/remote context; extract that into
a shared _seed_identity_baggage helper and run both MCP emitters through it.
Identity comes only from the authenticated payload, never the client carrier, so this
keeps the earlier spoofing fix intact while restoring attribution. Regression tests
assert the authenticated team lands on both MCP spans and that a spoofed
params._meta.baggage value can't override it.
* refactor(otel): model MCP spans as roots that link the transport in SPAN_REGISTRY
The proxy auth path calls phase_span() and seed_request_identity() in
litellm/integrations/otel/runtime.py on every request, each doing a
try/except lazy import of litellm.integrations.otel.logger. When the
OpenTelemetry SDK is not installed (the default), that import raises, and
CPython never caches a failed import, so every request re-scanned sys.path
and contended on the import lock. At 750 concurrent users this cost about
12% throughput versus v1.85.0.
Resolve the hooks once and cache the outcome, absence included, with
functools.cache, so the import is attempted a single time instead of per
request. Throughput returns to the v1.85.0 baseline.
* feat(proxy): type Customer Management response_model for OpenAPI coverage
Add response_model to the five remaining untyped /customer operations
(block, unblock, new, update, delete) so the generated OpenAPI schema
documents a concrete response body. new/update reuse the canonical
LiteLLM_EndUserTable (matching info/list); block, unblock, and delete
get small dedicated models in
litellm/types/proxy/management_endpoints/customer_endpoints.py.
Together with the already-typed info/list/daily-activity routes this
brings the Customer Management group to full response_model coverage.
Regression tests assert each public /customer/* route declares the
expected response_model and that /customer/new surfaces a typed schema
in app.openapi(), so dropping a response_model fails CI.
* fix(proxy): keep budget_id in typed customer responses
Address review feedback on the Customer Management response_model typing.
Greptile flagged that response_model=LiteLLM_EndUserTable on /customer/new
and /customer/update silently drops fields the raw Prisma model_dump()
echoed. Checking the schema, budget_id is the only such scalar column that
was missing from the Pydantic model (created_at/updated_at/tpm_limit do not
exist on litellm_endusertable), so add budget_id to LiteLLM_EndUserTable.
This restores budget_id on new/update and also fixes the pre-existing gap
where /customer/info and /customer/list (already typed) dropped it, which
the UI Customer type expects. A regression test pins budget_id surviving the
response_model filter on /customer/update.
Also document UnblockUsersResponse.blocked_users via a Field description: it
holds the users that remain blocked after the call. The key name predates
this PR and is kept to avoid a backwards-incompatible rename on a beta route.
* fix(proxy): keep nested budget fields in customer responses
response_model=LiteLLM_EndUserTable nests the budget as the narrow write
allowlist LiteLLM_BudgetTable, which silently drops the server-managed
fields the customer endpoints used to return (budget_reset_at, created_at).
Introduce CustomerResponse, a thin response model that nests
LiteLLM_BudgetTableFull (the repo's budget response model), and apply it on
/customer/new, /customer/update, /customer/info and /customer/list. list
also builds CustomerResponse so its budget isn't narrowed at construction
time. created_by/updated_at/updated_by remain omitted, matching how budgets
are returned elsewhere.
The shared LiteLLM_EndUserTable is left untouched: it's constructed in many
places that pass narrow budget instances, and pydantic v2 won't coerce a
budget instance into a wider nested model. Typing only at the response
boundary (where the handler hands FastAPI a dict) sidesteps that. A
regression test pins budget_reset_at + created_at through the filter and
asserts the internal audit fields stay out.
* test(proxy): add golden-master characterization tests for customer responses
Lock the exact JSON body each customer-object endpoint (info/list/new/update)
and delete return today, so the upcoming type-safety refactor of the handlers
is only allowed to land if it reproduces these byte for byte. Pins null-field
inclusion, the nested budget shape (server fields kept, audit fields dropped),
and object_permission reverse-relation stripping. Green against current code.
* refactor(proxy): make the customer response flow type-safe
Replace the untyped dict + bolt-on response_model pattern on the customer
object endpoints with explicit typed construction. A single mapper,
_to_customer_response, validates a DB row into CustomerResponse at one
Any -> typed seam; new/update/info/list now return it (or a list of it) and
carry real -> CustomerResponse / -> List[CustomerResponse] return
annotations, and delete returns DeleteCustomersResponse. basedpyright now
verifies the handlers' return shapes instead of a runtime filter doing it
silently.
This also deletes the four copy-pasted object_permission reverse-relation
cleanup loops: pydantic's extra=ignore drops those undeclared fields during
validation, so the loops were dead code (proven by the golden-master tests,
which stay byte-for-byte green). basedpyright errors on the file drop from
140 to 116, all from removed dict plumbing.
CustomerResponse stays a thin subclass of LiteLLM_EndUserTable so it inherits
the existing validators/config unchanged (behavior preservation); only the
nested budget type is widened.
* refactor(proxy): annotate customer response mapper param as BaseModel
Address review nit: the mapper's untyped `record` added an ANN001 violation.
The incoming rows are pydantic v2 models, so type the param as BaseModel
rather than object (object has no model_dump, which would just move the
problem to basedpyright). This clears the ANN001 and also drops three
basedpyright unknown-type violations the untyped param was adding.
* style(test): ruff format customer endpoint tests
* test(proxy): give customer budget test update mocks a valid model_dump
The type-safe response refactor validates the update result via
_to_customer_response (CustomerResponse.model_validate(record.model_dump())).
These budget tests mocked the end-user update to return a bare MagicMock,
so model_dump() yielded a MagicMock that fails validation. Give each update
mock a minimal valid dict; the tests assert on the prisma calls, not the body.
* chore(ui): regenerate API types from proxy OpenAPI spec
* fix(ui): make generated API types stable across Python versions
Python 3.13 strips a docstring's common leading indentation at compile
time while 3.12 keeps it, so app.openapi() emits differently-indented
description strings depending on the interpreter. The dashboard type
generator ran locally on 3.13 and in CI on 3.12, so schema.d.ts drifted
and the "Verify schema.d.ts matches the proxy OpenAPI spec" check failed
Normalize every description through inspect.cleandoc in the spec dump so
the output is identical regardless of interpreter, then regenerate
* chore: shift CI lint left with a pre-commit hook and CLAUDE.md rule
Add an opt-in pre-commit hook (.githooks/pre-commit, active after
make install-hooks) that runs the CI-equivalent checks against staged
files: make lint for Python, prettier plus eslint for the dashboard,
and a gen:api drift check for the proxy OpenAPI types. Document the
same expectation in CLAUDE.md so reds surface locally instead of in CI.
* fix: make `make lint` isomorphic to the CI lint job
`make lint` diverged from test-linting.yml in ways that produced both
false reds and false greens: its format-check ran over the whole repo
(CI scopes it to changed files vs the base), its ruff-strict budget ran
in absolute mode (CI runs it as a delta vs base), and it omitted the
type-discipline gate entirely. Recompose `lint` to replay CI's exact
sequence: diff-scoped ruff format check, whole-tree ruff check, the
strict / type-discipline / basedpyright budgets as a delta resolved the
same way CI resolves it (merge-base with origin/litellm_internal_staging),
then circular-import and import-safety. Factor the repeated base fetch
into one shared prerequisite so the chain hits the network once.
Align the pre-commit hook's eslint invocation with the CI frontend-lint
job (`--pass-on-unpruned-suppressions`) and fix the CLAUDE.md guidance to
point at the diff-scoped frontend commands instead of the whole-folder
npm scripts, which are broader than CI.
* fix(githooks): make pre-commit 1:1 with CI frontend-lint, lint, and type-gen
The shift-left pre-commit hook diverged from the CI jobs it claims to mirror, so a clean commit did not actually mean a green CI lint.
The dashboard block only ran prettier and eslint over js/jsx/ts/tsx/mjs/cjs, but CI's frontend-lint runs prettier over a wider set (also json, css, scss, md, mdx, yml, yaml, html) and additionally gates the whole-folder eslint lint budgets via scripts/check-lint-budgets.mjs. The hook now mirrors that split and runs the budget check, so a dashboard commit that passes locally passes the job.
The API-types block ran npm run gen:api without LITELLM_PYTHON, so it shelled out to the system python3 which has no litellm installed and always failed with a false 'could not regenerate API types' red. It now passes LITELLM_PYTHON="uv run --no-sync python" the way check-ui-api-types.yml does.
make lint format-checks the files in origin/base...HEAD, which at pre-commit time predates the staged change, so a brand-new commit's formatting went unchecked. The Python block now also runs ruff format --check over the staged litellm files directly to cover that case, and its trigger is scoped to staged litellm/ files (the only tree CI's lint job inspects) so a tests-only or scripts-only commit skips the slow make lint instead of wasting time on a run that could not catch anything.
CLAUDE.md's shift-left rule was cut off mid-sentence and understated the frontend checks; it now describes all three gates accurately and points agents at make install-hooks to run them automatically before each commit.
Co-authored-by: Mateo Wang <mateo-berri@users.noreply.github.com>
* fix(githooks): scope the API-types check to all of check-ui-api-types.yml's triggers
spec_files was filtered from the staged Python files, so the gen:api drift
check only fired for .py changes under litellm/proxy or litellm/types. CI's
check-ui-api-types.yml triggers on any file under those directories (Prisma
schema, configs) plus the generator script and the dashboard package files,
so a non-Python proxy/types change could pass the hook and still fail CI.
Match the workflow's full trigger set instead.
* fix(pre-commit): run prisma generate before gen:api to mirror CI
* refactor(githooks): run shift-left lint via on-demand make pre-commit, not an auto-firing hook
The pre-commit hook ran make lint plus the dashboard eslint budgets, which are minutes of work (basedpyright over litellm/, a whole-folder eslint . pass at ~40s). Wiring that into core.hooksPath via make install-hooks meant every human commit, not just an agent's, paid that cost, which is real friction for interactive committers.
Move the staged-file checks out of .githooks/ into scripts/pre_commit_lint.sh and expose them as make pre-commit, and keep .githooks/ to only the fast Conventional Commits / Branches hooks so make install-hooks no longer makes commits slow. Agents run make pre-commit right before each commit (CLAUDE.md instructs this), so the slow gates fire only for the commits an agent is making and never auto-fire for a human typing git commit. The script stays hook-compatible for anyone who still wants it to fire automatically via a symlink.
Preferred this over sniffing an agent env var to auto-fire only for agents: that is fragile (misses agents when the var is unset, fires on humans when it leaks into their shell, and silently no-ops a hook a human deliberately installed), whereas an on-demand command achieves the same humans-never, agents-per-commit outcome deterministically.
Co-authored-by: Mateo Wang <mateo-berri@users.noreply.github.com>
* fix(pre-commit): run make lint last so it can't prune the proxy deps gen:api needs
make lint's install-dev prerequisite runs uv sync --frozen, which prunes the proxy extras (prisma, websockets, ...) from the venv. With the Python block running first, the subsequent API-types block then failed: gen:api imports litellm.proxy.proxy_server, which needs those deps, so every litellm/proxy change (the main trigger for the API-types check) hit a false 'could not regenerate API types' red. Run the dashboard and API-types blocks before the Python block so gen:api sees an intact env; CI is unaffected because there the lint and check-ui-api-types jobs run in separate environments.
Co-authored-by: Mateo Wang <mateo-berri@users.noreply.github.com>
* fix: make CLAUDE.md more concise
* fix(makefile): give make lint the CI lint env and stop it pruning the venv
make lint diverged from test-linting.yml's lint job in two ways: it never generated the Prisma client (so basedpyright resolved the DB wrappers as Unknown, drifting from CI's counts), and its bare uv sync --frozen pruned the proxy extras (prisma, websockets, ...) out of the venv on every run, which broke the gen:api step that imports litellm.proxy.proxy_server and left a dev unable to run the proxy until re-syncing.
Add a lint-install target that mirrors the job's environment (the proxy-dev group plus prisma generate) and runs before the checks, and make both it and install-dev use uv sync --inexact so they top up the venv instead of tearing packages out. CI is unaffected since it installs its own env per job.
Because make lint no longer prunes, the pre-commit reorder that ran it last (to dodge the prune) is no longer needed, so restore the original block order.
Co-authored-by: Mateo Wang <mateo-berri@users.noreply.github.com>
* fix(makefile): drop lint-install so make lint matches CI's slimmer env
test-linting.yml's lint job installs deps with a bare uv sync --frozen
(default dev group only, no proxy-dev, no prisma generate), but the
lint-install target chained into make lint pulled in --group proxy-dev
and ran prisma generate. Because the basedpyright budget step compares
head and base counts against fixed thresholds, the extra symbols and
Prisma client locally resolved can shift error counts away from CI's,
producing false greens or false reds on the type-check gate.
Remove the lint-install target and its slot in lint. The remaining
sub-targets already chain install-dev, which now uses
uv sync --inexact --frozen, so the venv still isn't pruned but the
installed set stays aligned with what CI sees.
* ci(linting): install proxy-dev and generate prisma in lint job, matching make lint
make lint now installs the proxy-dev group and generates the Prisma client so basedpyright resolves the DB wrappers; the lint job here still installed only the base env, so a local pre-commit could pass while the required CI lint failed (or vice versa). Bring this job in line, which is the same environment litellm_internal_staging's lint job already uses.
Co-authored-by: Mateo Wang <mateo-berri@users.noreply.github.com>
* fix(makefile): keep make lint on the proxy-dev + prisma env to match CI
A concurrent change dropped lint-install to match what looked like CI's slim env, but test-linting.yml's lint job (and the merge ref this PR's CI actually runs) installs --group proxy-dev and generates the Prisma client. With make lint slim and CI fat, basedpyright resolves fewer symbols locally than CI, so a prisma-typed error can stay Unknown locally (green) while CI catches it (red). Restore lint-install so make lint installs the same env CI does; the previous commit also brought this PR's test-linting.yml in line with that env, so the two now match.
Co-authored-by: Mateo Wang <mateo-berri@users.noreply.github.com>
---------
Co-authored-by: Cursor Agent <cursoragent@cursor.com>
Co-authored-by: Mateo Wang <mateo-berri@users.noreply.github.com>
litellm_overhead_latency_metric only covers the SDK wrapper window and excludes
proxy guardrails. Add a histogram that sums SDK overhead plus pre/post-call
guardrail durations (during-call excluded since it runs concurrently with the LLM
call, alongside logging_only and MCP modes that never block the response),
recorded next to the existing overhead metric with the same labels and buckets.
No existing metric's value is changed.
* chore: remove _experimental/out
* fix(ci): recreate _experimental/out before copying UI build output
The build scripts cp the Next.js output into litellm/proxy/_experimental/out,
which was removed from git. cp failed because the target directory no longer
existed; mkdir -p recreates it before the copy.
* fix(proxy): make UI serving resilient to a missing _experimental/out
Removing the committed UI export means the source/test tree no longer
ships litellm/proxy/_experimental/out. Three things assumed it was always
present and broke once it was gone:
- get_favicon hard-coded the built favicon path and 404'd without it; it
now falls back to the bundled swagger/favicon.ico
- the /_next and /ui static mounts raised at construction when the export
was absent, so the whole UI-setup block was swallowed and no mounts
registered; they now use check_dir=False
- _restructure_ui_html_files was a nested function only exposed as a
module attribute when that block happened to succeed; it is now a real
module-level function
test_admin_ui_export_serves_nested_extensionless_routes validated the
committed artifact, whose premise this PR removes; it now drives the same
MCP OAuth callback restructure guarantee through a synthetic export.
* chore(greptile): ignore generated _experimental/out so review fits the file limit
* Revert "chore(greptile): ignore generated _experimental/out so review fits the file limit"
ignorePatterns is applied after Greptile counts the files changed, so it
does not bring the diff under the file limit; the config had no effect.
Docs live in BerriAI/litellm-docs; these four files were swept into the
repo by unrelated PRs. The Crusoe provider and XecGuard guardrail pages
are migrated to litellm-docs (BerriAI/litellm-docs#438); plugin_architecture.md
is already covered there by docs/proxy/plugins.md, and the orphaned image
was referenced by no doc.
Drop the explicit size on Request ID so it falls back to the default
width like the other reverted columns. Narrow Session ID from 160px to
120px since its truncated value needs less room
The explicit 90px/80px sizes were too narrow for the Duration (s) and
TTFT (s) headers once the sort arrows were factored in, cramping the
header labels. Dropping the size lets these two columns fall back to the
default width like before
* fix(anthropic): drop unsignable thinking blocks and allow null signature in logging (LIT-4007)
Open-source reasoning models (DeepSeek-R1 and distills, Qwen3/QwQ, IBM
Granite 3.2 via vLLM/Ollama/OpenRouter/DeepSeek) return reasoning_content
with no Anthropic-style signature, which LiteLLM represents as a thinking
block with a null signature.
Two failures resulted. First, ChatCompletionThinkingBlock.signature was a
required str, so building the StandardLoggingObject raised a ValidationError
on signature=None and the success log record was silently dropped while the
request still returned 200; relaxing it to Optional[str] lets the log build.
Second, replaying such a turn to a real Anthropic model forwarded the
null-signature thinking block unchanged and Anthropic rejected it with
400 thinking.signature.str; since Anthropic verifies the signature
cryptographically, a null, empty, or missing signature cannot be repaired,
so anthropic_messages_pt now drops the unsignable thinking block while
preserving the assistant text and keeping genuinely signed blocks.
* style: use builtin generics for thinking-block filter helpers
* fix(ui): regenerate schema.d.ts for nullable thinking-block signature
Reasoning-token cost was computed but folded into output_cost, and cache
cost was only populated from the top-level cache_read_input_tokens attribute,
so providers that report cache tokens under prompt_tokens_details (Gemini,
OpenAI, Vertex) never got a cache breakdown.
Adds a provider-agnostic get_token_type_cost_breakdown helper that derives
reasoning, cache-read and cache-creation cost from the normalized usage object
using the same rate-resolution primitives as the total-cost path, so the
breakdown reconciles with the totals. completion_cost stores these via
set_cost_breakdown, surfacing reasoning_cost (new), cache_read_cost and
cache_creation_cost in StandardLoggingPayload.cost_breakdown and the spend logs.
Co-authored-by: Kunal Nayyar <48790070+kunal2002@users.noreply.github.com>
The MCP gateway authenticated to upstream OAuth token endpoints only with
client_secret_post (client_secret placed in the POST body). Providers that
require HTTP Basic client authentication (client_secret_basic, the OIDC
default) reject that with invalid_client, which surfaced as a 500 on the
/<server>/token exchange and broke both the initial authorization_code
exchange and refresh.
Add a per-server token_endpoint_auth_method ("client_secret_basic" |
"client_secret_post") and a single helper that builds the right headers and
body for the configured method, then route every upstream token-endpoint POST
through it: the inbound exchange and refresh in discoverable_endpoints, the v1
per-user refresh in db, the v2 authorization_code refresher, the M2M
client_credentials fetch, and the RFC 8693 token exchange. The default stays
client_secret_post so existing servers are unaffected; basic sends
Authorization: Basic base64(form-urlencode(client_id):form-urlencode(secret))
per RFC 6749 section 2.3.1 and omits the secret from the body.
client_secret_basic is a confidential-client method, so a server configured for
it with a missing client_id/secret raises rather than silently downgrading to a
body request (no-silent-fallback); the inbound endpoint maps that to a 400 and
the refresh paths to a failed-refresh / needs-reauth. A secretless client_id
under the default method stays valid for public clients authenticating with PKCE.
Resolves LIT-4091
* fix(proxy): emit x-litellm-response-cost header on /messages and /generateContent (LIT-4076)
The Anthropic /v1/messages and Google native :generateContent routes return
TypedDict results (AnthropicMessagesResponse, GenerateContentResponseBody) that
are plain dicts at runtime and cannot hold a _hidden_params attribute. The cost
is computed by update_response_metadata, but ResponseMetadata.apply() only
persists _hidden_params back when the result object has that attribute, so for
those two routes the computed response_cost was dropped. The non-streaming
header build in base_process_llm_request then saw an empty response_cost and
get_custom_headers filtered the x-litellm-response-cost header out, even though
the other x-litellm-* headers still appeared.
The non-streaming success path now recovers the cost from the logging object
when the response cannot carry _hidden_params, preferring the value already
stored in model_call_details and recomputing from the same calculator only when
it has not been stored yet. Object responses (ModelResponse, ResponsesAPIResponse)
keep their existing behavior, so chat/completions, /responses, and the Anthropic
error path that intentionally emits a zero cost are unaffected. Streaming stays
out of scope because the header is emitted at stream start, before the cost is
known.
* fix(proxy): also recover response cost header for /generateContent responses with _hidden_params (LIT-4076)
* fix(proxy): compute generateContent response cost synchronously so cost header is emitted (LIT-4076)
* fix(lint): suppress BLE001 on generate_content cost normalization guard
The defensive blind except keeps cost normalization from ever breaking the
response path; mark it noqa so it does not breach the strict-rule budget.
* fix(bedrock): drop unmappable Responses tools instead of failing the request (LIT-3858)
When an OpenAI Responses request is routed to a Bedrock Converse Anthropic model,
litellm translates the tools array into Bedrock toolConfig. Responses built-in tool
types beyond function (web_search, image_generation, namespace, tool_search, custom)
have no Bedrock equivalent, and previously caused two failures.
A web_search tool is derived into a web_search_options param. Bedrock Anthropic
models do not list web_search_options in get_supported_openai_params, so the request
raised UnsupportedParamsError (HTTP 400) even though it never needed web search. The
derived param is now dropped on the Bedrock chat-completion bridge for models that
do not support it, scoped to Bedrock so other providers are untouched and without
requiring drop_params. Nova still keeps it since it maps to a nova_grounding systemTool.
The remaining non-function tools reached _bedrock_tools_pt and were emitted as junk
litellm_unnamed_tool_N toolSpecs with empty schemas, polluting toolConfig with tools
the model could hallucinate calls to. They are now dropped because they carry neither
an OpenAI function nor an Anthropic input_schema, while mappable function and
input_schema tools survive untouched.
* refactor(responses): drop derived web_search_options via provider config
Greptile flagged that the LIT-3858 fix put Bedrock-specific logic in the
generic Responses->Chat Completion bridge: it imported AmazonConverseConfig
and branched on custom_llm_provider.startswith("bedrock").
Read web_search_options support from each provider's own
get_supported_openai_params instead, so the bridge stays provider-agnostic
and Bedrock capability knowledge lives in the Bedrock config that already
owns it. Behavior is unchanged for the cases the PR targeted (Bedrock
Anthropic drops, Bedrock Nova and OpenAI keep) and now generalizes correctly
to any provider whose config does not support the derived param.
Add a Cohere regression test proving the drop is provider-agnostic; it fails
under the old bedrock-only check and passes now.
* fix(responses): drop derived web_search_options for bedrock_converse alias
Greptile/T-Rex caught that the provider-agnostic drop regressed the
bedrock_converse route: get_supported_openai_params did not map the
bedrock_converse alias (only "bedrock"), so it returned None (unmapped) and
the derived web_search_options was forwarded for
model="bedrock/converse/us.anthropic.claude-sonnet-4-6",
custom_llm_provider="bedrock_converse" instead of being dropped. The previous
startswith("bedrock") check happened to match the alias.
Map bedrock_converse through AmazonConverseConfig in get_supported_openai_params,
mirroring the existing ["bedrock", "bedrock_converse"] pairing in
_strip_model_name. Add regression tests at both levels: the alias now drops the
derived param for Anthropic Converse models, still keeps it for Nova, and the
helper resolves identically to "bedrock".
* fix: skip health check for semantic auto_router deployments
auto_router/<name> deployments are semantic meta-routers that select among
real LLM deployments at request time. They have no LLM endpoint to probe.
The health check was passing model=auto_router/router_1 to get_llm_provider(),
which raised BadRequestError: "Unmapped LLM provider for this endpoint" because
auto_router is not a real LLM provider, causing these deployments to always
appear unhealthy and curl requests to hang.
Detect semantic auto_router deployments in _run_model_health_check and return
{} (healthy) without calling litellm.ahealth_check. Sub-strategies
(complexity_router, adaptive_router, quality_router) are excluded from this
fast path and continue to be health-checked normally.
* ci: trigger circleci
The Model Armor guardrail only sent text extracted from user messages to
sanitizeUserPrompt, so harmful content inside attached PDFs, Office docs,
and CSVs reached the LLM unscanned. A file-only message had no extractable
text, so the pre-call and moderation hooks returned early and the document
was never submitted to Model Armor at all.
Wire inline document/file scanning into async_pre_call_hook and
async_moderation_hook. extract_file_attachments walks message content blocks
(OpenAI type:file file_data and Anthropic type:document source), decodes the
base64 bytes, maps the MIME type to a Model Armor byteDataType, and skips
remote URLs, bare file_id references, oversize files past the 4 MB limit, and
unsupported types. Each attachment is sent through the byte API and a
MATCH_FOUND blocks the request before it reaches the LLM.
Resolves LIT-4084
* fix(passthrough): drop top-level additional_drop_params on /v1/messages
On the Anthropic Messages pass-through path, additional_drop_params only
stripped nested dotted paths, so plain top-level keys like `thinking` and
`context_management` were forwarded to the provider. Bedrock rejects these
with "Extra inputs are not permitted", returning a 400 to Claude App/CLI
even when the user configured `additional_drop_params: ["thinking"]`.
delete_nested_value already handles plain top-level fields, so route every
drop param through it and remove the nested-only filter. Fixes#25931.
* fix(passthrough): drop thinking for bedrock inference-profile ARNs on /v1/messages
Opaque Bedrock Application Inference Profile ARNs contain neither "anthropic"
nor "claude", so is_anthropic_claude_model returned False and the thinking
param was rewritten to reasoning_effort before additional_drop_params ran.
That made additional_drop_params: ["thinking"] a no-op for the converse-ARN
form, and the Bedrock Converse transform re-expanded reasoning_effort back into
additionalModelRequestFields.thinking, so the request 400'd.
Extend the thinking-translation gates to also accept bedrock ARNs via the
existing is_bedrock_arn_model helper, mirroring the cache_control path, so
thinking is preserved as thinking and additional_drop_params can drop it.
* fix(mcp): resolve per-user OAuth identity authoritatively at the token endpoint
The OAuth token endpoint stored a user's per-server token under the identity
returned by _extract_user_id_from_request, which read only the Authorization
header and did getattr(cached, "user_id") on a raw user_api_key_cache lookup
with no model_type rehydration and no DB fallback. That silently returned None
in two common cases on a multi-replica gateway: the LiteLLM key arrives on
x-litellm-api-key (what MCP clients such as Claude Desktop and Claude Code
send) rather than Authorization, and a cross-replica cache hit deserializes to
a plain dict rather than a UserAPIKeyAuth, so getattr finds no attribute. When
it returned None the token was not persisted.
This was survivable until the authorization_code v2 migration began stripping
the caller's Authorization for migrated per-user OAuth servers and routing the
preemptive 401 existence check through the stored token, so a persist miss now
hard-fails: the egress challenges with 401 on every reconnect (the client sees
"rejected them on reconnect" or a successful connect with zero tools).
Resolve identity through get_key_object, the canonical resolver that reads the
cache with model_type and falls back to the DB, and accept the key from
x-litellm-api-key as well as Authorization. The silent persist skip is now a
warning. The caller-Authorization stripping stays as is, since reinstating it
would reopen the cross-user credential override it was added to prevent.
* fix(mcp): reject blocked or expired keys when resolving the token-endpoint identity
The OAuth token endpoint is unauthenticated, and get_key_object resolves a key row without the
blocked/expiry checks the main user_api_key_auth pipeline runs (that pipeline is bypassed here). So
a holder of a revoked or expired LiteLLM key could POST a valid upstream authorization code with
that key in x-litellm-api-key/Authorization and write or overwrite the stored per-user OAuth token
for that key's user. The cache-only resolver this replaced incidentally dropped blocked keys
(blocking purges the cache entry), so moving to the authoritative cache-then-DB resolution removed
that accidental shield.
Validate the resolved key before trusting its identity: return None when blocked or expired, so the
upsert is skipped. Deleted keys are already rejected, since get_key_object raises on a missing row.
Regression tests cover the blocked and expired cases and fail without the guard.
* fix(proxy): count only active users toward license seat limit
SCIM-deactivated users (metadata.scim_active == false) are kept in LiteLLM_UserTable for audit and reactivation, but they were still counted toward the per-user license limit, so deactivating a user never freed a seat. Okta never sends a SCIM DELETE and Entra only hard-deletes well after deactivation, so deactivation has to be what frees the seat
Add UserRepository.count_billable_users(), which counts every row except those where metadata.scim_active is false (absent, null, and true all count), and route the user-create license gate, the free-SSO 5-user cap, and the enterprise /user/available_users display through it. A separate litellm_active_users Prometheus gauge reports the billable count while litellm_total_users keeps its original meaning so existing dashboards are unaffected
* fix(proxy): floor billable user count at zero
count_billable_users() runs two separate count queries (total, then deactivated). Under a burst of deactivations between them, the deactivated count can momentarily exceed the earlier total and produce a negative result, which would flow into is_over_limit as a negative and show a negative seat count in the display and gauge. Clamp the result to zero so a transient race can never yield a nonsensical negative; the value self-corrects on the next call
Addresses Greptile P1 on the PR
* refactor(proxy): count teams via TeamRepository in available_users
* style: ruff format changed files at line-length 120
* fix(vertex_ai/files): upload batch files in a single media request to fix 499s on large uploads
PR #31036 switched the vertex batch file upload from a single GCS media
upload to a chunked resumable session. The resumable path sends the body as
many sequential PUTs, each waiting a full round-trip to GCS before the next,
so a multi-GB upload accumulates hundreds of round-trips and overruns the
client/load-balancer request timeout, surfacing as 499s (client closed
connection) on files as small as 500MB. This was a regression from the
last-known-good commit, where the upload completed as one continuous request.
Revert the batch upload to a single uploadType=media request, but stage the
transformed payload to a temp file first so peak memory stays bounded (the
goal of the resumable rewrite) without the per-chunk round-trips. The temp
file is closed deterministically (TemporaryFile unlinks on close), not left
to the GC. The now-unused resumable chunked-upload plumbing is removed.
Also swap the per-row transform's stdlib json for orjson (parse + serialize),
which is ~4x faster on this hot path; the streaming body now emits compact
orjson bytes.
The request stays synchronous, so the returned file object is real and
POST /v1/batches keeps working immediately against the uploaded object.
Tests: single media request carries the whole payload with a real
Content-Length (no chunked transfer-encoding); failed upload raises; the
staged temp file is closed deterministically; byte-for-byte transform parity.
* test(vertex_ai/files): mock single media upload POST instead of removed resumable method
test_avertex_batch_prediction patched BaseLLMHTTPHandler._aresumable_chunked_upload, which was removed when the batch jsonl upload moved from a chunked resumable GCS session to a single uploadType=media request. Patch the raw httpx.AsyncClient.post that _astage_and_upload_media issues so the real staging, upload and response transform run while the GCS object response is mocked, and assert the media URL and Content-Type.
* fix(vertex_ai/files): forward request timeout to media upload, drop orjson, sort imports
Forward the per-request timeout through _stage_and_upload_media /
_astage_and_upload_media to the GCS POST. Every other upload branch forwards
it; the new media path was dropping it, so a caller-provided timeout was
silently ignored (the files path passes 600s by default, but a custom
request_timeout would not have reached this upload). Regression test asserts
the resolved timeout reaches the request (mutation-verified).
Revert the orjson swap in the batch transform: importing orjson at module load
in this core-path file broke `import litellm` on environments without orjson
(the Windows import test). Back to stdlib json; the upload leg dominates large
uploads anyway, so the transform-side win was marginal.
Fix import ordering in llm_http_handler.py (I001) introduced by the new imports.
* fix(vertex_ai/files): stream batch upload to GCS instead of staging to a temp file
Addresses a disk-exhaustion concern: staging the full transformed batch body to
a local temp file before the GCS request meant an authenticated user could fill
the proxy's temp volume with large concurrent uploads (on top of Starlette's
input spool).
GCS's simple/media upload accepts chunked transfer-encoding, so stream the
transform straight to the single media request instead. Each block is produced
on a worker thread (the transform never runs on the event loop) and sent
chunked, so the body is neither buffered in memory nor written to disk, and the
upload is still one continuous request (no per-chunk round-trips, no 499). Drops
the temp-file staging, the tempfile/IO imports, and Content-Length computation.
Regression test asserts the upload streams (chunked transfer-encoding, no
Content-Length) and creates no temp file; mutation-verified that reintroducing
staging fails it.
Defines the TypedDict mirror of LiteLLM_ObjectPermissionBase under
litellm/types/object_permission.py so SDK-side modules can adopt the
type without violating the SDK-must-not-import-from-proxy layering
rule. litellm/proxy/_types.py re-exports it for existing callers.
Supports the validator-surface retypes in the preceding proxy refactor
commit.
Mirror the scalar `max_budget` guard in `_common_key_generation_helper`
for the per-window check: a CLI session token caller (carrying
`max_budget=None`) cannot set `budget_limits` on a personal key. Pass
`team_table` into the helper so it can detect the personal-key shape;
reject before the `delegation_ceiling is None` early return.
Four new regression tests cover the personal-key reject, the team-key
happy path, the team-key over-team-budget path, and the proxy-admin
exemption.
Enforce that every `budget_limits[*].max_budget` is a finite number;
applies to every caller including proxy admin and runs before the
role / ceiling checks. Six parametrized regression tests cover NaN /
+inf / -inf for both non-admin and admin callers.
increment_spend_counters was parallelized in #31578, but the dominant
per-request cost under high concurrency is the pre-call budget enforcement
in common_checks, which still ran a Redis-first get_current_spend per scope
(team, team windows, key windows, org, tag, user, team member, end user)
one sequential await after another inside the auth span.
The per-scope reads target distinct counter keys with no cross-scope
ordering dependency, so they now run concurrently under asyncio.gather.
Key metadata.tags injection still runs before the gather so the tag budget
check sees it, and every scope settles before the first error in
scope-priority order propagates, preserving the previous rejection semantics.
Resolves LIT-4090
/key/generate validated the caller's delegation ceiling against
data.max_budget only. The per-window entries in data.budget_limits
bypassed the check, so a non-admin caller could mint a key whose
1-day window vastly exceeded their own max_budget. The data.permissions
dict also went unvalidated for non-admin callers, so they could
self-grant capabilities like allow_pii_controls (and on Enterprise,
get_spend_routes).
Both gates now live in _common_key_generation_helper, covering
/key/generate and /key/service-account/generate. The existing empty
{} default on permissions still passes for non-admin callers.
* fix(databricks): split parallel tool calls so each tool message follows tool_calls
Databricks OpenAI-compatible serving (e.g. GPT models) 400s with "messages with
role 'tool' must be a response to a preceeding message with 'tool_calls'" when an
assistant turn makes parallel tool calls. LiteLLM faithfully sends one assistant
message holding all tool_calls followed by one 'tool' message per result, so every
result after the first is preceded by another 'tool' message rather than the
assistant tool_calls message, which Databricks rejects.
Re-emit each result immediately after an assistant message that carries only its
matching tool_call, turning assistant(tool_calls=[A, B]), tool(A), tool(B) into
assistant(tool_calls=[A]), tool(A), assistant(tool_calls=[B]), tool(B). The
rewrite is a no-op when the turn is already valid (single call), the group is
incomplete, or ids don't line up, so no tool call is ever dropped. Scoped to
non-Claude models, matching the existing OpenAI-shaped transformation path.
* style(databricks): use builtin list generics in parallel tool-call split
Switch the List[...] annotations introduced by _split_parallel_tool_calls
to lowercase list[...] so the UP006 strict-rule budget stays within its
ceiling.
* perf(ui): load virtual-keys team filter from the fast v2 endpoint
The virtual-keys table sourced all teams through fetchAllTeams, which
hits the unpaginated /team/list. On a proxy with 125 teams that call
takes ~9.5s, so the Team ID filter and the team-alias/budget columns sat
empty for that whole window. The key list itself does not carry
team_alias or team_max_budget, so the table genuinely needs a team
lookup and cannot just drop the fetch.
Add useAllTeams, which pages the fast /v2/team/list to completion (~0.6s
per 100-team page, so ~1.2s for 125 vs ~9.5s), and point VirtualKeysTable
at it instead of fetchAllTeams. The allTeams shape, the filter searchFn,
the column lookups, and the loading indicator are all unchanged; only the
source endpoint changes. fetchAllTeams stays for its other callers.
* test(ui): tighten team-filter test readability and robustness
Address adversarial review of the added tests. Rename the single-page
useAllTeams test to match what it asserts (one request for a one-page
result) rather than implying it guards the early-return, and drop the
unread, misleading total: 125 from the mock page response. Scope the
created_by alias-over-email assertion to the key's table row so it checks
the visible cell value; the hover popover that also holds the email is
portaled out of the row, so the previous document-wide negative assertion
was relying on antd's lazy popover mounting.
* fix(ui): scope useAllTeams cache by access token
The previous /team/list query keyed on accessToken, so a user switch in
the same SPA session produced a distinct cache entry. useAllTeams dropped
that, so team IDs and aliases could be briefly reused across users until
the staleTime expired. Put accessToken back in the query key to restore
per-identity isolation, and add a regression test that a token switch
triggers a refetch rather than serving the cached list.
* fix(ui): keep virtual-keys filters across delete and refresh (LIT-4080)
Filtering virtual keys by User ID and then deleting a key reset the filter to show all keys, and re-clicking Fetch did not re-apply it. The page ran two competing fetch paths: useKeys (React Query) fetched the page unfiltered while a separate useFilterLogic hook held its own filteredKeys list and, on any refresh, only re-applied Team and Organization client-side, silently dropping the User ID and Key Alias filters. Delete refreshed through the unfiltered useKeys path, so the filtered view collapsed back to everything
VirtualKeysTable now owns its filter state and feeds every filter (team, organization, key alias, user id, key hash) straight into the useKeys options, so the filters are part of the React Query key. Any refetch or invalidation re-runs the same filtered query, which makes the reset-on-delete bug structurally impossible. Free-text inputs are debounced with @tanstack/react-pacer, sorting and pagination are server-side, and changing a filter or sort resets to page 1
Delete now invalidates keyKeys.lists() from key_info_view, matching the create path, instead of prop-drilling a refetch; the window "storage" refetch effect is removed. The dual-path useFilterLogic hook (and its test) are deleted
Regression coverage: VirtualKeysTable threads an active User ID filter into the useKeys query and clears it on reset, useKeys encodes filter options in its query key so a filter change refetches, and key_info_view invalidates the keys list on delete
* refactor(ui): simplify virtual-keys table data flow
VirtualKeysTable now fetches its own teams and organizations via
useOrganizations and the existing all-teams query instead of taking them
as props, so the prop-drill through UserDashboard and the two page
callers (page.tsx, ApiKeysDashboard) is gone along with their redundant
organization state and fetch
Filter state collapses from a useState plus a useDebouncedState mirror
into a single source whose debounced copy is derived with
useDebouncedValue, and one typed toKeyListFilters adapter maps it to the
key/list query options. Behavior is unchanged; same 300ms debounce and
the same reset timing
The unused onSortChange/currentSort props and their sync effect are
removed since no caller passed them, leaving sorting fully internal
Adds a created_by_user alias-over-email regression test that fails if the
display precedence is swapped
* test(ui): add required last_active to useKeys mock fixtures
The KeyResponse type requires last_active, so the typed mockKeys
fixtures were missing it. Add it so the file type-checks cleanly.
* chore(ui): ratchet lint budgets after virtual-keys refactor
Deleting filter_logic.tsx and simplifying VirtualKeysTable lowered the
no-explicit-any (2026 to 2016) and complexity (128 to 127) counts, so the
eslint-metrics.json baseline was stale and failed the frontend-lint
budget gate. Regenerate it, and drop the now-dead filter_logic.tsx
suppression entry for the file this PR removed.
* fix(ui): show a loading state for data-backed filter dropdowns
The Team ID and Organization ID filters source their options from async
hooks (teams / organizations). While that data was still loading the
dropdowns rendered 'No results found', so they looked empty rather than
loading. Add an opt-in loading flag to FilterOption that the searchable
select surfaces as a spinner and a 'Loading...' empty state, and wire it
from the teams and organizations query loading states. While loading, the
filter no longer caches an empty initial-options list, so the real
options appear once the data arrives.