Commit graph

12328 commits

Author SHA1 Message Date
Devin AI
85c5153aec style: ruff format spend logs migration test
Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>
2026-08-18 23:27:40 +00:00
Devin AI
b6fe1d9866 fix(db): build the spend-logs api_key index without CONCURRENTLY so partitioned tables can migrate
Postgres rejects a concurrent index build on a partitioned parent, so the migration aborted on deployments that ran db_scripts/partition_spend_logs.sql. Also carry the index through the partition and unpartition runbooks, and cover the partitioned shape with a migration test that applies the shipped statement for real.

Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>
2026-08-18 23:23:59 +00:00
Shivam Rawat
ab3bcb95ef perf(db): index LiteLLM_SpendLogs on (api_key, startTime)
The Logs tab filters LiteLLM_SpendLogs by api_key inside a startTime
window. The table only had indexes on startTime, (startTime, request_id),
end_user and session_id, so the api_key predicate degraded to a full scan
of the window on both the pagination count and the paged select, re-run on
every page turn. On a multi-month table that pinned a customer's Aurora
writer near 100% CPU for ~50 minutes.

Built CONCURRENTLY: this is the largest table on most deployments, and a
blocking build would stall spend-log inserts for the whole build.
2026-08-18 14:17:17 -07:00
Mateo Wang
c6fad3683d
Merge pull request #36032 from Scott-Wilson-ZocDoc/litellm_fix_responses_tool_choice_auto
fix(responses): unwrap object-form tool_choice before calling the Responses API
2026-08-17 14:37:38 -07:00
Mateo Wang
c1fc5983ca
Merge pull request #36482 from irosh-colombage-ZocDoc2/fix/mcp-oauth-scoped-issuer
fix(mcp): scope authorization server issuer
2026-08-17 14:27:38 -07:00
yuneng-jiang
77c8a6452f
Merge pull request #36258 from BerriAI/litellm_/elastic-ishizaka-db31e6
feat(proxy): let USE_V2_MIGRATION_RESOLVER select the v2 migration resolver
2026-08-17 14:13:04 -07:00
mateo-berri
5ddff616dc fix(mcp): keep origin issuer on the openid-configuration alias 2026-08-17 13:41:36 -07:00
Mateo Wang
5c15c097b6
Merge pull request #36033 from Scott-Wilson-ZocDoc/litellm_fix_session_resume_thinking
fix(anthropic): stop emitting empty thinking blocks on the Responses adapter
2026-08-17 13:32:34 -07:00
Mateo Wang
e11fe1d6cc
Merge pull request #36781 from daniel-meismer-zocdoc/feature/request-logs-user-id-filter
feat(ui): add user ID request log filter
2026-08-17 13:32:26 -07:00
Mateo Wang
47eaff19d3
Merge pull request #36979 from Scott-Wilson-ZocDoc/fix/anthropic-responses-optional-tool-props
fix(anthropic): preserve optional Responses tool properties
2026-08-17 13:26:14 -07:00
Mateo Wang
cde5465e6f
Merge pull request #37203 from BerriAI/litellm_batches_metadata_type_400
fix(proxy): return 400 for non-object metadata and litellm_metadata instead of silent drop or 500
2026-08-17 13:19:33 -07:00
Mateo Wang
3b6e56716a
Merge pull request #36978 from Scott-Wilson-ZocDoc/fix/mcp-guardrail-usage-monitor
fix(guardrails): record MCP tool guardrail evaluations and blocks in …
2026-08-17 13:18:03 -07:00
mateo-berri
3cdf041827 Merge branch 'litellm_internal_staging' into fix/anthropic-responses-optional-tool-props 2026-08-17 13:10:29 -07:00
mateo-berri
98bfbb99d2 refactor(responses): drop commentary from the tool_choice fix 2026-08-17 13:04:45 -07:00
mateo-berri
11e2341fc9 Merge branch 'litellm_internal_staging' into feature/request-logs-user-id-filter 2026-08-17 13:01:19 -07:00
mateo-berri
c7b17b6615 Merge remote-tracking branch 'origin/litellm_internal_staging' into litellm_pr36482_head 2026-08-17 12:54:14 -07:00
mateo-berri
cba076715d Merge remote-tracking branch 'origin/litellm_internal_staging' into pr36032_drive 2026-08-17 12:52:12 -07:00
mateo-berri
ce4eaa16e8 Merge remote-tracking branch 'origin/litellm_internal_staging' into litellm_pr36033_drive 2026-08-17 12:50:32 -07:00
mateo-berri
a21de07b3f fix(proxy): drop every invalid metadata field before raising so failure hooks never see them 2026-08-17 12:44:46 -07:00
mateo-berri
f3b0cdca43 fix(proxy): keep ProxyException status codes on /v1/moderations instead of wrapping into 500 2026-08-17 12:32:30 -07:00
mateo-berri
c894697a2a fix(proxy): keep ProxyException status codes on /v1/messages instead of wrapping into 500 2026-08-17 12:21:12 -07:00
Mateo Wang
badd737526
Merge pull request #37199 from BerriAI/litellm_lit5657_batches_400_missing_fields
fix(proxy): return 400 naming the missing required param on POST /v1/batches
2026-08-17 12:20:25 -07:00
mateo-berri
57e1bf41c6 fix(proxy): return 400 for non-object metadata and litellm_metadata instead of silent drop or 500 2026-08-17 12:08:56 -07:00
mateo-berri
a9bb09905d test(batches): drop redundant section banner 2026-08-17 12:00:40 -07:00
Yassin Kortam
d5b91b94d3
test(e2e): replay a real tool-search assistant turn back to Bedrock Invoke (#36856)
The tool_search x bedrock_invoke cell only ever probed the first turn, so
nothing in the suite has sent a server_tool_use block back to a provider.
Every turn of a real Claude Code session after the first carries the
server_tool_use and tool_search_tool_result blocks the previous turn
produced, and that path was uncovered.

Adds probe_tool_search_multiturn, which takes the real assistant turn
back, answers any client-side tool_use with the id the model actually
emitted, and replays the whole thing as history with the tools still
declared. The assertion refuses to go green unless both server-tool
blocks made it into the replayed history, so a first turn truncated at
max_tokens reads as a failure instead of a vacuous pass.

The replay assertion's red paths never run in a green cell, so they get
markerless harness tests of their own alongside the existing
_builder_unit_tests tree.

No production code.
2026-08-17 11:59:26 -07:00
mateo-berri
e4ce526900 fix(proxy): return 400 naming the missing required param on POST /v1/batches 2026-08-17 11:55:06 -07:00
ryan-crabbe-berri
ad6a3a7b9e
fix(proxy): registry caches stop per-request tag and end-user Postgres reads in auth (#36801)
* fix(proxy): cache tag-name registry so unregistered request tags skip Postgres

Request tags are free-form attribution labels, so most have no LiteLLM_TagTable
row. get_tag_objects_batch never cached that absence: every tagged request ran
a find_many that came back empty, and under Prisma pool contention those
per-request queries queued for minutes inside user_api_key_auth.

Cache the bounded set of registered tag names under one aggregate key with the
management-object TTL. Uncached request tags are filtered against it before any
per-tag DB fetch, so unregistered tags cost zero DB reads on a warm path. An
empty registry is cached as a valid answer; DB errors are not cached and fall
back to the per-tag lookup; tables past TAG_REGISTRY_MAX_SIZE cache an overflow
sentinel that disables filtering. Tag create/update/delete endpoints now evict
the registry and per-tag keys and publish cross-worker invalidation (they
previously evicted nothing). The per-tag write-back also gains the management
TTL it was missing, and the hand-built tag:{name} key strings are replaced with
a shared builder.

* fix(proxy): skip per-request end-user DB reads via restricted-id registry

Every request carrying a user id ran get_end_user_object, and with high-cardinality
auto-created end-user rows (hundreds of thousands of ids, all restriction fields
NULL) the per-pod cache missed on nearly every request, so each one paid a Postgres
find_unique that queued behind the Prisma pool during background-job bursts. True
misses were never cached, and unknown ids paid the read twice per request.

Cache the bounded set of end-user ids that carry any restriction (blocked, budget,
region, default model, or object permission) under one aggregate key with the
management-object TTL. When an id misses the per-id cache and is absent from a
usable registry, get_end_user_object returns None with zero DB reads; restricted
ids keep today's fetch-and-cache path. The skip is bypassed whenever
litellm.max_end_user_budget_id is set (default budgets make unrestricted rows
behaviorally distinct from missing rows), validate_end_user_id_in_db is on
(existence checks need the row), or the token carries end_user_max_budget from
custom auth (the row's recorded spend seeds the budget counter). Empty registries
cache as a valid answer, DB errors are never cached, and oversized tables cache an
overflow sentinel that disables filtering. Customer create/update/block/delete now
evict the registry and per-id keys and publish cross-worker invalidation (they
previously evicted nothing), and the per-id write-back gains the management TTL it
was missing so Redis entries no longer live forever.

* refactor(proxy): single generic registry loader with error sentinel and single-flight

Code review follow-ups on the two registry caches. Registry DB errors now cache
the overflow sentinel for a short REGISTRY_ERROR_NEGATIVE_CACHE_TTL window and
log at warning, so a degraded Postgres stops paying the failing registry scan on
every request on top of the per-id fallback. Cold registry loads are single-flight
per worker behind per-registry locks with a recheck after acquire, so a TTL expiry
no longer fans out one full-table scan per in-flight request. The tag and end-user
loaders collapse into one _load_bounded_registry with per-entity fetch closures,
and the triplicated evict-then-broadcast protocol becomes one evict_and_broadcast
helper beside publish_auth_cache_invalidation, shared by the tag, customer, and
project eviction paths.

* chore(lint): suppress fail-safe registry excepts and ratchet BLE001 budget

* docs(proxy): trim registry cache commentary to single-line why docstrings

* fix(lint): move tag fetch return to else block to satisfy TRY300 budget
2026-08-17 18:52:13 +00:00
Brian Cox
3b6ae75f73
fix(bedrock): preserve cache token usage when invocationMetrics replace the usage block (#36878)
Bedrock invoke /v1/messages streaming reports cache_read_input_tokens and
cache_creation_input_tokens on message_stop.usage while attaching
amazon-bedrock-invocationMetrics to the same chunk. The stream decoder
rebuilt that chunk's usage block from inputTokenCount/outputTokenCount
alone, which exclude cache reads and writes, so the cache breakdown was
destroyed before _promote_message_stop_usage could surface it and cache
tokens were billed at $0. Merge instead of replace, and also map
cacheReadInputTokenCount/cacheWriteInputTokenCount when Bedrock reports
the cache itemization inside the invocation metrics.

Co-authored-by: Brian Cox <3924351+brian5021@users.noreply.github.com>
2026-08-17 11:28:35 -07:00
Yassin Kortam
9b7ed77fcc
fix(azure): rename max_tokens to max_completion_tokens for gpt-5-chat deployments (#36857)
Azure rejects the legacy `max_tokens` key for the whole gpt-5 name family, but
`AzureOpenAIGPT5Config.is_model_gpt_5_model` deliberately excludes `gpt-5-chat*`
so those deployments fall through to `AzureOpenAIConfig`, which sends `max_tokens`
verbatim and gets a 400 back on every request that carries it, `/health` probes
included.

One predicate was answering two independent questions. Split it: the new
`AzureOpenAIConfig.requires_max_completion_tokens` covers the whole gpt-5 name
family and drives only the rename, while `is_model_gpt_5_model` keeps keying
reasoning_effort, the temperature clamp and the dropped penalties off the
reasoning question, so #13781 stays fixed.
2026-08-17 11:27:55 -07:00
Yassin Kortam
ee08b63657
feat(bedrock): forward LiteLLM identity and metadata into Bedrock requestMetadata (#36861)
Adds an opt-in operator allow-list, litellm_settings::bedrock_request_metadata_fields, that forwards LiteLLM key, team and end-user identity plus client spend_logs_metadata into Bedrock request metadata so Bedrock spend can be grouped in AWS Cost Explorer.

Covers all three Bedrock surfaces: the Converse body requestMetadata field, and a signed X-Amzn-Bedrock-Request-Metadata header on Invoke chat completions and on Invoke /v1/messages, where the header is the only viable leg.

The resolver reads both metadata variable names, reserves the whole user_api_key_ prefix against caller-supplied keys, caps the client slot budget explicitly at 16 minus the reserved count, and drops rather than rejects auto-injected values that violate Bedrock constraints. Caller-supplied requestMetadata keeps its existing 400 semantics.

The request-metadata field and header are proxy-owned whenever forwarding is enabled. A caller-supplied value, reachable through the generic extra_headers passthrough, is dropped unconditionally and compared case-insensitively, and is replaced only by the proxy's own value, so identity in the AWS billing record cannot be forged. Absence of a resolved value still means absence on the wire rather than a fallback to the caller's. The guardrail headers keep their existing no-displace behaviour.
2026-08-17 11:27:46 -07:00
yucheng-berri
1139012b45
fix(guardrails): scan text on /guardrails/apply_guardrail for Azure Content Safety (#36894)
* fix(guardrails): scan text on /guardrails/apply_guardrail for Azure Content Safety

The two Azure Content Safety guardrails never implemented apply_guardrail, so the
endpoint fell through to the base no-op and answered 200 with the caller's text
echoed back, having scanned nothing.

Implementing that method also flips the proxy's unified-vs-native dispatch, which
would move request traffic off these guardrails' own hooks. Add an opt-out that
keeps every lifecycle event on the native hooks, so only the endpoint changes.

* test(guardrails): cover the remaining native-hook opt-out dispatch sites

Adds regression tests for the parallel post-call path, the MCP post-call hook, and
the policy engine step, so every read of the opt-out flag fails when removed.
2026-08-17 11:18:40 -07:00
Mateo Wang
2bc87ec3cc
Merge pull request #34067 from MUSE-CODE-SPACE/fix/batch-logging-null-output-file
fix(batches): don't crash logging when a completed batch has no output file
2026-08-17 10:01:02 -07:00
Mateo Wang
41c3133d0e
Merge pull request #34087 from ArjunPakhan/fix/bedrock-cancel-batch
fix(batches): support AWS Bedrock batch cancellation via `StopModelInvocationJob`
2026-08-17 10:00:57 -07:00
Mateo Wang
a1644eaf84
Merge pull request #36392 from BerriAI/devin_ai_fix_bedrock_batch_file_bytes_36388
fix(bedrock): report uploaded size in the FileObject returned by managed batch uploads
2026-08-17 10:00:53 -07:00
mateo-berri
5965648547 fix(proxy): close websocket cleanly when OpenAI credentials are missing 2026-08-16 14:40:37 -07:00
mateo-berri
4ba9d6b136 fix(proxy): expose url join helper at module level for websocket route 2026-08-16 14:29:21 -07:00
mateo-berri
81aefe4b3c Merge remote-tracking branch 'origin/litellm_internal_staging' into litellm_pr36151_ws_passthrough
# Conflicts:
#	litellm/proxy/pass_through_endpoints/pass_through_endpoints.py
#	tests/test_litellm/proxy/pass_through_endpoints/test_pass_through_endpoints.py
2026-08-16 14:21:57 -07:00
mateo-berri
862f33bbaa fix(proxy): negotiate client subprotocol on OpenAI websocket passthrough 2026-08-16 14:11:12 -07:00
mateo-berri
a258b2b130 fix(proxy): harden OpenAI websocket passthrough
- decode upstream first frame as utf-8 instead of ascii
- reject model-restricted keys at connect to match HTTP model enforcement
- log the actual request path for /openai_passthrough traffic
2026-08-16 13:51:50 -07:00
mateo-berri
1b13957776 fix(bedrock): treat ConflictException on stop as idempotent cancel 2026-08-16 13:44:27 -07:00
Yuneng Jiang
de2b220c36
test(e2e/ui): assert the log drawer chevrons by their lucide classes
The log details drawer moved off Ant Design in 03d2b16bc, so its section
header renders lucide ChevronUp/ChevronDown rather than antd's UpOutlined
and DownOutlined. The collapse test still waited on .anticon-up and
.anticon-down, which no longer exist anywhere under view_logs, so it
failed on every run and burned all three attempts identically.

Point the three assertions at .lucide-chevron-up and .lucide-chevron-down,
matching how the dashboard's other suites address lucide icons.
2026-08-15 17:50:40 -07:00
mateo-berri
904ff9efa7 Merge branch 'litellm_internal_staging' into devin_ai_fix_bedrock_batch_file_bytes_36388
Some checks failed
Terraform Modules / fmt, validate, test (aws) (push) Has been cancelled
Terraform Provider / gofmt, vet, build, test (push) Has been cancelled
Terraform Provider / Provider endpoints vs proxy OpenAPI schema (push) Has been cancelled
Resolves the transform_create_file_response conflict by keeping the
_uploaded_object_size handoff over the response Content-Length read,
and adds the rebind-ok justification LIT011 now requires for the
upload-size litellm_params handoff after the base budget ratcheted.
2026-08-15 17:31:52 -07:00
mateo-berri
4114f907ea fix(bedrock): reraise cancel validation errors for non-terminal jobs, allow bedrock in acancel_batch typing 2026-08-15 17:25:16 -07:00
mateo-berri
57e946f279 Merge origin/litellm_internal_staging into fix/bedrock-cancel-batch 2026-08-15 17:18:08 -07:00
yuneng-jiang
91aee78e78
Merge pull request #37065 from BerriAI/litellm_/nice-wilson-9fbed6
test(e2e): assert provider error shape instead of pinned prose
2026-08-15 17:15:30 -07:00
yuneng-jiang
1968562733
Merge pull request #37059 from BerriAI/litellm_/circleci-pipeline-triage-9b92e5
test: unstick the suites CircleCI is failing on
2026-08-15 17:12:46 -07:00
Yuneng Jiang
4df93c713c
Merge remote-tracking branch 'origin/litellm_internal_staging' into litellm_/nice-wilson-9fbed6 2026-08-15 16:59:36 -07:00
Mateo Wang
13d94ec546
Merge pull request #36869 from BerriAI/litellm_lit002_typeddict_dict_literals
feat(lint): exempt TypedDict-annotated dict literals from LIT002
2026-08-15 16:35:24 -07:00
devin-ai-integration[bot]
74a1beda77
fix(panw_prisma_airs): scan tool call args as plain text, not a tool_event (#37038)
* fix(panw_prisma_airs): scan tool call args as plain text, not a tool_event

Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>

* refactor(panw_prisma_airs): type the tool call argument extractor

Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>

* test(panw_prisma_airs): cover tool call error fallback and dict masking paths

Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>

* fix(panw_prisma_airs): scan tool names with args and tolerate custom tool calls

Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>

* fix(panw_prisma_airs): scan tool call arguments that arrive already parsed

The tool call slice types arguments as a string, so a client posting parsed JSON
failed validation and the whole tool call, name included, read as unscannable and
was skipped without ever reaching AIRS. The OpenAI request path forwards
client-supplied tool_calls verbatim, so that shape is reachable.

Coerce non-string arguments instead of rejecting them, so the content is scanned.

* fix(panw_prisma_airs): route tool-block masked data by scan side, not by key name

Merging #37036 (already on staging) with this PR produces no conflict and a
silent bug. #37036 withholds prompt_masked_data on response-side tool blocks,
which was right while tool calls went out as a request-side tool_event: AIRS
reported the model's arguments under that key. This PR scans tool calls as
ordinary prompt/response text, so the side of the scan now decides which key
holds what. The model's arguments arrive under response_masked_data, already
covered by _CLIENT_HIDDEN_SCAN_FIELDS, and prompt_masked_data goes back to
being the caller's own input -- one of the audit fields LIT-5638 asks for.

Left as merged, a response-side tool block drops that field with nothing to
flag it.

- Tool-path block branch calls _build_error_detail without also_hide
- also_hide parameter removed; after this change it has no callers
- Regression test asserts both directions: model output withheld, caller
  input preserved. It fails against the auto-merged combination.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>

* fix(panw_prisma_airs): a wrong-typed tool name must not suppress the scan

_ToolCallFunctionSlice types name as str, and _get_tool_call_function turns any
ValidationError into (None, None), which _scan_tool_calls_for_guardrail reads as
an unscannable tool call and skips. So a client posting "name": 123 keeps its
arguments off the wire to AIRS entirely -- no error, no log, no block. The
OpenAI request path forwards client tool_calls verbatim, so this is reachable by
any caller holding a valid key.

_coerce_arguments already existed for exactly this failure mode on the sibling
field. Widening it to cover name closes the gap:

  name='transfer_funds'   AIRS called: 1x   args scanned: True
  name=123 (int)          AIRS called: 0x   args scanned: False   <- before
  name=123 (int)          AIRS called: 1x   args scanned: True    <- after

Reported by Cursor Bugbot on fd9f6396e5.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>

---------

Co-authored-by: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>
Co-authored-by: Yucheng Zhu <yucheng@berri.ai>
Co-authored-by: Claude Opus 5 (1M context) <noreply@anthropic.com>
2026-08-15 23:14:43 +00:00
Yuneng Jiang
481ab07ca7
test(e2e): skip the bedrock web search cell the stack cannot provision
This cell needs the websearch_interception callback and a declared search
backend, both listed in its own module docstring. The ephemeral e2e stack
ships neither, so the request falls through to the bedrock transformation
and takes the by-design 400 that tells you to enable interception.

The cell has never been green here: the error path merged about an hour and
a half before the cell did, and the last full suite to pass predates the
cell entirely. Skip it with the reason recorded so the run reports honestly
instead of carrying a permanent red, and unskip once the stack ships the
config the docstring already spells out.
2026-08-15 16:09:02 -07:00