Commit graph

9968 commits

Author SHA1 Message Date
Yujong Lee
6805d01709 fix(vector-store): preserve aliases with embedding config 2026-09-01 15:33:18 -07:00
Yujong Lee
5635811726 fix(vector-store): route embeddings through router 2026-09-01 15:33:18 -07:00
Yujong Lee
bfa5eac76b fix(vector-store): resolve embedding aliases for search 2026-09-01 15:33:18 -07:00
ryan-crabbe-berri
9c417ba08b
Merge pull request #38969 from emerzon/litellm_strict_order_fallback
fix(router): keep order fallback on the requested order level
2026-09-01 15:21:59 -07:00
ryan-crabbe-berri
55d638412b fix: stop a cleared Team field from blocking personal key creation
Clearing the Team combobox in the Create Key modal left team_id set to an
empty string, so /key/generate treated the request as team key generation
and failed with a team-not-found error for non-admin members.

TeamDropdown now emits null on clear, and GenerateKeyRequest normalizes an
empty team_id to None so the request runs the personal key path.
2026-09-01 15:17:00 -07:00
ryan-crabbe-berri
45fa78470d test(router): inject the upstream client instead of mutating litellm.aclient_session
The text-completion wire test set litellm.aclient_session, which the
test-quality gate (TQ005) flags as a process-wide global write. Pass an
AsyncOpenAI client through the router's client kwarg instead, so the test
owns its transport and needs no cache flush or global restore.

Claude-Session: https://claude.ai/code/session_01XKkTFa6g7Rmd6vtHL91GMn
2026-09-01 15:14:56 -07:00
devin-ai-integration[bot]
97dbd8efcb
fix(docker): add public Wolfi apk repo to runtime image (#39033)
* fix(docker): add public Wolfi apk repo to runtime image

Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>

* test(docker): accept quote variants in Wolfi repo assertion

Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>

---------

Co-authored-by: yassin <yassin@berri.ai>
Co-authored-by: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>
2026-09-01 15:11:15 -07:00
devin-ai-integration[bot]
846900320e
feat(alerting): slack alerts for per-user daily/monthly spend thresholds and spend anomaly detection (#38438)
* feat(alerting): slack alerts for per-user daily/monthly spend thresholds and spend anomaly detection

Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>

* test(alerting): use specific ValidationError matches in config rejection test

Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>

* fix(proxy): tolerate mocked slack alerting args when scheduling user spend scan

Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>

* fix(alerting): reject non-finite values in user spend alert settings

Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>

---------

Co-authored-by: yassin <yassin@berri.ai>
Co-authored-by: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>
2026-09-01 15:09:03 -07:00
devin-ai-integration[bot]
5988d93fed
fix(logging): guarantee max_parallel_requests slot release when streaming logging fails (#39093)
* fix(logging): guarantee max_parallel_requests slot release when stream logging fails

Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>

* test(logging): cover guardrail branch of streaming logging hook failure isolation

Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>

---------

Co-authored-by: yassin <yassin@berri.ai>
Co-authored-by: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>
2026-09-01 15:08:11 -07:00
devin-ai-integration[bot]
f9c6eda909
fix(cli): quote the Claude Code apiKeyHelper for cmd.exe on Windows (#39174)
* fix(cli): quote the Claude Code apiKeyHelper for cmd.exe on Windows

lite up and lite login --config-claude wrote the helper command with
POSIX shlex quoting, so a backslashed Windows install path came out
wrapped in single quotes that cmd.exe and PowerShell take literally.
Quote every token with the cmd.exe rules already used for agent shims
when running on Windows, and keep the POSIX output unchanged elsewhere.

Resolves LIT-6627

Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>

* test(cli): split the Windows apiKeyHelper with cmd.exe and C runtime rules

The invocation test pulled tokens back out with a regex, which cannot see
the doubled quotes or the percent guard quote_for_cmd emits. Model the two
parsers that read the helper on Windows instead and check argv round
trips for backslashed, spaced, metacharacter, percent and quoted tokens

---------

Co-authored-by: yassin <yassin@berri.ai>
Co-authored-by: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>
2026-09-01 15:08:00 -07:00
ryan-crabbe-berri
0b89c59be2 fix(router): consume _target_order at deployment selection so it never reaches a provider
Reading _target_order with .get left it in the request kwargs after selection, and only
nine provider boundaries stripped it. _atext_completion and _aadapter_completion spread
the raw kwargs, so an order-2 hop on /completions sent _target_order upstream, which
real providers reject as an unknown argument. Popping at selection strips it for every
path in one place; the PR's retry-keeping test already passed with pop because each retry
hands the callee its own kwargs copy.

Claude-Session: https://claude.ai/code/session_01XKkTFa6g7Rmd6vtHL91GMn
2026-09-01 15:02:01 -07:00
mateo
4da12795fc fix: filter deployment default API key limits
Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>
2026-09-01 21:50:45 +00:00
mateo-berri
a38dfecd96 Merge remote-tracking branch 'origin/litellm_internal_staging' into litellm_fix_stream_modify_response_chunks 2026-09-01 14:49:06 -07:00
mateo-berri
d59fcda8af fix(rerank): adopt declared authenticating providers in arerank instead of resolving them
get_llm_provider runs the OAuth device flow for github_copilot and chatgpt,
so calling it on the event loop before the executor dispatch let an
authenticated caller block the loop for the length of the polling window.
Adopt the declared provider via declared_authenticating_provider, matching
the metadata callers in utils.py, and only resolve for everything else.
2026-09-01 14:47:59 -07:00
Mateo Wang
b52b5d9421
Merge pull request #38734 from BerriAI/litellm_fix_bedrock_buffered_responses_stream
fix(bedrock): route streamed responses-API output through the unified guardrail
2026-09-01 14:44:41 -07:00
Mateo Wang
d3a399f6a7
Merge pull request #38870 from BerriAI/litellm_fix_azure_chat_anyof_tool_schema
fix(azure): flatten top-level tool schema combinators on Azure chat completions
2026-09-01 14:43:04 -07:00
Mateo Wang
981c5924a3
Merge pull request #39000 from BerriAI/litellm_fix_supported_openai_params_router_alias
fix(proxy): resolve router model aliases in /utils/supported_openai_params
2026-09-01 14:40:04 -07:00
ryan-crabbe-berri
ac964918c5 fix(router): strip _target_order at every provider boundary via a shared helper 2026-09-01 14:34:48 -07:00
yucheng-berri
5767a2da0f
fix(mcp): follow tools/list pagination from upstream servers (#39172)
* fix(mcp): follow tools/list pagination from upstream servers

Adopts BerriAI/litellm#32244 by Jupiter363 onto litellm_internal_staging
with merge conflicts resolved

* fix(mcp): degrade buggy pagination to partial results and bound the preview walk

A repeated nextCursor now returns the tools collected so far instead of
discarding every page with a RuntimeError, an empty-string cursor is treated
as terminal, load_mcp_tools shares the same pagination walk instead of
returning only the first page, and the tools/list preview is bounded by the
listing timeout instead of only the per-request timeout times the page cap

* fix(mcp): annotate deliberate rebind for the preview timeout scope

* fix(mcp): bound the shared pagination walk with an overall listing deadline

The per-request session read timeout restarts on every page, so direct SDK
callers of list_tools and load_mcp_tools could run up to the page cap with
no overall bound. The walk now returns the tools collected so far when
max(MCP_CLIENT_TIMEOUT, MCP_TOOL_LISTING_TIMEOUT) expires

* fix(mcp): let a per-server timeout extend the pagination deadline

MCPClient carries a per-server timeout that can exceed the global default;
list_tools now passes max(self.timeout, MCP_TOOL_LISTING_TIMEOUT) into the
shared walk so a deliberately slow server is not silently truncated at the
global deadline

* fix(mcp): honor per-server timeouts in the preview deadline and test the walk sessionless

The preview deadline now extends with the created client's own timeout, and
the pagination walk's cap, repeated-cursor, and empty-cursor cases are tested
directly against the helper instead of through patched SDK internals

* fix(mcp): forward the preview request's per-server timeout to the temporary server model

The tools preview built its temporary MCPServer without the request's
timeout field, so the client factory always fell back to the global
default and a per-server timeout could never extend the preview's
listing deadline (or its per-request timeout).
2026-09-01 14:34:14 -07:00
devin-ai-integration[bot]
558f42e304
fix(proxy): default max_idle_connection_lifetime to 60s on DB URLs (#39134)
* fix(proxy): default max_idle_connection_lifetime to 60s on DB URLs

Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>

* fix(proxy): regenerate schema.d.ts for database_max_idle_connection_lifetime

Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>

* fix(proxy): keep URL-pinned max_idle_connection_lifetime over config value

Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>

---------

Co-authored-by: yassin <yassin@berri.ai>
Co-authored-by: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>
2026-09-01 14:16:51 -07:00
Mateo Wang
deb67ce6e2 Merge remote-tracking branch 'origin/litellm_internal_staging' into litellm_fix_bedrock_buffered_responses_stream
# Conflicts:
#	tests/test_litellm/proxy/guardrails/guardrail_hooks/test_bedrock_guardrails.py
2026-09-01 14:11:29 -07:00
mateo-berri
b0aa1506fc Merge branch 'litellm_internal_staging' into litellm_fix_supported_openai_params_router_alias 2026-09-01 14:11:17 -07:00
mateo-berri
b311c24253 Merge remote-tracking branch 'origin/litellm_internal_staging' into litellm_fix_azure_chat_anyof_tool_schema 2026-09-01 14:08:34 -07:00
mateo-berri
28d0ac5339 fix(router): guard the declared-provider check for requests without a model 2026-09-01 14:06:14 -07:00
mateo-berri
0f35328a21 Merge remote-tracking branch 'origin/litellm_internal_staging' into litellm_fix_azure_chat_anyof_tool_schema
# Conflicts:
#	tests/test_litellm/litellm_core_utils/prompt_templates/test_litellm_core_utils_prompt_templates_common_utils.py
2026-09-01 13:59:59 -07:00
mateo-berri
848a3edbe8 merge: sync litellm_internal_staging into litellm_rerank_provider_error_body 2026-09-01 13:55:04 -07:00
mateo-berri
018d8ae9eb Merge remote-tracking branch 'origin/litellm_internal_staging' into HEAD 2026-09-01 13:49:51 -07:00
Yassin Kortam
aab9abdd1d
fix: keep litellm_credential_name from LiteLLM Params JSON and gate stored credential attach to proxy admins (#39047)
* fix(ui): keep litellm_credential_name from LiteLLM Params JSON when no credential is selected

Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>

* test(ui): drop null litellm_credential_name from AddModelPanel payload fixture

Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>

* fix(ui): validate JSON litellm_credential_name against accessible credentials

Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>

* fix(proxy): enforce proxy-admin-only credential attachment on model create/update

Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>

* fix(proxy): raise ProxyException for unauthorized credential attach and gate /model/update

Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>

* refactor(proxy): fold credential-change detection into can_user_attach_credential to satisfy complexity budget

Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>

* fix(proxy): decrypt stored credential name before unchanged-credential comparison

Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>

* test(proxy): cover credential attach rejection on add_new_model and patch_model

Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>

* test(proxy): annotate proxy-global patches with test-quality suppressions

Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>

---------

Co-authored-by: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>
2026-09-01 13:46:18 -07:00
Mateo Wang
8ce02c0666
Merge pull request #39184 from BerriAI/litellm_fable_5_1_structured_output 2026-09-01 13:41:22 -07:00
Mateo Wang
44b595f0bb
Merge pull request #39136 from BerriAI/litellm_lit6611_requested_model_label_cap
fix(prometheus): bound requested_model label cardinality on client failure paths
2026-09-01 13:35:33 -07:00
yuneng-jiang
6028afa2ab
Merge pull request #39178 from BerriAI/litellm_revert_v2_migration_resolver_default
revert: default the proxy back to the v1 migration resolver
2026-09-01 13:32:18 -07:00
mateo-berri
95a0ca3ba5 Merge remote-tracking branch 'origin/litellm_internal_staging' into litellm_gemini_thinking_content 2026-09-01 13:31:55 -07:00
yucheng-berri
cdb1245e74
fix(s3): bound s3 object keys and download filenames for long Responses API ids (#39164)
* fix(s3): bound object keys and download filenames to s3 limits

Long OpenAI-compatible Responses API ids pushed the s3 object key past s3's
1024 UTF-8 byte cap, so the PUT failed with a 400 and the log record was
dropped. Keys that still fit are unchanged, byte for byte. An oversized one
now keeps a readable head of the file name and appends the sha256 of the full
name. A configured path/alias prefix that is long enough to overflow on its
own keeps whole leading path segments, so a prefix-scoped IAM policy or
lifecycle rule still matches, and ends in a short digest of the full
configured value so two operators do not land in the same folder.

The Content-Disposition filename carried the same unbounded id and hit s3's
2048 byte metadata-header cap, so the upload still failed with
MetadataTooLarge once the key was bounded. It is bounded the same way, head
plus digest, so two records downloaded from the console stay distinct files.

The full response id stays in the uploaded JSON payload.

* fix(s3): keep the configured prefix whole and spend the whole key budget

Shorten the response id first and only trim the operator's configured prefix
when the prefix itself is what does not fit, so prefix scoped IAM policies and
lifecycle rules keep matching. Trim by bytes rather than whole segments so the
longest possible string prefix survives, and route the audit log key through
the same shared builder.

* chore(s3): trim the comments and docstrings the review flagged

Keep the two external facts that are not visible from the code, the 1024 byte
object key cap and the 2048 byte metadata header cap, and drop the rest.
2026-09-01 13:30:02 -07:00
mateo-berri
8b0441a628 fix(vector_stores): block caller-supplied embedding selection params on query surfaces 2026-09-01 13:29:25 -07:00
Mateo Wang
286a754999
Merge pull request #38839 from BerriAI/litellm_fix_chat_anyof_tool_schema
fix(openai): flatten top-level tool schema combinators on chat completions
2026-09-01 13:25:03 -07:00
mateo
d568bbe58d fix(bedrock): use tool fallback without forced tool_choice for claude-fable-5-1 structured output
Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>
2026-09-01 20:19:46 +00:00
Mateo Wang
c913b09e66
Merge pull request #31725 from Srivatsa03/time-based-cost-pricing
feat(cost): support time-based off-peak pricing in cost calculation
2026-09-01 13:18:30 -07:00
mateo-berri
8b5ae3da9d test(vector_stores): package the suite dir to avoid test_main basename collision 2026-09-01 13:15:51 -07:00
mateo-berri
60de5468d7 fix(cost): inherit the backend's raw cost map entry for off-peak-only deployments
Filtering copied fields by name dropped companion billing rules like
web_search_billing_unit and the regional uplift multipliers, so grounding
and uplifts billed differently through the deployment entry. Copy the
backend's raw litellm.model_cost entry wholesale instead, which also
removes the synthesized-zero special case since the raw entry only holds
real values.
2026-09-01 13:05:25 -07:00
mateo-berri
ac53ea0756 Merge remote-tracking branch 'origin/litellm_internal_staging' into litellm_stream_usage_cost_default 2026-09-01 13:03:18 -07:00
mateo
3e3e4d6970 fix(anthropic): use native structured output for claude-fable-5-1 on Vertex AI and Bedrock Invoke
Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>
2026-09-01 20:01:03 +00:00
Mateo Wang
19ca4cd4a9
Merge branch 'litellm_internal_staging' into litellm_fix_supported_openai_params_router_alias 2026-09-01 12:57:36 -07:00
Yuneng Jiang
fd72a39b1b
revert: default the proxy back to the v1 migration resolver
This reverts merge commit 2b1bd20834 (#31125)

Two CircleCI jobs on the staging-to-main promotion went red the moment
that PR landed. proxy_multi_instance_tests boots two proxies against one
database, and both now race the same migration:

  Error: P3018 A migration failed to apply
  Database error code: 40P01, deadlock detected
  Process 73 waits for ShareLock on virtual transaction 4/11;
  blocked by process 75. Process 75 waits for ExclusiveLock on
  advisory lock [16384,0,72707369,1]; blocked by process 73

Neither proxy comes up, so the job times out after 300s waiting on
localhost:4000. The same wait took 36.5s on the last green run

Timeline: #31125 merged at 18:46:14Z and the failing run started at
18:49:59Z. The merge commit is not an ancestor of the last green
revision (194a3cc) and is an ancestor of the first failing one
(01de2837)

The v2 resolver was meant to avoid exactly this class of contention, so
the deadlock looks like a bug in it rather than a reason to abandon it.
Putting the default back to v1 buys time to fix it without holding up
the release
2026-09-01 12:56:08 -07:00
mateo-berri
babe7816ad fix(rag): store-wins merge, single lookup, allowlisted search params
rag_query reuses the store resolved during authorization instead of a
second registry lookup, merges registry data store-wins so callers
cannot override a managed store's provider or credentials, and logs ids
instead of the merged config, which can carry resolved credentials.
aquery forwards only allowlisted retrieval_config keys to vector store
search, keeping caller-supplied connection overrides like api_base and
api_key away from the search call
2026-09-01 12:52:45 -07:00
mateo-berri
3914de24ef fix(router): route model-less sync vector store calls to the SDK
_generic_api_call_with_fallbacks requires a model, so sync
vector_store_search and vector_store_create raised a TypeError whenever
the call carried no model. Model-less calls now go directly to the SDK
function, with the router injected for search, matching the async
wrapper's behavior
2026-09-01 12:52:45 -07:00
Mateo Wang
8ae072b501
Merge pull request #39147 from BerriAI/litellm_openai_drop_toolless_tool_choice
fix(openai): drop tool_choice when request has no tools on chat completions
2026-09-01 12:50:38 -07:00
Mateo Wang
ab8b8deb39
Merge pull request #39144 from BerriAI/litellm_fix_responses_tool_call_id_shape
fix(responses): tool call id shape breaks gpt-5 -> claude fallback conversations
2026-09-01 12:50:35 -07:00
mateo-berri
0ec3e936b7 fix(cost): inherit the backend's full price structure for off-peak-only deployments
Copying only the flat token rates dropped threshold, tiered, service-tier,
cache, character, and per-second rates from peak-hour billing once cost
lookup switched to the deployment entry, and get_model_info synthesizes
zero flat rates for backends without one, which would have marked
tiered-only backends explicitly priced free. Copy every price-bearing
field instead, deep-copied, rejecting the synthesized zeros the way
_inherit_builtin_tiered_output_rate already does.
2026-09-01 12:47:26 -07:00
mateo-berri
d3dab8e294 fix(rerank): map provider errors with the resolved provider on sync and async paths 2026-09-01 12:46:04 -07:00
mateo-berri
4567fc784c fix: close non-message open items as incomplete when a stream is blocked 2026-09-01 12:45:23 -07:00