Commit graph

46248 commits

Author SHA1 Message Date
devin-ai-integration[bot]
846900320e
feat(alerting): slack alerts for per-user daily/monthly spend thresholds and spend anomaly detection (#38438)
* feat(alerting): slack alerts for per-user daily/monthly spend thresholds and spend anomaly detection

Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>

* test(alerting): use specific ValidationError matches in config rejection test

Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>

* fix(proxy): tolerate mocked slack alerting args when scheduling user spend scan

Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>

* fix(alerting): reject non-finite values in user spend alert settings

Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>

---------

Co-authored-by: yassin <yassin@berri.ai>
Co-authored-by: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>
2026-09-01 15:09:03 -07:00
devin-ai-integration[bot]
5988d93fed
fix(logging): guarantee max_parallel_requests slot release when streaming logging fails (#39093)
* fix(logging): guarantee max_parallel_requests slot release when stream logging fails

Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>

* test(logging): cover guardrail branch of streaming logging hook failure isolation

Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>

---------

Co-authored-by: yassin <yassin@berri.ai>
Co-authored-by: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>
2026-09-01 15:08:11 -07:00
devin-ai-integration[bot]
f9c6eda909
fix(cli): quote the Claude Code apiKeyHelper for cmd.exe on Windows (#39174)
* fix(cli): quote the Claude Code apiKeyHelper for cmd.exe on Windows

lite up and lite login --config-claude wrote the helper command with
POSIX shlex quoting, so a backslashed Windows install path came out
wrapped in single quotes that cmd.exe and PowerShell take literally.
Quote every token with the cmd.exe rules already used for agent shims
when running on Windows, and keep the POSIX output unchanged elsewhere.

Resolves LIT-6627

Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>

* test(cli): split the Windows apiKeyHelper with cmd.exe and C runtime rules

The invocation test pulled tokens back out with a regex, which cannot see
the doubled quotes or the percent guard quote_for_cmd emits. Model the two
parsers that read the helper on Windows instead and check argv round
trips for backslashed, spaced, metacharacter, percent and quoted tokens

---------

Co-authored-by: yassin <yassin@berri.ai>
Co-authored-by: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>
2026-09-01 15:08:00 -07:00
ryan-crabbe-berri
0b89c59be2 fix(router): consume _target_order at deployment selection so it never reaches a provider
Reading _target_order with .get left it in the request kwargs after selection, and only
nine provider boundaries stripped it. _atext_completion and _aadapter_completion spread
the raw kwargs, so an order-2 hop on /completions sent _target_order upstream, which
real providers reject as an unknown argument. Popping at selection strips it for every
path in one place; the PR's retry-keeping test already passed with pop because each retry
hands the callee its own kwargs copy.

Claude-Session: https://claude.ai/code/session_01XKkTFa6g7Rmd6vtHL91GMn
2026-09-01 15:02:01 -07:00
Mateo Wang
bad55da9bf
Merge pull request #38872 from BerriAI/litellm_fix_viewer_add_model_tab
fix(ui): hide model write affordances from view-only admin sessions
2026-09-01 14:52:54 -07:00
ryan-crabbe-berri
af11db9fe5 test(e2e): cover retry-on-timeout and the context-window fallback
Two P0 rows in the reliability coverage registry had no test.

reliability.retry.timeout.succeeds_within_retries gets a new file. The model
group is a pair: an always-timing-out deployment holding all of the group's
shuffle weight, and a healthy backup at weight 0. The weighted pick always opens
on the timing-out one, its first Timeout benches it via an allowed_fails_policy
of TimeoutErrorAllowedFails 0, and the retry falls through to the only
deployment left, so the outcome is a completion plus a reported retry with no
random first pick in the middle.

reliability.fallback.context_window.routes_to_fallback joins the existing
fallbacks spec. It registers a genuinely small-context OpenAI deployment, sends
a prompt past its limit so the provider refuses it on length, and reroutes with
context_window_fallbacks, which is the setting that handles that refusal rather
than plain fallbacks.

Both drive real provider calls through router_settings_override, so no config
change and no second proxy is needed. Reliability & Performance goes 16/36 to
18/36.

Claude-Session: https://claude.ai/code/session_01QvQzYztinxj8ZuD5YxbVdL
2026-09-01 14:51:15 -07:00
mateo
4da12795fc fix: filter deployment default API key limits
Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>
2026-09-01 21:50:45 +00:00
mateo-berri
a38dfecd96 Merge remote-tracking branch 'origin/litellm_internal_staging' into litellm_fix_stream_modify_response_chunks 2026-09-01 14:49:06 -07:00
Mateo Wang
b52b5d9421
Merge pull request #38734 from BerriAI/litellm_fix_bedrock_buffered_responses_stream
fix(bedrock): route streamed responses-API output through the unified guardrail
2026-09-01 14:44:41 -07:00
Mateo Wang
d3a399f6a7
Merge pull request #38870 from BerriAI/litellm_fix_azure_chat_anyof_tool_schema
fix(azure): flatten top-level tool schema combinators on Azure chat completions
2026-09-01 14:43:04 -07:00
Mateo Wang
981c5924a3
Merge pull request #39000 from BerriAI/litellm_fix_supported_openai_params_router_alias
fix(proxy): resolve router model aliases in /utils/supported_openai_params
2026-09-01 14:40:04 -07:00
ryan-crabbe-berri
ac964918c5 fix(router): strip _target_order at every provider boundary via a shared helper 2026-09-01 14:34:48 -07:00
yucheng-berri
5767a2da0f
fix(mcp): follow tools/list pagination from upstream servers (#39172)
* fix(mcp): follow tools/list pagination from upstream servers

Adopts BerriAI/litellm#32244 by Jupiter363 onto litellm_internal_staging
with merge conflicts resolved

* fix(mcp): degrade buggy pagination to partial results and bound the preview walk

A repeated nextCursor now returns the tools collected so far instead of
discarding every page with a RuntimeError, an empty-string cursor is treated
as terminal, load_mcp_tools shares the same pagination walk instead of
returning only the first page, and the tools/list preview is bounded by the
listing timeout instead of only the per-request timeout times the page cap

* fix(mcp): annotate deliberate rebind for the preview timeout scope

* fix(mcp): bound the shared pagination walk with an overall listing deadline

The per-request session read timeout restarts on every page, so direct SDK
callers of list_tools and load_mcp_tools could run up to the page cap with
no overall bound. The walk now returns the tools collected so far when
max(MCP_CLIENT_TIMEOUT, MCP_TOOL_LISTING_TIMEOUT) expires

* fix(mcp): let a per-server timeout extend the pagination deadline

MCPClient carries a per-server timeout that can exceed the global default;
list_tools now passes max(self.timeout, MCP_TOOL_LISTING_TIMEOUT) into the
shared walk so a deliberately slow server is not silently truncated at the
global deadline

* fix(mcp): honor per-server timeouts in the preview deadline and test the walk sessionless

The preview deadline now extends with the created client's own timeout, and
the pagination walk's cap, repeated-cursor, and empty-cursor cases are tested
directly against the helper instead of through patched SDK internals

* fix(mcp): forward the preview request's per-server timeout to the temporary server model

The tools preview built its temporary MCPServer without the request's
timeout field, so the client factory always fell back to the global
default and a per-server timeout could never extend the preview's
listing deadline (or its per-request timeout).
2026-09-01 14:34:14 -07:00
ryan-crabbe-berri
c7212e7fe2 refactor(router): drop _target_order via pop to satisfy the mutable-collection budget 2026-09-01 14:16:57 -07:00
devin-ai-integration[bot]
558f42e304
fix(proxy): default max_idle_connection_lifetime to 60s on DB URLs (#39134)
* fix(proxy): default max_idle_connection_lifetime to 60s on DB URLs

Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>

* fix(proxy): regenerate schema.d.ts for database_max_idle_connection_lifetime

Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>

* fix(proxy): keep URL-pinned max_idle_connection_lifetime over config value

Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>

---------

Co-authored-by: yassin <yassin@berri.ai>
Co-authored-by: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>
2026-09-01 14:16:51 -07:00
Mateo Wang
deb67ce6e2 Merge remote-tracking branch 'origin/litellm_internal_staging' into litellm_fix_bedrock_buffered_responses_stream
# Conflicts:
#	tests/test_litellm/proxy/guardrails/guardrail_hooks/test_bedrock_guardrails.py
2026-09-01 14:11:29 -07:00
mateo-berri
b0aa1506fc Merge branch 'litellm_internal_staging' into litellm_fix_supported_openai_params_router_alias 2026-09-01 14:11:17 -07:00
ryan-crabbe-berri
2de33555ea style: run ruff format on fallback_event_handlers 2026-09-01 14:08:40 -07:00
mateo-berri
b311c24253 Merge remote-tracking branch 'origin/litellm_internal_staging' into litellm_fix_azure_chat_anyof_tool_schema 2026-09-01 14:08:34 -07:00
mateo-berri
28d0ac5339 fix(router): guard the declared-provider check for requests without a model 2026-09-01 14:06:14 -07:00
mateo-berri
0f35328a21 Merge remote-tracking branch 'origin/litellm_internal_staging' into litellm_fix_azure_chat_anyof_tool_schema
# Conflicts:
#	tests/test_litellm/litellm_core_utils/prompt_templates/test_litellm_core_utils_prompt_templates_common_utils.py
2026-09-01 13:59:59 -07:00
yuneng-jiang
8bc862f52c
Merge pull request #39129 from BerriAI/litellm_/litellm-issue-39078-c025be
fix(ui): render the logs Tools panel with theme tokens
2026-09-01 13:58:03 -07:00
Mateo Wang
2c5f429ad4
Merge pull request #39185 from BerriAI/litellm_fix_embedding_encoding_format_suite_break
test: exempt MockTransport request-shape embedding tests from VCR replay
2026-09-01 13:52:51 -07:00
Yassin Kortam
aab9abdd1d
fix: keep litellm_credential_name from LiteLLM Params JSON and gate stored credential attach to proxy admins (#39047)
* fix(ui): keep litellm_credential_name from LiteLLM Params JSON when no credential is selected

Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>

* test(ui): drop null litellm_credential_name from AddModelPanel payload fixture

Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>

* fix(ui): validate JSON litellm_credential_name against accessible credentials

Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>

* fix(proxy): enforce proxy-admin-only credential attachment on model create/update

Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>

* fix(proxy): raise ProxyException for unauthorized credential attach and gate /model/update

Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>

* refactor(proxy): fold credential-change detection into can_user_attach_credential to satisfy complexity budget

Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>

* fix(proxy): decrypt stored credential name before unchanged-credential comparison

Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>

* test(proxy): cover credential attach rejection on add_new_model and patch_model

Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>

* test(proxy): annotate proxy-global patches with test-quality suppressions

Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>

---------

Co-authored-by: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>
2026-09-01 13:46:18 -07:00
Mateo Wang
8ce02c0666
Merge pull request #39184 from BerriAI/litellm_fable_5_1_structured_output 2026-09-01 13:41:22 -07:00
ryan-crabbe-berri
6d6c9af4ab
Merge pull request #39155 from BerriAI/litellm_agent_hub_search
feat(ui): add search to the Agent Hub tab and admin agents table
2026-09-01 13:36:36 -07:00
Mateo Wang
44b595f0bb
Merge pull request #39136 from BerriAI/litellm_lit6611_requested_model_label_cap
fix(prometheus): bound requested_model label cardinality on client failure paths
2026-09-01 13:35:33 -07:00
yuneng-jiang
6028afa2ab
Merge pull request #39178 from BerriAI/litellm_revert_v2_migration_resolver_default
revert: default the proxy back to the v1 migration resolver
2026-09-01 13:32:18 -07:00
yucheng-berri
cdb1245e74
fix(s3): bound s3 object keys and download filenames for long Responses API ids (#39164)
* fix(s3): bound object keys and download filenames to s3 limits

Long OpenAI-compatible Responses API ids pushed the s3 object key past s3's
1024 UTF-8 byte cap, so the PUT failed with a 400 and the log record was
dropped. Keys that still fit are unchanged, byte for byte. An oversized one
now keeps a readable head of the file name and appends the sha256 of the full
name. A configured path/alias prefix that is long enough to overflow on its
own keeps whole leading path segments, so a prefix-scoped IAM policy or
lifecycle rule still matches, and ends in a short digest of the full
configured value so two operators do not land in the same folder.

The Content-Disposition filename carried the same unbounded id and hit s3's
2048 byte metadata-header cap, so the upload still failed with
MetadataTooLarge once the key was bounded. It is bounded the same way, head
plus digest, so two records downloaded from the console stay distinct files.

The full response id stays in the uploaded JSON payload.

* fix(s3): keep the configured prefix whole and spend the whole key budget

Shorten the response id first and only trim the operator's configured prefix
when the prefix itself is what does not fit, so prefix scoped IAM policies and
lifecycle rules keep matching. Trim by bytes rather than whole segments so the
longest possible string prefix survives, and route the audit log key through
the same shared builder.

* chore(s3): trim the comments and docstrings the review flagged

Keep the two external facts that are not visible from the code, the 1024 byte
object key cap and the 2048 byte metadata header cap, and drop the rest.
2026-09-01 13:30:02 -07:00
mateo-berri
6a9dcb5ce6 test: allow dashscope domain in qwen alias default api_base check 2026-09-01 13:29:12 -07:00
mateo-berri
f4347f25de test: exempt MockTransport request-shape embedding tests from VCR replay 2026-09-01 13:29:12 -07:00
Mateo Wang
286a754999
Merge pull request #38839 from BerriAI/litellm_fix_chat_anyof_tool_schema
fix(openai): flatten top-level tool schema combinators on chat completions
2026-09-01 13:25:03 -07:00
mateo
d568bbe58d fix(bedrock): use tool fallback without forced tool_choice for claude-fable-5-1 structured output
Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>
2026-09-01 20:19:46 +00:00
Mateo Wang
c913b09e66
Merge pull request #31725 from Srivatsa03/time-based-cost-pricing
feat(cost): support time-based off-peak pricing in cost calculation
2026-09-01 13:18:30 -07:00
yuneng-jiang
6c26345ab5
Merge pull request #39175 from BerriAI/litellm_ui_select_popup_race
test(ui): pick select options by role instead of by text
2026-09-01 13:05:54 -07:00
mateo-berri
60de5468d7 fix(cost): inherit the backend's raw cost map entry for off-peak-only deployments
Filtering copied fields by name dropped companion billing rules like
web_search_billing_unit and the regional uplift multipliers, so grounding
and uplifts billed differently through the deployment entry. Copy the
backend's raw litellm.model_cost entry wholesale instead, which also
removes the synthesized-zero special case since the raw entry only holds
real values.
2026-09-01 13:05:25 -07:00
Yuneng Jiang
529ac12ba5
test(ui): drop the helper docblock
The repo does not take explanatory comments. The reason the helper queries by
role lives in the commit that introduced it and in the PR description.
2026-09-01 13:01:46 -07:00
mateo
3e3e4d6970 fix(anthropic): use native structured output for claude-fable-5-1 on Vertex AI and Bedrock Invoke
Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>
2026-09-01 20:01:03 +00:00
Mateo Wang
19ca4cd4a9
Merge branch 'litellm_internal_staging' into litellm_fix_supported_openai_params_router_alias 2026-09-01 12:57:36 -07:00
Yuneng Jiang
fd72a39b1b
revert: default the proxy back to the v1 migration resolver
This reverts merge commit 2b1bd20834 (#31125)

Two CircleCI jobs on the staging-to-main promotion went red the moment
that PR landed. proxy_multi_instance_tests boots two proxies against one
database, and both now race the same migration:

  Error: P3018 A migration failed to apply
  Database error code: 40P01, deadlock detected
  Process 73 waits for ShareLock on virtual transaction 4/11;
  blocked by process 75. Process 75 waits for ExclusiveLock on
  advisory lock [16384,0,72707369,1]; blocked by process 73

Neither proxy comes up, so the job times out after 300s waiting on
localhost:4000. The same wait took 36.5s on the last green run

Timeline: #31125 merged at 18:46:14Z and the failing run started at
18:49:59Z. The merge commit is not an ancestor of the last green
revision (194a3cc) and is an ancestor of the first failing one
(01de2837)

The v2 resolver was meant to avoid exactly this class of contention, so
the deadlock looks like a bug in it rather than a reason to abandon it.
Putting the default back to v1 buys time to fix it without holding up
the release
2026-09-01 12:56:08 -07:00
Mateo Wang
8ae072b501
Merge pull request #39147 from BerriAI/litellm_openai_drop_toolless_tool_choice
fix(openai): drop tool_choice when request has no tools on chat completions
2026-09-01 12:50:38 -07:00
Mateo Wang
ab8b8deb39
Merge pull request #39144 from BerriAI/litellm_fix_responses_tool_call_id_shape
fix(responses): tool call id shape breaks gpt-5 -> claude fallback conversations
2026-09-01 12:50:35 -07:00
mateo-berri
0ec3e936b7 fix(cost): inherit the backend's full price structure for off-peak-only deployments
Copying only the flat token rates dropped threshold, tiered, service-tier,
cache, character, and per-second rates from peak-hour billing once cost
lookup switched to the deployment entry, and get_model_info synthesizes
zero flat rates for backends without one, which would have marked
tiered-only backends explicitly priced free. Copy every price-bearing
field instead, deep-copied, rejecting the synthesized zeros the way
_inherit_builtin_tiered_output_rate already does.
2026-09-01 12:47:26 -07:00
mateo-berri
4567fc784c fix: close non-message open items as incomplete when a stream is blocked 2026-09-01 12:45:23 -07:00
Yuneng Jiang
116efee3f7
Merge remote-tracking branch 'origin/litellm_internal_staging' into litellm_ui_select_popup_race 2026-09-01 12:42:33 -07:00
Yuneng Jiang
eb53639ecb
test(ui): pick select options by role instead of by text
Clicking a Base UI select entry found by text or by a title attribute is a
race. The text node exists one render before the popup finishes entering,
and until then the positioner still carries pointer-events: none, so
user-event refuses the click and the test throws. Querying by role only
matches once the popup is exposed to the accessibility tree, which is after
that window closes.

Route the 37 remaining select interactions through chooseSelectOption, which
does the role query. Instrumenting the converted files shows the text query
resolving while the popup was still pointer-blocked on 6 of 41 samples; the
role query was never blocked.

Seven files kept their text queries because their popup entries carry no
accessible role, so there is nothing to query by.
2026-09-01 12:42:27 -07:00
yuneng-jiang
a3e115f4cd
fix(ui): render the guardrail garden detail page with theme tokens (#39131)
The page set its headings, table borders, sidebar labels and tag pills
inline with a fixed light palette (#202124, #5f6368, #dadce0, #f8f9fa,
#fff), so in dark mode it drew dark text on hardcoded white surfaces.

Move those to the foreground/muted/border/card/info tokens, matching
the back link and Create Guardrail button that already used them.
2026-09-01 12:39:43 -07:00
mateo-berri
4c7dd0b522 Merge remote-tracking branch 'origin/litellm_internal_staging' into litellm_fix_stream_modify_response_chunks
# Conflicts:
#	type-discipline-budget.json
2026-09-01 12:36:48 -07:00
yuneng-jiang
4c86b1d58c
test(e2e/ui): cover the Usage page activity tabs (#39061)
* test(e2e/ui): cover the Usage page activity tabs

Usage had one test, on Top Virtual Keys. The Key, Model and Endpoint Activity
tabs are the ones an admin reads to answer where the spend went, and none of
them was covered.

Also fixes waitForKeyInDailyActivity, which only read the first page of
/user/daily/activity. The route paginates, so once a run generates more keys
than one page holds, the helper spins for its full 120 seconds and then blames
the rollup for a key the rollup wrote correctly. The Usage page itself already
walks every page; the helper now matches it.

* test(e2e/ui): route the user-creation call through SERVER_ROOT_PATH

Review caught /user/new posting to the server root, which misses the proxy
when it is mounted under a prefix. traffic.ts already had the helper for
this; it is now exported so specs making their own management calls can use
it too.

Also drops the mutable accumulators from the daily-activity paging, which
the repo conventions ask for.

---------

Co-authored-by: Claude <noreply@anthropic.com>
2026-09-01 12:36:17 -07:00
yuneng-jiang
284e96cfe7
test(e2e/ui): cover the team Settings tab (#39058)
The Teams tests covered creating, deleting and membership, but nothing on the
Settings tab, which is the form that posts the whole team back. That is the
shape behind the reports of a team losing its metadata or its model aliases
after an unrelated edit.

Each test creates its own team rather than editing a seeded one. The limits
test pins the models and members the edit had no business touching, and the
alias test calls the new alias with a team key instead of trusting the
readback, since an alias the router never resolves reads the same either way.

Co-authored-by: Claude <noreply@anthropic.com>
2026-09-01 12:35:45 -07:00