Commit graph

46489 commits

Author SHA1 Message Date
tin-berri
502b3a2f79
feat(ui): auto-router controls for context-window escalation (#39054)
* feat(ui): auto-router controls for context-window escalation

Adds an Advanced: Context Window Escalation section to the auto-router
form, both create and edit arms, with the toggle for
enable_context_window_escalation and a clamped decimal input for
context_window_escalation_buffer. An untouched control keeps both keys
out of the payload so the router tracks the backend defaults; an
explicit opt-out (false) survives the edit round-trip through the
managed-keys projection and the hydrator, and preset prefill maps both
keys straight through so a preset cannot silently drop them

Resolves LIT-6601

* fix(ui): clearing the context-window buffer removes it from the payload

Both review bots converged on the same defect: an emptied buffer field
early-returned in commitBuffer, the draft was discarded on blur, and the
stale number reappeared and stayed in the saved config, contradicting
the copy that an empty field tracks the backend default. An empty commit
now removes the key, which the managed-keys projection propagates as a
real deletion on edit. Also trims the narrative comments the review
flagged as restating behavior
2026-08-31 20:02:08 -07:00
ryan-crabbe-berri
ca1f69fb73 Merge remote-tracking branch 'origin/litellm_internal_staging' into pr37044 2026-08-31 19:58:43 -07:00
tin-berri
0565d33fa5
fix(ui): let the auto-router scoring tier list follow the theme (#39040)
* fix(ui): let the auto-router scoring tier list follow the theme

* test(ui): assert the tier list carries the muted-foreground token
2026-08-31 19:55:16 -07:00
tin-berri
8d6d7f9ce9
feat(complexity_router): opt-in modality-based capability routing for image requests (#39032) 2026-08-31 19:51:41 -07:00
yucheng-berri
ccd76dac50
fix(proxy): wire team-level logging callbacks into passthrough endpoints (#38979)
* fix(proxy): wire team-level logging callbacks into passthrough endpoints

LIT-5152: passthrough routes now wire dynamic team-level callbacks
(success_callback, failure_callback, callback_vars) into Logging constructor,
mirroring the add_litellm_data_to_request behavior. Three hardening fixes:

1. Catch TypeError/AttributeError in _get_validated_callback_metadata when
   team logging metadata has wrong shape (e.g., logging list instead of dict),
   preventing HTTP 500 on passthrough routes with malformed config.

2. Wrap websocket passthrough logging initialization in try/except, since
   the socket is already accepted at that point; errors after accept() yield
   abrupt close (1006/1011) rather than clean HTTP error response.

3. Handle malformed deprecated callback_settings gracefully with try/except.

4. Wrap HTTP passthrough callback resolution in try/except to prevent 500 on
   malformed team metadata (backward-compatibility fix).

Changes:
- pass_through_endpoints.py: wire dynamic callbacks in HTTP+WS paths, handle
  malformed metadata gracefully with try/except fallbacks
- litellm_pre_call_utils.py: expand exception handling in validators
- test file: regression test for happy-path team callback wiring

* refactor(proxy): share passthrough team-callback resolution and cover its fail-open path

Collapse the duplicated callback wiring on the HTTP and websocket passthrough
paths into one helper that returns a frozen wiring value, log resolution
failures at error level so a broken logging config stays visible, and add
regression tests for malformed team metadata and an operational lookup failure.

Reverts the _get_validated_callback_metadata except widening: it changed
behavior for normal LLM routes, which is outside this ticket's scope.

* fix(proxy): keep passthrough alive when team callback vars hold env references

The deprecated team_metadata.callback_settings branch builds
TeamCallbackMetadata directly, skipping the AddTeamCallback validation
that strips os.environ/ references from the newer logging list. Stamping
those vars onto the Logging object made its constructor raise, so a team
on the legacy shape got HTTP 500 on every passthrough call. Validate the
resolved vars inside the fail-open boundary instead, so the request goes
through with dynamic callbacks skipped and the reason logged.

* fix(proxy): lint violations in team callback wiring helper

* style: format lint
2026-09-01 02:38:39 +00:00
ryan-crabbe-berri
f7accc4e29 test(e2e): drop the two mgmt registry cells no shared-proxy test can cover
mgmt.cache_settings.update.happy_path and
mgmt.config_override.hashicorp_vault.happy_path were the last two uncovered
Management/UI cells, and neither can be covered against the shared proxy the
e2e suites run on. Both routes reconfigure the whole process rather than a
resource the test owns.

/cache/settings persists whatever it receives into a row that outranks the YAML
cache_params and is re-applied on a timer, so a partial write downgrades a TLS
cluster to a plaintext standalone node and every later Redis call hangs. That is
what took out 60 of 72 tests on 2026-07-25 and got the original test removed in
PR #34664.

/config_overrides/hashicorp_vault has the same shape: a POST sets the HCP_VAULT_*
env vars, swaps litellm.secret_manager_client process-wide, and writes a row the
config-reload poll re-applies, so every os.environ/ lookup on the pod resolves
against the test's Vault until the DELETE lands. Its constructor also never
dials Vault, so a POST to a bogus address still returns 200 and a smoke test
built on it would pass for the wrong reason.

Keeping rows we have decided not to cover only inflates the denominator, so drop
them and record the reasoning where someone would go to write the test. Filing
the isolated-proxy harness they both need separately; the cells come back with
it.

Management/UI goes 75/77 to 75/75, headline 402/544 to 402/542.
2026-08-31 19:33:04 -07:00
Yuneng Jiang
d26a960190
Merge remote-tracking branch 'origin/litellm_internal_staging' into litellm_/non-admin-key-type-change-e76bec 2026-08-31 19:12:20 -07:00
yucheng-berri
3fadcd7155
fix(auth): quiet malformed virtual key rejections to stdout (#38838)
* fix(auth): quiet malformed virtual key rejections to stdout

Reduce noisy invalid-api-key error logs by classifying malformed virtual
keys and routing their rejections to stdout as WARNING instead of stderr
as ERROR. Suppressible via LITELLM_LOG=ERROR or log_client_error_tracebacks=true.

Changes:
- auth_utils: is_invalid_virtual_key_error() classifier and marker functions
- auth_exception_handler: log invalid keys as WARNING to child logger before
  identity seeding and callbacks, escalate non-401 transforms to ERROR
- user_api_key_auth: websocket early-raise WebSocketException(1008) to avoid
  double-logging at HTTP layer
- _logging: child logger verbose_proxy_stdout_logger with no handler/level;
  LevelRoutingStreamHandler routes its WARNING records to stdout; handler
  setLevel in _turn_on_json() closes JSON config handler level leak
- test_auth_exception_handler: new test case verifying malformed-key logs
  at WARNING with marker retention through transformations

Fixes LIT-5362

* fix(auth): classify malformed-key 401 by raise-site marker, not message text

Review round 1 (Greptile P2, veria Low):
- Move the marker attribute name to litellm/constants.py per the shared
  sentinel convention
- Stamp the marker on the malformed-key 401 where it is raised and classify
  only by it. Message text is caller-influenceable on other 401s (vector
  store ids, organization ids are interpolated into their messages), so a
  phrase match would let a request body demote an authorization failure to
  the quiet log path
- Regression test: a 401 carrying the phrase but not the marker stays at
  ERROR on stderr
2026-08-31 18:19:40 -07:00
Yuneng Jiang
4a163f1a6a
test(e2e-ui): assert user-observable behavior instead of DOM structure in audit fixes
Replace table tbody and data-slot locators with getByRole, restore prior
public MCP hub entries instead of clearing the whitelist on cleanup, seed
the public agent via the append-semantics per-agent route, and rework
mutable cleanup state into const-scoped try/finally blocks
2026-08-31 18:15:18 -07:00
Yassin Kortam
b473339ac0
Revert "fix(ui): keep litellm_credential_name from LiteLLM Params JSON when n…" (#39046)
This reverts commit 33cc9c1c48.
2026-08-31 18:00:29 -07:00
devin-ai-integration[bot]
33cc9c1c48
fix(ui): keep litellm_credential_name from LiteLLM Params JSON when no credential is selected (#39005)
* fix(ui): keep litellm_credential_name from LiteLLM Params JSON when no credential is selected

Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>

* test(ui): drop null litellm_credential_name from AddModelPanel payload fixture

Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>

---------

Co-authored-by: yassin <yassin@berri.ai>
Co-authored-by: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>
2026-08-31 17:55:53 -07:00
yujonglee
97cac0b8ed
Merge pull request #39020 from BerriAI/litellm_rust_native_build_profiles
build(rust): configure native extension profiles
2026-08-31 17:54:38 -07:00
Yuneng Jiang
c818aa153d
Merge remote-tracking branch 'origin/litellm_internal_staging' into litellm_/jovial-archimedes-1d743b
# Conflicts:
#	tests/e2e/ui/tests/internal-user/internalUserWithTeams.spec.ts
2026-08-31 17:45:19 -07:00
yuneng-jiang
68cfe1697b
Merge pull request #39016 from BerriAI/litellm_/flaky-e2e-tests-d022f6
test(e2e): assert user-observable behavior instead of DOM structure
2026-08-31 17:34:30 -07:00
Yassin Kortam
e34f43328c
fix(proxy): ship psycopg so partitioned SpendLogs detection actually runs (#38994)
ProxyExtrasDBManager.spend_logs_is_partitioned() (#38452) silently returns
False when psycopg can't be imported, and psycopg was never added to the
extra_proxy install, so every production image lacks it. Schema
reconciliation then generates the unfiltered primary-key rewrite against a
genuinely partitioned LiteLLM_SpendLogs and Postgres rejects it, exactly the
failure the fix was meant to prevent. Ships psycopg via extra_proxy and logs
a warning when it's still missing instead of failing silently.
2026-08-31 17:33:12 -07:00
mateo-berri
329654765a test(guardrails): update chat eos block tests for finish-chunk withholding 2026-08-31 17:31:59 -07:00
mateo-berri
76cfa6339b test: give the mocked prepared request real headers for the masked debug log 2026-08-31 17:28:12 -07:00
Yujong Lee
9b774d3dcf
fix(ci): isolate editable Cargo cache namespace 2026-08-31 17:25:28 -07:00
Yujong Lee
849269d52d
build(rust): configure native extension profiles 2026-08-31 17:25:28 -07:00
ryan-crabbe-berri
50eed7efa2
Merge pull request #38851 from BerriAI/litellm_window_spend_seed_by_time
refactor(proxy): bound the budget window seed by time instead of request ids
2026-08-31 17:21:45 -07:00
mateo-berri
38825cf9c6 fix(guardrails): match Responses stream event types by value so enum-typed events close the open item 2026-08-31 17:19:55 -07:00
mateo-berri
e601114383 fix(bedrock): mask signed request headers in guardrail debug log 2026-08-31 17:15:54 -07:00
mateo-berri
9b15febe13 Merge remote-tracking branch 'origin/litellm_internal_staging' into litellm_fix_stream_modify_response_chunks
# Conflicts:
#	type-discipline-budget.json
2026-08-31 17:10:36 -07:00
mateo-berri
78b57fb427 fix(guardrails): withhold chat finish chunk in end_of_stream_only mode and close open Responses items before a mid-stream block 2026-08-31 17:05:45 -07:00
mateo-berri
3a750cbf92 Merge remote-tracking branch 'origin/litellm_internal_staging' into litellm_decrease_anys_opus5_r2
# Conflicts:
#	basedpyright-code-budget.json
#	type-discipline-budget.json
2026-08-31 17:01:33 -07:00
Mateo Wang
99703a30f0
Merge pull request #36008 from nuernber/litellm_bedrock_messages_disconnect_billing
fix(anthropic_messages): drain upstream in a detached pump so client …
2026-08-31 16:54:41 -07:00
mateo-berri
a99f62d1bc fix(anthropic_messages): park deferred billing before the end-of-stream sentinel
At end of drain the pump enqueued the sentinel first and picked the
billing mode from client_detached afterward, so a client that consumed
the sentinel and tore the relay down before the pump resumed (possible
whenever the sentinel enqueue hit a full queue) had its fully delivered
response billed through the teardown path, skipping the proxy's
post-response hook. Bill or park before the sentinel goes out, and let
an unconsumed sentinel fall back to dispatching the parked billing.
2026-08-31 16:46:12 -07:00
mateo-berri
92b46538e1 fix(policy_engine): restore request guardrails list after pipeline allow 2026-08-31 16:38:44 -07:00
tin-berri
3829418878
feat(shadow_eval): target teams and users so JWT-auth traffic can be evaluated (#39015)
Shadow eval jobs previously targeted only virtual keys, so deployments on
pure JWT auth (which present no key at all) could never sample their
traffic. Jobs now carry a typed (target_type, target_id) pair covering
keys, teams, and users; sampling matches the identity every request
resolves to at auth time, so team and user jobs cover JWT traffic with
no client changes.

Resolves LIT-6578
2026-08-31 16:37:38 -07:00
mateo-berri
31a9f7e6ad test(guardrails): type streaming-block test helpers and drop mutable accumulators 2026-08-31 16:34:21 -07:00
mateo-berri
158220f151 Merge remote-tracking branch 'origin/litellm_internal_staging' into litellm_fix_stream_modify_response_chunks 2026-08-31 16:20:29 -07:00
tin-berri
f93d9b6b67
feat(complexity_router): escalate oversized prompts to a tier that fits before dispatch (#38844)
* feat(complexity_router): escalate oversized prompts to a tier that fits before dispatch

The classifier scores complexity and never prompt size, so a long agentic
session whose newest ask is trivial classifies SIMPLE onto a small-window
tier and the provider rejects it with a context-window 400 that nothing
retries. The gate runs after classification on every decision path
(classify tail and session-affinity pin), estimates prompt tokens
including the out-of-band carriers (top-level system, tools,
instructions), and when the decided tier provably cannot hold the prompt
moves the request to the lowest configured tier with a model whose
declared window fits, restricting the pick to fitting models when the
decided tier can keep it. Models with no resolvable window are never
escalated away from or onto, escalated decisions are never written as
session pins, and the decision records context_escalated plus the
original tier in spend logs.

Resolves LIT-6503

* fix(complexity_router): judge groups by smallest window, bound skips by bytes, filter adaptive picks

Review-round rework, one mechanism per finding. A group is judged by its
smallest resolvable deployment window, since the core router picks within
a group with no fit check. The counting skip is gated on UTF-8 byte
length, which BPE token counts can never exceed, so token-dense scripts
cannot slip past it; only a real tokenizer count ever moves a request and
a failed count leaves the placement alone. The fit facts now filter every
adaptive phase including cold start and the tier fallbacks. Window
questions adopt the declared provider and never resolve authenticating
providers, and a router instance without get_model_list degrades the gate
to a no-op. Tests rebuilt on real Router instances resolving deployment
model_info end to end, plus a full-path test through
async_get_available_deployment
2026-08-31 16:12:59 -07:00
mateo-berri
38dc51cc28 Merge branch 'litellm_internal_staging' of https://github.com/BerriAI/litellm into litellm_decrease_anys_opus5_r2
# Conflicts:
#	tests/logging_callback_tests/gcs_pub_sub_body/spend_logs_payload.json
2026-08-31 16:12:29 -07:00
mateo-berri
7edf5b36cf fix(guardrails): deliver modify_response block as valid SSE on streaming chat and Responses
A guardrail modify_response verdict on a streaming request only produced a
proper replacement on /v1/messages: the chat completions and Responses API
translations had no build_block_sse_chunks, so the ModifyResponseException
re-raised and surfaced as an in-stream 500 error frame (or a whole-request
500 in buffered mode) instead of the documented 200 replacement.

Implement build_block_sse_chunks for both OpenAI translations: chat emits a
content delta plus a finish_reason content_filter chunk with real usage;
Responses emits the typed event sequence (standalone via
build_synthetic_response_events pre-stream, or an output-item continuation
under the in-progress response id mid-stream) ending in response.completed.
2026-08-31 16:01:39 -07:00
Yuneng Jiang
3eb1eee1fa
Merge remote-tracking branch 'origin/litellm_internal_staging' into litellm_/flaky-e2e-tests-d022f6 2026-08-31 15:57:51 -07:00
Yuneng Jiang
859bd01dda
fix(e2e): assert the users table's own empty-state copy
UsersTable overrides DataTable's default noDataMessage with its own
EmptyState, so the row reads "No users found" rather than "No results".
Assert that, and pair it with the seeded user being absent so the check
cannot pass while the filter silently does nothing.
2026-08-31 15:57:49 -07:00
Mateo Wang
81c8c93bef
Merge pull request #38873 from BerriAI/litellm_fix_model_block_response_500
fix(proxy): return 200 from /model/block and /model/unblock instead of 500
2026-08-31 15:56:21 -07:00
Mateo Wang
d60e77ae8c
Merge pull request #38819 from BerriAI/litellm_fix_gemini_tts_response_format
fix(speech): stop forwarding response_format as a chat param for Gemini TTS
2026-08-31 15:56:18 -07:00
Mateo Wang
9b87413540
Merge pull request #38878 from BerriAI/litellm_fix_master_key_rotation_blocked
fix(proxy): preserve model table columns on master key rotation
2026-08-31 15:56:03 -07:00
Yuneng Jiang
8ab132b8be
fix(key_management): allow non-admin key_type preset transitions on /key/update
A non-admin switching an existing key's type between the safe preset
buckets (llm_api_routes, info_routes, and empty = full access) got a 403
from the allowed_routes admin gate, because /key/update, unlike
/key/generate and /key/regenerate, had no carve-out for preset-derived
values. Skip the gate only when both the incoming and the stored
allowed_routes consist entirely of safe presets, so clearing an
admin-set custom route restriction still requires proxy admin.
2026-08-31 15:54:20 -07:00
tin-berri
296bde0d0d
feat(complexity-router): add classification_mode to skip classifier on continuation turns (#38861) 2026-08-31 15:50:04 -07:00
ryan-crabbe-berri
b5c156a7d5 style(proxy): drop the explicit return None update_database no longer needs
Reverting the function to -> None left two bare `return None` statements that
RET501 rejects now that None is the only value it can return.

Claude-Session: https://claude.ai/code/session_01QvQzYztinxj8ZuD5YxbVdL
2026-08-31 15:49:49 -07:00
ryan-crabbe-berri
5eff708d0f fix(proxy): keep persisted spend another pod has not incremented in the window seed
The batch start told the seed which LiteLLM_SpendLogs rows were its own, but
using it as a hard cutoff also dropped rows another pod had already persisted.
Those rows are only repaid by that pod's own increment, so if it died first the
window row stayed permanently under the recorded spend.

The seed now reads both sums in one scan and takes off this batch's own spend,
flooring at the pre-batch total for the case where its log rows have not landed
yet. Redis payloads keep an empty request_ids so a leader from before the field
was dropped can still merge what it pops during a rolling deploy.

Claude-Session: https://claude.ai/code/session_01QvQzYztinxj8ZuD5YxbVdL
2026-08-31 15:49:03 -07:00
Mateo Wang
7a02e4163f
Merge pull request #38913 from BerriAI/litellm_gigachat_passthrough_25886
feat(gigachat): add native API passthrough routes with spend logging
2026-08-31 15:47:27 -07:00
yuneng-jiang
68c8d5ca28
Merge pull request #39027 from BerriAI/litellm_/qa-checklist-e2e-audit-952d1d
test(e2e): cover SCIM token creation and SCIM API auth in the Admin UI suite
2026-08-31 15:45:23 -07:00
mateo-berri
bab347a28e test(gcs_pubsub): expect router_metadata key in spend logs fixture 2026-08-31 15:37:56 -07:00
Yuneng Jiang
94f6827530
test(e2e): drop redundant SCIM key cleanup, the throwaway db is the teardown 2026-08-31 15:36:35 -07:00
Yuneng Jiang
c8b82aa0b6
Merge remote-tracking branch 'origin/litellm_internal_staging' into litellm_/flaky-e2e-tests-d022f6 2026-08-31 15:29:06 -07:00
Yuneng Jiang
abd8beec01
fix(e2e): measure the clipped popup and assert the table's empty state
Two assertions were checking the wrong thing. The anchoring tests read
getByRole("listbox"), which resolves to SelectPrimitive.List; that sits at
full content height inside the popup that clips and scrolls it, so the box
overlapped the trigger even when nothing visible did. Measure the popup.

The SSO-ID search expected zero rows, but DataTable renders a "No results"
message row when a filter matches nothing, so the count is one. Assert the
empty state the user actually sees.
2026-08-31 15:28:34 -07:00
mateo-berri
eb00986f18 Merge remote-tracking branch 'origin/litellm_internal_staging' into litellm_gigachat_passthrough_25886
# Conflicts:
#	osv-scanner.toml
2026-08-31 15:25:10 -07:00