Commit graph

47584 commits

Author SHA1 Message Date
moe-berri
901e312b17 style(ui): position Auto Setup before templates 2026-09-03 19:47:44 -07:00
moe-berri
7e2f345a86 style(ui): place Auto Setup under templates 2026-09-03 19:46:07 -07:00
moe-berri
2cb27985d9 fix(ui): refresh Auto Setup model ladders 2026-09-03 19:43:55 -07:00
moe-berri
78c40ed6a7 style(ui): compact Auto Setup control 2026-09-03 19:30:04 -07:00
moe-berri
c0a401947a test(router): cover retry policy opt-out 2026-09-03 19:29:54 -07:00
moe-berri
8acb8de997 refactor(ui): simplify Auto Setup model selection 2026-09-03 19:28:54 -07:00
moe-berri
d671e0ea5d fix(router): honor explicit retry opt-out 2026-09-03 19:20:47 -07:00
moe-berri
ddc5d8dc37 fix(router): bound auto-router classifier latency 2026-09-03 19:10:53 -07:00
moe-berri
a2e5e7e066 copy(ui): describe recommended Auto Setup models 2026-09-03 19:03:07 -07:00
moe-berri
72ddd699dd feat(ui): prefer proven models in Auto Setup fallback 2026-09-03 19:00:26 -07:00
moe-berri
a48afd2242 fix(ui): exclude existing Auto Routers from auto setup 2026-09-03 18:48:40 -07:00
mateo-berri
b93cf2e20a Merge branch 'litellm_fix-batch-spend-key-double-hash-bcae' of https://github.com/BerriAI/litellm into litellm_fix-batch-spend-key-double-hash-bcae 2026-09-03 18:46:25 -07:00
mateo-berri
2b7e14872f fix(spend-tracking): hand plain dict rows to polars in the CloudZero and Focus exports 2026-09-03 18:46:01 -07:00
devin-ai-integration[bot]
16db51e2cf
feat(caching): add semantic_cache_scope to isolate semantic cache hits per end user (#39590)
Semantic cache keys omit the prompt, so every end user behind one virtual key
shares a bucket and can be served another user's semantically similar response.
Add an opt-in cache_params.semantic_cache_scope (key | end_user) that appends the
authenticated end-user id to the tenant scope, read from metadata and
litellm_metadata so /v1/chat/completions, /v1/responses and /v1/messages are all
covered, falling back to the key scope when no end-user id is present. Expose the
setting in the cache settings API and the Admin UI cache settings form

Co-authored-by: yassin <yassin@berri.ai>
Co-authored-by: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>
2026-09-03 18:44:54 -07:00
Cursor Agent
048499cdf5
fix(spend): only reverse-hash export rows whose key alias join missed
Team and service keys often have no user_email after a successful token
join. Treating empty email as a miss hashed every verification token on
routine CloudZero and Focus exports.

Co-authored-by: Mateo Wang <mateo-berri@users.noreply.github.com>
2026-09-04 01:44:18 +00:00
mateo-berri
1763052ff5 Merge branch 'litellm_internal_staging' of https://github.com/BerriAI/litellm into litellm_govcloud_profiles_lit6421 2026-09-03 18:41:00 -07:00
mateo-berri
43f31b0a4b feat(pricing): add GovCloud Claude Opus 5 and us-gov. inference profile rows 2026-09-03 18:40:59 -07:00
Yassin Kortam
b7f53ce9a9
fix(mcp): pre-flight the ID-JAG credential at the transport edge (#35392) 2026-09-04 01:39:37 +00:00
Cursor Agent
33815682bc
fix(spend): keep CloudZero and Focus spend-user email when filling alias
Export rows already join user_email from DailyUserSpend.user_id. Recovered
key-owner email must not replace that when only the alias join missed.

Co-authored-by: Mateo Wang <mateo-berri@users.noreply.github.com>
2026-09-04 01:38:26 +00:00
devin-ai-integration[bot]
7ae352e5cf
fix(model_checks): drop wildcard routes like bedrock/* from /v1/models (#31731)
* fix: remove wildcard routes from /v1/models response

Wildcard routes like bedrock/* were leaking into the /v1/models response
because _get_wildcard_models only removed them from unique_models in the
fallback branches (no router or no deployment), but not when the router
had a matching deployment. Now wildcards are always removed from the base
list; they are only re-added to the result when return_wildcard_routes=True
is explicitly passed.

Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>

* refactor(model_checks): collapse wildcard expansion branches and tighten regression tests

Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>

---------

Co-authored-by: Krrish Dholakia <krrishdholakia@berri.ai>
Co-authored-by: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>
Co-authored-by: yassin <yassin@berri.ai>
2026-09-04 01:36:46 +00:00
moe-berri
940fdfb26b feat(ui): add one-click Auto Router setup 2026-09-03 18:36:28 -07:00
yucheng-berri
4e18c0f63a
fix(azure): restrict the storage credential chain to deployment identities (#39637)
* fix(azure): restrict the storage credential chain to deployment identities

The keyless Azure Storage path walks the full DefaultAzureCredential chain, so a
proxy with no storage service principal authenticates as whichever identity the
host happens to carry: an operator's az login on a workstation, or the
AZURE_CLIENT_ID/AZURE_CLIENT_SECRET service principal set for Azure OpenAI.
Neither is the identity granted Storage Blob Data Contributor.

Narrow the chain to workload identity and managed identity, the two credentials
a deployment legitimately holds. Azure OpenAI, Postgres IAM auth and the other
callers of get_azure_ad_token_provider keep the full chain.

* test(azure): read the credential chain off the mock instead of an accumulator

* chore: drop a stray launch traceback committed at the repo root

* fix(azure): let the storage chain reach a system assigned managed identity

DefaultAzureCredential keeps one managed identity link and pins it to
AZURE_CLIENT_ID, so a host that sets that variable for Azure OpenAI and runs as
a system assigned identity never got asked for a storage token. Build the chain
from the three credentials a deployment can carry instead of subtracting the
ones it cannot.
2026-09-03 18:29:32 -07:00
mateo-berri
2f3d1ca575 Merge branch 'litellm_internal_staging' of https://github.com/BerriAI/litellm into litellm_fix_v1_messages_midstream_timeout_failure_logging
# Conflicts:
#	litellm/litellm_core_utils/litellm_logging.py
2026-09-03 18:27:43 -07:00
devin-ai-integration[bot]
4a537e2c19
fix(proxy): emit SSE keepalives on queue, rag, azure passthrough, usage chat and policy enrich streams (#39273)
Co-authored-by: yassin <yassin@berri.ai>
Co-authored-by: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>
2026-09-03 18:23:37 -07:00
devin-ai-integration[bot]
aec083cdac
feat(proxy): per-worker admission control that rejects excess requests with 503 (#39352)
* feat(proxy): reject excess per-worker requests with 503

Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>

* refactor(proxy): drop redundant suppressions in admission middleware

Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>

* fix(proxy): exempt the /metrics/ redirect target from admission control

Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>

* ci: allowlist live Granian saturation benchmark

Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>

* fix(proxy): regenerate dashboard API types

Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>

* fix(proxy): normalize root_path for admission exemptions, validate settings, inject state

Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>

* test(proxy): cover prometheus metric factory, lifespan scope, and prefix lookalike paths

Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>

* fix(proxy): queue behind pending waiters, cache admission settings parsing, log invalid limits

Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>

* refactor(proxy): simplify invalid admission settings handling

Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>

---------

Co-authored-by: yassin <yassin@berri.ai>
Co-authored-by: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>
2026-09-03 18:19:04 -07:00
mateo-berri
ce95afe2bd fix(spend-tracking): reverse-hash dirty spend keys in Postgres instead of paging token tables 2026-09-03 17:58:17 -07:00
tin-berri
bd10977a9a
fix(snowflake): normalize Cortex Claude request shapes (#39453)
* fix(snowflake): normalize Cortex Claude request shapes

Co-authored-by: Kamron Javaherpour <kamron@kargo.com>

Co-authored-by: Oleksandr Kononov <oleks.konov@kargo.com>

* style(snowflake): format Cortex request transformations

* fix(snowflake): annotate Cortex wire payloads

* fix(snowflake): route Cortex content through the shared Anthropic converters

* fix(snowflake): surface Cortex prompt-cache usage and thinking blocks

Parse Cortex's Anthropic-dialect responses and SSE with Anthropic's own parser so cache_creation/cache_read counts, thinking blocks and signatures reach the caller. Restore thinking for every Claude model: Cortex documents extended thinking broadly and only adaptive thinking is 4.6-gated.

* fix(snowflake): echo signed thinking blocks on every assistant turn

The reference converter extends signed thinking blocks on each assistant turn, not just tool-call turns, so a replayed thinking-plus-text response keeps its signed block. Content-less thinking turns send no empty text block.

* fix(snowflake): preserve thinking list content

---------

Co-authored-by: Oleksandr Kononov <oleks.konov@kargo.com>
2026-09-03 17:44:07 -07:00
mateo-berri
533a1c959c fix(proxy): reserve budget for the Bedrock Converse prompt, not the context window
Resolving the model from the /bedrock path sends passthrough calls through
optimistic budget reservation, whose tokenizer cannot walk Converse content
blocks and so fell back to the model's max_input_tokens. Count those messages
as text and read inferenceConfig.maxTokens so a budgeted key reserves the
request's cost.
2026-09-03 17:42:46 -07:00
mubashir1osmani
18373f8f51 fix(vertex_ai): route endpoint resolution through safe_get to guard caller-supplied api_base 2026-09-03 20:38:24 -04:00
yucheng-berri
a06d63f99e
fix(logging): blocked requests no longer report guardrail_status=success in multi-guardrail configs (#39596)
* fix(logging): aggregate guardrail_status by severity across guardrail entries

A pre_call guardrail that passed (e.g. hide-secrets recording a mask)
appends its entry before a later guardrail's block, and the first-wins
reader reported the blocked request as guardrail_status=success in
StandardLoggingPayload.status_fields. Take the most severe status
across all entries instead: guardrail_intervened >
guardrail_failed_to_respond > success > not_run.

* refactor(logging): express guardrail status severity as an immutable order

Replace the precedence dict and rebinding loop with a severity-ordered
tuple and a max() aggregation, per the repo's no-mutation and
mutable-collection lint gates; parametrize the severity test cases.
No behavior change.

* style(logging): apply ruff format to entries binding
2026-09-03 17:33:31 -07:00
devin-ai-integration[bot]
fe770700f4
fix(caching): keep a node timeout from forcing a cluster-wide topology reinit on redis-py 8.x (#39349)
* fix(caching): keep node timeout from forcing cluster reinit

Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>

* docs(caching): describe the 8.x timeout-tolerant wrapper in the module docstring

Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>

* style(caching): format redis cluster isolation wrapper

Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>

* fix(caching): keep concurrent reinit requests when tolerating a node timeout

Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>

* test(caching): cover redis cluster redirect branches

Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>

* fix(caching): let overlapping tolerated timeouts release their own reinit requests

Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>

---------

Co-authored-by: yassin <yassin@berri.ai>
Co-authored-by: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>
2026-09-03 17:31:58 -07:00
devin-ai-integration[bot]
5dd3fdbc3d
fix(router): evict stale global pattern_router entries on upsert/delete (#39664)
* fix(router): evict stale global pattern_router entries on upsert/delete

upsert_deployment and delete_deployment cleaned team_pattern_routers but left
the outgoing deployment in the global pattern_router, so wildcard requests kept
round-robining onto the stale entry after a PATCH /model/{id}/update.

Fixes #29064

Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>

* test(router): dedupe test_get_configured_mode_reads_deployment_model_info name

Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>

* fix(router): restore global pattern_router eviction dropped by previous commit

Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>

---------

Co-authored-by: yassin <yassin@berri.ai>
Co-authored-by: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>
2026-09-03 17:31:05 -07:00
ryan-crabbe-berri
b7d1e89667
Merge pull request #39684 from BerriAI/litellm_lit_4738_table_scrolling
fix(ui): scroll admin table rows inside the table instead of the page
2026-09-03 17:30:53 -07:00
Mateo Wang
6071a0767d
Merge pull request #39678 from BerriAI/litellm_fix_search_users_spec_placeholder
test(e2e): match the Internal Users search placeholder shipped by #39604
2026-09-03 17:27:40 -07:00
devin-ai-integration[bot]
1add1b4655
perf(mcp): cache SSO identity assertion reads on the ID-JAG path (#39348)
* perf(mcp): cache SSO identity assertion reads on the ID-JAG path

Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>

* fix(mcp): guard sso assertion cache against stale relogin reads

Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>

* fix(mcp): keep sso assertion cache entries and generation markers in separate namespaces

Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>

* test(router): rename duplicate get_configured_mode test so ruff F811 passes

Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>

* fix(mcp): use a process-wide epoch for sso assertion cache invalidation

Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>

---------

Co-authored-by: yassin <yassin@berri.ai>
Co-authored-by: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>
2026-09-03 17:25:37 -07:00
ryan-crabbe-berri
ff97e71652 refactor(ui): type and de-mutate the table scrolling spec, drop CSS narration
DataTable loses the comment that narrated its sticky header classes. The
table scrolling e2e spec now types every management API response it
reads, seeds rows through an immutable reduce instead of pushing into
arrays, and deletes what it seeded in each test's finally block instead
of draining a shared mutable list in afterEach.

Refs LIT-4738

Claude-Session: https://claude.ai/code/session_018yW93iDaEMhoQUXcYjus7D
2026-09-03 17:25:26 -07:00
tin-berri
e26d607f5d
feat(ui): configure auto-router affinity idle TTL (#39679) 2026-09-03 17:22:47 -07:00
ryan
c911740d82 test(ui): update useModelsInfo call assertions for the new filter arguments
Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>
2026-09-04 00:19:13 +00:00
Mateo Wang
c4e9076267
Merge pull request #39037 from BerriAI/litellm_fix_anthropic_messages_error_envelope
fix(anthropic_endpoints): return Anthropic type:error envelope for /v1/messages errors
2026-09-03 17:16:32 -07:00
Mateo Wang
5f1c63a9c7
Merge pull request #39662 from BerriAI/litellm_lit6905_vertex_passthrough_default_location
fix(proxy): apply default_vertex_config location before building the Vertex passthrough base URL
2026-09-03 17:14:48 -07:00
mubashir1osmani
54fd69beb2 fix(vertex_ai): build a well-formed endpoint-resolution url for path-mounted custom api_base 2026-09-03 20:13:51 -04:00
ryan-crabbe-berri
dc9f40c11f fix(ui): scroll admin table rows inside the table instead of the page
Virtual Keys, Teams, Request Logs and Tags now hand DataTable a bounded
flex chain and use fillHeight, so the app shell main stays the only page
scroller, the rows scroll under a pinned header and the pagination footer
sits at the bottom of the page. DataTable keeps the sticky header inside
its own scroller in maxBodyHeight mode too, which is what let the header
scroll away with the rows on Keys, Teams and Models. Model Hub, Vector
Stores and the team detail keys tab drop their 75vh boxes and flow with
the page scroller.

Adds an e2e spec that fails on the merge base for every one of those
pages and passes at this tip.

Refs LIT-4738

Claude-Session: https://claude.ai/code/session_018yW93iDaEMhoQUXcYjus7D
2026-09-03 17:13:18 -07:00
ryan
c9f9b4ae9c test(ui): hoist mock responses to named variables to stay within lint budget
Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>
2026-09-04 00:11:00 +00:00
mubashir1osmani
988ae7ca80 test(vertex_ai): assert the publisher-model batch payload instead of only mock calls 2026-09-03 20:04:05 -04:00
ryan
9ba6cab889 fix(ui): make Admin UI table pagination honor the selected page size
All Models now pushes the model group, access group and wildcard filters into
/v2/model/info (new optional access_group and wildcard_only params) so the
server total_count matches the rendered rows. Request Logs defaults to 25,
uses the shared page size options and counts rendered rows in the footer.
Deleted Teams gets the shared DataTable server pagination footer instead of a
hard-coded page size of 100. Per-user usage and the remaining unbounded list
tables get paginationMode so the size selector renders.

Resolves LIT-4738

Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>
2026-09-03 23:59:58 +00:00
mubashir1osmani
14d5ff5c21 fix(vertex_ai): keep new batch handler code within the mutable-collection budget 2026-09-03 19:58:37 -04:00
mateo-berri
c1a607f90f test(e2e): match the Internal Users search placeholder shipped by #39604
PR #39604 renamed the Internal Users search box placeholder to "Search by email or ID…" but left searchUsers.spec.ts looking for the old "Search by email…" copy, so e2e_ui_testing has been red on litellm_internal_staging since it merged. Point the locator at the shipped placeholder
2026-09-03 16:57:10 -07:00
devin-ai-integration[bot]
cf3af0f486
perf(spend): group /spend/logs summary by day in Postgres instead of per-row Prisma group_by (#39351)
* perf(spend): group /spend/logs summary by day in Postgres instead of per-row Prisma group_by

Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>

* refactor(spend): simplify /spend/logs daily summary aggregation

Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>

* fix(spend): preserve spend logs response schema

Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>

* chore: ratchet lint budgets after merge

Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>

* chore: ratchet lint budgets after merge

Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>

* fix(spend): compare spend log range bounds as naive UTC timestamps

Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>

* chore: ratchet lint budgets after merge

Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>

* chore: ratchet lint budgets after merge

Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>

* chore: ratchet lint budgets after merge

Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>

* chore: ratchet lint budgets after merge

Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>

* test: cover spend logs summary edge cases

Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>

* chore: ratchet lint budgets after merge

Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>

* test: cover spend summary request filters

Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>

* chore: ratchet lint budgets after merge

Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>

* chore: ratchet lint budgets after merge

Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>

---------

Co-authored-by: yassin <yassin@berri.ai>
Co-authored-by: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>
2026-09-03 23:53:59 +00:00
mubashir1osmani
1d2ed0bdac fix(vertex_ai): address batch review findings
Derive the GCS batch object path from the deployment's configured model
when present, so a user-crafted JSONL body.model cannot redirect an
authorized deployment's credentials to a different endpoint; the JSONL
value remains the fallback for direct SDK calls with no deployment
config.

Route the fine-tuned endpoint resolution GET through _check_custom_proxy
so custom api_base deployments do not contact Google directly.

Prefer the publisher model path over an endpoints/ segment when parsing
GCS uris, and use the last endpoints/ occurrence, so a bucket prefix
containing endpoints/<digits> cannot shadow the real model path.

Move the custom_endpoint rejection from the batches dispatcher into the
Vertex batch handler so the provider policy lives in the provider
module.
2026-09-03 19:49:37 -04:00
mubashir1osmani
39a17898ff
test(proxy-extras): repoint the migrate-deploy harness at the run_prisma seam (#39673) 2026-09-03 16:45:09 -07:00