Commit graph

5176 commits

Author SHA1 Message Date
Yucheng He
97a6b920ff Merge remote-tracking branch 'origin/litellm_internal_staging' into litellm_lit3371_proxy_metadata_forwarding 2026-09-04 19:15:21 -07:00
ryan-crabbe-berri
2151dcbd73
Merge pull request #39822 from BerriAI/litellm_lit_6594_access_group_resource_names
feat(access-groups): resolve resource names on access group responses
2026-09-04 19:03:20 -07:00
tin-berri
d8ca43a800
feat(complexity_router): let the LLM classifier see request images (#39825)
The classifier scores extracted text, so a turn whose complexity lives in
its image is invisible to it: a screenshot of a stack trace classifies on
its caption, and an image-only turn flattens to empty text and never
reaches the classifier at all.

classifier_llm_config.vision opts in, off by default, with max_images
bounding what one turn can add. Images are still dropped when the
classifier model is declared supports_vision false. Anthropic and
Responses image parts are rewritten into chat-completions dialect before
they reach the classifier call, since /v1/messages hands the pre-routing
hook its own dialect untranslated.

The local scorer no longer short-circuits heuristic_first or hybrid on a
turn carrying forwarded images, because it reads text alone and its
confidence describes a request it has only partly seen.
2026-09-04 18:59:50 -07:00
moe-berri
639b3f4f62
Merge pull request #39809 from BerriAI/litellm_stall_escalation
feat(router): auto-escalate stalled complexity-router tasks
2026-09-04 18:44:25 -07:00
yuneng-jiang
e733ca1065
Merge pull request #39811 from BerriAI/litellm_/mongodb-vector-store-e4ff63
feat(vector_stores): add a MongoDB vector store provider for Atlas and self-managed deployments
2026-09-04 18:19:03 -07:00
ryan-crabbe-berri
be76dfad9c
Merge pull request #35853 from BerriAI/litellm_retry_policy_503
fix(router): resolve retry_policy by exception hierarchy, add ServiceUnavailableErrorRetries and DefaultRetries
2026-09-04 17:34:00 -07:00
Yucheng He
396bcca423 Merge origin/litellm_internal_staging into litellm_lit3371_proxy_metadata_forwarding 2026-09-04 17:27:11 -07:00
tin-berri
b3c867c7b2
fix(auto_router): derive tier definitions in prompt editor (#39688) 2026-09-04 16:48:33 -07:00
moe-berri
544822b1a8 fix: escalate a stalled keyword-forced tier, and let a blocked toggle clear
Two issues Bugbot found on #39809.

A keyword_tier_rule forces its tier and returns before any classification
runs, so stall escalation never reached that path even though keyword
escalation did. That left the one path that can pin a weak model to a
whole conversation as the one path a stall could not lift. Stall
detection now resolves before the override branch and both paths bump.

The dashboard switch disabled itself whenever session pinning or
user-turn classification was on, including for a router that already had
stall escalation enabled. The conflicting keys stayed set, the backend
rejected the save, and the disabled switch was the only way to clear
them. It now disables only the off-to-on direction.
2026-09-04 16:14:28 -07:00
ryan-crabbe-berri
b29f9a94bc refactor(router): resolve retry policy by exception MRO and add DefaultRetries
Replace the hand-ordered isinstance ladder in get_num_retries_from_retry_policy
with a class-to-field mapping walked along the exception's MRO, most specific
class first. A RetryPolicy field can no longer go silently dead the way
InternalServerErrorRetries did, and subclasses such as
ContentPolicyViolationError or MidStreamFallbackError pick up their parent's
field when they have none of their own.

Add a DefaultRetries catch-all so errors without a dedicated field
(BadGatewayError, APIConnectionError, NotFoundError, ...) can be governed by the
policy too. Specific fields still win over DefaultRetries.

Wiring the previously dead InternalServerErrorRetries changes one test
expectation: a policy of 2 now overrides a per-deployment num_retries of 5, so
the amplification test sees 3 upstream requests instead of 6.

Expose DefaultRetries as "All other errors" in the Admin UI retry settings tab
and ratchet the lint budgets down by the violations this branch fixed.
2026-09-04 16:09:01 -07:00
ryan-crabbe-berri
c5c10bc91f feat(access-groups): resolve resource names on access group responses
The access group detail page rendered MCP servers, agents, attached teams and keys as bare ids, so an admin had to look each one up elsewhere to audit a group

Every access group response now also carries access_mcp_servers, access_agents, assigned_teams and assigned_keys as {id, name} pairs. Names come from the DB rows first and fall back to config-declared MCP servers and agents (including legacy agent ids), resolved with one query per table across all groups in a list call. The existing *_ids columns are unchanged

The UI renders the name with the id in a tooltip, links teams and keys to their detail pages, and shows the raw id only when nothing resolves
2026-09-04 16:08:50 -07:00
moe-berri
939039f492 fix(router): anchor stall detection on the newest tool call
Counting whichever pattern was most common across the window escalated a
task that had already recovered: three identical failures stay in the
window for a few turns after the model breaks out of them, and on their
own they met the threshold.

Both tests now anchor on the newest call. The repeat test counts calls
matching the newest one, and the error test only runs while the newest
call is itself an error, so a window whose recent calls are healthy no
longer escalates. The matches still do not have to be adjacent, so a
retry loop broken up by an unrelated lookup keeps counting.

Found by Greptile on #39809.
2026-09-04 16:03:20 -07:00
ryan-crabbe-berri
2de21d0695 Merge remote-tracking branch 'origin/litellm_internal_staging' into litellm_retry_policy_503
# Conflicts:
#	litellm/types/router.py
#	tests/test_litellm/test_router.py
2026-09-04 15:58:04 -07:00
moe-berri
01b55daee9 feat(ui): add stalled task escalation controls to the auto-router form
Adds an "Advanced: Stalled Task Escalation" section to the complexity
router config: a toggle plus the repeat threshold and the window of recent
tool calls to examine. Both knobs are seeded on enable and cleared on
disable, so an off router sends none of the three keys, which is what the
backend requires next to session pinning and a custom tier set.

The toggle locks out with an explanation when "How often to classify" is
set to once-per-session or new-user-message, since both replay a held
routing decision instead of classifying and a stall would never reach the
classifier. The keys join the custom-tier restriction registry, which both
strips them from a custom-tier save and marks the section restricted.

ResponseFormatControls moves into its own file to keep
ComplexityRouterConfig.tsx under the 800-line lint ceiling, matching the
one-file-per-control layout its siblings already use.
2026-09-04 15:31:54 -07:00
ryan-crabbe-berri
d23bec84c4
Merge pull request #39196 from BerriAI/litellm_guardrail_usage_cost_rollup
feat(guardrails): roll up Bedrock guardrail cost per usage counter
2026-09-04 15:21:16 -07:00
tin-berri
6dff3a5f72
fix(complexity_router): fall back to a live peer when the decided tier model is fully cooled down (#39675)
A complexity tier can name several model groups, but the pool pick and the session-pin
replay both returned a group without consulting deployment health, so a group whose every
deployment was in cooldown was still routed to and the request died at the router's
zero-deployment check while a healthy peer sat in the same tier.

Gate the decided response at the pre-routing hook's exits, the seam the modality gate
already occupies, so every arm that can place a request is covered by one owner: a fresh
classification, a replayed or escalated pin, a plan-mode floor, a context-window
escalation, an adaptive pick, and whatever arm is added next.

Peers come from the decided tier only. Climbing to a higher tier costs more than the
classifier asked for and is left to a follow-up. The gate fails open on every uncertainty:
an unreadable cooldown view, a decision carrying no tier, a group the router knows no
deployments for, or a tier whose peers are all cooling.
2026-09-04 22:11:50 +00:00
ryan-crabbe-berri
9bd34adb6d Merge remote-tracking branch 'origin/litellm_internal_staging' into litellm_guardrail_usage_cost_rollup
# Conflicts:
#	type-discipline-budget.json
2026-09-04 14:41:55 -07:00
ryan-crabbe-berri
f59021f0e0
Merge pull request #39807 from BerriAI/litellm_lit_6690_relax_routing_group_name
fix(ui): accept any routing group name the backend accepts
2026-09-04 14:39:37 -07:00
ryan-crabbe-berri
6c81a5c423 feat(guardrails): store untracked units on the rollup row instead of nulling cost
A row that received both priced and unpriced increments used to collapse
to cost NULL, throwing away the priced subtotal and making every unit on
it read as untracked. The rollup now carries a second column,
untracked_units, that the aggregator increments for units with no known
price while cost keeps accruing for the rest, so cost covers exactly
units - untracked_units. Rows written before the migration keep cost
NULL and still read as untracked in full

The endpoints read untracked units off the column (or the whole row for
a legacy NULL) rather than from a NULL filter, and the policies overview
now fills totalUntrackedUsageUnits, which the previous commit missed

Claude-Session: https://claude.ai/code/session_01EX13mWex6RaBo9PYnkAtFW
2026-09-04 14:38:08 -07:00
moe-berri
336891bbc7
Merge pull request #39696 from BerriAI/litellm_bound_classifier_timeout
fix(router): bound auto-router classifier latency
2026-09-04 14:37:52 -07:00
Yuneng Jiang
38cd1bff7b
Merge remote-tracking branch 'origin/litellm_internal_staging' into litellm_/mongodb-vector-store-e4ff63 2026-09-04 14:20:31 -07:00
yuneng-jiang
2849aee57d
fix(health): probe test_connection with the credential the request names (#39801)
* fix(health): probe test_connection with the credential the request names

/health/test_connection matches the request's model string against the
configured deployments and merges the match's litellm_params underneath the
request. A request that named a stored credential but no key of its own
still satisfied the "request sets no connection fields" test, so it inherited
the matched deployment's api_key and api_base, and load_credentials_from_list
then skipped the named credential because api_key was already set.

A wildcard route covering the model is enough to match, so the Add Model
page's Test Connect probed with an unrelated deployment's key while echoing
back the credential that was selected.

Naming a credential the configuration does not name now withholds the
configuration's credential fields, the same set already withheld from a
request that supplies its own endpoint. Naming no credential still inherits
them, as documented.

* test(health): drop test docstrings that restate their own names

* test(health): assert the credential probe on the wire, not on the call args

The connection-test regressions patched litellm.ahealth_check and read the
params handed to it. Driving the endpoint through the app with respx faking
the upstream instead lets the real credential resolution run, so the tests
assert the key and host that actually go out, which is what the bug was about.

It also drops three of the five patched proxy internals; the two that are left
are proxy-global wiring with no injection seam, the same ones the image_edit
connection test already has to reach for.

* chore(ui): regenerate schema.d.ts for the test_connection docs change
2026-09-04 14:18:41 -07:00
ryan-crabbe-berri
1548be8235 feat(guardrails): report the usage units a guardrail's cost leaves out
A row's cost sums only the daily rows that carry a tracked cost, so it
silently under-reports whenever some rows are NULL (pre-migration days,
old pods mid-rollout, an unpriced counter). Both usage endpoints now
return the per-counter units behind those NULL rows next to the cost
(untrackedUsageUnits / totalUntrackedUsageUnits on the overview,
untracked_usage_units on the detail), so a partial cost is never mistaken
for a complete one and the reader can see exactly what it excludes

Claude-Session: https://claude.ai/code/session_01EX13mWex6RaBo9PYnkAtFW
2026-09-04 14:09:59 -07:00
ryan-crabbe-berri
6f18a4d81e fix(ui): accept any routing group name the backend accepts
The create form rejected names with slashes or spaces even though the
proxy stores and routes any non-empty string. Drop the client-only
character pattern and trim the name before the required check so a
whitespace-only name is still refused

Claude-Session: https://claude.ai/code/session_01HkaXiD6gssHnx3kqu1rR8C
2026-09-04 14:06:34 -07:00
Yuneng Jiang
431579dc16
Merge remote-tracking branch 'origin/litellm_internal_staging' into litellm_/mongodb-vector-store-e4ff63
# Conflicts:
#	.github/workflows/_test-unit-base.yml
#	litellm/litellm_core_utils/sensitive_data_masker.py
#	tests/test_litellm/litellm_core_utils/test_sensitive_data_masker.py
#	uv.lock
2026-09-04 14:00:10 -07:00
tin-berri
8beca1d58d
fix(auto-router): route 1M complex tier to GPT Sol (#39797)
* feat(ui): add 1M context auto-router preset

* feat(ui): use heuristic v2 for 1M preset

* fix(ui): keep 1M preset test within lint budget

* fix(auto-router): route 1M complex tier to GPT Sol

* test(auto-router): update 1M complex tier expectation
2026-09-04 13:33:06 -07:00
ryan-crabbe-berri
3bdc5ecd0e refactor(ui): drop the per-user usage page clamp now handled by the shared DataTable
The shared DataTable clamps a server-mode page index whenever rowCount no
longer reaches it (#39776), including the empty-dataset case this table's
own clamp skipped because it required total_pages > 0. Remove the local
clamp and cover the empty case through the component so the wiring into
the shared behavior is what the tests prove
2026-09-04 12:47:37 -07:00
ryan
7fde31fe08 fix(ui): fall back to the last page when per-user usage shrinks under the current page
Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>
2026-09-04 12:45:26 -07:00
ryan
fd42bddee6 fix(ui): reset per-user usage page in the same render as the tag filter change
Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>
2026-09-04 12:45:26 -07:00
ryan
11f272e08b fix(ui): paginate per-user usage with the shared server-side DataTable footer
Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>
2026-09-04 12:45:26 -07:00
ryan-crabbe-berri
62087c5d5f
Merge pull request #39776 from BerriAI/litellm_datatable_server_page_clamp
fix(ui): clamp server-paginated DataTable page index when rowCount shrinks
2026-09-04 12:43:38 -07:00
devin-ai-integration[bot]
dd01abc439
feat(team): report per-user spend within a team for JWT traffic (#39771)
* feat(team): report per-user spend within a team for JWT traffic

Add GET /team/spend/by_user, which groups raw spend logs by (team_id, user)
so JWT/SSO requests with no virtual key are attributed to the user inside
each selected team. Team admins see every member, plain members see only
their own row. The Team Usage page gets a Spend Per User Within Team card
with CSV export backed by the same endpoint.

Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>

* test(team): cover /team/spend/by_user in behavior suite, tf audit allowlist and EntityUsage unit test

Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>

* refactor(team): drop explanatory docstrings from /team/spend/by_user and regen schema.d.ts

Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>

---------

Co-authored-by: yassin <yassin@berri.ai>
Co-authored-by: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>
2026-09-04 12:00:47 -07:00
ryan-crabbe-berri
d05d2a6f05 fix(ui): clamp server-paginated DataTable page index when rowCount shrinks
Server-mode tables kept whatever page index the user was on after the
server's total dropped below it, for example after deleting the last
rows of the final page or when a refetch came back empty. The footer
then read "Page 2 of 1" and "Showing 26-25 of 25" with Previous and
First enabled over an empty body, and every one of the 13 server-mode
consumers was exposed since none of them clamped

The shared DataTable now snaps the controlled page index to the last
valid page as soon as a non-loading rowCount no longer reaches it, so
the fix applies to every consumer without per-table clamps. Loading
responses are ignored so a pending fetch never bounces the user to
page 1
2026-09-04 11:33:26 -07:00
Mateo Wang
04a198e3e3
Merge pull request #39568 from BerriAI/litellm_fix-batch-spend-key-double-hash-bcae
fix(spend-tracking): keep batch spend keys joinable after v1.99 provenance gate
2026-09-04 10:47:34 -07:00
yuneng-jiang
c8635ecc67
feat: page the public model hub table off /public/v1/model_hub, keeping every filter (#39691)
* feat(ui): page the public model hub table off /public/v1/model_hub

The public Model Hub page loaded every published model group in one call and
did all of its searching, sorting and filtering in the browser, so a proxy with
a few thousand groups sent megabytes to render one screen.

The models table now asks /public/v1/model_hub for one page at a time. Paging,
sorting, search and the provider and mode filters are query parameters on that
route, and the pagination footer counts from the response envelope's
total_count rather than the rows on screen. Column sortability is derived from
the fields the route declares sortable, so a header can no longer ask it for a
sort it answers with a 400.

The feature filter is dropped: supports_* are booleans and the route has no
boolean filter, so it could only ever have filtered the page in view.

* fix(ui): offer the model modes litellm actually prices in the hub filter

The mode filter listed 'moderations', which no model group's mode is ever set
to, so picking it could only ever return nothing; 'anthropic_messages' was
dead the same way. Four real modes the catalogue does use, search, ocr,
guardrail and vector_store, were missing entirely.

The list is now the mode vocabulary in model_prices_and_context_window.json,
and a test reads that file so an option that matches nothing, or a mode with
no option, fails instead of silently filtering to an empty table. A failed
page fetch logs the route's error detail again, as it did before the table
moved to the paginated route.

* test(ui): keep the model hub health rows out of the inline-object budget

frontend-lint's local/no-large-inline-object-arg budget went 555 to 556: the
health check rows became arguments to the row helper. They are plain literals
spreading a shared default again, which is what they were before, and the gate
reports 554 against a max of 555.

* feat: keep every model hub filter when the table pages

Moving the table onto /public/v1/model_hub cost it two controls the route
could not serve: the provider filter fell back to one substring because
providers only declared contains, and the feature filter went away entirely
because supports_* are booleans with no filter at all. Options for the
dropdowns went with them, since a page of rows only knows the values on that
page.

The route now declares providers in, a features field whose value is the
capability names a row has, and providers, rpm and tpm as sortable. Features
is one repeated field rather than a boolean per flag so selecting two of them
matches either, which is what the multi-select has always meant. Three facet
routes serve the distinct providers, modes and features across the published
groups, carrying the parent's filters, per section 12 of the list design.

All of it is additive: the route rejects unknown parameters, so no request
that worked before changes, and the design's stability policy calls new
filters and parameters safe within a version.

Health status stays unsortable. Health is read for the rows on the page, and
ordering the match set by it would mean reading it for every published group,
which is the cost the paging exists to avoid.

* fix(ui): put the model hub facet types where the generator emits them

The generated file lists paths in sorted order and operations in path order.
Both new blocks were spliced in one entry too late, after
/queue/chat/completions rather than before it, so the schema.d.ts sync check
regenerated the file and found them misplaced. Same blocks, byte for byte,
moved to the position the generator gives them.

* fix(proxy): type a facet payload as the sequence the framework hands it

The lint job's basedpyright gate flagged one new reportArgumentType: handle_facet
passes a tuple, and FacetListResponse declared data as list[str]. A list would
have traded that error for an LIT002 mutable construction, and both budgets are
already at their ceiling on the base.

Sequence[str] is what the framework actually produces and what the model always
accepted: pydantic emits the same array schema either way, verified against
model_json_schema, so the OpenAPI spec and schema.d.ts are unchanged, and the
existing list-passing caller in spend_logs still type checks.

* test(proxy): pin the facet route's rejection contract

handle_facet answers six ways before it ever reaches the executor, and
none of them was covered: a denied scope, a filter operator the spec does
not offer, a repeated parameter, a non-positive page or page_size, and the
where clause those last two feed. Every one is a 400 or 403 an
unauthenticated caller can reach, so each gets a test that fails when the
branch stops firing.
2026-09-03 22:36:16 -07:00
moe-berri
04c6ce9ac3
Merge pull request #39693 from BerriAI/litellm_auto_setup_simple
feat(ui): add one-click Auto Router setup
2026-09-03 20:28:01 -07:00
moe-berri
d5481ca037 fix(ui): respect reported reasoning efforts 2026-09-03 20:01:44 -07:00
moe-berri
510424c86c feat(router): add classifier circuit breaker 2026-09-03 19:59:14 -07:00
moe-berri
901e312b17 style(ui): position Auto Setup before templates 2026-09-03 19:47:44 -07:00
moe-berri
7e2f345a86 style(ui): place Auto Setup under templates 2026-09-03 19:46:07 -07:00
moe-berri
2cb27985d9 fix(ui): refresh Auto Setup model ladders 2026-09-03 19:43:55 -07:00
moe-berri
78c40ed6a7 style(ui): compact Auto Setup control 2026-09-03 19:30:04 -07:00
moe-berri
8acb8de997 refactor(ui): simplify Auto Setup model selection 2026-09-03 19:28:54 -07:00
moe-berri
a2e5e7e066 copy(ui): describe recommended Auto Setup models 2026-09-03 19:03:07 -07:00
moe-berri
72ddd699dd feat(ui): prefer proven models in Auto Setup fallback 2026-09-03 19:00:26 -07:00
moe-berri
a48afd2242 fix(ui): exclude existing Auto Routers from auto setup 2026-09-03 18:48:40 -07:00
devin-ai-integration[bot]
16db51e2cf
feat(caching): add semantic_cache_scope to isolate semantic cache hits per end user (#39590)
Semantic cache keys omit the prompt, so every end user behind one virtual key
shares a bucket and can be served another user's semantically similar response.
Add an opt-in cache_params.semantic_cache_scope (key | end_user) that appends the
authenticated end-user id to the tenant scope, read from metadata and
litellm_metadata so /v1/chat/completions, /v1/responses and /v1/messages are all
covered, falling back to the key scope when no end-user id is present. Expose the
setting in the cache settings API and the Admin UI cache settings form

Co-authored-by: yassin <yassin@berri.ai>
Co-authored-by: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>
2026-09-03 18:44:54 -07:00
moe-berri
940fdfb26b feat(ui): add one-click Auto Router setup 2026-09-03 18:36:28 -07:00
devin-ai-integration[bot]
aec083cdac
feat(proxy): per-worker admission control that rejects excess requests with 503 (#39352)
* feat(proxy): reject excess per-worker requests with 503

Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>

* refactor(proxy): drop redundant suppressions in admission middleware

Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>

* fix(proxy): exempt the /metrics/ redirect target from admission control

Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>

* ci: allowlist live Granian saturation benchmark

Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>

* fix(proxy): regenerate dashboard API types

Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>

* fix(proxy): normalize root_path for admission exemptions, validate settings, inject state

Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>

* test(proxy): cover prometheus metric factory, lifespan scope, and prefix lookalike paths

Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>

* fix(proxy): queue behind pending waiters, cache admission settings parsing, log invalid limits

Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>

* refactor(proxy): simplify invalid admission settings handling

Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>

---------

Co-authored-by: yassin <yassin@berri.ai>
Co-authored-by: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>
2026-09-03 18:19:04 -07:00
ryan-crabbe-berri
b7d1e89667
Merge pull request #39684 from BerriAI/litellm_lit_4738_table_scrolling
fix(ui): scroll admin table rows inside the table instead of the page
2026-09-03 17:30:53 -07:00