Commit graph

48845 commits

Author SHA1 Message Date
yujonglee
7276caecd4
refactor(rust): extract config crate (#39706)
* refactor(rust): extract config crate

* refactor(config): split crate modules

* refactor(gateway): remove gil health counter
2026-09-04 08:17:07 -07:00
mateo
0c29f510bc fix(registry): drop gpt-image-2 text output price, add openrouter minimax-m3 and qwen3.7-plus
Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>
2026-09-04 15:03:29 +00:00
mateo
c707f2fe5d fix(model_prices): databricks gpt-5-3-codex is served via the Responses API
Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>
2026-09-04 14:32:11 +00:00
mateo
8e83d6d63d fix(model_prices): add Databricks Sep-2026 catalog, Azure gpt-realtime-2.x, per-token realtime image pricing
Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>
2026-09-04 14:06:54 +00:00
mateo
af4340dce3 Merge remote-tracking branch 'origin/litellm_internal_staging' into litellm_registry_audit_2026_09_02 2026-09-04 13:03:22 +00:00
Atharva-Kanherkar
7399b3844d fix(anthropic): satisfy lint budget gates for refusal translation 2026-09-04 17:33:06 +05:30
Atharva-Kanherkar
0b34abe8fe fix(anthropic): harden refusal translation 2026-09-04 17:02:07 +05:30
amasen02
9c8594c7b8 style(proxy): format reset_budget_job with ruff 2026-09-04 16:37:58 +05:30
amasen02
3623aecc64 style(proxy): add Final type annotations to enduser budget reset variables 2026-09-04 16:33:57 +05:30
amasen02
daced81f20 fix(proxy): invalidate end-user spend counter and cache on budget reset (#39726)
Signed-off-by: amasen02 <amasen02@users.noreply.github.com>
2026-09-04 15:37:53 +05:30
mateo
03a82823fc test: deflake redis loop-stall burst test and pre-commit interrupt cleanup
The redis breaker test raced the event loop: the fake call had to still be
pending when a real time.sleep stall began, which needs the loop to get from
scheduling to the stall in under 1ms. The fake now holds its answer behind an
asyncio.Event so the whole burst times out deterministically.

The pre-commit interrupt test found a real leak: lint_dashboard creates its
eslint report with mktemp and only removed it on the happy path, so an
interrupt landing during the whole-folder eslint run left the file behind.
The subshell now removes it from an EXIT trap, and the test drives the
interrupt while that eslint run is in flight.

Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>
2026-09-04 10:07:34 +00:00
mateo
9835883e03 Merge remote-tracking branch 'origin/litellm_internal_staging' into litellm_deflake_20260902
Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>
2026-09-04 09:20:07 +00:00
Atharva-Kanherkar
200e2901d6 fix(anthropic_responses): preserve Responses refusal blocks in Anthropic translation
When OpenAI Responses returns a refusal content block, Anthropic /v1/messages
erased the refusal text into an empty content array and emitted stop_reason 'end_turn'.
Translate refusal blocks to Anthropic text blocks, map stop_reason to 'refusal',
and add 'refusal' to AnthropicFinishReason.

Fixes #39721
2026-09-04 14:47:56 +05:30
mateo
8b37de14b1 refactor: drop fresh Any annotations and suppressions from admission control, spend summary, and dual cache
Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>
2026-09-04 08:05:04 +00:00
Ninad Phalak
76ac63abd4
test(guardrails): assert the depth bound instead of only reaching the end
The depth test asserted nothing, so it passed whether or not the bound held,
and the test-quality gate counted it as a zero-assert test. It now sends a
shallow value alongside a 200-deep chain and asserts the shallow one is
collected while the value past the bound is not.
2026-09-04 02:47:19 -05:00
mateo
b287d915b0 merge: bring litellm_internal_staging into the rolling tech debt branch
Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>
2026-09-04 07:45:29 +00:00
Ninad Phalak
119ec62952
fix(guardrails): redact Responses PromptObject variables
A Responses request can send `prompt` as a PromptObject rather than a string.
Its `variables` are substituted into the stored prompt on the provider side, so
they are caller text, and the dict shape was falling through untouched.

`id` and `version` pick which stored prompt to run and are left unchanged.
2026-09-04 02:35:32 -05:00
Ninad Phalak
ce921f49e4
fix(guardrails): walk nested tool results iteratively, with a depth bound
CI flagged _collect_content as recursive. It was, and worse, it was unbounded:
a tool_result nests its own content, the nesting is caller controlled, and the
descent had nothing to stop it. That is a JSON bomb, not a style issue.

Now an explicit queue with a depth bound of 8. Real payloads nest one or two
deep. The queue is walked in document order because the shield maps its replies
back by position, so collection order is part of the contract.
2026-09-04 02:31:20 -05:00
ryan
7783efc3b1 chore: merge litellm_internal_staging into litellm_lit_4738_table_pagination
Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>
2026-09-04 05:39:32 +00:00
yuneng-jiang
c8635ecc67
feat: page the public model hub table off /public/v1/model_hub, keeping every filter (#39691)
* feat(ui): page the public model hub table off /public/v1/model_hub

The public Model Hub page loaded every published model group in one call and
did all of its searching, sorting and filtering in the browser, so a proxy with
a few thousand groups sent megabytes to render one screen.

The models table now asks /public/v1/model_hub for one page at a time. Paging,
sorting, search and the provider and mode filters are query parameters on that
route, and the pagination footer counts from the response envelope's
total_count rather than the rows on screen. Column sortability is derived from
the fields the route declares sortable, so a header can no longer ask it for a
sort it answers with a 400.

The feature filter is dropped: supports_* are booleans and the route has no
boolean filter, so it could only ever have filtered the page in view.

* fix(ui): offer the model modes litellm actually prices in the hub filter

The mode filter listed 'moderations', which no model group's mode is ever set
to, so picking it could only ever return nothing; 'anthropic_messages' was
dead the same way. Four real modes the catalogue does use, search, ocr,
guardrail and vector_store, were missing entirely.

The list is now the mode vocabulary in model_prices_and_context_window.json,
and a test reads that file so an option that matches nothing, or a mode with
no option, fails instead of silently filtering to an empty table. A failed
page fetch logs the route's error detail again, as it did before the table
moved to the paginated route.

* test(ui): keep the model hub health rows out of the inline-object budget

frontend-lint's local/no-large-inline-object-arg budget went 555 to 556: the
health check rows became arguments to the row helper. They are plain literals
spreading a shared default again, which is what they were before, and the gate
reports 554 against a max of 555.

* feat: keep every model hub filter when the table pages

Moving the table onto /public/v1/model_hub cost it two controls the route
could not serve: the provider filter fell back to one substring because
providers only declared contains, and the feature filter went away entirely
because supports_* are booleans with no filter at all. Options for the
dropdowns went with them, since a page of rows only knows the values on that
page.

The route now declares providers in, a features field whose value is the
capability names a row has, and providers, rpm and tpm as sortable. Features
is one repeated field rather than a boolean per flag so selecting two of them
matches either, which is what the multi-select has always meant. Three facet
routes serve the distinct providers, modes and features across the published
groups, carrying the parent's filters, per section 12 of the list design.

All of it is additive: the route rejects unknown parameters, so no request
that worked before changes, and the design's stability policy calls new
filters and parameters safe within a version.

Health status stays unsortable. Health is read for the rows on the page, and
ordering the match set by it would mean reading it for every published group,
which is the cost the paging exists to avoid.

* fix(ui): put the model hub facet types where the generator emits them

The generated file lists paths in sorted order and operations in path order.
Both new blocks were spliced in one entry too late, after
/queue/chat/completions rather than before it, so the schema.d.ts sync check
regenerated the file and found them misplaced. Same blocks, byte for byte,
moved to the position the generator gives them.

* fix(proxy): type a facet payload as the sequence the framework hands it

The lint job's basedpyright gate flagged one new reportArgumentType: handle_facet
passes a tuple, and FacetListResponse declared data as list[str]. A list would
have traded that error for an LIT002 mutable construction, and both budgets are
already at their ceiling on the base.

Sequence[str] is what the framework actually produces and what the model always
accepted: pydantic emits the same array schema either way, verified against
model_json_schema, so the OpenAPI spec and schema.d.ts are unchanged, and the
existing list-passing caller in spend_logs still type checks.

* test(proxy): pin the facet route's rejection contract

handle_facet answers six ways before it ever reaches the executor, and
none of them was covered: a denied scope, a filter operator the spec does
not offer, a repeated parameter, a non-positive page or page_size, and the
where clause those last two feed. Every one is a 400 or 403 an
unauthenticated caller can reach, so each gets a test that fails when the
branch stops firing.
2026-09-03 22:36:16 -07:00
Ninad Phalak
403b06ea76
fix(guardrails): flush every held choice, and cover tool results and suffix
Three review findings.

The trailing flush walked the last chunk's choices, so a choice that finished
earlier and stopped appearing lost whatever text was still held for it and its
answer was truncated. It is now driven by the windows themselves and emits one
chunk per choice, synthesising the choice when the terminal chunk omits it.
That was data loss, not just under-redaction.

An Anthropic tool_result carries its own content, as a string or as further
blocks, and only each part's `text` was being collected. Handled recursively;
image and audio parts still fall through untouched.

The legacy completions `suffix` is forwarded to providers that support it and
was never collected. Note the placement: it has to be gathered before the
string-prompt early return, which is what the new test pins.
2026-09-03 23:28:26 -05:00
mateo
08bfdadb10 chore: merge litellm_internal_staging into litellm_registry_audit_2026_09_02, drop the deleted ocr ledger
Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>
2026-09-04 04:18:09 +00:00
yujonglee
eb67e5402b
test(ocr): complete Rust unit test parity (#39689)
* feat(ocr): complete Rust unit test parity

* test(ocr): keep parity changes harness-only

* test(ocr): share azure DI native fixture across response tests

* refactor(ocr): colocate gateway unit tests and extract lifecycle integration tests
2026-09-03 21:15:01 -07:00
yujonglee
ee08c36fc0
refactor(tests): restructure rust python harness around strategy definitions (#39628)
* wip

* refactor(tests): move sdk function tracing into rust python harness

* dead code

* fix: handle harness keyboard interrupts

* refactor(tests): deduplicate rust python harness helpers

* fix(harness): expose validated strategy choices

* wip

* refactor(harness): let strategies own parity reports

* docs(harness): update strategy structure

* refactor(harness): localize strategy report views

* wip

* fix(harness): satisfy mapping runner type checks

* fix(harness): clarify trace parity output

* wip

* fix(harness): clarify unit mapping report

* fix(harness): finalize trace parity contracts

* refactor(harness): structure parity contracts

* feat: derive unit test mapping from traces

* feat(harness): map rstest test families

* feat(ocr): port Azure document intelligence tests

* feat(harness): enforce complete unit mappings

* feat(ocr): add reducto core transforms

* feat(harness): classify host-only unit tests

* fix(ocr): complete Rust provider plumbing

* fix(harness): reuse OCR parity workers
2026-09-03 21:15:01 -07:00
yujonglee
8fc0663198
fix(ci): pin setup-uv to v10.0.1 (#39111) 2026-09-03 20:44:24 -07:00
moe-berri
04c6ce9ac3
Merge pull request #39693 from BerriAI/litellm_auto_setup_simple
feat(ui): add one-click Auto Router setup
2026-09-03 20:28:01 -07:00
moe-berri
81dd911bdc fix(router): recognize asyncio classifier timeouts 2026-09-03 20:19:13 -07:00
moe-berri
5a2845d183 fix(router): preserve classifier breaker state under concurrency 2026-09-03 20:10:02 -07:00
moe-berri
d5481ca037 fix(ui): respect reported reasoning efforts 2026-09-03 20:01:44 -07:00
moe-berri
510424c86c feat(router): add classifier circuit breaker 2026-09-03 19:59:14 -07:00
moe-berri
901e312b17 style(ui): position Auto Setup before templates 2026-09-03 19:47:44 -07:00
moe-berri
7e2f345a86 style(ui): place Auto Setup under templates 2026-09-03 19:46:07 -07:00
moe-berri
2cb27985d9 fix(ui): refresh Auto Setup model ladders 2026-09-03 19:43:55 -07:00
moe-berri
78c40ed6a7 style(ui): compact Auto Setup control 2026-09-03 19:30:04 -07:00
moe-berri
c0a401947a test(router): cover retry policy opt-out 2026-09-03 19:29:54 -07:00
moe-berri
8acb8de997 refactor(ui): simplify Auto Setup model selection 2026-09-03 19:28:54 -07:00
Ninad Phalak
6536d61adf
feat(guardrails): redact the participant name on a message
`name` on a user or assistant turn identifies a person and was going to the
provider intact. The proxy this integrates with already redacts it, so the
integration was the weaker of the two.

On a tool or function turn the same field carries the function's name, which
has to arrive unchanged or the call stops routing. That case is skipped, and
a test asserts the value is never even sent to the shield.
2026-09-03 21:22:30 -05:00
moe-berri
d671e0ea5d fix(router): honor explicit retry opt-out 2026-09-03 19:20:47 -07:00
moe-berri
ddc5d8dc37 fix(router): bound auto-router classifier latency 2026-09-03 19:10:53 -07:00
Ninad Phalak
a0abb9a499
refactor(guardrails): name the guardrail llm_shield_proxy throughout
The integration was called llm_shield in code, llm-shield in the example
config, and LLM Shield in the dashboard, while the product and its PyPI
package are both llm-shield-proxy. An operator who saw the guardrail in
LiteLLM could not tell what to install.

One identifier now: llm_shield_proxy for the enum value, module, directory,
class, config model, logo and environment variables, with LLM Shield Proxy
as the display name. That matches `pip install llm-shield-proxy`.

Renames only; no behaviour change.
2026-09-03 21:09:48 -05:00
moe-berri
a2e5e7e066 copy(ui): describe recommended Auto Setup models 2026-09-03 19:03:07 -07:00
moe-berri
72ddd699dd feat(ui): prefer proven models in Auto Setup fallback 2026-09-03 19:00:26 -07:00
Ninad Phalak
46438d7cf7
fix(guardrails): restore every streaming choice, not just the first
Streaming rehydration read and rewrote choices[0] only, so with n>1 every
later choice went back to the caller still holding its placeholders.

Each choice is its own token stream, so the sliding window is now tracked
per choice index rather than once per stream. A single shared window would
have been worse than the bug: it would splice the characters held back for
one choice onto the next one's delta.

The final flush walks every choice the same way, and the two helpers that
only ever looked at choices[0] are gone.

Adds a test that both choices come back restored, and one that each choice
gets its own window handed back rather than its neighbour's.
2026-09-03 20:48:50 -05:00
moe-berri
a48afd2242 fix(ui): exclude existing Auto Routers from auto setup 2026-09-03 18:48:40 -07:00
mateo-berri
b93cf2e20a Merge branch 'litellm_fix-batch-spend-key-double-hash-bcae' of https://github.com/BerriAI/litellm into litellm_fix-batch-spend-key-double-hash-bcae 2026-09-03 18:46:25 -07:00
mateo-berri
2b7e14872f fix(spend-tracking): hand plain dict rows to polars in the CloudZero and Focus exports 2026-09-03 18:46:01 -07:00
devin-ai-integration[bot]
16db51e2cf
feat(caching): add semantic_cache_scope to isolate semantic cache hits per end user (#39590)
Semantic cache keys omit the prompt, so every end user behind one virtual key
shares a bucket and can be served another user's semantically similar response.
Add an opt-in cache_params.semantic_cache_scope (key | end_user) that appends the
authenticated end-user id to the tenant scope, read from metadata and
litellm_metadata so /v1/chat/completions, /v1/responses and /v1/messages are all
covered, falling back to the key scope when no end-user id is present. Expose the
setting in the cache settings API and the Admin UI cache settings form

Co-authored-by: yassin <yassin@berri.ai>
Co-authored-by: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>
2026-09-03 18:44:54 -07:00
Cursor Agent
048499cdf5
fix(spend): only reverse-hash export rows whose key alias join missed
Team and service keys often have no user_email after a successful token
join. Treating empty email as a miss hashed every verification token on
routine CloudZero and Focus exports.

Co-authored-by: Mateo Wang <mateo-berri@users.noreply.github.com>
2026-09-04 01:44:18 +00:00
mateo-berri
1763052ff5 Merge branch 'litellm_internal_staging' of https://github.com/BerriAI/litellm into litellm_govcloud_profiles_lit6421 2026-09-03 18:41:00 -07:00
mateo-berri
43f31b0a4b feat(pricing): add GovCloud Claude Opus 5 and us-gov. inference profile rows 2026-09-03 18:40:59 -07:00