Commit graph

47584 commits

Author SHA1 Message Date
mateo
93abc3a0cd Merge remote-tracking branch 'origin/litellm_internal_staging' into litellm_registry_audit_2026_09_02 2026-09-04 19:02:44 +00:00
devin-ai-integration[bot]
dd01abc439
feat(team): report per-user spend within a team for JWT traffic (#39771)
* feat(team): report per-user spend within a team for JWT traffic

Add GET /team/spend/by_user, which groups raw spend logs by (team_id, user)
so JWT/SSO requests with no virtual key are attributed to the user inside
each selected team. Team admins see every member, plain members see only
their own row. The Team Usage page gets a Spend Per User Within Team card
with CSV export backed by the same endpoint.

Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>

* test(team): cover /team/spend/by_user in behavior suite, tf audit allowlist and EntityUsage unit test

Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>

* refactor(team): drop explanatory docstrings from /team/spend/by_user and regen schema.d.ts

Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>

---------

Co-authored-by: yassin <yassin@berri.ai>
Co-authored-by: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>
2026-09-04 12:00:47 -07:00
Yuneng Jiang
2042364fc2
fix(proxy): strip every TypedDict qualifier before numeric form-field detection
_numeric_form_type only peeled a single ReadOnly layer, so a field still
wrapped in Required/NotRequired was read as non-numeric and dropped from the
mapping. Which qualifiers survive get_type_hints varies by interpreter version
and by include_extras, so on Python 3.10 NotRequired[ReadOnly[int]] reached the
check intact and the field was silently skipped, which is what turns the mapped
test red on the 3.10 leg only.

Peel Required/NotRequired/ReadOnly/Annotated in any order and nesting instead.
The one production caller feeds a schema with no qualifiers, so the resulting
mapping is unchanged on every interpreter in the matrix, but a field written the
house-convention way stops being dropped.
2026-09-04 11:50:28 -07:00
ryan-crabbe-berri
d05d2a6f05 fix(ui): clamp server-paginated DataTable page index when rowCount shrinks
Server-mode tables kept whatever page index the user was on after the
server's total dropped below it, for example after deleting the last
rows of the final page or when a refetch came back empty. The footer
then read "Page 2 of 1" and "Showing 26-25 of 25" with Previous and
First enabled over an empty body, and every one of the 13 server-mode
consumers was exposed since none of them clamped

The shared DataTable now snaps the controlled page index to the last
valid page as soon as a non-loading rowCount no longer reaches it, so
the fix applies to every consumer without per-table clamps. Loading
responses are ignored so a pending fetch never bounces the user to
page 1
2026-09-04 11:33:26 -07:00
ryan-crabbe-berri
caa1ab0e60 fix(guardrails): store an unpriced Bedrock counter as unknown, not free
A counter missing from the cost map entry was priced at 0.0 per unit, so
the rollup recorded it as known-free usage. It now stamps None for that
counter and the rollup writes NULL, while the per-request guardrail_cost
that feeds spend and budgets still sums only the known prices.

Claude-Session: https://claude.ai/code/session_01EX13mWex6RaBo9PYnkAtFW
2026-09-04 11:27:57 -07:00
ryan-crabbe-berri
4f3b02360e test(organization): type the legacy update helper's request body precisely 2026-09-04 11:21:37 -07:00
Yuneng Jiang
7d3b03d006
test(caching): drive the redis stall burst off the clock, not asyncio.wait_for
test_event_loop_stall_timeout_burst_keeps_breaker_closed built its timeout
burst by wrapping a healthy fake call in asyncio.wait_for. Before 3.12,
wait_for returns the inner result when the inner future also completed while
the loop was blocked, so no call timed out, the burst never materialised, and
the test's own liveness guard failed with 0 >= 3.

The fake now checks its own client deadline against the clock, the way a client
library does, so the stall produces a real redis TimeoutError burst on every
interpreter. The breaker itself is unchanged: its duration gate is plain
time.time() bookkeeping and never depended on the version.
2026-09-04 11:09:01 -07:00
Yuneng Jiang
4bd3cd9e06
ci: report every failing test in a job instead of stopping at the first
Drops `-x` from all 24 pytest invocations in .circleci/config.yml. With
`-x`, a job stops at its first failure, so a second broken test in the
same suite stays invisible until the first is fixed and CI is re-run.
That turns one round trip into N when a job has several broken tests.

This is exactly what happened in #39770: fixing
test_missing_model_parameter_curl in
tests/store_model_in_db_tests/test_openai_error_handling.py immediately
unmasked test_chat_completion_bad_model_with_spend_logs in the same
file, which had been failing for a long time without ever being
reported.

Only `-x` is removed; -v/-vv/-s/-n/--reruns and every other flag are
untouched.
2026-09-04 11:05:58 -07:00
yuneng-jiang
f74bc9427b
Merge pull request #39770 from BerriAI/litellm_/chronic-test-failures-e0994f
test: repair four chronically failing CI tests
2026-09-04 11:01:52 -07:00
Mateo Wang
04a198e3e3
Merge pull request #39568 from BerriAI/litellm_fix-batch-spend-key-double-hash-bcae
fix(spend-tracking): keep batch spend keys joinable after v1.99 provenance gate
2026-09-04 10:47:34 -07:00
Yuneng Jiang
a5a78670d0
test: address review notes on the chronic-test repairs
Drop the two new docstrings, annotate the new locals Final, and replace the
mutable call recorder with a rebuild stub that fails the test if it is ever
reached.
2026-09-04 10:18:15 -07:00
Yuneng Jiang
dbf8fe0f4e
test: repair four chronically failing CI tests
test_no_linear_scans_in_router: #39468 added config_deployments() and
heuristic_v2_router_limit_violation(), which both scan the whole model_list
from admin-only paths (model add/upsert), so add them to the allowlist. The
allowlist becomes a mapping so each exemption carries its reason as data.

test_missing_model_parameter_curl: a request with no model is rejected by the
proxy when nothing can serve it and by the router when a wildcard or default
deployment exists, and by the upstream provider when a wildcard forwards it,
so the message text is not a stable contract. Assert the contract that holds
in every case: HTTP 400 with a non-empty error message.

test_model_group_info_e2e: /model_group/info resolves wildcards, so it can
never return "anthropic/*" verbatim. cc3f9cd65b rewrote the assertion to
expect the raw pattern after claude-3-5-haiku-20241022 left the price map,
which made it unsatisfiable. Assert the expansion instead.

test_should_derive_ocr_mapping_status_from_live_tests: the audit needs a
native bridge built with the trace-parity feature, which CI never builds, so
skip with the harness's own diagnostic instead of erroring. Extract that
check out of ensure_trace_bridge as trace_bridge_error so a pytest run
reports the state without kicking off a maturin rebuild.
2026-09-04 10:10:47 -07:00
Krrish Dholakia
dad1b13225 style(fireworks_ai): format capability fallback
Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>
2026-09-04 17:07:59 +00:00
Krrish Dholakia
2f1da035ae fix(fireworks_ai): keep generic capability fallback for reasoning and tool_choice
Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>
2026-09-04 17:06:01 +00:00
mateo-berri
bef3585d82 feat(pricing): add GovCloud rows for every live but unpriced Bedrock model
Every model bedrock list-foundation-models and list-inference-profiles
report as live in us-gov-west-1 or us-gov-east-1 now has a priced row:
Claude Fable 5.1 (profile plus in-region), Nemotron Nano 9B (profile plus
in-region), Grok 4.6 (profile plus Mantle in both regions), the us-gov.
Claude 3 Haiku profile in the east, Nova Lite, Micro and the Nova 2
multimodal embeddings in the west, and the Gemma 4 and gpt-oss Mantle
SKUs the GovCloud offer files price. Offer-file rates are used where AWS
publishes them; Claude rows carry the 1.2x GovCloud premium.
2026-09-04 10:03:43 -07:00
mateo
e7c29351e8 chore: ratchet lint budgets after merging litellm_internal_staging
Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>
2026-09-04 16:35:26 +00:00
mateo-berri
46974fe46e Merge branch 'litellm_internal_staging' of https://github.com/BerriAI/litellm into litellm_govcloud_profiles_lit6421 2026-09-04 09:28:33 -07:00
mateo-berri
6c27754455 fix(anthropic): bill an uncostable partial pass-through stream at zero cost instead of dropping its usage 2026-09-04 09:26:51 -07:00
Krrish Dholakia
788efea7b3 fix(fireworks_ai): resolve tool_choice/reasoning support for short model names
Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>
2026-09-04 16:25:15 +00:00
mateo
41cb19df55 chore: merge litellm_internal_staging into litellm_techdebt_20260903
Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>
2026-09-04 16:23:52 +00:00
Mateo Wang
2e734004f9
Merge pull request #39399 from BerriAI/litellm_python_version_ci
fix: restore Python compatibility and test 3.10 through 3.14
2026-09-04 09:21:52 -07:00
Yujong Lee
6675e1fd9c fix: preserve Python 3.10 harness compatibility 2026-09-04 09:13:05 -07:00
Yujong Lee
fae3d224eb Merge remote-tracking branch 'origin/litellm_internal_staging' into litellm_python_version_ci
# Conflicts:
#	basedpyright-code-budget.json
#	tests/sdk_function_trace/profiler.py
#	tests/sdk_function_trace/test_profiler.py
2026-09-04 09:01:13 -07:00
yujonglee
b75ac5cf52
feat(python): rename Rust rollout API (#39704) 2026-09-04 08:40:44 -07:00
yujonglee
7276caecd4
refactor(rust): extract config crate (#39706)
* refactor(rust): extract config crate

* refactor(config): split crate modules

* refactor(gateway): remove gil health counter
2026-09-04 08:17:07 -07:00
mateo
0c29f510bc fix(registry): drop gpt-image-2 text output price, add openrouter minimax-m3 and qwen3.7-plus
Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>
2026-09-04 15:03:29 +00:00
mateo
c707f2fe5d fix(model_prices): databricks gpt-5-3-codex is served via the Responses API
Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>
2026-09-04 14:32:11 +00:00
mateo
8e83d6d63d fix(model_prices): add Databricks Sep-2026 catalog, Azure gpt-realtime-2.x, per-token realtime image pricing
Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>
2026-09-04 14:06:54 +00:00
mateo
af4340dce3 Merge remote-tracking branch 'origin/litellm_internal_staging' into litellm_registry_audit_2026_09_02 2026-09-04 13:03:22 +00:00
Atharva-Kanherkar
7399b3844d fix(anthropic): satisfy lint budget gates for refusal translation 2026-09-04 17:33:06 +05:30
Atharva-Kanherkar
0b34abe8fe fix(anthropic): harden refusal translation 2026-09-04 17:02:07 +05:30
amasen02
9c8594c7b8 style(proxy): format reset_budget_job with ruff 2026-09-04 16:37:58 +05:30
amasen02
3623aecc64 style(proxy): add Final type annotations to enduser budget reset variables 2026-09-04 16:33:57 +05:30
amasen02
daced81f20 fix(proxy): invalidate end-user spend counter and cache on budget reset (#39726)
Signed-off-by: amasen02 <amasen02@users.noreply.github.com>
2026-09-04 15:37:53 +05:30
mateo
03a82823fc test: deflake redis loop-stall burst test and pre-commit interrupt cleanup
The redis breaker test raced the event loop: the fake call had to still be
pending when a real time.sleep stall began, which needs the loop to get from
scheduling to the stall in under 1ms. The fake now holds its answer behind an
asyncio.Event so the whole burst times out deterministically.

The pre-commit interrupt test found a real leak: lint_dashboard creates its
eslint report with mktemp and only removed it on the happy path, so an
interrupt landing during the whole-folder eslint run left the file behind.
The subshell now removes it from an EXIT trap, and the test drives the
interrupt while that eslint run is in flight.

Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>
2026-09-04 10:07:34 +00:00
mateo
9835883e03 Merge remote-tracking branch 'origin/litellm_internal_staging' into litellm_deflake_20260902
Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>
2026-09-04 09:20:07 +00:00
Atharva-Kanherkar
200e2901d6 fix(anthropic_responses): preserve Responses refusal blocks in Anthropic translation
When OpenAI Responses returns a refusal content block, Anthropic /v1/messages
erased the refusal text into an empty content array and emitted stop_reason 'end_turn'.
Translate refusal blocks to Anthropic text blocks, map stop_reason to 'refusal',
and add 'refusal' to AnthropicFinishReason.

Fixes #39721
2026-09-04 14:47:56 +05:30
mateo
8b37de14b1 refactor: drop fresh Any annotations and suppressions from admission control, spend summary, and dual cache
Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>
2026-09-04 08:05:04 +00:00
mateo
b287d915b0 merge: bring litellm_internal_staging into the rolling tech debt branch
Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>
2026-09-04 07:45:29 +00:00
ryan
7783efc3b1 chore: merge litellm_internal_staging into litellm_lit_4738_table_pagination
Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>
2026-09-04 05:39:32 +00:00
yuneng-jiang
c8635ecc67
feat: page the public model hub table off /public/v1/model_hub, keeping every filter (#39691)
* feat(ui): page the public model hub table off /public/v1/model_hub

The public Model Hub page loaded every published model group in one call and
did all of its searching, sorting and filtering in the browser, so a proxy with
a few thousand groups sent megabytes to render one screen.

The models table now asks /public/v1/model_hub for one page at a time. Paging,
sorting, search and the provider and mode filters are query parameters on that
route, and the pagination footer counts from the response envelope's
total_count rather than the rows on screen. Column sortability is derived from
the fields the route declares sortable, so a header can no longer ask it for a
sort it answers with a 400.

The feature filter is dropped: supports_* are booleans and the route has no
boolean filter, so it could only ever have filtered the page in view.

* fix(ui): offer the model modes litellm actually prices in the hub filter

The mode filter listed 'moderations', which no model group's mode is ever set
to, so picking it could only ever return nothing; 'anthropic_messages' was
dead the same way. Four real modes the catalogue does use, search, ocr,
guardrail and vector_store, were missing entirely.

The list is now the mode vocabulary in model_prices_and_context_window.json,
and a test reads that file so an option that matches nothing, or a mode with
no option, fails instead of silently filtering to an empty table. A failed
page fetch logs the route's error detail again, as it did before the table
moved to the paginated route.

* test(ui): keep the model hub health rows out of the inline-object budget

frontend-lint's local/no-large-inline-object-arg budget went 555 to 556: the
health check rows became arguments to the row helper. They are plain literals
spreading a shared default again, which is what they were before, and the gate
reports 554 against a max of 555.

* feat: keep every model hub filter when the table pages

Moving the table onto /public/v1/model_hub cost it two controls the route
could not serve: the provider filter fell back to one substring because
providers only declared contains, and the feature filter went away entirely
because supports_* are booleans with no filter at all. Options for the
dropdowns went with them, since a page of rows only knows the values on that
page.

The route now declares providers in, a features field whose value is the
capability names a row has, and providers, rpm and tpm as sortable. Features
is one repeated field rather than a boolean per flag so selecting two of them
matches either, which is what the multi-select has always meant. Three facet
routes serve the distinct providers, modes and features across the published
groups, carrying the parent's filters, per section 12 of the list design.

All of it is additive: the route rejects unknown parameters, so no request
that worked before changes, and the design's stability policy calls new
filters and parameters safe within a version.

Health status stays unsortable. Health is read for the rows on the page, and
ordering the match set by it would mean reading it for every published group,
which is the cost the paging exists to avoid.

* fix(ui): put the model hub facet types where the generator emits them

The generated file lists paths in sorted order and operations in path order.
Both new blocks were spliced in one entry too late, after
/queue/chat/completions rather than before it, so the schema.d.ts sync check
regenerated the file and found them misplaced. Same blocks, byte for byte,
moved to the position the generator gives them.

* fix(proxy): type a facet payload as the sequence the framework hands it

The lint job's basedpyright gate flagged one new reportArgumentType: handle_facet
passes a tuple, and FacetListResponse declared data as list[str]. A list would
have traded that error for an LIT002 mutable construction, and both budgets are
already at their ceiling on the base.

Sequence[str] is what the framework actually produces and what the model always
accepted: pydantic emits the same array schema either way, verified against
model_json_schema, so the OpenAPI spec and schema.d.ts are unchanged, and the
existing list-passing caller in spend_logs still type checks.

* test(proxy): pin the facet route's rejection contract

handle_facet answers six ways before it ever reaches the executor, and
none of them was covered: a denied scope, a filter operator the spec does
not offer, a repeated parameter, a non-positive page or page_size, and the
where clause those last two feed. Every one is a 400 or 403 an
unauthenticated caller can reach, so each gets a test that fails when the
branch stops firing.
2026-09-03 22:36:16 -07:00
mateo
08bfdadb10 chore: merge litellm_internal_staging into litellm_registry_audit_2026_09_02, drop the deleted ocr ledger
Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>
2026-09-04 04:18:09 +00:00
yujonglee
eb67e5402b
test(ocr): complete Rust unit test parity (#39689)
* feat(ocr): complete Rust unit test parity

* test(ocr): keep parity changes harness-only

* test(ocr): share azure DI native fixture across response tests

* refactor(ocr): colocate gateway unit tests and extract lifecycle integration tests
2026-09-03 21:15:01 -07:00
yujonglee
ee08c36fc0
refactor(tests): restructure rust python harness around strategy definitions (#39628)
* wip

* refactor(tests): move sdk function tracing into rust python harness

* dead code

* fix: handle harness keyboard interrupts

* refactor(tests): deduplicate rust python harness helpers

* fix(harness): expose validated strategy choices

* wip

* refactor(harness): let strategies own parity reports

* docs(harness): update strategy structure

* refactor(harness): localize strategy report views

* wip

* fix(harness): satisfy mapping runner type checks

* fix(harness): clarify trace parity output

* wip

* fix(harness): clarify unit mapping report

* fix(harness): finalize trace parity contracts

* refactor(harness): structure parity contracts

* feat: derive unit test mapping from traces

* feat(harness): map rstest test families

* feat(ocr): port Azure document intelligence tests

* feat(harness): enforce complete unit mappings

* feat(ocr): add reducto core transforms

* feat(harness): classify host-only unit tests

* fix(ocr): complete Rust provider plumbing

* fix(harness): reuse OCR parity workers
2026-09-03 21:15:01 -07:00
yujonglee
8fc0663198
fix(ci): pin setup-uv to v10.0.1 (#39111) 2026-09-03 20:44:24 -07:00
moe-berri
04c6ce9ac3
Merge pull request #39693 from BerriAI/litellm_auto_setup_simple
feat(ui): add one-click Auto Router setup
2026-09-03 20:28:01 -07:00
moe-berri
81dd911bdc fix(router): recognize asyncio classifier timeouts 2026-09-03 20:19:13 -07:00
moe-berri
5a2845d183 fix(router): preserve classifier breaker state under concurrency 2026-09-03 20:10:02 -07:00
moe-berri
d5481ca037 fix(ui): respect reported reasoning efforts 2026-09-03 20:01:44 -07:00
moe-berri
510424c86c feat(router): add classifier circuit breaker 2026-09-03 19:59:14 -07:00