Commit graph

16900 commits

Author SHA1 Message Date
Yuneng Jiang
9e0659212a
test(e2e): repair two suites broken by intentional behaviour changes
Both of these are e2e assumptions that PRs #31731 and #39532 invalidated, not
product regressions. They have been red in litellm-e2e builds 119-123.

Wildcard readiness probe (6 errors in test_model_access_group_e2e.py)

#31731 made _get_wildcard_models drop a wildcard route from /v1/models
unconditionally; before it, a wildcard with a matching router deployment stayed
in the list and only the no-router / no-deployment fallbacks removed it. The
shared readiness helper polls /v1/models for an exact id match, so registering
openai/gpt-5.4* now times out at model_servable_timeout every run and every
test in the class errors in setup.

return_wildcard_routes=True still re-adds the route, so the poll asks for it.
The flag is a no-op for a concrete model name -- it only ever adds wildcard
entries -- so it is set unconditionally rather than sniffing the name.

Semantic auto-router spend assertion

#39532 bills the routing embedding to the caller's key on purpose, so the key's
spend logs now legitimately carry an openai/text-embedding-3-small row and
_assert_served_only_by rejects it.

Widening the allowlist would have weakened the assertion this test exists for --
that the request reached the target deployment. Instead the embedding row is
split off and asserted separately, which turns the break into coverage for
#39532. The poll gains a predicate so it waits for the embedding row rather than
racing whichever row is written first.
2026-09-04 13:48:42 -07:00
Mateo Wang
300d335255
Merge pull request #39361 from BerriAI/litellm_fix_mantle_host_re_anchor
fix(bedrock_mantle): anchor MANTLE_HOST_RE so custom Mantle hosts are honored
2026-09-04 13:40:02 -07:00
Yuneng Jiang
50fb35e17e
fix(vector_stores): make MongoDB errors actionable on self-managed deployments
mongod serves $vectorSearch identically whether mongot runs under Atlas or beside
a self-managed deployment, so the provider already worked against on-prem. The
guidance did not: a refused connection told the operator to check their project's
IP access list and whether the cluster was paused, neither of which exists outside
Atlas, and the index errors claimed an "Atlas Vector Search index" they do not have.
Every message now names a remedy for both, keeping the Atlas-specific hint labelled
as such.

Also diagnoses unescaped credentials, which self-managed deployments hit more often
because the password is usually generated. pymongo reports those three different
ways and none of them mentions the password: '@', ':' and '%' raise an RFC 3986
complaint, '/' is read as the database separator and surfaces as Bad database name,
and an unescaped ':' looks like a bad port and comes back as a plain ValueError.
All three now point at the credentials. The ValueError branch's comment claimed it
fired on an unescaped '/', which pymongo actually reports as InvalidURI; corrected
to the port parse it really catches.

Verified against a self-managed mongod 8.0 with mongot, reached over plain
mongodb:// with no SRV and no TLS: 13 cases with live OpenAI embeddings, and 4
credential cases against an auth-enabled instance whose password holds % @ / and :.
list_search_indexes returns the same queryable and status fields there as on Atlas,
so the index-readiness check needed no change.
2026-09-04 13:39:54 -07:00
tin-berri
8beca1d58d
fix(auto-router): route 1M complex tier to GPT Sol (#39797)
* feat(ui): add 1M context auto-router preset

* feat(ui): use heuristic v2 for 1M preset

* fix(ui): keep 1M preset test within lint budget

* fix(auto-router): route 1M complex tier to GPT Sol

* test(auto-router): update 1M complex tier expectation
2026-09-04 13:33:06 -07:00
ryan-crabbe-berri
189bd857f4
Merge pull request #39794 from BerriAI/litellm_v2_org_update_public
feat(organization): expose PATCH /v2/organization/{organization_id} in the OpenAPI schema
2026-09-04 13:32:52 -07:00
ryan-crabbe-berri
38a1b44993
Merge pull request #39793 from BerriAI/litellm_v2_org_update_validation
fix(organization): reject negative limits and unparseable budget_duration on PATCH /v2/organization
2026-09-04 13:32:07 -07:00
ryan-crabbe-berri
a502c728ac
Merge pull request #39670 from BerriAI/litellm_fix_org_update_null_budget_limits
fix(organization): clear org budget limits when PATCH /organization/update sends null
2026-09-04 13:31:56 -07:00
Mateo Wang
338a37d8cd
Merge pull request #39632 from BerriAI/litellm_lit6874_fireworks_perplexity_off_peak_pricing
fix(cost): honor off_peak_pricing in the fireworks_ai and perplexity cost calculators
2026-09-04 13:20:52 -07:00
Mateo Wang
44b1cc7b0f
Merge pull request #39589 from BerriAI/litellm_fix_v1_messages_midstream_timeout_failure_logging
fix(proxy): log mid-stream /v1/messages failures as failures with partial usage
2026-09-04 13:20:30 -07:00
moe-berri
6234399f9e fix(router): keep circuit-open fallbacks out of session pins
An open classifier circuit routed through the ordinary heuristic or
classifier_fallback path, and both causes are pin-worthy, so a session
whose turn landed on the cooldown fallback held that model for the whole
session_affinity TTL and never reclassified after the breaker closed.

The circuit-open signal now blocks the pin, and _classifier_failure_outcome
tags its outcomes through one helper instead of reassigning a Final.
2026-09-04 13:16:34 -07:00
ryan-crabbe-berri
7a717740dd test(organization): assert rejected values write nothing to the DB 2026-09-04 13:11:31 -07:00
mateo
50d6b26a86 fix(registry): mark baseten GLM-5.3 as vision-capable per Baseten vision docs
Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>
2026-09-04 20:10:58 +00:00
ryan-crabbe-berri
0eb2363074 test(organization): assert route publicity through the production OpenAPI generator 2026-09-04 13:10:37 -07:00
ryan-crabbe-berri
3e4a884b25 feat(organization): expose PATCH /v2/organization/{organization_id} in the OpenAPI schema 2026-09-04 13:03:18 -07:00
ryan-crabbe-berri
976ff0a785 fix(organization): 422 on negative limits and unparseable budget_duration in v2 update 2026-09-04 12:58:33 -07:00
mateo
f5157a63eb test: allow 128k and 256k tiered cache fields in registry schema test
Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>
2026-09-04 19:56:30 +00:00
devin-ai-integration[bot]
205a5e9d6c
feat(mcp): use x-mcp-<access_group>-* headers as default upstream credentials for group members (#39717)
Co-authored-by: yassin <yassin@berri.ai>
Co-authored-by: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>
2026-09-04 12:45:24 -07:00
mateo-berri
8debf8294c fix(batches): resolve a legacy row's org from the key's organization_id and keep the Batch label on grouped rows 2026-09-04 12:42:39 -07:00
yuneng-jiang
4ad4db22f1
Merge pull request #39773 from BerriAI/litellm_/pr-39770-test-failure-c4aacd
test(caching): drive the redis stall burst off the clock, not asyncio.wait_for
2026-09-04 12:33:01 -07:00
mateo
93abc3a0cd Merge remote-tracking branch 'origin/litellm_internal_staging' into litellm_registry_audit_2026_09_02 2026-09-04 19:02:44 +00:00
devin-ai-integration[bot]
dd01abc439
feat(team): report per-user spend within a team for JWT traffic (#39771)
* feat(team): report per-user spend within a team for JWT traffic

Add GET /team/spend/by_user, which groups raw spend logs by (team_id, user)
so JWT/SSO requests with no virtual key are attributed to the user inside
each selected team. Team admins see every member, plain members see only
their own row. The Team Usage page gets a Spend Per User Within Team card
with CSV export backed by the same endpoint.

Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>

* test(team): cover /team/spend/by_user in behavior suite, tf audit allowlist and EntityUsage unit test

Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>

* refactor(team): drop explanatory docstrings from /team/spend/by_user and regen schema.d.ts

Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>

---------

Co-authored-by: yassin <yassin@berri.ai>
Co-authored-by: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>
2026-09-04 12:00:47 -07:00
mateo-berri
7e51fbc819 Merge remote-tracking branch 'origin/litellm_internal_staging' into litellm_batch_ui_logs
# Conflicts:
#	tests/proxy_unit_tests/test_check_batch_cost.py
2026-09-04 12:00:14 -07:00
Yuneng Jiang
2042364fc2
fix(proxy): strip every TypedDict qualifier before numeric form-field detection
_numeric_form_type only peeled a single ReadOnly layer, so a field still
wrapped in Required/NotRequired was read as non-numeric and dropped from the
mapping. Which qualifiers survive get_type_hints varies by interpreter version
and by include_extras, so on Python 3.10 NotRequired[ReadOnly[int]] reached the
check intact and the field was silently skipped, which is what turns the mapped
test red on the 3.10 leg only.

Peel Required/NotRequired/ReadOnly/Annotated in any order and nesting instead.
The one production caller feeds a schema with no qualifiers, so the resulting
mapping is unchanged on every interpreter in the matrix, but a field written the
house-convention way stops being dropped.
2026-09-04 11:50:28 -07:00
ryan-crabbe-berri
caa1ab0e60 fix(guardrails): store an unpriced Bedrock counter as unknown, not free
A counter missing from the cost map entry was priced at 0.0 per unit, so
the rollup recorded it as known-free usage. It now stamps None for that
counter and the rollup writes NULL, while the per-request guardrail_cost
that feeds spend and budgets still sums only the known prices.

Claude-Session: https://claude.ai/code/session_01EX13mWex6RaBo9PYnkAtFW
2026-09-04 11:27:57 -07:00
ryan-crabbe-berri
4f3b02360e test(organization): type the legacy update helper's request body precisely 2026-09-04 11:21:37 -07:00
Yuneng Jiang
7d3b03d006
test(caching): drive the redis stall burst off the clock, not asyncio.wait_for
test_event_loop_stall_timeout_burst_keeps_breaker_closed built its timeout
burst by wrapping a healthy fake call in asyncio.wait_for. Before 3.12,
wait_for returns the inner result when the inner future also completed while
the loop was blocked, so no call timed out, the burst never materialised, and
the test's own liveness guard failed with 0 >= 3.

The fake now checks its own client deadline against the clock, the way a client
library does, so the stall produces a real redis TimeoutError burst on every
interpreter. The breaker itself is unchanged: its duration gate is plain
time.time() bookkeeping and never depended on the version.
2026-09-04 11:09:01 -07:00
yuneng-jiang
f74bc9427b
Merge pull request #39770 from BerriAI/litellm_/chronic-test-failures-e0994f
test: repair four chronically failing CI tests
2026-09-04 11:01:52 -07:00
Mateo Wang
04a198e3e3
Merge pull request #39568 from BerriAI/litellm_fix-batch-spend-key-double-hash-bcae
fix(spend-tracking): keep batch spend keys joinable after v1.99 provenance gate
2026-09-04 10:47:34 -07:00
Yuneng Jiang
a5a78670d0
test: address review notes on the chronic-test repairs
Drop the two new docstrings, annotate the new locals Final, and replace the
mutable call recorder with a rebuild stub that fails the test if it is ever
reached.
2026-09-04 10:18:15 -07:00
Yuneng Jiang
dbf8fe0f4e
test: repair four chronically failing CI tests
test_no_linear_scans_in_router: #39468 added config_deployments() and
heuristic_v2_router_limit_violation(), which both scan the whole model_list
from admin-only paths (model add/upsert), so add them to the allowlist. The
allowlist becomes a mapping so each exemption carries its reason as data.

test_missing_model_parameter_curl: a request with no model is rejected by the
proxy when nothing can serve it and by the router when a wildcard or default
deployment exists, and by the upstream provider when a wildcard forwards it,
so the message text is not a stable contract. Assert the contract that holds
in every case: HTTP 400 with a non-empty error message.

test_model_group_info_e2e: /model_group/info resolves wildcards, so it can
never return "anthropic/*" verbatim. cc3f9cd65b rewrote the assertion to
expect the raw pattern after claude-3-5-haiku-20241022 left the price map,
which made it unsatisfiable. Assert the expansion instead.

test_should_derive_ocr_mapping_status_from_live_tests: the audit needs a
native bridge built with the trace-parity feature, which CI never builds, so
skip with the harness's own diagnostic instead of erroring. Extract that
check out of ensure_trace_bridge as trace_bridge_error so a pytest run
reports the state without kicking off a maturin rebuild.
2026-09-04 10:10:47 -07:00
Krrish Dholakia
2f1da035ae fix(fireworks_ai): keep generic capability fallback for reasoning and tool_choice
Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>
2026-09-04 17:06:01 +00:00
mateo-berri
bef3585d82 feat(pricing): add GovCloud rows for every live but unpriced Bedrock model
Every model bedrock list-foundation-models and list-inference-profiles
report as live in us-gov-west-1 or us-gov-east-1 now has a priced row:
Claude Fable 5.1 (profile plus in-region), Nemotron Nano 9B (profile plus
in-region), Grok 4.6 (profile plus Mantle in both regions), the us-gov.
Claude 3 Haiku profile in the east, Nova Lite, Micro and the Nova 2
multimodal embeddings in the west, and the Gemma 4 and gpt-oss Mantle
SKUs the GovCloud offer files price. Offer-file rates are used where AWS
publishes them; Claude rows carry the 1.2x GovCloud premium.
2026-09-04 10:03:43 -07:00
mateo-berri
8b3faa6ed8 fix(proxy): keep /health inside the caller's team for unrestricted non-admin keys 2026-09-04 09:53:53 -07:00
mateo-berri
69a45e81cc fix(proxy): refuse /health targets outside the caller's scope instead of probing the rest 2026-09-04 09:40:00 -07:00
mateo-berri
471f51cc4d fix(proxy): keep another team's deployment out of /health for keys with no team 2026-09-04 09:29:09 -07:00
mateo-berri
46974fe46e Merge branch 'litellm_internal_staging' of https://github.com/BerriAI/litellm into litellm_govcloud_profiles_lit6421 2026-09-04 09:28:33 -07:00
mateo-berri
6c27754455 fix(anthropic): bill an uncostable partial pass-through stream at zero cost instead of dropping its usage 2026-09-04 09:26:51 -07:00
Krrish Dholakia
788efea7b3 fix(fireworks_ai): resolve tool_choice/reasoning support for short model names
Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>
2026-09-04 16:25:15 +00:00
Yujong Lee
6675e1fd9c fix: preserve Python 3.10 harness compatibility 2026-09-04 09:13:05 -07:00
Yujong Lee
fae3d224eb Merge remote-tracking branch 'origin/litellm_internal_staging' into litellm_python_version_ci
# Conflicts:
#	basedpyright-code-budget.json
#	tests/sdk_function_trace/profiler.py
#	tests/sdk_function_trace/test_profiler.py
2026-09-04 09:01:13 -07:00
yujonglee
b75ac5cf52
feat(python): rename Rust rollout API (#39704) 2026-09-04 08:40:44 -07:00
mateo
0c29f510bc fix(registry): drop gpt-image-2 text output price, add openrouter minimax-m3 and qwen3.7-plus
Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>
2026-09-04 15:03:29 +00:00
mateo
8e83d6d63d fix(model_prices): add Databricks Sep-2026 catalog, Azure gpt-realtime-2.x, per-token realtime image pricing
Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>
2026-09-04 14:06:54 +00:00
mateo
af4340dce3 Merge remote-tracking branch 'origin/litellm_internal_staging' into litellm_registry_audit_2026_09_02 2026-09-04 13:03:22 +00:00
Atharva-Kanherkar
0b34abe8fe fix(anthropic): harden refusal translation 2026-09-04 17:02:07 +05:30
amasen02
3623aecc64 style(proxy): add Final type annotations to enduser budget reset variables 2026-09-04 16:33:57 +05:30
amasen02
daced81f20 fix(proxy): invalidate end-user spend counter and cache on budget reset (#39726)
Signed-off-by: amasen02 <amasen02@users.noreply.github.com>
2026-09-04 15:37:53 +05:30
mateo
03a82823fc test: deflake redis loop-stall burst test and pre-commit interrupt cleanup
The redis breaker test raced the event loop: the fake call had to still be
pending when a real time.sleep stall began, which needs the loop to get from
scheduling to the stall in under 1ms. The fake now holds its answer behind an
asyncio.Event so the whole burst times out deterministically.

The pre-commit interrupt test found a real leak: lint_dashboard creates its
eslint report with mktemp and only removed it on the happy path, so an
interrupt landing during the whole-folder eslint run left the file behind.
The subshell now removes it from an EXIT trap, and the test drives the
interrupt while that eslint run is in flight.

Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>
2026-09-04 10:07:34 +00:00
mateo
9835883e03 Merge remote-tracking branch 'origin/litellm_internal_staging' into litellm_deflake_20260902
Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>
2026-09-04 09:20:07 +00:00
Atharva-Kanherkar
200e2901d6 fix(anthropic_responses): preserve Responses refusal blocks in Anthropic translation
When OpenAI Responses returns a refusal content block, Anthropic /v1/messages
erased the refusal text into an empty content array and emitted stop_reason 'end_turn'.
Translate refusal blocks to Anthropic text blocks, map stop_reason to 'refusal',
and add 'refusal' to AnthropicFinishReason.

Fixes #39721
2026-09-04 14:47:56 +05:30