Commit graph

18968 commits

Author SHA1 Message Date
mateo-berri
ce722ab1b3 fix(proxy): evict the member's cached user row on team member add
JWT auth caches the user row before it adds the user to the JWT's team, and admission checks the credential's team against the cached row on whichever worker takes the next request. On a two-worker gateway the credential minted for a newly joined team answered 403 "not in your team memberships" until the management-object TTL ran out, because /team/member_add only evicted the membership spend sentinel. The add now evicts the added members' cached user rows and broadcasts the eviction to the other workers, the way /team/member_delete already did

The mint test now also covers a user SCIM deactivated after the cache last saw them active: the database read refuses the mint while the cached row still says active
2026-09-16 16:31:28 -07:00
Mateo Wang
390ab45448
Merge pull request #37506 from BerriAI/litellm_dashscope_reasoning_effort
Some checks are pending
LiteLLM Rust / rust-wheel (push) Waiting to run
Unit Tests: Documentation Validation / documentation (push) Waiting to run
Unit Tests: Proxy DB Operations / budgets (push) Blocked by required conditions
Unit Tests: Proxy DB Operations / assert-shard-coverage (push) Waiting to run
Unit Tests: Proxy DB Operations / auth-checks (push) Blocked by required conditions
Unit Tests: Proxy DB Operations / custom-logging (push) Blocked by required conditions
Unit Tests: Proxy DB Operations / db-and-spend (push) Blocked by required conditions
Unit Tests: Proxy DB Operations / endpoints-and-responses (push) Blocked by required conditions
Unit Tests: Proxy DB Operations / guardrails-hooks (push) Blocked by required conditions
Unit Tests: Proxy DB Operations / jwt-and-keys (push) Blocked by required conditions
Unit Tests: Proxy DB Operations / key-generation (push) Blocked by required conditions
Unit Tests: Proxy DB Operations / logging-misc (push) Blocked by required conditions
Unit Tests: Proxy DB Operations / proxy-runtime (push) Blocked by required conditions
Unit Tests: Proxy DB Operations / proxy-server-core (push) Blocked by required conditions
Unit Tests: Proxy DB Operations / proxy-utils (push) Blocked by required conditions
Unit Tests / caching-local (push) Waiting to run
Unit Tests / core-utils (push) Waiting to run
Unit Tests / enterprise-package (push) Waiting to run
Unit Tests / enterprise-routing (push) Waiting to run
Unit Tests / integrations (push) Waiting to run
Unit Tests / All Other Providers (push) Waiting to run
Unit Tests / Vertex AI (push) Waiting to run
Unit Tests / misc (push) Waiting to run
Unit Tests / proxy-auth (push) Waiting to run
Unit Tests / proxy-endpoints (push) Waiting to run
Unit Tests / proxy-extras (push) Waiting to run
Unit Tests / proxy-infra (push) Waiting to run
Unit Tests / proxy-server (push) Waiting to run
Unit Tests / responses-caching-types (push) Waiting to run
GitHub Actions Security Analysis / zizmor (push) Waiting to run
fix(dashscope): forward reasoning_effort to the provider
2026-09-16 16:29:25 -07:00
mateo-berri
2f186055e6 fix(logging): decide keep-or-scrub for a log extra by comparing it to its scrubbed copy
Some checks are pending
ai-gateway image / ai-gateway release image (push) Waiting to run
The code-quality check refuses recursive functions and the walk that inspected
extras was one, so the filter no longer walks anything itself. safe_dumps now
builds its JSON-native structure through safe_json_structure, the filter scrubs
the extra through that, and the original object is kept only when the scrubbed
copy compares equal to it. Anything the serializer skipped (non-string keys,
nests past its depth, fields a repr hides) makes the copy differ, so the copy
wins. A host object whose equality raises, as numpy arrays and torch tensors
do, counts as changed instead of breaking the caller's log call
2026-09-16 16:29:23 -07:00
yassin
bddb64ddc5 test(vertex_ai): cover unterminated last row, unparseable rows, and sync stream rejection
Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>
2026-09-16 23:28:39 +00:00
ryan-crabbe-berri
14fbd623d7 fix(proxy): clear both cache partitions and carry the walk position as a value
UserApiKeyCache's batch delete ran the two partitions in sequence, so a Redis
failure on the hashed token partition returned before the ordinary management
keys were touched. Both partitions are attempted now and the first failure is
re-raised for the caller to report.

The customer walk kept its position in two locals it reassigned each page. It
now mirrors the window walk in the same file: a page helper returns where the
walk goes next, and the driver rebinds one value.
2026-09-16 16:28:16 -07:00
kerry
884e96958c fix(streaming): cap estimated reasoning tokens to the provider total and cover dict chunks
Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>
2026-09-16 23:27:48 +00:00
Moe Khalil
1c9525fdab test(router): cover deployment replacement during metadata discovery
Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>
2026-09-16 23:27:31 +00:00
kerry
15c18ad6cd fix(bedrock_mantle): accept verbosity on gpt-5.x chat completions
Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>
2026-09-16 23:27:07 +00:00
ryan-crabbe-berri
ab92a6637d test(e2e): cover team admin editable fields on /team/update
Team admins are refused until a proxy admin enables a field, then limited to the enabled fields, and resending unchanged budget settings keeps the team's budget reset times
2026-09-16 16:26:52 -07:00
joshua-berri
41410e9556
Merge pull request #41364 from BerriAI/litellm_fix_mcp_auth_fail_closed_4501
fix(mcp): fail closed on missing upstream credentials
2026-09-16 23:26:38 +00:00
Moe Khalil
b7c6befb37 feat(router): discover token limits for hosted OpenAI-compatible models
Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>
2026-09-16 23:19:54 +00:00
mateo-berri
3712d8de92 Merge remote-tracking branch 'origin/main' into litellm_model_group_info_proxy_admin_all_models
Some checks failed
LiteLLM Rust / rust-wheel (push) Has been cancelled
LiteLLM Rust / rust-lint (push) Has been cancelled
LiteLLM Rust / rust-test (push) Has been cancelled
Terraform Provider / gofmt, vet, build, test (push) Has been cancelled
Terraform Provider / Provider endpoints vs proxy OpenAPI schema (push) Has been cancelled
2026-09-16 16:18:45 -07:00
mateo-berri
2a9fa48730 fix: expand wildcard deployments for proxy admins on /model_group/info 2026-09-16 16:18:44 -07:00
kerry
813d96f26e fix(e2e): resolve remaining merge markers in e2e_config
Some checks failed
LiteLLM Rust / rust-lint (push) Has been cancelled
LiteLLM Rust / rust-test (push) Has been cancelled
LiteLLM Rust / rust-wheel (push) Has been cancelled
Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>
2026-09-16 23:18:26 +00:00
kerry
3a5121e126 Merge remote-tracking branch 'origin/main' into litellm_e2e_cost_calculation_scripted_provider
Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>

# Conflicts:
#	tests/e2e/conftest.py
#	tests/e2e/e2e_config.py
2026-09-16 23:18:20 +00:00
Yujong Lee
13cb739089 fix(rust_bridge): qualify runtime calls in dispatch and drop OCR transport rows from wheel matrix
Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>
2026-09-16 23:16:31 +00:00
ryan-crabbe-berri
cdf0142f4a fix(proxy): isolate each cache and each page in the budget reset invalidation
Greptile review follow-ups on the paged end-user cache invalidation.

UserApiKeyCache keeps hashed token keys in a second in-memory partition, and
routes delete_cache / async_delete_cache there. It inherited the new batch
delete unchanged, so a budget cascade cleared the main partition and left the
key object sitting on its pre-reset spend. Override it the way
async_set_cache_pipeline already partitions its entries.

The spend counters and the management cache shared one exception handler, so a
Redis failure on the counters returned before the management cache was touched
at all. Each cache gets its own await and its own handler now.

A failed page read returned the same empty tuple that ends the walk normally,
so a truncated pass was reported as a complete one. The window is advanced by
then and no later tick comes back for the customers past that page, so the walk
now says it was cut short and the service log carries it.
2026-09-16 16:14:08 -07:00
mateo-berri
a1ad95dbbd fix(gemini): read the minimal thinking floor from the cost map and cover the /v1/messages bridge
Some checks failed
ai-gateway image / ai-gateway release image (push) Has been cancelled
LiteLLM Rust / rust-lint (push) Has been cancelled
LiteLLM Rust / rust-test (push) Has been cancelled
2026-09-16 16:13:41 -07:00
kerry
43514a7ffe Merge remote-tracking branch 'origin/main' into litellm_fix_interrupted_anthropic_reasoning_usage 2026-09-16 23:13:06 +00:00
kerry
02bccfd89f fix(streaming): fill text_tokens when reasoning is counted from stream content
Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>
2026-09-16 23:13:05 +00:00
yuneng-jiang
765e6e498d
Merge pull request #41402 from BerriAI/litellm_/buildkite-litellm-e2e-setup-ff714d
feat(e2e): make the provider cache reusable across builds and mount Bedrock behind it
2026-09-16 16:12:29 -07:00
ryan-crabbe-berri
81ae5caa7e fix(proxy): pass only a team admin's changed fields on to the team update 2026-09-16 16:11:52 -07:00
yassin
acd4f0eb04 fix(spend_tracking): attribute router-rejected requests to the model group provider
A request for a configured model group that the router rejects before picking a deployment (all deployments in cooldown, no healthy deployment) never gets a custom_llm_provider in its logging kwargs. The spend log payload persisted an empty provider, the daily spend tables carried it through, and the Admin UI Usage page rendered those requests under unknown even though every model in the group has a provider

get_logging_payload now takes the proxy router and, when the logged provider is missing, infers it from the model group's deployments. It only attributes when every deployment in the group resolves to the same provider; mixed groups, unknown groups and a missing router leave the value empty as before. Explicitly logged providers keep precedence

Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>
2026-09-16 23:10:47 +00:00
mateo-berri
610249c663 fix(bedrock): carry the tool_call_id into the rewritten tool call text 2026-09-16 16:08:46 -07:00
yuneng-jiang
030901211e
Merge pull request #41360 from BerriAI/litellm_unpin_together_ai_serverless_model
test(together_ai): move request-shape checks to the mapped file, drop the live ones
2026-09-16 16:07:33 -07:00
yuneng-jiang
fddf83a2ac
Merge pull request #41487 from BerriAI/litellm_isolate_generic_api_ndjson_test
test(logging): pick this test's own records out of the shared log batch
2026-09-16 16:06:47 -07:00
ryan-crabbe-berri
be6e7ef173
Merge pull request #41345 from BerriAI/litellm_remove_duplicate_user_budget_hook
fix(proxy): remove duplicate user budget hook that 429'd zero-cost models
2026-09-16 16:06:41 -07:00
Yujong Lee
c85acc8d28 test(rust_bridge): cover binding validation, async upstream errors, and OCR preparation failures
Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>
2026-09-16 23:04:38 +00:00
yassin
e2e7f5f879 feat(vertex_ai): stream GCS batch output files from /v1/files/{id}/content
Vertex AI file content retrieval downloaded the whole GCS object into memory
before responding, which made large batch output files (hundreds of MB, image
generation JSONL past 4 GiB) impractical to fetch through the proxy.

Add BaseLLMHTTPHandler.async_retrieve_file_content_streaming, an httpx
stream=True path that hands the byte iterator to the provider config through
the new BaseFilesConfig.transform_file_content_stream hook and closes the
response on completion, early close, and HTTP error. VertexAIFilesConfig peeks
at the first JSONL row: Generate Content batch output is converted to OpenAI
batch format one row at a time (content-length dropped since it changes),
embeddings output stays buffered so fanned-out rows can be regrouped, and
anything else passes through with the upstream content-type and content-length.

vertex_ai joins FILE_CONTENT_STREAMING_PROVIDERS, so the proxy returns a
StreamingResponse for it while OpenAI-compatible providers and the buffered
Vertex path are unchanged.

Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>
2026-09-16 23:03:22 +00:00
yucheng-berri
659c4ce5a1
Merge pull request #40669 from BerriAI/litellm_otel_passthrough_trace_propagation
fix(otel): propagate W3C trace context on HTTP and WebSocket passthrough
2026-09-16 16:02:46 -07:00
yucheng
db97616149 chore: merge main into litellm_lit5285_login_rate_limit_v2
Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>
2026-09-16 23:02:23 +00:00
mateo-berri
02d9aae8c8 fix(proxy): stop forwarding LiteLLM credential headers on Bedrock agent-runtime passthrough
The agent-runtime branch of /bedrock/{endpoint} (agents, knowledgebases, flows,
retrieveAndGenerate, rerank, generateQuery, optimize-prompt) forwarded every
caller header to AWS next to the SigV4 signature, so a LiteLLM key presented in
x-api-key or x-litellm-api-key reached bedrock-agent-runtime verbatim. Build the
upstream header set explicitly: drop LiteLLM credential headers by name and any
authenticated secret by value, keep the remaining caller headers, and let the
signed headers win on collisions.
2026-09-16 16:01:52 -07:00
Yuneng Jiang
afb28540bb
fix(e2e): keep the CLI determinism test out of the in-cluster suite
It drives the real CLI for several seconds. The edge stamps every
upstream call with PYTEST_CURRENT_TEST, a process-global that names
whichever test the worker is in when the call arrives rather than the one
that made it, so a test that holds a worker that long collects other
tests' in-flight calls. Build 234's key report credits this test with 20
Bedrock and 7 Anthropic misses, and it makes no provider call at all.

Those misattributed calls take the wrong test id into the cache key and
write recordings under it, so the test was polluting the shared corpus it
exists to protect.

Deselected unless E2E_CLI_DETERMINISM is set, the same opt-in shape the
managed-files, prompt-caching and redis-chaos markers already use. The
attribution bug itself is older than this branch and is reported, not
fixed here.
2026-09-16 16:01:32 -07:00
mateo-berri
8045794c37 Merge remote-tracking branch 'origin/main' into litellm_dashscope_reasoning_effort 2026-09-16 15:59:08 -07:00
kerry
2466975d29 test(e2e): clean cost map decimals and simplify scripted wire helpers
Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>
2026-09-16 22:57:23 +00:00
yuneng-jiang
ceaa0adef8
Merge branch 'main' into litellm_unpin_together_ai_serverless_model 2026-09-16 15:57:17 -07:00
yuneng-jiang
7f83650a41
Merge branch 'main' into litellm_isolate_generic_api_ndjson_test 2026-09-16 15:56:53 -07:00
ryan-crabbe-berri
36eb9cdb35 test(proxy): drop the route-list membership test that the PATCH gate tests already cover 2026-09-16 15:56:27 -07:00
Yujong Lee
c635399f6e merge: origin/main into litellm_rust_bridge_declarative_route_catalog
Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>
2026-09-16 22:54:51 +00:00
kerry
9885dc8962 test(e2e): add azure, bedrock converse and vertex wires to the cost suite
Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>
2026-09-16 22:52:42 +00:00
ryan-crabbe-berri
c37d0a2e66 fix(proxy): read general_settings without a cast and test the /team/update gate by behavior only 2026-09-16 15:51:04 -07:00
kerry
4b70696afa fix(streaming): estimate interrupted Anthropic stream usage from reasoning_content
Interrupted Anthropic streams that die before message_delta were billed at
the message_start placeholder (any value above 1 was trusted) or at 0 when
the partial response was reasoning-only, because the token_counter fallback
only looked at visible text. Reset the placeholder whenever no finish_reason
or second usage event arrived, fold the already-counted reasoning tokens into
the fallback estimate, and drop the stale completion_tokens_details so cost
is computed from the recovered count

Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>
2026-09-16 22:49:25 +00:00
Yujong Lee
9617312ab2 add PublicDispatch 2026-09-16 15:44:46 -07:00
ryan
dae16264c1 test(auth): cover over-budget user on zero-cost vs paid model in common_checks
Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>
2026-09-16 15:42:15 -07:00
ryan
287bbaa6c1 fix(proxy): remove duplicate user budget hook that 429'd zero-cost models
_PROXY_MaxBudgetLimiter re-checked spend:user:{id} against user_max_budget in
async_pre_call_hook without the zero-cost model exemption that
_user_max_budget_check applies in auth, so free models were rejected with
"Max budget limit reached." once a user was over budget. Auth already owns
this check, so the hook is deleted rather than taught the exemption again

Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>
2026-09-16 15:42:15 -07:00
ryan-crabbe-berri
6e2ae19670 fix(proxy): enforce org budget ceilings on /team/update
update_team loaded the org without its budget row, so the org max_budget,
tpm_limit and rpm_limit checks silently passed. It now loads the budget the
same way /team/new does
2026-09-16 15:38:30 -07:00
ryan-crabbe-berri
a44a58e91a feat(proxy): let team admins edit tpm_limit when a proxy admin enables it
tpm_limit is the first field in the team-admin allow-list registry. The team
settings tab gives a team admin a form with only the enabled fields and sends
only those on save, and UI Settings labels the checkbox the same way
2026-09-16 15:38:30 -07:00
ryan-crabbe-berri
d233043b05 test(proxy): expect team-admin budget changes to stop at the allow-list
max_budget is not a team-admin editable field yet, so the behavior suite now
pins the 403 in both directions instead of the old lower-is-allowed rule
2026-09-16 15:38:30 -07:00
ryan-crabbe-berri
66519da9b6 fix(proxy): report the caller's team edit access on /team/info
The dashboard gated the team settings form on a role it guessed from the
is_* props, the members list and the org list. The org list is premium
gated and empty while loading, so a team admin who is also an org admin
was told team admins cannot edit, although /team/update accepts them as
an org admin

/team/info now returns caller_edit_access, resolved by the same helper
/team/update uses, and TeamInfo keys the form and the toast off that
field. The org list is only read for the organization dropdown now, and
general_settings is read through one validated accessor in both handlers
2026-09-16 15:38:30 -07:00
ryan-crabbe-berri
50f890d02d fix(proxy): keep the team-admin field allow-list importable on python 3.10
assert_never landed in typing in 3.11, so importing it from typing broke the
3.10 import smoke job and the py310 typing gate. Take it from typing_extensions
like the rest of the proxy does.

Also drop /team/update from test_neighbouring_team_routes_stay_closed. The route
is self-managed now, so the coarse gate admits the caller and update_team decides,
which the sibling test in the same file already asserts.

Claude-Session: https://claude.ai/code/session_018PUCupsaarVLJy4iDFx256
2026-09-16 15:38:06 -07:00