Commit graph

18680 commits

Author SHA1 Message Date
Joshua Valluru
92e182b898 fix(mcp): persist OAuth credentials for validated JWT users 2026-09-15 12:29:03 -07:00
Devin AI
2f33727cc9 fix(proxy): reset budgets by decrementing pre-reset spend instead of zeroing rows
The budget reset job read a row's spend, reset it in place, then wrote
spend: 0 (or decremented by max_budget under rollover) when committing.
Any spend the batch writer incremented into the row between the read and
the commit was erased while LiteLLM_DailyUserSpend kept it, so the daily
rollup permanently exceeded the counters.

Capture each row's spend before _reset_budget_common mutates it and write
a decrement of pre_spend - post_spend, which equals max_budget in the
rollover-over-cap case it replaces. Rows with no spend still get an
absolute spend: 0.

Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>
2026-09-15 19:24:33 +00:00
yucheng
d5e056491c fix(guardrails): keep the assistant turn when scoping empties the request history
A request whose turns all fall outside the guardrail's scope, such as a user-only
request under scan_only_tool_results, still supplied a conversation, so the response
scan now carries the reply as the sole assistant turn instead of dropping
structured_messages. Response-only behavior stays when no conversation was supplied

Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>
2026-09-15 19:12:15 +00:00
Devin AI
f219535717 test(azure): avoid rebinding messages in managed file id regression test
Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>
2026-09-15 19:06:31 +00:00
Tin Chi Lo
cadb7ee44d fix(router): preserve native encrypted capability tasks 2026-09-15 12:05:30 -07:00
Devin AI
7566164e46 Merge remote-tracking branch 'origin/main' into litellm_registry_audit_2026_09_14 2026-09-15 19:02:41 +00:00
Devin AI
81524d212f fix(azure): rebuild content dicts instead of mutating, drop test docstrings
Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>
2026-09-15 18:54:35 +00:00
Mateo Wang
d3929287fe
Merge pull request #41189 from BerriAI/litellm_per_turn_control_beta
fix(anthropic): add the per-turn-control beta when a message carries output_config
2026-09-15 11:53:01 -07:00
Tin Chi Lo
896f35c751 fix(router): extract capability tasks with request scoped markers 2026-09-15 11:52:20 -07:00
Yassin Kortam
501be3143d fix(proxy): enforce organization budgets when max_budget is 0
_organization_max_budget_check returned early whenever org_max_budget
was <= 0, so an organization with an explicit max_budget of 0 was
treated as unlimited instead of zero allowance. Key, team, and user
budget checks already skip only on None; align organization budgets
with that convention.

validate_team_org_change had the same defect in a different shape: it
used a truthy check on the org's max_budget when validating a team
move, so an explicit 0 there silently skipped the guard too.

Co-Authored-By: Claude Sonnet 5 <noreply@anthropic.com>
2026-09-15 11:49:05 -07:00
Tin Chi Lo
e62f0d0376 fix(router): reject unknown capability policy fields 2026-09-15 11:43:04 -07:00
Devin AI
f925c1d1e6 fix(azure): strip litellm format field from file and image content parts
Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>
2026-09-15 18:41:05 +00:00
mateo-berri
f490172338 test(anthropic): drop docstrings and wrap a long line in the per-turn-control tests 2026-09-15 11:39:01 -07:00
Mateo Wang
2e06d195b2
Merge pull request #39857 from BerriAI/litellm_e2e_reliability_module_cells
test(e2e): cover the reliability retry, cooldown, fallback, and routing-strategy cells
2026-09-15 11:38:51 -07:00
yassin
fb00567e4c fix(proxy): track per-member organization spend so the Organizations UI shows member spend
Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>
2026-09-15 18:38:07 +00:00
Yassin Kortam
41b5d47c71
Merge pull request #41144 from BerriAI/litellm_responses_bridge_filters_unknown_params
fix(responses): filter bridged kwargs like the native Responses path
2026-09-15 11:32:38 -07:00
Devin AI
8978b4562f test(main): drop unrelated reformatting
Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>
2026-09-15 18:26:54 +00:00
Yassin Kortam
837423237a
Merge pull request #36775 from MvdB/litellm_presidio_new_entities
feat(guardrails): add new upstream presidio pii entities including german set
2026-09-15 11:26:21 -07:00
yujonglee
33000d7e25
Merge pull request #41180 from BerriAI/litellm_rust_bridge_native_stub
build(rust-bridge): add typed _native stub and validate it with mypy.stubtest
2026-09-15 11:25:56 -07:00
Devin AI
fd2fb4c44e fix(http): address review on outbound HTTP/2
Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>
2026-09-15 18:25:26 +00:00
Devin AI
da7853c20a test: drop tests that pin vendor facts and add the CLAUDE.md rule
Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>
2026-09-15 18:24:30 +00:00
Tin Chi Lo
2da9bbfc0f chore: merge main into capability classifier copy 2026-09-15 11:24:01 -07:00
tin-berri
3ac79757f4
Merge pull request #41175 from BerriAI/litellm_team_member_auto_routers
feat(auto-router): allow opted-in team members to manage their routers
2026-09-15 11:19:58 -07:00
Yassin Kortam
dfcefd8298
Merge pull request #41256 from BerriAI/litellm_team_model_access_error_lists_all_models
fix(proxy): list directly assigned team models in model access errors
2026-09-15 11:16:25 -07:00
yassin
cbd72ff6d8 Merge remote-tracking branch 'origin/main' into litellm_responses_bridge_filters_unknown_params 2026-09-15 18:15:19 +00:00
ryan-crabbe-berri
8b6c398b92
Merge pull request #41039 from BerriAI/litellm_bulk_user_delete
feat(proxy): add POST /user/bulk_delete and POST /team/bulk_member_delete
2026-09-15 11:12:40 -07:00
mateo-berri
ce40b5773d Merge remote-tracking branch 'origin/main' into litellm_per_turn_control_beta 2026-09-15 11:11:52 -07:00
ryan
6ea1085bc3 Merge branch 'main' into litellm_bulk_user_delete
Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>
2026-09-15 18:03:15 +00:00
mateo-berri
c0fd8f6012 fix(anthropic): merge a case-variant Anthropic-Beta client header instead of clobbering it 2026-09-15 11:03:10 -07:00
Emerson Gomes
4595b4f62f
fix(azure-ai): coerce FLUX controls and preserve response dimensions 2026-09-15 13:01:12 -05:00
ryan-crabbe-berri
08b433267e
Merge pull request #40917 from BerriAI/litellm_credential_conflict_409
fix(credentials): answer 409 on a credential name collision, make Terraform adoption opt-in
2026-09-15 10:58:26 -07:00
ryan-crabbe-berri
6dcca8c4ae
Merge pull request #41028 from BerriAI/litellm_bulk_new_user
feat(proxy): add POST /management/v1/users/bulk for batched user and team membership creation
2026-09-15 10:57:46 -07:00
Devin AI
a426df108a feat(http): opt-in outbound HTTP/2 for httpx clients
Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>
2026-09-15 17:54:03 +00:00
ryan-crabbe-berri
81806f33cf fix(credentials): answer 409 on a name collision, let PATCH resolve values from model_id
POST /credentials let a duplicate name hit the unique index and handed back
Prisma's "Unique constraint failed" as a 500, so callers string-matched that
message to tell a caller mistake from a server fault. The unique violation now
maps to a 409 whose message names the PATCH route, two concurrent creates of
one name agree on it, and the detection lives in a repository helper the five
hand-rolled copies can move onto later

PATCH /credentials/{name} took a CredentialItem body, so the model_id the
Terraform adopt path sent was dropped. It now accepts UpdateCredentialItem and
shares the deployment lookup with create. Both handlers take the router as a
FastAPI dependency instead of reading the proxy global, which is what the
tests override
2026-09-15 10:41:43 -07:00
Emerson Gomes
911f66aff6
feat(azure_ai): support FLUX.2 flex images 2026-09-15 12:25:19 -05:00
Devin AI
9fce6e1989 test(pricing): let synced GovCloud Bedrock rows cite the AWS price list
Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>
2026-09-15 17:09:12 +00:00
Devin AI
79450121f8 test(proxy): document access group test seam
Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>
2026-09-15 17:00:11 +00:00
tin-berri
3ad9a7f336
Merge pull request #41174 from BerriAI/litellm_tier_model_affinity
fix(router): preserve session model choice within each complexity tier
2026-09-15 09:54:53 -07:00
Devin AI
56d0f953f5 fix(proxy): list directly assigned team models in model access errors
Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>
2026-09-15 16:33:41 +00:00
mateo-berri
f8fe704516 chore: merge main into fix/batch-retrieve-model-group 2026-09-15 07:02:59 -07:00
mateo-berri
2bbf34c652 fix(utils): keep the tool_choice validator's model argument optional 2026-09-15 07:00:33 -07:00
Devin AI
631ba7f9e3 fix(models): dedupe merged keys, price gemini *-latest aliases at live targets, add Nova cache read prices and Fireworks deprecation dates
Absorbs #41148 and #41152 into the rolling registry PR.

Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>
2026-09-15 13:36:11 +00:00
Mateo Wang
c8114ba41f
Merge pull request #40921 from BerriAI/litellm_unified_key_policy_hook
feat(proxy): unified custom_key_policy hook for key generate, update and regenerate
2026-09-15 06:34:28 -07:00
mateo-berri
76488beaf8 fix(utils): reject an untranslatable tool_choice with a 400 instead of a 500 2026-09-15 05:55:25 -07:00
Marty Sullivan
785c6cffc4 fix(cost): carry image and video input tokens through the Responses usage bridge
Realtime cost is computed from *_tokens_details after the usage round-trips
through the Responses shape, and the input half of that shape carried audio
only, so image and video prompt tokens stopped being billable as themselves.

Vertex splits prompt tokens by modality, so a session sending camera frames
arrives with image_tokens set. Those were folded into text_tokens and lost
their attribution. The amount happens not to move today, because the
calculator falls back to input_cost_per_token when no per-modality rate is
set, but the tokens have to survive before any such rate can ever apply.

InputTokensDetails now declares image_tokens and video_tokens instead of
leaning on pydantic extras, the repeated per-field copying is a loop over the
modality names so adding a modality no longer adds a branch, and the read-back
in ResponseAPILoggingUtils picks up video_tokens, which
PromptTokensDetailsWrapper already declared.

The output half of the original change is dropped: 449c091391 landed the same
OutputTokensDetails.audio_tokens fix upstream, with its own coverage in
test_gemini_realtime_transformation.py, and it always sets
output_tokens_details rather than only when non-empty. That structure is kept
as upstream wrote it.
2026-09-15 05:52:54 -07:00
Mateo Wang
9496f16f12
Merge pull request #41173 from BerriAI/litellm_realtime_health_check_credential_name
fix(health): resolve litellm_credential_name in realtime health checks
2026-09-15 05:24:29 -07:00
mateo-berri
1f2d050386 Merge remote-tracking branch 'origin/main' into litellm_unified_key_policy_hook 2026-09-15 05:11:29 -07:00
mateo-berri
7c0cca4d36 Merge remote-tracking branch 'origin/main' into litellm_fix_ocr_native_multipage_pdf
Some checks failed
ai-gateway image / ai-gateway release image (push) Waiting to run
LiteLLM Rust / rust-lint (push) Has been cancelled
LiteLLM Rust / rust-test (push) Has been cancelled
# Conflicts:
#	litellm/_logging.py
2026-09-15 04:49:14 -07:00
yucheng
43ae9aff3d fix(guardrails): tolerate a model-less request when translating Anthropic response context
The proxy-endpoints shard failed with KeyError: 'model' because the new Anthropic post-call context translation reached translate_anthropic_to_openai with request data that only carried messages and guardrail metadata.

Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>
2026-09-15 10:22:35 +00:00
yucheng
8e25720c08 fix(guardrails): give post-call scans the scoped request conversation and tools
Response-side guardrail scans on OpenAI Chat Completions, Anthropic Messages, and OpenAI Responses now carry structured_messages (the request turns scoped exactly like the pre-call scan, closed by the model's reply as an assistant turn) and tools (the request's function definitions), in addition to texts, images, and tool_calls.

Guardrails that used structured_messages or tools as a response-side signal (akto, crowdstrike_aidr, hiddenlayer, openai moderations, promptguard, qualifire, straiker) keep their previous response payloads. Logging-only scans whose output translation differs from the input translation get a chat-shaped request so the context survives.

Resolves LIT-6628

Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>
2026-09-15 09:58:53 +00:00