Commit graph

18968 commits

Author SHA1 Message Date
yucheng-berri
e99902c4dd
Merge pull request #41569 from BerriAI/litellm_azure_ptu_spillover_cost
fix(cost): price Azure PTU spillover requests at standard token rates
2026-09-17 17:09:40 -07:00
ryan
e79d03e604 fix(proxy): evict jwt key mapping cache on bulk user and team member deletion
Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>
2026-09-18 00:00:17 +00:00
yucheng-berri
4c70cb4815
Merge pull request #41578 from BerriAI/litellm_grafana_all_prometheus_metrics_dashboard
feat(grafana): add all-metrics dashboard and fix stale dashboard_v2 gauges
2026-09-17 16:57:32 -07:00
kerry-berri
9e651c8fe3
Merge pull request #41536 from BerriAI/litellm_aws_govcloud_partition_gate
test(e2e): cover bedrock batch file upload and create in the us-gov-west-1 partition
2026-09-17 16:55:53 -07:00
mateo-berri
064e49810a fix(bedrock): keep the tool search rule off azure_ai and pin dotted ids and Vertex fills 2026-09-17 16:48:15 -07:00
mateo-berri
c29d1c2813 chore: merge main into litellm_foundry_a2a_entra_agents 2026-09-17 16:45:07 -07:00
yassin
d0591665b5 Merge remote-tracking branch 'origin/main' into HEAD
Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>

# Conflicts:
#	litellm/proxy/pass_through_endpoints/success_handler.py
#	tests/test_litellm/proxy/pass_through_endpoints/test_llm_pass_through_endpoints.py
2026-09-17 23:44:13 +00:00
Devin AI
569ef3ea25 merge: resolve conflict with main in member budget seeding
Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>
2026-09-17 23:42:41 +00:00
ryan
1d50d1ad3b fix(proxy): rewrite every model allowlist in one statement on rename
Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>
2026-09-17 23:40:59 +00:00
mateo-berri
99b83a2d52 fix(fireworks_ai): restore supports_vision on minimax-m3 in the cost map 2026-09-17 16:39:31 -07:00
mateo-berri
8e3742a5f3 fix(proxy): answer 503 temporarily_unavailable when the token exchange cannot verify the subject token
Some checks are pending
LiteLLM Rust / rust-lint (push) Waiting to run
LiteLLM Rust / rust-test (push) Waiting to run
LiteLLM Rust / rust-wheel (push) Waiting to run
A subject token JWT auth could not check, because the IdP's JWKS was unreachable with no cached copy or the auth database was down, came back as 400 invalid_request with the same fixed message a bad token gets, so clients re-logged in instead of retrying the way they already do for a mint-time 503. Those checks now answer 503 temporarily_unavailable and log the reason, while real rejections stay 400 invalid_request.
2026-09-17 16:39:02 -07:00
mateo-berri
42f2978061 Merge remote-tracking branch 'origin/main' into HEAD 2026-09-17 16:38:54 -07:00
mateo-berri
1e7e5b695f fix(router): reject a blank Jev api_key so it cannot pair with a caller-chosen api_base 2026-09-17 16:38:54 -07:00
Mateo Wang
deb9d8aedd
Merge pull request #41607 from BerriAI/litellm_typesafe_passthrough
Some checks are pending
LiteLLM Rust / rust-test (push) Waiting to run
Unit Tests: Documentation Validation / documentation (push) Waiting to run
Unit Tests: Proxy DB Operations / assert-shard-coverage (push) Waiting to run
Unit Tests: Proxy DB Operations / auth-checks (push) Blocked by required conditions
Unit Tests: Proxy DB Operations / budgets (push) Blocked by required conditions
Unit Tests: Proxy DB Operations / custom-logging (push) Blocked by required conditions
Unit Tests: Proxy DB Operations / db-and-spend (push) Blocked by required conditions
Unit Tests: Proxy DB Operations / endpoints-and-responses (push) Blocked by required conditions
Unit Tests: Proxy DB Operations / guardrails-hooks (push) Blocked by required conditions
Unit Tests: Proxy DB Operations / jwt-and-keys (push) Blocked by required conditions
Unit Tests: Proxy DB Operations / key-generation (push) Blocked by required conditions
Unit Tests: Proxy DB Operations / logging-misc (push) Blocked by required conditions
Unit Tests: Proxy DB Operations / proxy-runtime (push) Blocked by required conditions
Unit Tests: Proxy DB Operations / proxy-server-core (push) Blocked by required conditions
Unit Tests: Proxy DB Operations / proxy-utils (push) Blocked by required conditions
Unit Tests / caching-local (push) Waiting to run
Unit Tests / core-utils (push) Waiting to run
Unit Tests / enterprise-package (push) Waiting to run
Unit Tests / enterprise-routing (push) Waiting to run
Unit Tests / integrations (push) Waiting to run
Unit Tests / All Other Providers (push) Waiting to run
Unit Tests / Vertex AI (push) Waiting to run
Unit Tests / misc (push) Waiting to run
Unit Tests / proxy-auth (push) Waiting to run
Unit Tests / proxy-endpoints (push) Waiting to run
Unit Tests / proxy-extras (push) Waiting to run
Unit Tests / proxy-infra (push) Waiting to run
Unit Tests / proxy-server (push) Waiting to run
Unit Tests / responses-caching-types (push) Waiting to run
GitHub Actions Security Analysis / zizmor (push) Waiting to run
feat(proxy): add TypeSafe AI Jev evaluate passthrough with registry-priced spend tracking
2026-09-17 16:37:33 -07:00
ryan
884a467a7f fix(proxy): evict jwt_key_mapping cache when user, team, or org deletion removes mapped keys
Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>
2026-09-17 23:33:59 +00:00
ryan-crabbe-berri
e5aa10a1ea
Merge pull request #41686 from BerriAI/litellm_member_budget_clone_reset_and_audit
fix(team): keep a forked member budget's reset window and audit bulk member budget writes
2026-09-17 16:32:31 -07:00
ryan-crabbe-berri
89289d4f77 fix(proxy): surface a failed project spend enqueue at error level
Some checks are pending
LiteLLM Rust / rust-lint (push) Waiting to run
LiteLLM Rust / rust-test (push) Waiting to run
LiteLLM Rust / rust-wheel (push) Waiting to run
The call-site guard kept a project enqueue failure from skipping the sibling
spend writes, but logged it at debug only, so a dropped project charge was
invisible on a default log level. Log it through spend_log_error inside the
helper and re-raise, the way the org and agent helpers already do.
2026-09-17 16:26:26 -07:00
ryan-crabbe-berri
e4778cd2d1 test(proxy): drive the real spend queue in the project enqueue isolation test
Substitutes a queue that rejects project items instead of replacing writer
methods with mocks, so the test asserts the sibling key, team and tag spend
actually landed in the queue rather than that a mock was awaited.
2026-09-17 16:24:51 -07:00
ryan
de0047c802 fix(proxy): skip allowlist rewrite when the model name is unchanged
Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>
2026-09-17 23:24:26 +00:00
yassin
33531649c3 perf(mcp): count gateway session groups with Counter and pin the oversized initialize peek invariant
Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>
2026-09-17 23:20:47 +00:00
ryan
70164dabd2 fix(proxy): isolate project spend enqueue failures from sibling spend writes
Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>
2026-09-17 23:20:10 +00:00
yassin
c230393731 fix(mcp): refuse sessionless and stale-session POSTs that skip initialize while mcp_allowed_clients is set
Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>
2026-09-17 23:18:54 +00:00
ryan
d3f5cde530 fix(proxy): propagate db model renames to key, team, org, project and user model allowlists
Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>
2026-09-17 23:11:52 +00:00
Mateo Wang
57d41bda29
Merge pull request #41684 from BerriAI/litellm_wildcard_license_auto_router
fix(license): let a wildcard allowed_features license grant the auto_router feature
2026-09-17 16:11:40 -07:00
Mateo Wang
07b5051c0d
Merge pull request #41689 from BerriAI/litellm_responses_bridge_strip_internal_kwargs
fix(responses): keep the addressed response id off bridged provider requests
2026-09-17 16:11:04 -07:00
mateo-berri
1c40e6034d fix(a2a): Entra credentials own the chat route bearer over a stored api_key or authorization header 2026-09-17 16:09:36 -07:00
mateo-berri
7f581f6bc7 fix(router): keep the TypeSafe key off caller-chosen Jev endpoints 2026-09-17 16:09:31 -07:00
yassin
912edaa8cc Merge remote-tracking branch 'origin/main' into litellm_transcribe_passthrough 2026-09-17 23:07:57 +00:00
Mateo Wang
0da001901b
Merge pull request #41672 from BerriAI/litellm_autoroute_start_stop
feat(cli): rename lite autoroute up/down to start/stop, keeping the old names as deprecated aliases
2026-09-17 16:07:00 -07:00
Devin AI
9d5f65c26c merge: resolve conflict with main, move temp budget patch fields to shared member_budget_patch
Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>
2026-09-17 23:04:25 +00:00
ryan-crabbe-berri
29a959b3e8
Merge pull request #41632 from BerriAI/litellm_bulk_team_member_budget_update
feat(management_v1): bulk update team member budgets
2026-09-17 15:58:32 -07:00
yassin
ce735f586c fix(proxy): scope Transcribe jobs to the key that started them and charge rewritten media the maximum
Standard jobs are tagged litellm-owner on StartTranscriptionJob so GetTranscriptionJob
and DeleteTranscriptionJob only work for the owner or a proxy admin, and account-wide
operations need a proxy admin. Media rewritten after job creation is charged the eight
hour maximum, and the success handler takes an injected log dispatch instead of tests
patching its private method

Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>
2026-09-17 22:58:15 +00:00
mateo-berri
d3963e5d63 test(cli): pin the up alias port forwarding and the removed-settings stop message 2026-09-17 15:55:50 -07:00
mateo-berri
79029d89f9 test(responses): drop the history docstrings from the bridge regression tests 2026-09-17 15:51:35 -07:00
kerry
91619376d2 test: restore azure ai cached-token billing coverage with derived rates
Some checks failed
LiteLLM Rust / rust-lint (push) Has been cancelled
LiteLLM Rust / rust-test (push) Has been cancelled
LiteLLM Rust / rust-wheel (push) Has been cancelled
Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>
2026-09-17 22:49:02 +00:00
Mateo Wang
289c52bcd6
Merge pull request #39512 from BerriAI/litellm_fix_image_edits_bracketed_alias
fix(images): stop forwarding the raw image[] and mask[] form keys
2026-09-17 15:48:32 -07:00
Mateo Wang
fec8231b83
Merge pull request #41419 from BerriAI/litellm_bedrock_openai_no_cachepoint
fix(bedrock): never emit Converse cachePoint for OpenAI-family models
2026-09-17 15:43:04 -07:00
ryan-crabbe-berri
f04f0258f7 refactor(team): model the bulk budget audit payload as frozen types 2026-09-17 15:41:06 -07:00
yucheng
e9825f1d26 test(proxy): drive the heuristics responsiveness check without mutable state
Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>
2026-09-17 22:39:15 +00:00
mateo-berri
1feaa48705 fix(proxy): log TypeSafe calls that name no model as unknown 2026-09-17 15:37:00 -07:00
mateo-berri
ca91751d5b fix(responses): keep the addressed response id off bridged provider requests
The Responses id security hook keeps the id a client addressed under
`_litellm_addressed_response_id` in the request body so internal retries can
re-authorize it. On a model without a native Responses config that body is
bridged into `completion()` kwargs, the key was treated as a provider param,
and providers rejected it, so every follow-up turn carrying
`previous_response_id` returned 400.

Register the key in `all_litellm_params` so it is dropped before any provider
request, and share one constant between the hook and the param list.
2026-09-17 15:35:11 -07:00
jesus
1b69a5b0a4 test(proxy): model missing organizations in MCP auth fixtures
Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>
2026-09-17 22:34:02 +00:00
ryan
7e5b3b49d4 fix(proxy): treat non-positive project max_budget as unbudgeted
Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>
2026-09-17 22:33:54 +00:00
kerry
b054d54bee Merge remote-tracking branch 'origin/litellm_remove_brittle_price_pinning_tests' into litellm_remove_brittle_price_pinning_tests
Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>

# Conflicts:
#	tests/test_litellm/llms/parallel_ai/test_parallel_ai_search.py
#	tests/test_litellm/proxy/common_utils/test_prompt_cache_pricing.py
#	tests/test_litellm/proxy/test_proxy_utils.py
2026-09-17 22:33:26 +00:00
kerry
9eb6fbc572 test: read cost-map keys the implementation resolves and isolate the tariff test's model_cost copy
Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>
2026-09-17 22:32:39 +00:00
ryan
bde5593523 Merge remote-tracking branch 'origin/main' into litellm_lit_3269_project_spend_tracking
Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>

# Conflicts:
#	litellm/proxy/common_utils/reset_budget_job.py
2026-09-17 22:29:16 +00:00
jesus
c4ad6194a0 fix(auth): only fail closed on DB outages during org lookup
Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>
2026-09-17 22:28:11 +00:00
yucheng
36844ef301 fix(proxy): clamp prompt injection heuristics worker count to at least one
Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>
2026-09-17 22:26:00 +00:00
mateo-berri
4371ddb620 feat(cli): keep lite autoroute up and down as hidden deprecated aliases 2026-09-17 15:25:12 -07:00
yassin
52914a06d2 Merge remote-tracking branch 'origin/main' into litellm_mcp_client_allowlist 2026-09-17 22:22:28 +00:00