Commit graph

53363 commits

Author SHA1 Message Date
Moe Khalil
d0ed8145b0 chore: merge latest main into model info discovery branch 2026-09-17 00:19:16 +00:00
ryan-crabbe-berri
3dfd24a8da
Merge pull request #41445 from BerriAI/litellm_ui_url_state_orgs-projects
feat(ui): persist organizations and projects list, detail tab and key table state in the URL
2026-09-16 17:17:01 -07:00
Yuneng Jiang
185a712d24
test(e2e): validate the captured /model/new body with its pydantic model 2026-09-16 17:16:22 -07:00
ryan-crabbe-berri
a36d2de5c9 fix(proxy): read the allowed request off the verdict and type the empty metadata set
CodeQL flagged the match capture as possibly uninitialized
2026-09-16 17:13:18 -07:00
Mateo Wang
d4a22acb66
Merge pull request #41513 from BerriAI/litellm_internal_copy_31400
fix(bedrock): neutralize orphaned tool blocks instead of raising or injecting a dummy tool (internal copy of #31400)
2026-09-16 17:12:35 -07:00
mateo-berri
dee5724c21 fix(a2a): read a stored card's capabilities the way the spec does and keep one Authorization line 2026-09-16 17:12:08 -07:00
kerry
9e84d8a9c0 fix(ci): drop the pull_request_review trigger so the auto-merge workflow only runs from main
Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>
2026-09-17 00:11:31 +00:00
Moe Khalil
39286245b5 chore: merge main into model info discovery branch
Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>
2026-09-17 00:10:00 +00:00
Yuneng Jiang
dcde8395ec
Merge remote-tracking branch 'origin/main' into litellm_/attribution-investigation-a81211 2026-09-16 17:08:17 -07:00
Yuneng Jiang
c63d0e6922
fix(e2e): bind provider-cache recordings to the deployment's test, not the serving process
The cache edge keyed every recording on its own process's PYTEST_CURRENT_TEST.
Under xdist that names whatever test the serving worker is in, which is
unrelated to the caller: the proxy is a separate pod, and the Claude Code compat
matrix registered its shared aliases from every worker, each pointing at that
worker's edge, so the router spread one worker's calls across all eight edges.
Builds 234 and 235 of litellm-e2e, same commit, credited the same Bedrock
request to unrelated tests 92% of the time, and Bedrock never converged past a
~20% hit rate while OpenAI, whose deployments are per test, sat at 90%.

A deployment registered from inside a test now carries its test's slug in the
edge URL it is pointed at, `{edge}/{mount}/t/{slug}`, and the edge reads that
segment off every request before forwarding. A request without one is forwarded
live and never cached, and the edge no longer falls back to process state. The
compat aliases are registered with provider_live=True and stay on their real
provider path: no single test owns them, and the matrix exists to prove the real
CLI against real providers.
2026-09-16 17:08:17 -07:00
Devin AI
2b8e17d581 refactor(router): move model info discovery provider set into openai_like module
Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>
2026-09-17 00:07:28 +00:00
ryan-crabbe-berri
43713f7508
Merge pull request #39996 from BerriAI/litellm_team_admin_editable_fields
feat(proxy): let proxy admins choose which team fields team admins may edit
2026-09-16 17:07:13 -07:00
kerry
0a47fe160d Merge remote-tracking branch 'origin/main' into litellm_mantle_gpt5_verbosity 2026-09-17 00:06:50 +00:00
Mateo Wang
09a188b583
Merge pull request #41094 from BerriAI/litellm_model_group_info_proxy_admin_all_models
fix(proxy): show all model groups to proxy admins in /model_group/info
2026-09-16 17:06:44 -07:00
yassin
5ddc96e560 fix(vertex_ai): drop stale transfer headers when GCS serves an encoded file body
Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>
2026-09-17 00:06:43 +00:00
Mateo Wang
913ef6ed49
Merge pull request #33856 from BerriAI/litellm_azure_ai_responses_native
fix(azure_ai): route Responses API to native /openai/v1/responses for Foundry Models
2026-09-16 17:06:29 -07:00
kerry-berri
f93b31679b
Merge pull request #41503 from BerriAI/litellm_fix_interrupted_anthropic_reasoning_usage
fix(streaming): estimate interrupted Anthropic stream usage from reasoning_content
2026-09-16 17:04:59 -07:00
yassin
b35ca7d2c3 fix(spend_tracking): honour the global litellm_proxy override when inferring a model group provider
Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>
2026-09-17 00:03:01 +00:00
Mateo Wang
6cdf398bea
Merge pull request #41504 from BerriAI/litellm_bedrock_agent_runtime_strip_virtual_key
fix(proxy): stop forwarding LiteLLM credential headers on Bedrock agent-runtime passthrough
2026-09-16 17:02:17 -07:00
Yujong Lee
6d30006592 refactor(rust_bridge): share call_hook instead of per-route native lambdas
Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>
2026-09-17 00:01:52 +00:00
Yujong Lee
43c50325a6 refactor(rust_bridge): pass dispatch context functions directly 2026-09-16 16:59:57 -07:00
kerry
a3d3fe4ada fix(ci): pin the auto-merge request to the evaluated head sha
Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>
2026-09-16 23:57:44 +00:00
yucheng
8d054ba230 test(otel): cover the routed tracer budget for spans opened at the pre_call boundary
Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>
2026-09-16 23:56:57 +00:00
kerry
42bf9f0ece Merge remote-tracking branch 'origin/main' into litellm_auto_merge_price_sync 2026-09-16 23:56:48 +00:00
Devin AI
75eec8712c fix(a2a): keep the upstream status on card discovery failures and inject the card client in tests
Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>
2026-09-16 23:55:57 +00:00
kerry
bba15b1382 style(responses_bridge): apply ruff format
Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>
2026-09-16 23:52:53 +00:00
yassin
d6b13f938d test(proxy): cover guardrail tag budget edge cases
Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>
2026-09-16 23:52:39 +00:00
yassin
488666ccae fix(proxy): enforce tag budgets for tags added by guardrails
Auth runs the tag budget check before pre_call_hook, so a tag that a custom guardrail adds is attributed spend but never budget checked. After the pre-call hook, budget check only the newly added tags with the same exemptions auth applied (budget-free routes, zero-cost models), keep the pre-guardrail tag baseline across fallback retries, and surface an over-budget tag as the same budget_exceeded 429 auth returns

Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>
2026-09-16 23:52:39 +00:00
Devin AI
bc8e28cfcf test(a2a): inject a fake httpx client into card resolver tests
Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>
2026-09-16 23:52:16 +00:00
Mateo Wang
319f427c40
Merge pull request #41201 from BerriAI/litellm_gemini_37_38_flash_no_minimal_thinking
fix(gemini): map minimal thinking to low for Gemini 3.7 and 3.8 Flash
2026-09-16 16:52:08 -07:00
yassin
5f1d87911a test(spend_tracking): cover unresolvable deployment leaving provider empty
Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>
2026-09-16 23:50:22 +00:00
mateo-berri
ada0a1ad3a fix(azure_ai): strip the azure_ai/ prefix when a Responses call is remapped to azure
A catalog OpenAI name on an .openai.azure.com host (or with AZURE_AI_API_BASE set to one) is remapped from azure_ai to azure before the Responses request is built, and the azure_ai/ prefix stayed in the wire model, so Azure answered DeploymentNotFound. The Azure Responses config now strips azure_ai/ next to responses/ and o_series/.
2026-09-16 16:49:29 -07:00
yassin
8bd598f13c feat(proxy): add Amazon Transcribe SigV4 pass-through routes
Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>
2026-09-16 23:48:15 +00:00
yucheng-berri
c5325b1492
Merge pull request #40596 from BerriAI/litellm_lit_7470_rate_limit_fallback_pristine_data
fix(proxy): retry rate-limit fallbacks from a pristine request snapshot
2026-09-16 16:47:35 -07:00
ryan-crabbe-berri
fd411373fc style(proxy): collapse the enduser page signature onto one line 2026-09-16 16:44:39 -07:00
yucheng-berri
672f43fd54
Merge pull request #41356 from BerriAI/litellm_lit7836_call_id_endpoint_logs
fix(proxy): carry litellm_call_id through endpoint specific error logs and failure responses
2026-09-16 16:43:12 -07:00
kerry
b7a9042f1b style(responses_bridge): suppress type-discipline flags with reasons
Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>
2026-09-16 23:43:04 +00:00
mateo-berri
574ea15b8f fix(mcp): count admin static headers as api_key credential slots
An api_key server whose key lives in static_headers, the documented
shape for upstreams that expect a custom header name, dispatched fine
before the fail-closed check and was rejected as misconfigured after
it. The check now treats every static header the admin configured as a
credential slot for api_key mode, on both the MCP client path and the
OpenAPI tool path, with regression tests at all three layers.
2026-09-16 16:42:20 -07:00
ryan-crabbe-berri
0d8b46b88c refactor(proxy): trim the invalidation docstrings and inject the page read failure
Cuts the new docstrings back to the parts a reader cannot get from the code,
and fixes a stale reference: the walk this one is modelled on is
_reset_windows_for, not _reset_windows_for_source.

The truncation test reached in and replaced MockTable.find_many. The mock takes
a scheduled read failure instead, the way it already takes canned rows.
2026-09-16 16:41:13 -07:00
kerry
9212124266 style(responses_bridge): use dict.get in text merge helper
Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>
2026-09-16 23:37:45 +00:00
kerry
255ef3bc26 refactor(bedrock_mantle): build supported params without mutation
Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>
2026-09-16 23:37:45 +00:00
ryan-crabbe-berri
3ee8d43fdd fix(ui): send only a changed TPM limit from the team admin settings form
Some checks failed
LiteLLM Rust / rust-lint (push) Has been cancelled
LiteLLM Rust / rust-test (push) Has been cancelled
LiteLLM Rust / rust-wheel (push) Has been cancelled
Save stays disabled until the value differs from the team's, so an unchanged form never reaches /team/update
2026-09-16 16:37:35 -07:00
mateo-berri
41737aeda8 test(proxy): drop the docstring from the agent-runtime passthrough regression class 2026-09-16 16:37:32 -07:00
Moe Khalil
b95dbb41ae fix(router): scope discovered limits to active deployments
Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>
2026-09-16 23:37:28 +00:00
yassin
9af014d75a fix(spend_tracking): keep inferred provider out of model reconstruction and OAuth provider lookup
Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>
2026-09-16 23:36:54 +00:00
mateo-berri
e243237a7c feat(a2a): reach Microsoft Foundry agents with Entra auth and versioned card discovery
Foundry serves its agent card only at agentCard/v1.0, accepts only an Entra ID
bearer, and defaults to a non-blocking send, so the A2A relay and the chat
completions route could not use it.

The relay gains an agent_card_path litellm_param plus agentCard/v1.0 as a third
discovery probe, mints a bearer from flat Entra fields on the agent
(tenant_id, client_id, client_secret, azure_ad_token, azure_username,
azure_password, azure_scope) for https://ai.azure.com/.default, and sends it on
the card fetch, message/send, message/stream, tasks/* and the chat bridge.

Chat completions look the registered agent up by its provider-stripped name so
its api_key and headers reach the request, tag every message with its kind, ask
for a blocking send, fall back to a blocking send when the registered card says
streaming: false, and fail the call on a JSON-RPC error inside a stream instead
of yielding an empty one. Entra fields stay out of the chat bridge's logged
parameters.

Resolves LIT-5122
2026-09-16 16:36:47 -07:00
yucheng
d3f4a8b983 fix(otel): read the attribute budget from the span's own provider limits so routed tracers fit correctly
Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>
2026-09-16 23:34:43 +00:00
kerry
438681be1a refactor(responses_bridge): share text merge helper
Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>
2026-09-16 23:33:28 +00:00
kerry
e8f246bb6b fix(responses_bridge): forward verbosity as text.verbosity
Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>
2026-09-16 23:32:32 +00:00
mateo-berri
ce722ab1b3 fix(proxy): evict the member's cached user row on team member add
JWT auth caches the user row before it adds the user to the JWT's team, and admission checks the credential's team against the cached row on whichever worker takes the next request. On a two-worker gateway the credential minted for a newly joined team answered 403 "not in your team memberships" until the management-object TTL ran out, because /team/member_add only evicted the membership spend sentinel. The add now evicts the added members' cached user rows and broadcasts the eviction to the other workers, the way /team/member_delete already did

The mint test now also covers a user SCIM deactivated after the cache last saw them active: the database read refuses the mint while the cached row still says active
2026-09-16 16:31:28 -07:00