Commit graph

18968 commits

Author SHA1 Message Date
kerry
a1560936f7 fix(timing): use epoch math for detailed pre-processing and drop client-supplied timing windows
Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>
2026-09-19 01:06:53 +00:00
yassin
453eccb2fa test(router): drop docstrings from the shared tpm regression tests
Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>
2026-09-19 01:06:50 +00:00
Yassin Kortam
15f63c33bf
Merge pull request #41911 from BerriAI/litellm_rate_limit_reset_time_utc
fix(rate_limiter): render the 429 reset time in UTC as labelled
2026-09-18 18:05:45 -07:00
ryan
f3bbeed82f feat(proxy): let team admins manage projects via team_admin_editable_team_fields
Adds a projects entry to the team_admin_editable_team_fields setting. When set, team admins (legacy admins list or members_with_roles role admin) can call /project/new and /project/update for the teams they administer. The two routes join self_managed_routes so the endpoint check runs instead of the route gate's blanket 401. /project/delete stays proxy admin only

Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>
2026-09-19 01:02:50 +00:00
kerry
80b0ea6a2f test(xai): narrow raises match for missing api key
Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>
2026-09-19 01:00:53 +00:00
ryan
8bd9d356dc test(team): drop docstrings that restate the budget source and reset assertions
Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>
2026-09-19 01:00:31 +00:00
ryan
52a71ff681 fix(team): let team admins reach the member reset_budget route and cover it in the behavior suite
Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>
2026-09-19 00:58:39 +00:00
Yujong Lee
8d2476465f fix(rust): honor environment proxies by default and name the cause in transport errors
Python's aiohttp transport reads HTTP(S)_PROXY on every request unless disable_aiohttp_trust_env is set, so the Rust clients now do the same instead of requiring aiohttp_trust_env. Transport error messages include reqwest's source chain, so a rejected certificate or refused connection is no longer reported as just 'error sending request'
2026-09-18 17:58:07 -07:00
kerry
a15b0fa6d2 test(integration): tidy xdist collection bookkeeping
Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>
2026-09-19 00:54:45 +00:00
kerry-berri
f71129f65b
Merge pull request #41847 from BerriAI/litellm_lit_8128_off_peak_pricing_schema
fix(schema): classify off_peak_pricing as a structured object in the model prices schema generator
2026-09-18 17:53:13 -07:00
kerry
6f54ad5166 fix(xai): parse integer speaker ids and simplify stt form build
Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>
2026-09-19 00:53:02 +00:00
Devin AI
6247b75543 test(mcp): keep the import block as merged on main
Co-Authored-By: bot_apk <apk@cognition.ai>
2026-09-19 00:51:51 +00:00
kerry
6eb67a8423 test(integration): run the cost shard with xdist workers
Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>
2026-09-19 00:50:29 +00:00
kerry
e52eea84e6 test(integration): serve scripted wires from the shared upstream
Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>
2026-09-19 00:50:26 +00:00
Devin AI
e3755a88e7 test(mcp): assert the misconfigured credential message on fail-closed rejections
Co-Authored-By: bot_apk <apk@cognition.ai>
2026-09-19 00:50:00 +00:00
ryan
d3a364d74f fix(team): report no budget source when the team default row was deleted
Derive budget_source from the budget row /team/info actually loaded, so a
metadata id whose row was removed via /budget/delete reads as none instead
of team_default. Share the /team/info test scaffolding so the added patch
calls stay within the TQ008 budget, and allowlist the imperative
reset_budget route in the provider endpoint audit

Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>
2026-09-19 00:49:58 +00:00
kerry
6b0ad3bed3 feat(xai): add speech-to-text via /v1/audio/transcriptions
Route xai audio transcription through a provider config hitting POST
https://api.x.ai/v1/stt instead of the openai-compatible chat handler
which targets /audio/transcriptions. Supports language, diarize,
keyterm, filler_words and other provider fields as passthrough kwargs

Resolves LIT-8153

Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>
2026-09-19 00:49:34 +00:00
jesus
3301fdafaf Merge remote-tracking branch 'origin/main' into litellm_org_alias_from_team 2026-09-19 00:48:54 +00:00
Yassin Kortam
6d8a960e1d
Merge pull request #41667 from BerriAI/litellm_mcp_client_allowlist
feat(mcp): allowlist MCP client applications at the gateway
2026-09-18 17:48:36 -07:00
kerry
d74e1bb445 fix(timing): union provider timing windows and anchor detailed pre-processing at receive time
Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>
2026-09-19 00:48:17 +00:00
mateo-berri
f213855558 refactor: drop the docstrings from the websocket relay and its tests 2026-09-18 17:44:32 -07:00
mateo-berri
3911d62bbe fix(vertex_ai): prune a discarded turn's id once its marker is delivered 2026-09-18 17:44:28 -07:00
Mateo Wang
a6e3a72ed8
Merge pull request #41870 from BerriAI/litellm_bedrock_openai_gpt_min_max_tokens
fix(bedrock): clamp maxTokens to the 16-token minimum for OpenAI GPT and xAI Grok models on Converse
2026-09-18 17:44:18 -07:00
kerry
26bcc537a5 fix(proxy): allow latest release info and reset banner dismissal
Some checks failed
LiteLLM Rust / rust-lint (push) Waiting to run
LiteLLM Rust / rust-test (push) Waiting to run
LiteLLM Rust / rust-wheel (push) Waiting to run
Terraform Provider / gofmt, vet, build, test (push) Waiting to run
Terraform Provider / Provider endpoints vs proxy OpenAPI schema (push) Waiting to run
Terraform Modules / fmt, validate, test (aws) (push) Has been cancelled
Terraform Modules / fmt, validate, test (gcp) (push) Has been cancelled
Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>
2026-09-19 00:44:10 +00:00
kerry
99659e9e7e test(bedrock): drop redundant recorder docstring
Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>
2026-09-19 00:43:51 +00:00
yassin
75b290969b fix(router): enforce model tpm limits against shared redis usage across replicas
The model tpm pre-call check read only the in-memory counter, so each proxy replica enforced the limit against its own traffic and the deployment admitted up to N times the configured tpm across N replicas. Read the shared Redis counter when the local counter is under the limit, keep the local counter authoritative when it is already at the limit, and fall back to local usage when Redis is unavailable

Supersedes #40854, Fixes #40291

Co-authored-by: Jahanzeb-git <jahanzebahmed2002@gmail.com>
Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>
2026-09-19 00:41:55 +00:00
yassin
aa0fb915d0 fix(rate_limiter): render the 429 reset time in UTC as labelled
The proxy rate limiters formatted the reset epoch with a naive datetime.fromtimestamp, which reads the process timezone, and then appended a literal UTC suffix. A proxy running outside UTC returned a local wall-clock time labelled as UTC in the 429 body and reset_at header. Convert with tz=timezone.utc in both the request limiter and the batch limiter so the label is true

Co-authored-by: Priyansh Nandwana <nandwana.priyansh103@gmail.com>
Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>
2026-09-19 00:40:51 +00:00
Devin AI
890e5feabe test(e2e): keep e2e_config formatting untouched
Co-Authored-By: bot_apk <apk@cognition.ai>
2026-09-19 00:38:57 +00:00
Devin AI
c1bd5ba91d test(e2e): share the Linear readonly tool constant and fail fast on unexpected consent
Co-Authored-By: bot_apk <apk@cognition.ai>
2026-09-19 00:38:28 +00:00
Yujong Lee
a3aceec2f8 fix(rust): match Python proxy, ssl_verify and client expiry behavior in the http pool
Honor environment proxies whenever Python would use httpx (sync calls, HTTP/2, aiohttp disabled), apply the per-call ssl_verify argument, ignore empty or missing SSL env values the way http_handler.py does, expire pooled clients after an hour so rotated certificates reload, keep the client certificate off media downloads, and decline instead of raising when a litellm global has an unexpected type
2026-09-18 17:37:08 -07:00
joshua
a873ead5d3 test(mcp): read SDK2 snake_case fields on CallToolResult
Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>
2026-09-19 00:36:03 +00:00
Devin AI
6c8f1c22e0 test(e2e): cover MCP OAuth happy path through gateway
Co-Authored-By: bot_apk <apk@cognition.ai>
2026-09-19 00:35:56 +00:00
yucheng
95c1d5b0a6 feat(otel v2): name and re-root the kept generation under llm_only, widening to the account's widest scope
Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>
2026-09-19 00:33:46 +00:00
ryan
d611fd33d9 Merge remote-tracking branch 'origin/main' into litellm_team_member_budget_source_reset 2026-09-19 00:31:55 +00:00
ryan
b32d1112a6 feat(team): show whether a member follows the team default budget and allow resetting to it
Adds budget_source (team_default, custom, none) to each membership in /team/info and a
POST /team/{team_id}/member/{user_id}/reset_budget route that relinks a member to the team's
shared team_member_budget row without touching their spend. The Admin UI team members table
shows a Team default or Custom badge next to each member's budget and offers a
Use team default action on customized members

Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>
2026-09-19 00:31:55 +00:00
mateo-berri
8b1f78fa08 fix(vertex_ai): drop a cleared turn's queued transcripts and carry its billed seconds 2026-09-18 17:30:57 -07:00
mateo-berri
50629ff5ca Merge branch 'main' of https://github.com/BerriAI/litellm into litellm_jwt_token_exchange_grant
# Conflicts:
#	tests/test_litellm/proxy/auth/test_auth_checks.py
2026-09-18 17:30:38 -07:00
Mateo Wang
cda022ca68
Merge pull request #40243 from zoroyihan7/fix-responses-stream-error-events
fix(responses): emit typed streaming failure events
2026-09-18 17:29:57 -07:00
Yujong Lee
542ad7dbac fix(ocr): forward the supplied client on the Python path and build pooled clients outside the lock
Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>
2026-09-19 00:27:54 +00:00
kerry
e91f17ac3a Merge remote-tracking branch 'origin/main' into litellm_lit_8128_off_peak_pricing_schema 2026-09-19 00:27:52 +00:00
Mateo Wang
f6d4766ebe
Merge pull request #41878 from BerriAI/litellm_requeue_daily_spend_without_redis_buffer
fix(proxy): requeue daily spend rows when the commit fails without the Redis buffer
2026-09-18 17:27:36 -07:00
kerry-berri
92d01fa568
Merge pull request #41901 from BerriAI/litellm_off_peak_pricing_integration_tests
test(integration): cover off-peak pricing on a live proxy
2026-09-18 17:26:51 -07:00
kerry
6b082d3a01 test(bedrock): type the SigV4 request recorder and drop caller-owned mutation
Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>
2026-09-19 00:25:01 +00:00
mateo-berri
cdc0e57e93 fix(websearch_interception): surface a failed search as a web_search_tool_result_error block and end the turn 2026-09-18 17:24:52 -07:00
mateo-berri
59f7a00cf6 fix(claude_code_gateway): scope the protobuf body skip to the OTLP routes and match the metrics middleware on the route path 2026-09-18 17:23:59 -07:00
mateo-berri
5030214671 Merge origin/main into litellm_fix_responses_ws_encrypted_content_affinity
Resolves the conflicts with the WebSocket request defaults from main (PR #41881):
the relay keeps both custom_llm_provider and request_defaults, and a masked
response.create frame is re-serialized when the defaults changed it.

Keeps the first-frame routing hints (input, previous_response_id) out of the
deployment request defaults so they never get injected into later frames on the
same connection, with a regression test.
2026-09-18 17:23:59 -07:00
Yujong Lee
b2d6cd1fcf refactor(rust): read litellm HTTP globals through one Python shim and tighten the http pool
Drop the core ocr() facade so VertexAuth and the http pool stay out of litellm-core's
public API, move the http Error enum to error.rs, and inject the media DNS resolver into
HttpClientPool instead of a per-call builder hook the cache key ignored.

The bridge now reads litellm.* HTTP settings only through litellm/rust_bridge/settings.py,
pinned by python_settings.json, while env overrides stay in Rust. This adds the Python
default User-Agent, parses string ssl_verify globals like get_ssl_verify, drops per-call
ssl_verify that Python OCR never honored, and removes the unused request_timeout.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
2026-09-18 17:23:26 -07:00
Mateo Wang
9b342cdd40
Merge pull request #41868 from BerriAI/litellm_config_update_rejects_config_owned_keys
fix(proxy): refuse config-owned keys on POST /config/update
2026-09-18 17:23:23 -07:00
kerry
4036b769a9 fix(proxy): count unprefixed release bullets and coalesce concurrent latest release fetches
Unprefixed release bullets now count as other_updates, concurrent cache misses share one upstream GitHub request through an injected asyncio.Lock, and the dashboard upgrade banner is announced as status rather than alert

Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>
2026-09-19 00:22:15 +00:00
kerry
f836bb481d test(integration): keep cost diagnostics and widen shard timeout
Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>
2026-09-19 00:19:11 +00:00