Commit graph

52143 commits

Author SHA1 Message Date
ryan-crabbe-berri
8b64f1ef03
Merge pull request #41525 from BerriAI/litellm_team_admin_rpm_budget_fields
feat(proxy): let team admins edit rpm_limit and max_budget when enabled
2026-09-16 20:39:05 -07:00
yucheng
2cbfd280e9 Merge remote-tracking branch 'origin/main' into litellm_prompt_injection_async_llm_check
Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>

# Conflicts:
#	litellm/proxy/proxy_server.py
2026-09-17 03:32:34 +00:00
yassin
f2305879d0 feat(proxy): price Azure Speech short audio pass-through from the recognized duration
Short audio responses carry Offset and Duration in 100ns ticks; convert their sum to seconds and price it with the existing azure/speech/azure-stt entry through transcription_cost. Batch calls and responses without an integer duration stay at zero cost

Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>
2026-09-17 03:31:29 +00:00
tin-berri
d18e06f736
Merge pull request #41508 from BerriAI/litellm_1789600151_discover_context_limits
feat(router): discover token limits for hosted OpenAI-compatible models
2026-09-16 20:29:57 -07:00
Yujong Lee
e0ce998091 fmt 2026-09-16 20:21:51 -07:00
yassin
6be9c4a978 refactor(router): compose DeploymentSemaphore over asyncio.Semaphore instead of subclassing it
Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>
2026-09-17 03:17:49 +00:00
yassin
6e1b4959d1 feat(passthrough): deepgram streaming /v1/listen WebSocket passthrough with duration-based cost tracking
Adds authenticated /deepgram/v1/listen and /deepgram/listen WebSocket routes that resolve the Deepgram
credential through the pass-through router, inject Authorization: Token upstream, default the model to
nova-3 when the client passes none, and relay audio and transcript frames unchanged. The shared WebSocket
relay no longer assumes the first upstream frame is JSON and forwards every frame as received, keeping
the Vertex AI Live setup handling on Vertex routes only. A Deepgram logging handler bills the call on
Metadata.duration, falling back to the furthest Results start + duration, at the deepgram/<model>
per-second rate from the model cost map

Resolves LIT-7937

Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>
2026-09-17 03:14:13 +00:00
yassin
2e8dc0a627 feat(proxy): add Azure AI Speech pass-through route
Adds /azure_speech/{endpoint:path}, an authenticated pass-through for the Azure AI Speech REST APIs: short-audio recognition on <region>.stt.speech.microsoft.com and batch transcription on <region>.api.cognitive.microsoft.com. The proxy resolves the subscription key through PassthroughEndpointRouter (AZURE_SPEECH_API_KEY or an Admin UI credential), picks the host from AZURE_SPEECH_REGION or AZURE_SPEECH_API_BASE, injects Ocp-Apim-Subscription-Key, strips the caller's Authorization and subscription-key headers, forwards the raw audio body byte for byte, and records a zero-cost SpendLogs row tagged azure_speech since the price map has no Azure Speech STT entry

Resolves LIT-7939

Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>
2026-09-17 03:03:19 +00:00
yassin
fe0eee6451 fix(anthropic): keep prompt cache prediction supported for queue-bounded deployments
Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>
2026-09-17 03:02:42 +00:00
yassin
446fadc4c7 feat(router): bound the max_parallel_requests wait queue and return 429 on overflow
Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>
2026-09-17 02:50:45 +00:00
Yujong Lee
edfa01da81 refactor(ocr): mirror Python provider layout and preserve tests 2026-09-16 19:48:26 -07:00
yuneng-berri
44a0e16c81
test(e2e): read a deleted key back as deleted, not as a 404
/key/info now serves a deleted key from the archive with status deleted
instead of answering 404, so the delete test's convergence predicate never
settled and the read timed out against a 200 it kept discarding.

The predicate now waits for status deleted through the same
_key_info_everywhere helper the rest of the file uses, and KeyInfo carries
the status field. The chat-rejection assertion after it is unchanged, so
the test still proves the key stops serving.
2026-09-17 02:38:27 +00:00
yucheng
3c000e4ffb fix(proxy): run prompt injection heuristics on a dedicated bounded executor
Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>
2026-09-17 02:38:13 +00:00
yucheng
d50bac391e test(proxy): cover startup router wiring for registered prompt injection detectors
Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>
2026-09-17 02:26:12 +00:00
yassin
8691a1e190 fix(vault): cache the secret body per url so mutations evict every field
Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>
2026-09-17 02:23:50 +00:00
yassin
4694bd0c63 fix(vault): key the secret cache by url and data field
Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>
2026-09-17 02:08:58 +00:00
Yujong Lee
03cd00fbb1 refactor(rust): standardize messages errors
Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>
2026-09-17 02:08:26 +00:00
yucheng
fc77914df3 test(proxy): type the moderation override stub in hook detection tests
Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>
2026-09-17 01:59:04 +00:00
yassin
bb9ff8cb2c fix(bedrock): keep realtime SDK error range inside websocket close reason
Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>
2026-09-17 01:57:33 +00:00
Devin AI
ed18edbbdd chore(proxy): drop a comment that restated the NUM_WORKERS assignment
Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>
2026-09-17 01:56:05 +00:00
kerry
ef1f306a7d test(e2e): emit gemini stream usage only on the final chunk
Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>
2026-09-17 01:53:52 +00:00
kerry
6a9ae2bba2 ci(aws-partition): count allowlisted literal occurrences so duplicates in allowed files fail
Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>
2026-09-17 01:47:33 +00:00
yassin
15bfe8f28a feat(vault): add separate login and secret namespaces for HashiCorp Vault
Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>
2026-09-17 01:41:48 +00:00
yassin
0259e8c7d5 fix(bedrock): support aws-sdk-bedrock-runtime 0.10 and 0.11 in the realtime handler
The bedrock-realtime extra pinned aws-sdk-bedrock-runtime 0.7.x, whose Config and BedrockRuntimeClient surface is gone in 0.11. The handler now resolves AsyncBedrockRuntimeConfig, builds AsyncBedrockRuntimeClient with the awscrt duplex transport, closes the client when the session ends, and tells an absent SDK apart from an installed but unsupported version. Moves the pin to >=0.10.0,<0.12.0 with the awscrt extra

Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>
2026-09-17 01:40:52 +00:00
yucheng
b1255a6f2c fix(proxy): run prompt injection heuristics off the event loop and dispatch llm_api_check moderation
Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>
2026-09-17 01:33:54 +00:00
kerry
302394edff ci: gate hardcoded commercial AWS partition literals and test us-gov endpoint builders
Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>
2026-09-17 01:33:43 +00:00
Yujong Lee
4e5a9efd9d feat(rust): map anthropic messages transforms 2026-09-16 18:25:40 -07:00
mateo-berri
47b2479c94 fix(bedrock): gate Invoke tool search on the model map's supports_tool_search flag
The Bedrock InvokeModel transformations decided whether to send the
tool-search-tool-2025-10-19 beta from hardcoded model name lists (a pattern
list on the messages path, an "opus-4" substring on the chat path), so Opus 4.8,
Opus 5 and Sonnet 5 never got the beta on the messages path, Opus 5 and Sonnet 5
never got it on the chat path, Opus 4.1 got it without support, and
/v1/model/info reported supports_tool_search as unset for all three.

Both paths now read the model map through one shared helper: the Bedrock
entries for Opus 4.8, Opus 5 and Sonnet 5 carry supports_tool_search
explicitly, and a claude-tool-search fallback rule flags Claude 4.5 and newer
for unmapped ids, inference-profile ARNs and mapped entries with no opinion,
so the next Claude gets the beta with no code change. An explicit false on a
resolved entry still wins.
2026-09-16 18:23:14 -07:00
kerry
1de633ac36 test(e2e): move matrix data freshness checks to collection time
Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>
2026-09-17 01:22:01 +00:00
Yassin Kortam
351a54e849
Merge pull request #41507 from BerriAI/litellm_attribute_router_rejected_spend_provider
fix(spend_tracking): attribute router-rejected requests to the model group provider
2026-09-16 18:21:44 -07:00
kerry
072b32baf2 test(e2e): derive goldens from first-principles rate selection
Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>
2026-09-17 01:18:02 +00:00
mateo-berri
d8c3a38a51 Merge remote-tracking branch 'origin/main' into litellm_jwt_token_exchange_grant
# Conflicts:
#	tests/test_litellm/proxy/auth/test_auth_checks.py
2026-09-16 18:17:57 -07:00
kerry
fc0cce553a test(e2e): derive cache rates from first principles and ungate all_components cases
Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>
2026-09-17 01:13:36 +00:00
Devin AI
b77f866dbb refactor(mistral): drop client_metadata without mutating optional_params
Some checks failed
LiteLLM Rust / rust-lint (push) Has been cancelled
LiteLLM Rust / rust-test (push) Has been cancelled
LiteLLM Rust / rust-wheel (push) Has been cancelled
Terraform Provider / gofmt, vet, build, test (push) Has been cancelled
Terraform Provider / Provider endpoints vs proxy OpenAPI schema (push) Has been cancelled
Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>
2026-09-17 01:13:13 +00:00
ryan-crabbe-berri
fc13cea479 fix(proxy): refuse a team admin's budget write when the budget changed mid-request
The keep-or-lower check compares against the budget update_team read, so the write now only lands while the stored max_budget still matches it and answers 409 otherwise. A concurrent proxy admin cut can no longer be overwritten with a higher value.
2026-09-16 18:11:53 -07:00
Yujong Lee
2414d1f028 feat(rust): scaffold anthropic stream transformation 2026-09-16 18:10:28 -07:00
yassin
65d0f3a03d fix(terraform): mirror /transcribe into the AWS and GCP gateway prefix lists
Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>
2026-09-17 01:07:57 +00:00
yassin
8533dc9673 fix(helm): route /transcribe to the gateway and drop pinned botocore operation from test
Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>
2026-09-17 01:06:15 +00:00
yucheng
7815719de7 fix(guardrails): stream Prompt Security post_call redactions in incremental_diff mode
Forward streaming_transform_mode from guardrail litellm_params into PromptSecurityGuardrail so incremental_diff is reachable from config; the default stays block_only. In incremental_diff the guardrail now returns stream_holdback_chars alongside the rewritten texts so that a value split across streamed chunks (or across an abbreviation period) is never partially released before the vendor rewrite arrives. Each response text gets its own protect call so modified_text maps back to the right choice when n > 1, and custom_guardrail no longer logs a clean response as mask just because the guardrail attached holdback metadata

Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>
2026-09-17 01:06:12 +00:00
Mateo Wang
2445bdd2b5
Merge pull request #40934 from BerriAI/litellm_fix_ocr_native_multipage_pdf
fix(logging): scan each log record once and collapse base64 payloads before the secret regex
2026-09-16 18:06:11 -07:00
kerry
3e11c98676 test(e2e): satisfy pyright in cost matrix derivation and golden generator
Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>
2026-09-17 01:02:41 +00:00
mateo-berri
f93d80ea84 fix(mistral): forward reasoning_effort only on models that accept it 2026-09-16 18:02:06 -07:00
yassin
e06c81665f refactor(spend_tracking): type the get_logging_payload parameters
Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>
2026-09-17 01:02:03 +00:00
mateo-berri
c2f77fd358 refactor(a2a): resolve the relay's Entra hop bearer inside the a2a provider helper 2026-09-16 17:56:42 -07:00
kerry
e1c9ae5ae4 test(e2e): drop needless sys.path bootstrap from golden generator
Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>
2026-09-17 00:56:39 +00:00
yuneng-jiang
38676aa599
Merge pull request #41078 from BerriAI/litellm_integration_extensions
test: add extension and browser integration contracts
2026-09-16 17:55:25 -07:00
mateo-berri
65f4d31bb8 Merge remote-tracking branch 'origin/main' into litellm_mistral_codex_reasoning_effort_client_metadata 2026-09-16 17:53:59 -07:00
yuneng-jiang
e927b63211
Merge pull request #41520 from BerriAI/litellm_/attribution-investigation-a81211
fix(e2e): bind provider-cache recordings to the deployment's test, not the serving process
2026-09-16 17:52:03 -07:00
kerry
522a7f5692 test(e2e): gate all_components cases by rates and tidy cost matrix names
Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>
2026-09-17 00:51:43 +00:00
ryan-crabbe-berri
d3f0607820 test(ui): find the max_budget checkbox by its new label 2026-09-16 17:50:57 -07:00