Commit graph

18680 commits

Author SHA1 Message Date
Yuneng Jiang
a0869fe835
test(budgets): avoid mutable fixture state 2026-09-16 22:06:31 -07:00
Devin AI
d4e54a0f34 fix(responses): keep sync text deltas and give the message item its own output index
Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>
2026-09-17 05:05:54 +00:00
yucheng
719d7a1983 test(proxy): cover inherited moderation overrides through during_call_hook
Replaces the capability flag assertion with a behavioral test that dispatches an
async_moderation_hook inherited from a parent class, and drops the dispatch
docstring that restated the code

Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>
2026-09-17 05:00:01 +00:00
Devin AI
441021fc96 fix(responses): announce message item before text events in the chat completions bridge
Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>
2026-09-17 04:46:26 +00:00
Yuneng Jiang
5c41e0b8dc
test(budgets): cover management null handling 2026-09-16 21:42:55 -07:00
yuneng-jiang
5ef40a630b
Merge pull request #41527 from BerriAI/litellm_/monitor-cci-failures-ac97a9
test: fix seven tests left stale by #41311, #41337, #39996, #41310, #41289 and #41315
2026-09-16 21:25:51 -07:00
yassin
1d8f19e4fd fix(proxy): keep Azure Speech multipart bodies intact through auth
user_api_key_auth called request.form() on multipart Azure Speech batch uploads, consuming the Starlette stream before the pass-through handler could read the raw bytes. The opaque body predicate now covers multipart on the whole /azure_speech prefix so auth caches an empty parsed body and the upload is forwarded byte for byte

Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>
2026-09-17 04:14:26 +00:00
yassin
45ceb56110 fix(router): validate max_parallel_requests_queue_size as a non-negative integer
Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>
2026-09-17 04:05:41 +00:00
yassin
6e56ba86c5 test(proxy): clear leaked auth dependency override before Azure Speech real-auth tests
Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>
2026-09-17 04:03:22 +00:00
yucheng-berri
821bcf5d78
Merge pull request #41140 from BerriAI/litellm_otel_v2_langfuse_user_session_tags 2026-09-16 20:55:42 -07:00
yassin
585c32d3f5 refactor(deepgram): move listen frame parsing into llms/deepgram and drop routine docstrings
Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>
2026-09-17 03:53:28 +00:00
Yuneng Jiang
c6023b4eec
test: pin the post-#41289 cooldown contract and scroll the auto-router select spec
test_router_fallbacks_with_cooldowns_and_dynamic_credentials expected a
caller-supplied credential to register its own deployment and cool it down.
#41289 stopped registering it, so cooldown logic skips that id and the
assertion can never hold. The test now asserts what the router guarantees
today: a 429 to a forwarded credential cools down none of the shared
deployments, the next credential is still served, and a 429 owned by a shared
deployment still cools it down. The final live OpenAI call becomes a mock

The auto-router template spec assumed the Add Auto Router form left room
below the Template select at 1280x900. #41315 added classifier fields above
it, so the options opened upward. The spec now scrolls the trigger to the top
of the dialog and asserts it sits in the upper half before checking placement
2026-09-16 20:52:19 -07:00
yuneng-jiang
54fa790e20
Merge pull request #41551 from BerriAI/litellm_cadence_319f427_key_lifecycle_delete
test(e2e): read a deleted key back as deleted, not as a 404
2026-09-16 20:48:55 -07:00
yuneng-jiang
3c34b92594
Merge pull request #41524 from BerriAI/litellm_aws_rotation_values
test(aws): verify rotated secret value
2026-09-16 20:41:12 -07:00
yucheng-berri
375cd4a668
Merge pull request #41498 from BerriAI/litellm_otel_indexed_messages_span_headroom 2026-09-16 20:40:24 -07:00
yassin
a0425a1f99 Merge remote-tracking branch 'origin/main' into litellm_azure_speech_passthrough 2026-09-17 03:40:09 +00:00
ryan-crabbe-berri
8b64f1ef03
Merge pull request #41525 from BerriAI/litellm_team_admin_rpm_budget_fields
feat(proxy): let team admins edit rpm_limit and max_budget when enabled
2026-09-16 20:39:05 -07:00
yucheng
2cbfd280e9 Merge remote-tracking branch 'origin/main' into litellm_prompt_injection_async_llm_check
Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>

# Conflicts:
#	litellm/proxy/proxy_server.py
2026-09-17 03:32:34 +00:00
yassin
f2305879d0 feat(proxy): price Azure Speech short audio pass-through from the recognized duration
Short audio responses carry Offset and Duration in 100ns ticks; convert their sum to seconds and price it with the existing azure/speech/azure-stt entry through transcription_cost. Batch calls and responses without an integer duration stay at zero cost

Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>
2026-09-17 03:31:29 +00:00
tin-berri
d18e06f736
Merge pull request #41508 from BerriAI/litellm_1789600151_discover_context_limits
feat(router): discover token limits for hosted OpenAI-compatible models
2026-09-16 20:29:57 -07:00
yassin
6be9c4a978 refactor(router): compose DeploymentSemaphore over asyncio.Semaphore instead of subclassing it
Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>
2026-09-17 03:17:49 +00:00
yassin
6e1b4959d1 feat(passthrough): deepgram streaming /v1/listen WebSocket passthrough with duration-based cost tracking
Adds authenticated /deepgram/v1/listen and /deepgram/listen WebSocket routes that resolve the Deepgram
credential through the pass-through router, inject Authorization: Token upstream, default the model to
nova-3 when the client passes none, and relay audio and transcript frames unchanged. The shared WebSocket
relay no longer assumes the first upstream frame is JSON and forwards every frame as received, keeping
the Vertex AI Live setup handling on Vertex routes only. A Deepgram logging handler bills the call on
Metadata.duration, falling back to the furthest Results start + duration, at the deepgram/<model>
per-second rate from the model cost map

Resolves LIT-7937

Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>
2026-09-17 03:14:13 +00:00
yassin
2e8dc0a627 feat(proxy): add Azure AI Speech pass-through route
Adds /azure_speech/{endpoint:path}, an authenticated pass-through for the Azure AI Speech REST APIs: short-audio recognition on <region>.stt.speech.microsoft.com and batch transcription on <region>.api.cognitive.microsoft.com. The proxy resolves the subscription key through PassthroughEndpointRouter (AZURE_SPEECH_API_KEY or an Admin UI credential), picks the host from AZURE_SPEECH_REGION or AZURE_SPEECH_API_BASE, injects Ocp-Apim-Subscription-Key, strips the caller's Authorization and subscription-key headers, forwards the raw audio body byte for byte, and records a zero-cost SpendLogs row tagged azure_speech since the price map has no Azure Speech STT entry

Resolves LIT-7939

Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>
2026-09-17 03:03:19 +00:00
yassin
fe0eee6451 fix(anthropic): keep prompt cache prediction supported for queue-bounded deployments
Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>
2026-09-17 03:02:42 +00:00
yassin
446fadc4c7 feat(router): bound the max_parallel_requests wait queue and return 429 on overflow
Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>
2026-09-17 02:50:45 +00:00
yuneng-berri
44a0e16c81
test(e2e): read a deleted key back as deleted, not as a 404
/key/info now serves a deleted key from the archive with status deleted
instead of answering 404, so the delete test's convergence predicate never
settled and the read timed out against a 200 it kept discarding.

The predicate now waits for status deleted through the same
_key_info_everywhere helper the rest of the file uses, and KeyInfo carries
the status field. The chat-rejection assertion after it is unchanged, so
the test still proves the key stops serving.
2026-09-17 02:38:27 +00:00
yucheng
3c000e4ffb fix(proxy): run prompt injection heuristics on a dedicated bounded executor
Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>
2026-09-17 02:38:13 +00:00
yucheng
d50bac391e test(proxy): cover startup router wiring for registered prompt injection detectors
Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>
2026-09-17 02:26:12 +00:00
yassin
8691a1e190 fix(vault): cache the secret body per url so mutations evict every field
Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>
2026-09-17 02:23:50 +00:00
yassin
4694bd0c63 fix(vault): key the secret cache by url and data field
Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>
2026-09-17 02:08:58 +00:00
yucheng
fc77914df3 test(proxy): type the moderation override stub in hook detection tests
Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>
2026-09-17 01:59:04 +00:00
yassin
bb9ff8cb2c fix(bedrock): keep realtime SDK error range inside websocket close reason
Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>
2026-09-17 01:57:33 +00:00
kerry
ef1f306a7d test(e2e): emit gemini stream usage only on the final chunk
Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>
2026-09-17 01:53:52 +00:00
kerry
6a9ae2bba2 ci(aws-partition): count allowlisted literal occurrences so duplicates in allowed files fail
Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>
2026-09-17 01:47:33 +00:00
yassin
15bfe8f28a feat(vault): add separate login and secret namespaces for HashiCorp Vault
Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>
2026-09-17 01:41:48 +00:00
yassin
0259e8c7d5 fix(bedrock): support aws-sdk-bedrock-runtime 0.10 and 0.11 in the realtime handler
The bedrock-realtime extra pinned aws-sdk-bedrock-runtime 0.7.x, whose Config and BedrockRuntimeClient surface is gone in 0.11. The handler now resolves AsyncBedrockRuntimeConfig, builds AsyncBedrockRuntimeClient with the awscrt duplex transport, closes the client when the session ends, and tells an absent SDK apart from an installed but unsupported version. Moves the pin to >=0.10.0,<0.12.0 with the awscrt extra

Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>
2026-09-17 01:40:52 +00:00
yucheng
b1255a6f2c fix(proxy): run prompt injection heuristics off the event loop and dispatch llm_api_check moderation
Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>
2026-09-17 01:33:54 +00:00
kerry
302394edff ci: gate hardcoded commercial AWS partition literals and test us-gov endpoint builders
Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>
2026-09-17 01:33:43 +00:00
mateo-berri
47b2479c94 fix(bedrock): gate Invoke tool search on the model map's supports_tool_search flag
The Bedrock InvokeModel transformations decided whether to send the
tool-search-tool-2025-10-19 beta from hardcoded model name lists (a pattern
list on the messages path, an "opus-4" substring on the chat path), so Opus 4.8,
Opus 5 and Sonnet 5 never got the beta on the messages path, Opus 5 and Sonnet 5
never got it on the chat path, Opus 4.1 got it without support, and
/v1/model/info reported supports_tool_search as unset for all three.

Both paths now read the model map through one shared helper: the Bedrock
entries for Opus 4.8, Opus 5 and Sonnet 5 carry supports_tool_search
explicitly, and a claude-tool-search fallback rule flags Claude 4.5 and newer
for unmapped ids, inference-profile ARNs and mapped entries with no opinion,
so the next Claude gets the beta with no code change. An explicit false on a
resolved entry still wins.
2026-09-16 18:23:14 -07:00
kerry
1de633ac36 test(e2e): move matrix data freshness checks to collection time
Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>
2026-09-17 01:22:01 +00:00
Yassin Kortam
351a54e849
Merge pull request #41507 from BerriAI/litellm_attribute_router_rejected_spend_provider
fix(spend_tracking): attribute router-rejected requests to the model group provider
2026-09-16 18:21:44 -07:00
kerry
072b32baf2 test(e2e): derive goldens from first-principles rate selection
Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>
2026-09-17 01:18:02 +00:00
mateo-berri
d8c3a38a51 Merge remote-tracking branch 'origin/main' into litellm_jwt_token_exchange_grant
# Conflicts:
#	tests/test_litellm/proxy/auth/test_auth_checks.py
2026-09-16 18:17:57 -07:00
kerry
fc0cce553a test(e2e): derive cache rates from first principles and ungate all_components cases
Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>
2026-09-17 01:13:36 +00:00
Devin AI
b77f866dbb refactor(mistral): drop client_metadata without mutating optional_params
Some checks failed
LiteLLM Rust / rust-lint (push) Has been cancelled
LiteLLM Rust / rust-test (push) Has been cancelled
LiteLLM Rust / rust-wheel (push) Has been cancelled
Terraform Provider / gofmt, vet, build, test (push) Has been cancelled
Terraform Provider / Provider endpoints vs proxy OpenAPI schema (push) Has been cancelled
Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>
2026-09-17 01:13:13 +00:00
ryan-crabbe-berri
fc13cea479 fix(proxy): refuse a team admin's budget write when the budget changed mid-request
The keep-or-lower check compares against the budget update_team read, so the write now only lands while the stored max_budget still matches it and answers 409 otherwise. A concurrent proxy admin cut can no longer be overwritten with a higher value.
2026-09-16 18:11:53 -07:00
yassin
8533dc9673 fix(helm): route /transcribe to the gateway and drop pinned botocore operation from test
Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>
2026-09-17 01:06:15 +00:00
yucheng
7815719de7 fix(guardrails): stream Prompt Security post_call redactions in incremental_diff mode
Forward streaming_transform_mode from guardrail litellm_params into PromptSecurityGuardrail so incremental_diff is reachable from config; the default stays block_only. In incremental_diff the guardrail now returns stream_holdback_chars alongside the rewritten texts so that a value split across streamed chunks (or across an abbreviation period) is never partially released before the vendor rewrite arrives. Each response text gets its own protect call so modified_text maps back to the right choice when n > 1, and custom_guardrail no longer logs a clean response as mask just because the guardrail attached holdback metadata

Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>
2026-09-17 01:06:12 +00:00
Mateo Wang
2445bdd2b5
Merge pull request #40934 from BerriAI/litellm_fix_ocr_native_multipage_pdf
fix(logging): scan each log record once and collapse base64 payloads before the secret regex
2026-09-16 18:06:11 -07:00
kerry
3e11c98676 test(e2e): satisfy pyright in cost matrix derivation and golden generator
Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>
2026-09-17 01:02:41 +00:00