Commit graph

18968 commits

Author SHA1 Message Date
yassin
6be9c4a978 refactor(router): compose DeploymentSemaphore over asyncio.Semaphore instead of subclassing it
Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>
2026-09-17 03:17:49 +00:00
yassin
6e1b4959d1 feat(passthrough): deepgram streaming /v1/listen WebSocket passthrough with duration-based cost tracking
Adds authenticated /deepgram/v1/listen and /deepgram/listen WebSocket routes that resolve the Deepgram
credential through the pass-through router, inject Authorization: Token upstream, default the model to
nova-3 when the client passes none, and relay audio and transcript frames unchanged. The shared WebSocket
relay no longer assumes the first upstream frame is JSON and forwards every frame as received, keeping
the Vertex AI Live setup handling on Vertex routes only. A Deepgram logging handler bills the call on
Metadata.duration, falling back to the furthest Results start + duration, at the deepgram/<model>
per-second rate from the model cost map

Resolves LIT-7937

Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>
2026-09-17 03:14:13 +00:00
yassin
2e8dc0a627 feat(proxy): add Azure AI Speech pass-through route
Adds /azure_speech/{endpoint:path}, an authenticated pass-through for the Azure AI Speech REST APIs: short-audio recognition on <region>.stt.speech.microsoft.com and batch transcription on <region>.api.cognitive.microsoft.com. The proxy resolves the subscription key through PassthroughEndpointRouter (AZURE_SPEECH_API_KEY or an Admin UI credential), picks the host from AZURE_SPEECH_REGION or AZURE_SPEECH_API_BASE, injects Ocp-Apim-Subscription-Key, strips the caller's Authorization and subscription-key headers, forwards the raw audio body byte for byte, and records a zero-cost SpendLogs row tagged azure_speech since the price map has no Azure Speech STT entry

Resolves LIT-7939

Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>
2026-09-17 03:03:19 +00:00
yassin
fe0eee6451 fix(anthropic): keep prompt cache prediction supported for queue-bounded deployments
Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>
2026-09-17 03:02:42 +00:00
yassin
446fadc4c7 feat(router): bound the max_parallel_requests wait queue and return 429 on overflow
Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>
2026-09-17 02:50:45 +00:00
yuneng-berri
44a0e16c81
test(e2e): read a deleted key back as deleted, not as a 404
/key/info now serves a deleted key from the archive with status deleted
instead of answering 404, so the delete test's convergence predicate never
settled and the read timed out against a 200 it kept discarding.

The predicate now waits for status deleted through the same
_key_info_everywhere helper the rest of the file uses, and KeyInfo carries
the status field. The chat-rejection assertion after it is unchanged, so
the test still proves the key stops serving.
2026-09-17 02:38:27 +00:00
yucheng
3c000e4ffb fix(proxy): run prompt injection heuristics on a dedicated bounded executor
Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>
2026-09-17 02:38:13 +00:00
yucheng
d50bac391e test(proxy): cover startup router wiring for registered prompt injection detectors
Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>
2026-09-17 02:26:12 +00:00
yassin
8691a1e190 fix(vault): cache the secret body per url so mutations evict every field
Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>
2026-09-17 02:23:50 +00:00
yassin
4694bd0c63 fix(vault): key the secret cache by url and data field
Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>
2026-09-17 02:08:58 +00:00
yucheng
fc77914df3 test(proxy): type the moderation override stub in hook detection tests
Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>
2026-09-17 01:59:04 +00:00
yassin
bb9ff8cb2c fix(bedrock): keep realtime SDK error range inside websocket close reason
Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>
2026-09-17 01:57:33 +00:00
kerry
ef1f306a7d test(e2e): emit gemini stream usage only on the final chunk
Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>
2026-09-17 01:53:52 +00:00
kerry
6a9ae2bba2 ci(aws-partition): count allowlisted literal occurrences so duplicates in allowed files fail
Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>
2026-09-17 01:47:33 +00:00
yassin
15bfe8f28a feat(vault): add separate login and secret namespaces for HashiCorp Vault
Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>
2026-09-17 01:41:48 +00:00
yassin
0259e8c7d5 fix(bedrock): support aws-sdk-bedrock-runtime 0.10 and 0.11 in the realtime handler
The bedrock-realtime extra pinned aws-sdk-bedrock-runtime 0.7.x, whose Config and BedrockRuntimeClient surface is gone in 0.11. The handler now resolves AsyncBedrockRuntimeConfig, builds AsyncBedrockRuntimeClient with the awscrt duplex transport, closes the client when the session ends, and tells an absent SDK apart from an installed but unsupported version. Moves the pin to >=0.10.0,<0.12.0 with the awscrt extra

Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>
2026-09-17 01:40:52 +00:00
yucheng
b1255a6f2c fix(proxy): run prompt injection heuristics off the event loop and dispatch llm_api_check moderation
Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>
2026-09-17 01:33:54 +00:00
kerry
302394edff ci: gate hardcoded commercial AWS partition literals and test us-gov endpoint builders
Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>
2026-09-17 01:33:43 +00:00
mateo-berri
47b2479c94 fix(bedrock): gate Invoke tool search on the model map's supports_tool_search flag
The Bedrock InvokeModel transformations decided whether to send the
tool-search-tool-2025-10-19 beta from hardcoded model name lists (a pattern
list on the messages path, an "opus-4" substring on the chat path), so Opus 4.8,
Opus 5 and Sonnet 5 never got the beta on the messages path, Opus 5 and Sonnet 5
never got it on the chat path, Opus 4.1 got it without support, and
/v1/model/info reported supports_tool_search as unset for all three.

Both paths now read the model map through one shared helper: the Bedrock
entries for Opus 4.8, Opus 5 and Sonnet 5 carry supports_tool_search
explicitly, and a claude-tool-search fallback rule flags Claude 4.5 and newer
for unmapped ids, inference-profile ARNs and mapped entries with no opinion,
so the next Claude gets the beta with no code change. An explicit false on a
resolved entry still wins.
2026-09-16 18:23:14 -07:00
kerry
1de633ac36 test(e2e): move matrix data freshness checks to collection time
Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>
2026-09-17 01:22:01 +00:00
Yassin Kortam
351a54e849
Merge pull request #41507 from BerriAI/litellm_attribute_router_rejected_spend_provider
fix(spend_tracking): attribute router-rejected requests to the model group provider
2026-09-16 18:21:44 -07:00
kerry
072b32baf2 test(e2e): derive goldens from first-principles rate selection
Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>
2026-09-17 01:18:02 +00:00
mateo-berri
d8c3a38a51 Merge remote-tracking branch 'origin/main' into litellm_jwt_token_exchange_grant
# Conflicts:
#	tests/test_litellm/proxy/auth/test_auth_checks.py
2026-09-16 18:17:57 -07:00
kerry
fc0cce553a test(e2e): derive cache rates from first principles and ungate all_components cases
Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>
2026-09-17 01:13:36 +00:00
Devin AI
b77f866dbb refactor(mistral): drop client_metadata without mutating optional_params
Some checks failed
LiteLLM Rust / rust-lint (push) Has been cancelled
LiteLLM Rust / rust-test (push) Has been cancelled
LiteLLM Rust / rust-wheel (push) Has been cancelled
Terraform Provider / gofmt, vet, build, test (push) Has been cancelled
Terraform Provider / Provider endpoints vs proxy OpenAPI schema (push) Has been cancelled
Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>
2026-09-17 01:13:13 +00:00
ryan-crabbe-berri
fc13cea479 fix(proxy): refuse a team admin's budget write when the budget changed mid-request
The keep-or-lower check compares against the budget update_team read, so the write now only lands while the stored max_budget still matches it and answers 409 otherwise. A concurrent proxy admin cut can no longer be overwritten with a higher value.
2026-09-16 18:11:53 -07:00
yassin
8533dc9673 fix(helm): route /transcribe to the gateway and drop pinned botocore operation from test
Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>
2026-09-17 01:06:15 +00:00
yucheng
7815719de7 fix(guardrails): stream Prompt Security post_call redactions in incremental_diff mode
Forward streaming_transform_mode from guardrail litellm_params into PromptSecurityGuardrail so incremental_diff is reachable from config; the default stays block_only. In incremental_diff the guardrail now returns stream_holdback_chars alongside the rewritten texts so that a value split across streamed chunks (or across an abbreviation period) is never partially released before the vendor rewrite arrives. Each response text gets its own protect call so modified_text maps back to the right choice when n > 1, and custom_guardrail no longer logs a clean response as mask just because the guardrail attached holdback metadata

Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>
2026-09-17 01:06:12 +00:00
Mateo Wang
2445bdd2b5
Merge pull request #40934 from BerriAI/litellm_fix_ocr_native_multipage_pdf
fix(logging): scan each log record once and collapse base64 payloads before the secret regex
2026-09-16 18:06:11 -07:00
kerry
3e11c98676 test(e2e): satisfy pyright in cost matrix derivation and golden generator
Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>
2026-09-17 01:02:41 +00:00
mateo-berri
f93d80ea84 fix(mistral): forward reasoning_effort only on models that accept it 2026-09-16 18:02:06 -07:00
mateo-berri
c2f77fd358 refactor(a2a): resolve the relay's Entra hop bearer inside the a2a provider helper 2026-09-16 17:56:42 -07:00
kerry
e1c9ae5ae4 test(e2e): drop needless sys.path bootstrap from golden generator
Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>
2026-09-17 00:56:39 +00:00
yuneng-jiang
38676aa599
Merge pull request #41078 from BerriAI/litellm_integration_extensions
test: add extension and browser integration contracts
2026-09-16 17:55:25 -07:00
mateo-berri
65f4d31bb8 Merge remote-tracking branch 'origin/main' into litellm_mistral_codex_reasoning_effort_client_metadata 2026-09-16 17:53:59 -07:00
yuneng-jiang
e927b63211
Merge pull request #41520 from BerriAI/litellm_/attribution-investigation-a81211
fix(e2e): bind provider-cache recordings to the deployment's test, not the serving process
2026-09-16 17:52:03 -07:00
kerry
522a7f5692 test(e2e): gate all_components cases by rates and tidy cost matrix names
Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>
2026-09-17 00:51:43 +00:00
yucheng-berri
0add8c0083
Merge pull request #41495 from BerriAI/litellm_converted_stream_post_call_hook
fix(utils): run post-call deployment hook on converted chat streams
2026-09-16 17:49:14 -07:00
kerry
bdfff602fb test(e2e): drive the cost matrix from cases.json and expected.json goldens
Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>
2026-09-17 00:47:55 +00:00
Yuneng Jiang
02ced74540
test: fix five tests left stale by #41311, #41337, #39996 and #41310
Every one of these fails on main's own scheduled CircleCI run with the same
assertion as on any PR, and each traces to a merged behavior change that
never updated the test that pinned the old behavior

- tests/integration/_support/client.py: #41311 made /key/info serve deleted
  keys from the archive with status deleted, so the scenario teardown asserts
  the live row is gone and the readback reports deleted instead of a 404.
  This alone accounts for nine integration-management and one
  integration-providers failure
- tests/integration/authorization/test_warmed_policy.py: #39996 made team
  admins unable to edit any team field unless a proxy admin allow-lists it,
  and tpm_limit is the only field it accepts today. The demotion test now
  enables tpm_limit for the scenario and edits that instead of team_alias
- tests/llm_responses_api_testing/test_base_responses_api_streaming_iterator.py:
  #41337 reads usage off the terminal response and copies the event when it
  is missing, which a Mock(spec=ResponsesAPIResponse) cannot survive. The
  four mocks now carry a usage object
- tests/test_openai_endpoints.py: #41310 lengthened the access-denied
  message, and the test matched against the ExceptionInfo repr, which
  saferepr truncates in the middle. It now matches the exception text
- tests/local_testing/test_text_completion.py: Together no longer serves
  Qwen2-1.5B serverless, the cheapest cost-map row. The test mocks the
  completions call and asserts the request litellm builds, so a vendor
  catalog rotation cannot fail it again

test_router_fallbacks_with_cooldowns_and_dynamic_credentials is deliberately
untouched: it passes and fails on main with identical code, and the failing
path is a product question about whether dynamic-credential 429s cool down
2026-09-16 17:47:27 -07:00
Yuneng Jiang
86f709d7c9
test(aws): preserve rotation response coverage 2026-09-16 17:47:09 -07:00
Yassin Kortam
617a40bb1c
Merge pull request #40842 from BerriAI/litellm_guardrail_tag_budget_enforcement
fix(proxy): enforce tag budgets for tags added by guardrails
2026-09-16 17:44:29 -07:00
yassin
291a43c409 Merge remote-tracking branch 'origin/main' into litellm_transcribe_passthrough 2026-09-17 00:44:17 +00:00
yucheng
126e257062 merge: resolve conflict with main in otel metadata module
Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>
2026-09-17 00:38:50 +00:00
yucheng
8a645bcc00 test(proxy): stub DATABASE_URL in the hold-pool regression test so it passes off the dev box
Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>
2026-09-17 00:36:12 +00:00
kerry-berri
db408f68ae
Merge pull request #41494 from BerriAI/litellm_auto_merge_price_sync
ci: auto-merge provider-info-sync PRs when CI, Greptile and Bugbot are clean
2026-09-16 17:36:02 -07:00
ryan-crabbe-berri
e3a82f2f66 fix(proxy): stop team admins raising an org team's max_budget under the org cap
The keep-or-lower budget rule only ran for standalone teams, so once max_budget is enabled a team admin on an org team could grow its own budget up to the organization's. It now applies to team admins on every team; org admins keep editing within the org cap.
2026-09-16 17:35:18 -07:00
Yuneng Jiang
a0a006f248
fix(e2e): own a shared fixture's deployment by the fixture's node, not the first test
A deployment registered while a module- or class-scoped fixture is being set up
was bound to whichever test asked for the fixture first, so every later test in
the module shared that partition. A session-scoped fixture is set up by every
xdist worker, so its deployment could never have one owner at all.

The e2e conftest now wraps pytest_fixture_setup and records the node the fixture
is scoped to: registrations made during a module or class fixture's setup carry
that node's slug, and a session- or package-scoped one has no owner and stays
live. The registration seam test moves from tests/e2e to the cache harness tests
beside the rest of the attribution coverage.
2026-09-16 17:35:05 -07:00
yucheng
e1cce943de Merge remote-tracking branch 'origin/main' into litellm_converted_stream_post_call_hook 2026-09-17 00:31:24 +00:00
mateo-berri
55cf4c43ed fix(logging): stamp scrubbed records with a private sentinel a caller cannot supply
A record stamped litellm_redacted=True skips the secret filter and both
formatters, and extra={"litellm_redacted": True} on any log call put that
stamp on a fresh record before the filter ran. The stamp is now a private
object compared by identity, so only the filter's own pass marks a record
scrubbed.
2026-09-16 17:30:49 -07:00