Commit graph

3448 commits

Author SHA1 Message Date
mateo-berri
cb80e8773e fix(policy_engine): fail open when an unended Messages stream has no text delta to carry the rewrite 2026-09-18 23:24:49 -07:00
mateo-berri
b3d9ba9e7b fix(policy_engine): deliver guardrail text rewrites on multi-choice, unfinished, and envelope-less streams
Post-call pipeline rewrites on buffered streams failed open on three shapes:
chat streams with n > 1 (the rebuilt response collapsed every choice into
index 0), streams that ended without a finish marker, and Responses streams
whose final event carried no response envelope.

The chat handler now rebuilds the ended stream one choice index at a time and
writes each choice's rewrite back to that choice's buffered deltas. The
Anthropic handler writes an unended stream's rewrite across its text deltas.
The Responses handler spreads an envelope-less rewrite over the buffered
output_text events, still failing open when a scanned event cannot be placed.

Tool-call rewrites on n > 1 chat streams keep failing open.
2026-09-18 22:59:52 -07:00
Mateo Wang
faed57f92c
Merge pull request #41918 from BerriAI/litellm_websearch_followup_api_base
fix(websearch): forward the deployment api_base to agentic follow-up calls on /v1/messages
2026-09-18 21:41:41 -07:00
Mateo Wang
9486caf584
Merge pull request #41721 from BerriAI/litellm_vertex_chirp3_streaming_stt
feat(vertex_ai): stream Chirp speech-to-text over /v1/realtime
2026-09-18 20:34:56 -07:00
kerry-berri
f26afabe6d
Merge pull request #41914 from BerriAI/litellm_xai_audio_transcription
feat(xai): add speech-to-text (Grok Voice Transcribe) via /v1/audio/transcriptions
2026-09-18 20:03:00 -07:00
kerry-berri
c4ddea7cfc
Merge pull request #41917 from BerriAI/litellm_fireworks_cache_read_default_discount
fix(cost_calc): default fireworks cached input to the documented 50% discount when the map has no cache-read rate
2026-09-18 18:48:58 -07:00
Mateo Wang
29bdd1ab47
Merge pull request #41905 from BerriAI/litellm_websearch_failed_search_error_block
fix(websearch_interception): surface a failed search as a web_search_tool_result_error block and end the turn
2026-09-18 18:34:11 -07:00
kerry
05cefb1480 fix(cost_calc): coerce string fireworks rates and drop the match fall-through
Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>
2026-09-19 01:33:41 +00:00
kerry-berri
fc0b37ff5d
Merge pull request #41904 from BerriAI/litellm_bedrock_batch_retrieve_sigv4_over_env_bearer
fix(bedrock): sign batch retrieve and cancel with deployment credentials when AWS_BEARER_TOKEN_BEDROCK is set
2026-09-18 18:33:10 -07:00
kerry
99d91d7205 fix(xai): reject non-success stt responses before parsing
Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>
2026-09-19 01:22:25 +00:00
mateo-berri
b1b7af884a fix(websearch): forward the deployment api_base to agentic follow-up calls on /v1/messages 2026-09-18 18:19:48 -07:00
kerry
c9cd666b36 refactor(cost_calc): move the fireworks cache-read default under litellm/llms
Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>
2026-09-19 01:18:09 +00:00
mateo-berri
17c519c40a test(custom_httpx): pass the token resolver and two-argument client factory in the realtime bridge test 2026-09-18 18:09:33 -07:00
kerry
1b305cd6b9 fix(cost_calc): default fireworks cached input to the documented 50% discount when the map has no cache-read rate
Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>
2026-09-19 01:07:07 +00:00
kerry
80b0ea6a2f test(xai): narrow raises match for missing api key
Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>
2026-09-19 01:00:53 +00:00
kerry
6f54ad5166 fix(xai): parse integer speaker ids and simplify stt form build
Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>
2026-09-19 00:53:02 +00:00
kerry
6b0ad3bed3 feat(xai): add speech-to-text via /v1/audio/transcriptions
Route xai audio transcription through a provider config hitting POST
https://api.x.ai/v1/stt instead of the openai-compatible chat handler
which targets /audio/transcriptions. Supports language, diarize,
keyterm, filler_words and other provider fields as passthrough kwargs

Resolves LIT-8153

Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>
2026-09-19 00:49:34 +00:00
mateo-berri
3911d62bbe fix(vertex_ai): prune a discarded turn's id once its marker is delivered 2026-09-18 17:44:28 -07:00
Mateo Wang
a6e3a72ed8
Merge pull request #41870 from BerriAI/litellm_bedrock_openai_gpt_min_max_tokens
fix(bedrock): clamp maxTokens to the 16-token minimum for OpenAI GPT and xAI Grok models on Converse
2026-09-18 17:44:18 -07:00
kerry
99659e9e7e test(bedrock): drop redundant recorder docstring
Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>
2026-09-19 00:43:51 +00:00
mateo-berri
8b1f78fa08 fix(vertex_ai): drop a cleared turn's queued transcripts and carry its billed seconds 2026-09-18 17:30:57 -07:00
kerry
6b082d3a01 test(bedrock): type the SigV4 request recorder and drop caller-owned mutation
Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>
2026-09-19 00:25:01 +00:00
mateo-berri
cdc0e57e93 fix(websearch_interception): surface a failed search as a web_search_tool_result_error block and end the turn 2026-09-18 17:24:52 -07:00
mateo-berri
12120fe59b refactor(bedrock): inline maxTokens clamp and cover inference-profile ARNs in tests 2026-09-18 17:11:29 -07:00
kerry
37da5b6f4d fix(bedrock): sign batch retrieve and cancel with deployment credentials when AWS_BEARER_TOKEN_BEDROCK is set
Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>
2026-09-19 00:10:17 +00:00
mateo-berri
dc11c34e3a Merge remote-tracking branch 'origin/main' into litellm_vertex_chirp3_streaming_stt 2026-09-18 16:56:02 -07:00
mateo-berri
3d7a771ea7 fix(vertex_ai): apply finals before interims and refresh the token per stream 2026-09-18 16:53:06 -07:00
Mateo Wang
ff7dc86947
Merge pull request #41892 from BerriAI/litellm_gemini_contentless_candidate_finish_reason
Some checks failed
Unit Tests: Proxy DB Operations / custom-logging (push) Blocked by required conditions
Unit Tests: Proxy DB Operations / db-and-spend (push) Blocked by required conditions
Unit Tests: Proxy DB Operations / endpoints-and-responses (push) Blocked by required conditions
Unit Tests: Proxy DB Operations / proxy-server-core (push) Blocked by required conditions
Unit Tests: Proxy DB Operations / guardrails-hooks (push) Blocked by required conditions
Unit Tests: Proxy DB Operations / jwt-and-keys (push) Blocked by required conditions
Unit Tests: Proxy DB Operations / key-generation (push) Blocked by required conditions
Unit Tests: Proxy DB Operations / logging-misc (push) Blocked by required conditions
Unit Tests: Proxy DB Operations / proxy-runtime (push) Blocked by required conditions
Unit Tests: Proxy DB Operations / proxy-utils (push) Blocked by required conditions
Unit Tests / caching-local (push) Waiting to run
Unit Tests / core-utils (push) Waiting to run
Unit Tests / enterprise-package (push) Waiting to run
Unit Tests / enterprise-routing (push) Waiting to run
Unit Tests / integrations (push) Waiting to run
Unit Tests / All Other Providers (push) Waiting to run
Unit Tests / Vertex AI (push) Waiting to run
Unit Tests / misc (push) Waiting to run
Unit Tests / proxy-auth (push) Waiting to run
Unit Tests / proxy-endpoints (push) Waiting to run
Unit Tests / proxy-extras (push) Waiting to run
Unit Tests / proxy-infra (push) Waiting to run
Unit Tests / proxy-server (push) Waiting to run
Unit Tests / responses-caching-types (push) Waiting to run
GitHub Actions Security Analysis / zizmor (push) Waiting to run
Issue label sync / sync-issue-labels-tests (push) Has been cancelled
Issue label sync / sync-issue-labels (push) Has been cancelled
Terraform Modules / fmt, validate, test (aws) (push) Has been cancelled
Terraform Modules / fmt, validate, test (gcp) (push) Has been cancelled
VS Code Extension / vscode-extension (push) Has been cancelled
fix(gemini): preserve candidates with finishReason and no content (#40477)
2026-09-18 16:34:46 -07:00
Mateo Wang
ec05cd0128
Merge pull request #41887 from BerriAI/litellm_gemma_4_26b_maas_context_window
fix: set vertex gemma-4-26b-a4b-it-maas context window to 262144
2026-09-18 16:14:56 -07:00
Mateo Wang
3490754e65
Merge pull request #41871 from BerriAI/litellm_bedrock_eager_input_streaming
feat: honor eager_input_streaming on Bedrock and Anthropic Claude tools
2026-09-18 16:05:10 -07:00
mateo-berri
a4624b6c6c fix(gemini): derive the finish reason key set from the Candidates type
Candidates.finishReason listed eleven values while the mapping key set
carried twenty-one, so typed fixtures could not spell the reasons this
PR handles. GeminiFinishReason is now the one list, the key set derives
from it, and a test checks every documented reason has an explicit
mapping instead of falling through to "stop"
2026-09-18 16:03:45 -07:00
mateo-berri
44034c1d5e test: source the gemma context window limits and isolate the cost map cache 2026-09-18 15:43:07 -07:00
mateo-berri
f8f165d416 Merge remote-tracking branch 'origin/main' into litellm_vertex_chirp3_streaming_stt
# Conflicts:
#	uv.lock
2026-09-18 15:38:33 -07:00
mateo-berri
7f9db61528 fix(vertex_ai): carry turns across stream rotation and route by model info
Rotating the Speech-to-Text stream at 240 s no longer ends the active turn:
the turn and its billed seconds continue on the new stream, forced at 280 s.
Bound the request and event queues (64 and 256) so a slow peer applies
backpressure instead of growing memory. Route a model to the Chirp realtime
path from its cost-map entry (mode audio_transcription plus /v1/realtime)
instead of a hardcoded name. Return on every branch of the recv and
transform helpers (CodeQL mixed returns), have the shared protocol helper take
the provider's error class so the Meta tests assert MuseProtocolError again,
and pin google-cloud-speech in the ci group so unit shards import it.
2026-09-18 15:35:59 -07:00
Devin AI
665327249d fix(bedrock): extend maxTokens minimum clamp to xAI Grok models on Converse
Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>
2026-09-18 22:06:09 +00:00
Mateo Wang
7283293d83
Merge pull request #41138 from BerriAI/litellm_bedrock_files_s3_endpoint_url
fix(bedrock): carry s3_endpoint_url and s3_region_name into file content downloads
2026-09-18 15:04:01 -07:00
mateo-berri
55249c7128 fix: set vertex gemma-4-26b-a4b-it-maas context window to 262144 2026-09-18 14:58:58 -07:00
mateo-berri
cf05466a27 fix(gemini): map every documented finishReason and reset per-candidate state
A content-less candidate is now kept as a choice whenever it carries a
finishReason, with the raw value on the choice's provider_specific_fields.
NO_IMAGE, IMAGE_RECITATION, IMAGE_OTHER and ESCALATION map to content_filter;
UNEXPECTED_TOOL_CALL and MISSING_THOUGHT_SIGNATURE map to stop. The
/v1/responses bridge reports content_filter and refusal as incomplete with
incomplete_details, and tool calls and reasoning no longer leak from one
candidate into the next.
2026-09-18 14:54:18 -07:00
Yassin Kortam
52d6aab421
Merge pull request #41554 from BerriAI/litellm_deepgram_listen_websocket_passthrough
feat(passthrough): deepgram streaming /v1/listen WebSocket passthrough with duration-based cost tracking
2026-09-18 14:48:37 -07:00
Mateo Wang
d45e04a9fd
Merge pull request #41062 from BerriAI/litellm_mistral_codex_reasoning_effort_client_metadata
fix(mistral): accept reasoning_effort on all models and drop client_metadata for Codex compatibility
2026-09-18 14:43:58 -07:00
Mateo Wang
a6bd779bd1
Merge pull request #39424 from emerzon/litellm_azure_ai_flux_2_flex
feat(azure_ai): support FLUX.2 flex images
2026-09-18 14:30:25 -07:00
yassin
0b5b69ea3a fix(deepgram): forward only the first model and language values to /listen
Authorization and pricing read the first model and language query value, but the raw query was forwarded, so Deepgram (which honours the last repeated value) could be sent a model the key was never allowed. Later duplicates of those two keys are now dropped before the upstream URL is built

Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>
2026-09-18 21:29:29 +00:00
yassin
0897b663c5 Merge remote-tracking branch 'origin/litellm_deepgram_listen_websocket_passthrough' into litellm_deepgram_listen_websocket_passthrough 2026-09-18 21:08:45 +00:00
yassin
93d61abfa5 fix(deepgram): refuse /listen sessions that have no streaming price
A caller could pick a model with only a pre-recorded registry row, or no row at all, and the session would be billed at the pre-recorded rate or logged at zero cost, so budgets did not apply. The route now closes the WebSocket with 1008 before dialing Deepgram unless deepgram/streaming/<model> (or the -multilingual row for language=multi) is an exact registry hit, and the logging handler applies the same check so a registry change under a live session records the duration with no cost instead of a substitute rate

Regression tests cover the route refusal, an operator-supplied streaming row for another model being accepted, and the handler never substituting the pre-recorded rate

Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>
2026-09-18 21:08:14 +00:00
mateo-berri
9d39d7efaa chore: merge main into fix/gemini-candidate-finish-reason-no-content 2026-09-18 14:05:25 -07:00
Mateo Wang
20e1e6f2a9
Merge pull request #41384 from BerriAI/litellm_fix_azure_vector_store_search_url
fix(azure): keep api-version query after vector store search path
2026-09-18 13:56:05 -07:00
yassin
b7d808e416 Merge remote-tracking branch 'origin/main' into litellm_deepgram_listen_websocket_passthrough
Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>
2026-09-18 20:52:42 +00:00
kerry-berri
6fa34a299b
Merge pull request #41754 from BerriAI/litellm_qwen3_8_omni_flash
feat(models): add qwen3.8 flash rows, fix Cohere embed v3 context, Bedrock Mantle and OpenRouter pricing
2026-09-18 13:49:35 -07:00
yassin
1f3bee2e10 Merge remote-tracking branch 'origin/main' into litellm_deepgram_listen_websocket_passthrough
Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>

# Conflicts:
#	uv.lock
2026-09-18 20:32:05 +00:00
mateo-berri
a37d1e8b9b test: type the eager_input_streaming test helpers and mark constants Final 2026-09-18 13:30:08 -07:00