Mateo Wang
8c4c394ecc
Merge pull request #41960 from BerriAI/litellm_deepseek_off_peak_pricing
...
fix(cost): bill DeepSeek V4.1 Flash and V4 Pro at off-peak rates outside peak hours
2026-09-19 06:45:46 -07:00
Mateo Wang
b8d837b2ef
Merge pull request #40147 from abhirup7/fix/azure-image-generation-entra-id-auth
...
fix(azure): send the resolved Entra ID token on image generation requests
2026-09-19 04:44:07 -07:00
Mateo Wang
dad8c32d23
Merge pull request #41940 from BerriAI/litellm_rag_ingest_registry_store
...
fix(rag): resolve registry stores on /v1/rag/ingest and reject providers without ingestion
2026-09-19 04:42:54 -07:00
mateo-berri
e0db862781
fix(cost): bill DeepSeek V4.1 Flash and V4 Pro at their off-peak rates outside peak hours
...
DeepSeek charges half the listed rate outside 01:00-04:00 and 06:00-10:00 UTC
Monday to Friday, so every deepseek-flash, deepseek-v4-flash,
deepseek-v4-flash-vision-exp, and deepseek-v4-pro entry now carries an
off_peak_pricing block with those windows and the halved input, output, and
cache-hit rates. The generated cost map schema picks up the block, and the
regression tests pin the peak and off-peak cost of one call at fixed moments.
2026-09-19 04:31:55 -07:00
mateo-berri
e4d01d1d78
fix(s3_vectors): reject a store id with an empty bucket or index part
...
A "bucket:" or ":index" id split into an empty name, so ingestion silently
generated a fresh index and search sent the empty name to AWS. Both sides now
raise the existing format error through the shared helper.
2026-09-19 03:52:48 -07:00
Mateo Wang
db04e7909e
Merge pull request #41934 from BerriAI/litellm_mistral_ocr_batches
...
feat(batches): support Mistral files/batches and per-page OCR batch cost tracking (internal copy of #40484 )
2026-09-19 03:51:27 -07:00
mateo-berri
217ff78ae7
fix(mistral): read back files whose purpose Mistral never lets us upload as user_data
2026-09-19 03:04:45 -07:00
Mateo Wang
5fc510a6fd
Merge pull request #41938 from BerriAI/litellm_gemini_cache_control_messages
...
fix(anthropic): keep cache_control for Gemini targets on /v1/messages and normalize Anthropic ttl units
2026-09-19 01:20:48 -07:00
mateo-berri
5ad1847835
fix(gemini): drop ttl values whose expiry Google cannot store (past the year 9999)
2026-09-19 01:00:27 -07:00
mateo-berri
3684e5cbcb
fix(gemini): drop ttl values outside the protobuf Duration range and the explanatory docstrings
2026-09-19 00:50:37 -07:00
mateo-berri
2303379c20
fix(gemini): keep the 2048 cache minimum on Gemini 2.5 Pro only, per Google's live cachedContents API
2026-09-19 00:07:22 -07:00
mateo-berri
3289e22834
fix(anthropic): keep cache_control for Gemini targets on /v1/messages and normalize Anthropic ttl units
2026-09-19 00:01:42 -07:00
mateo-berri
f89ca64481
fix(batches): honor deployment OCR page pricing in batch cost and answer 400 for unsupported Mistral file purposes
2026-09-18 23:57:41 -07:00
mateo-berri
6019e451ce
Merge branch 'main' into feat/gemini-cache-control-pass-through
2026-09-18 23:42:58 -07:00
mateo-berri
cb80e8773e
fix(policy_engine): fail open when an unended Messages stream has no text delta to carry the rewrite
2026-09-18 23:24:49 -07:00
mateo-berri
b3d9ba9e7b
fix(policy_engine): deliver guardrail text rewrites on multi-choice, unfinished, and envelope-less streams
...
Post-call pipeline rewrites on buffered streams failed open on three shapes:
chat streams with n > 1 (the rebuilt response collapsed every choice into
index 0), streams that ended without a finish marker, and Responses streams
whose final event carried no response envelope.
The chat handler now rebuilds the ended stream one choice index at a time and
writes each choice's rewrite back to that choice's buffered deltas. The
Anthropic handler writes an unended stream's rewrite across its text deltas.
The Responses handler spreads an envelope-less rewrite over the buffered
output_text events, still failing open when a scanned event cannot be placed.
Tool-call rewrites on n > 1 chat streams keep failing open.
2026-09-18 22:59:52 -07:00
mateo-berri
f3b198c1b7
fix(mistral): accept user_data as the OCR file purpose and keep OCR cost warnings single-line
2026-09-18 21:56:43 -07:00
mateo-berri
fb76b67e78
Merge remote-tracking branch 'origin/main' into litellm_mistral_ocr_batches
...
# Conflicts:
# litellm/batches/batch_utils.py
2026-09-18 21:51:33 -07:00
Mateo Wang
faed57f92c
Merge pull request #41918 from BerriAI/litellm_websearch_followup_api_base
...
fix(websearch): forward the deployment api_base to agentic follow-up calls on /v1/messages
2026-09-18 21:41:41 -07:00
Mateo Wang
9486caf584
Merge pull request #41721 from BerriAI/litellm_vertex_chirp3_streaming_stt
...
feat(vertex_ai): stream Chirp speech-to-text over /v1/realtime
2026-09-18 20:34:56 -07:00
kerry-berri
f26afabe6d
Merge pull request #41914 from BerriAI/litellm_xai_audio_transcription
...
feat(xai): add speech-to-text (Grok Voice Transcribe) via /v1/audio/transcriptions
2026-09-18 20:03:00 -07:00
mubashir1osmani
8cf2606e2d
fix(batches): mask pre-signed request auth headers before raw-request logging
...
A pre-signed batch/file request (Mistral, Bedrock) carries its auth header
inside the transformed request body, which pre_call logs verbatim into
raw_request_typed_dict and raw-request callbacks, leaking the provider key.
Mask the nested headers channel before handing the request to pre_call.
Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
2026-09-18 22:50:19 -04:00
kerry-berri
c4ddea7cfc
Merge pull request #41917 from BerriAI/litellm_fireworks_cache_read_default_discount
...
fix(cost_calc): default fireworks cached input to the documented 50% discount when the map has no cache-read rate
2026-09-18 18:48:58 -07:00
Mateo Wang
29bdd1ab47
Merge pull request #41905 from BerriAI/litellm_websearch_failed_search_error_block
...
fix(websearch_interception): surface a failed search as a web_search_tool_result_error block and end the turn
2026-09-18 18:34:11 -07:00
kerry
05cefb1480
fix(cost_calc): coerce string fireworks rates and drop the match fall-through
...
Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>
2026-09-19 01:33:41 +00:00
kerry-berri
fc0b37ff5d
Merge pull request #41904 from BerriAI/litellm_bedrock_batch_retrieve_sigv4_over_env_bearer
...
fix(bedrock): sign batch retrieve and cancel with deployment credentials when AWS_BEARER_TOKEN_BEDROCK is set
2026-09-18 18:33:10 -07:00
kerry
99d91d7205
fix(xai): reject non-success stt responses before parsing
...
Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>
2026-09-19 01:22:25 +00:00
mateo-berri
b1b7af884a
fix(websearch): forward the deployment api_base to agentic follow-up calls on /v1/messages
2026-09-18 18:19:48 -07:00
kerry
c9cd666b36
refactor(cost_calc): move the fireworks cache-read default under litellm/llms
...
Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>
2026-09-19 01:18:09 +00:00
mateo-berri
17c519c40a
test(custom_httpx): pass the token resolver and two-argument client factory in the realtime bridge test
2026-09-18 18:09:33 -07:00
kerry
1b305cd6b9
fix(cost_calc): default fireworks cached input to the documented 50% discount when the map has no cache-read rate
...
Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>
2026-09-19 01:07:07 +00:00
kerry
80b0ea6a2f
test(xai): narrow raises match for missing api key
...
Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>
2026-09-19 01:00:53 +00:00
kerry
6f54ad5166
fix(xai): parse integer speaker ids and simplify stt form build
...
Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>
2026-09-19 00:53:02 +00:00
kerry
6b0ad3bed3
feat(xai): add speech-to-text via /v1/audio/transcriptions
...
Route xai audio transcription through a provider config hitting POST
https://api.x.ai/v1/stt instead of the openai-compatible chat handler
which targets /audio/transcriptions. Supports language, diarize,
keyterm, filler_words and other provider fields as passthrough kwargs
Resolves LIT-8153
Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>
2026-09-19 00:49:34 +00:00
mateo-berri
3911d62bbe
fix(vertex_ai): prune a discarded turn's id once its marker is delivered
2026-09-18 17:44:28 -07:00
Mateo Wang
a6e3a72ed8
Merge pull request #41870 from BerriAI/litellm_bedrock_openai_gpt_min_max_tokens
...
fix(bedrock): clamp maxTokens to the 16-token minimum for OpenAI GPT and xAI Grok models on Converse
2026-09-18 17:44:18 -07:00
kerry
99659e9e7e
test(bedrock): drop redundant recorder docstring
...
Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>
2026-09-19 00:43:51 +00:00
mateo-berri
8b1f78fa08
fix(vertex_ai): drop a cleared turn's queued transcripts and carry its billed seconds
2026-09-18 17:30:57 -07:00
kerry
6b082d3a01
test(bedrock): type the SigV4 request recorder and drop caller-owned mutation
...
Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>
2026-09-19 00:25:01 +00:00
mateo-berri
cdc0e57e93
fix(websearch_interception): surface a failed search as a web_search_tool_result_error block and end the turn
2026-09-18 17:24:52 -07:00
mateo-berri
12120fe59b
refactor(bedrock): inline maxTokens clamp and cover inference-profile ARNs in tests
2026-09-18 17:11:29 -07:00
kerry
37da5b6f4d
fix(bedrock): sign batch retrieve and cancel with deployment credentials when AWS_BEARER_TOKEN_BEDROCK is set
...
Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>
2026-09-19 00:10:17 +00:00
mateo-berri
dc11c34e3a
Merge remote-tracking branch 'origin/main' into litellm_vertex_chirp3_streaming_stt
2026-09-18 16:56:02 -07:00
mateo-berri
3d7a771ea7
fix(vertex_ai): apply finals before interims and refresh the token per stream
2026-09-18 16:53:06 -07:00
Mateo Wang
ff7dc86947
Merge pull request #41892 from BerriAI/litellm_gemini_contentless_candidate_finish_reason
...
Unit Tests: Proxy DB Operations / custom-logging (push) Blocked by required conditions
Unit Tests: Proxy DB Operations / db-and-spend (push) Blocked by required conditions
Unit Tests: Proxy DB Operations / endpoints-and-responses (push) Blocked by required conditions
Unit Tests: Proxy DB Operations / proxy-server-core (push) Blocked by required conditions
Unit Tests: Proxy DB Operations / guardrails-hooks (push) Blocked by required conditions
Unit Tests / misc (push) Waiting to run
Unit Tests: Proxy DB Operations / jwt-and-keys (push) Blocked by required conditions
Unit Tests: Proxy DB Operations / key-generation (push) Blocked by required conditions
Unit Tests: Proxy DB Operations / logging-misc (push) Blocked by required conditions
Unit Tests: Proxy DB Operations / proxy-runtime (push) Blocked by required conditions
Unit Tests: Proxy DB Operations / proxy-utils (push) Blocked by required conditions
Unit Tests / caching-local (push) Waiting to run
Unit Tests / core-utils (push) Waiting to run
Unit Tests / enterprise-package (push) Waiting to run
Unit Tests / enterprise-routing (push) Waiting to run
Unit Tests / integrations (push) Waiting to run
Unit Tests / All Other Providers (push) Waiting to run
Unit Tests / Vertex AI (push) Waiting to run
Unit Tests / proxy-auth (push) Waiting to run
Unit Tests / proxy-endpoints (push) Waiting to run
Unit Tests / proxy-extras (push) Waiting to run
Unit Tests / proxy-infra (push) Waiting to run
Unit Tests / proxy-server (push) Waiting to run
Unit Tests / responses-caching-types (push) Waiting to run
GitHub Actions Security Analysis / zizmor (push) Waiting to run
Issue label sync / sync-issue-labels-tests (push) Has been cancelled
Issue label sync / sync-issue-labels (push) Has been cancelled
Terraform Modules / fmt, validate, test (aws) (push) Has been cancelled
Terraform Modules / fmt, validate, test (gcp) (push) Has been cancelled
VS Code Extension / vscode-extension (push) Has been cancelled
fix(gemini): preserve candidates with finishReason and no content (#40477 )
2026-09-18 16:34:46 -07:00
Mateo Wang
ec05cd0128
Merge pull request #41887 from BerriAI/litellm_gemma_4_26b_maas_context_window
...
fix: set vertex gemma-4-26b-a4b-it-maas context window to 262144
2026-09-18 16:14:56 -07:00
Mateo Wang
3490754e65
Merge pull request #41871 from BerriAI/litellm_bedrock_eager_input_streaming
...
feat: honor eager_input_streaming on Bedrock and Anthropic Claude tools
2026-09-18 16:05:10 -07:00
mateo-berri
a4624b6c6c
fix(gemini): derive the finish reason key set from the Candidates type
...
Candidates.finishReason listed eleven values while the mapping key set
carried twenty-one, so typed fixtures could not spell the reasons this
PR handles. GeminiFinishReason is now the one list, the key set derives
from it, and a test checks every documented reason has an explicit
mapping instead of falling through to "stop"
2026-09-18 16:03:45 -07:00
mateo-berri
44034c1d5e
test: source the gemma context window limits and isolate the cost map cache
2026-09-18 15:43:07 -07:00
mateo-berri
f8f165d416
Merge remote-tracking branch 'origin/main' into litellm_vertex_chirp3_streaming_stt
...
# Conflicts:
# uv.lock
2026-09-18 15:38:33 -07:00