Mateo Wang
58065d46fd
Merge pull request #42071 from BerriAI/litellm_remove_dead_telemetry_flag
2026-09-19 21:48:02 -07:00
mateo-berri
19d77e2442
test(guardrails): type the recorder hook's request_data as a Mapping
2026-09-19 20:21:37 -07:00
mateo-berri
2e83871d54
test(guardrails): type the recorder hook's logging_obj as object
2026-09-19 18:52:28 -07:00
mateo
f820472488
chore: remove the dead telemetry flag from the SDK, proxy CLI and configs
...
Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>
2026-09-20 01:44:10 +00:00
mateo-berri
2a35dc5217
fix(guardrails): keep the undeliverable rewrite reason through copies and name the responses mismatch
2026-09-19 18:00:03 -07:00
mateo-berri
114fc16554
merge: origin/main into litellm_lit_7346_multi_choice_stream_guardrails
2026-09-19 17:34:43 -07:00
tin-berri
1944d40097
Merge pull request #41177 from BerriAI/litellm_autorouter_baseline_cache
...
fix(proxy): estimate auto-router baseline costs from durable cache history
2026-09-19 16:16:29 -07:00
Mateo Wang
56408b3915
Merge pull request #42031 from BerriAI/litellm_internal_copy_41781
...
fix(azure): drop tool_choice when the request has no tools (internal copy of #41781 )
2026-09-19 16:03:32 -07:00
Mateo Wang
7651d6b550
Merge pull request #41942 from BerriAI/litellm_vllm_batch_runner
...
feat(batches): run hosted_vllm batches inside LiteLLM
2026-09-19 15:44:26 -07:00
mateo-berri
0ae5d7c2fa
Merge remote-tracking branch 'origin/main' into litellm_pr41781_azure_tool_choice
...
# Conflicts:
# tests/test_litellm/llms/azure/chat/test_azure_chat_gpt_transformation.py
2026-09-19 14:15:19 -07:00
Tin Chi Lo
ed40241d26
fix(proxy): estimate auto-router baseline costs from durable cache history
2026-09-19 12:44:47 -07:00
mateo-berri
9f0eb5082a
fix(batches): authorize executed upload targets before the files api probe
...
A batch upload naming a model on a LiteLLM-executed provider now checks
that the key may call that model before the upstream server is probed for
a Files API, matching the order batch create already uses. Only targets on
an executed provider are checked here, so provider-model uploads keep
their existing behavior.
File content reads and writes move out of the storage backend into
ManagedFileContentRepository, so the backend no longer queries Prisma
directly.
2026-09-19 12:26:11 -07:00
yucheng-berri
4301ac4940
Merge pull request #41986 from BerriAI/litellm_revert_post_call_guardrail_context
...
revert(guardrails): drop the scoped request conversation and tools from post-call scans (#41220 )
2026-09-19 12:06:30 -07:00
mateo-berri
1aca37e513
Merge remote-tracking branch 'origin/main' into litellm_vllm_batch_runner
...
# Conflicts:
# tests/test_litellm/proxy/openai_files_endpoint/test_files_common_utils.py
2026-09-19 11:51:54 -07:00
yucheng
537cdaf487
Revert "Merge pull request #41220 from BerriAI/litellm_post_call_guardrail_context"
...
This reverts commit e40b90bbfa , reversing
changes made to d8d5437f55 .
2026-09-19 17:45:15 +00:00
Joshua Valluru
fb56a14cd4
chore(mcp): merge main with unit test timeout safeguards
2026-09-19 09:42:08 -07:00
Mateo Wang
8c4c394ecc
Merge pull request #41960 from BerriAI/litellm_deepseek_off_peak_pricing
...
fix(cost): bill DeepSeek V4.1 Flash and V4 Pro at off-peak rates outside peak hours
2026-09-19 06:45:46 -07:00
Mateo Wang
b8d837b2ef
Merge pull request #40147 from abhirup7/fix/azure-image-generation-entra-id-auth
...
fix(azure): send the resolved Entra ID token on image generation requests
2026-09-19 04:44:07 -07:00
Mateo Wang
dad8c32d23
Merge pull request #41940 from BerriAI/litellm_rag_ingest_registry_store
...
fix(rag): resolve registry stores on /v1/rag/ingest and reject providers without ingestion
2026-09-19 04:42:54 -07:00
mateo-berri
e0db862781
fix(cost): bill DeepSeek V4.1 Flash and V4 Pro at their off-peak rates outside peak hours
...
DeepSeek charges half the listed rate outside 01:00-04:00 and 06:00-10:00 UTC
Monday to Friday, so every deepseek-flash, deepseek-v4-flash,
deepseek-v4-flash-vision-exp, and deepseek-v4-pro entry now carries an
off_peak_pricing block with those windows and the halved input, output, and
cache-hit rates. The generated cost map schema picks up the block, and the
regression tests pin the peak and off-peak cost of one call at fixed moments.
2026-09-19 04:31:55 -07:00
mateo-berri
608f8e2184
Merge remote-tracking branch 'origin/main' into litellm_vllm_batch_runner
...
# Conflicts:
# tests/test_litellm/proxy/batches_endpoints/test_endpoints.py
2026-09-19 04:23:28 -07:00
mateo-berri
e4d01d1d78
fix(s3_vectors): reject a store id with an empty bucket or index part
...
A "bucket:" or ":index" id split into an empty name, so ingestion silently
generated a fresh index and search sent the empty name to AWS. Both sides now
raise the existing format error through the shared helper.
2026-09-19 03:52:48 -07:00
Mateo Wang
db04e7909e
Merge pull request #41934 from BerriAI/litellm_mistral_ocr_batches
...
feat(batches): support Mistral files/batches and per-page OCR batch cost tracking (internal copy of #40484 )
2026-09-19 03:51:27 -07:00
mateo-berri
217ff78ae7
fix(mistral): read back files whose purpose Mistral never lets us upload as user_data
2026-09-19 03:04:45 -07:00
mateo-berri
fd45412c89
feat(batches): run hosted_vllm batches inside LiteLLM
...
vLLM serves no /v1/files or /v1/batches, so a hosted_vllm deployment can never
host a batch. Batch inputs for such a deployment now land in a LiteLLM-owned
storage backend, the batch is executed line by line through the deployment's
own chat, completion, embedding, or responses route, and the batch plus its
output and error files are served back from the database under the creating key
2026-09-19 01:47:43 -07:00
Mateo Wang
5fc510a6fd
Merge pull request #41938 from BerriAI/litellm_gemini_cache_control_messages
...
fix(anthropic): keep cache_control for Gemini targets on /v1/messages and normalize Anthropic ttl units
2026-09-19 01:20:48 -07:00
mateo-berri
5ad1847835
fix(gemini): drop ttl values whose expiry Google cannot store (past the year 9999)
2026-09-19 01:00:27 -07:00
mateo-berri
3684e5cbcb
fix(gemini): drop ttl values outside the protobuf Duration range and the explanatory docstrings
2026-09-19 00:50:37 -07:00
mateo-berri
2303379c20
fix(gemini): keep the 2048 cache minimum on Gemini 2.5 Pro only, per Google's live cachedContents API
2026-09-19 00:07:22 -07:00
mateo-berri
3289e22834
fix(anthropic): keep cache_control for Gemini targets on /v1/messages and normalize Anthropic ttl units
2026-09-19 00:01:42 -07:00
mateo-berri
f89ca64481
fix(batches): honor deployment OCR page pricing in batch cost and answer 400 for unsupported Mistral file purposes
2026-09-18 23:57:41 -07:00
mateo-berri
6019e451ce
Merge branch 'main' into feat/gemini-cache-control-pass-through
2026-09-18 23:42:58 -07:00
mateo-berri
cb80e8773e
fix(policy_engine): fail open when an unended Messages stream has no text delta to carry the rewrite
2026-09-18 23:24:49 -07:00
mateo-berri
b3d9ba9e7b
fix(policy_engine): deliver guardrail text rewrites on multi-choice, unfinished, and envelope-less streams
...
Post-call pipeline rewrites on buffered streams failed open on three shapes:
chat streams with n > 1 (the rebuilt response collapsed every choice into
index 0), streams that ended without a finish marker, and Responses streams
whose final event carried no response envelope.
The chat handler now rebuilds the ended stream one choice index at a time and
writes each choice's rewrite back to that choice's buffered deltas. The
Anthropic handler writes an unended stream's rewrite across its text deltas.
The Responses handler spreads an envelope-less rewrite over the buffered
output_text events, still failing open when a scanned event cannot be placed.
Tool-call rewrites on n > 1 chat streams keep failing open.
2026-09-18 22:59:52 -07:00
Joshua Valluru
f5ab563499
fix(mcp): preserve session expiry signals and scope dependency CI
2026-09-18 22:52:10 -07:00
mateo-berri
f3b198c1b7
fix(mistral): accept user_data as the OCR file purpose and keep OCR cost warnings single-line
2026-09-18 21:56:43 -07:00
mateo-berri
fb76b67e78
Merge remote-tracking branch 'origin/main' into litellm_mistral_ocr_batches
...
# Conflicts:
# litellm/batches/batch_utils.py
2026-09-18 21:51:33 -07:00
Mateo Wang
faed57f92c
Merge pull request #41918 from BerriAI/litellm_websearch_followup_api_base
...
fix(websearch): forward the deployment api_base to agentic follow-up calls on /v1/messages
2026-09-18 21:41:41 -07:00
Mateo Wang
9486caf584
Merge pull request #41721 from BerriAI/litellm_vertex_chirp3_streaming_stt
...
feat(vertex_ai): stream Chirp speech-to-text over /v1/realtime
2026-09-18 20:34:56 -07:00
kerry-berri
f26afabe6d
Merge pull request #41914 from BerriAI/litellm_xai_audio_transcription
...
feat(xai): add speech-to-text (Grok Voice Transcribe) via /v1/audio/transcriptions
2026-09-18 20:03:00 -07:00
mubashir1osmani
8cf2606e2d
fix(batches): mask pre-signed request auth headers before raw-request logging
...
A pre-signed batch/file request (Mistral, Bedrock) carries its auth header
inside the transformed request body, which pre_call logs verbatim into
raw_request_typed_dict and raw-request callbacks, leaking the provider key.
Mask the nested headers channel before handing the request to pre_call.
Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
2026-09-18 22:50:19 -04:00
kerry-berri
c4ddea7cfc
Merge pull request #41917 from BerriAI/litellm_fireworks_cache_read_default_discount
...
fix(cost_calc): default fireworks cached input to the documented 50% discount when the map has no cache-read rate
2026-09-18 18:48:58 -07:00
Mateo Wang
29bdd1ab47
Merge pull request #41905 from BerriAI/litellm_websearch_failed_search_error_block
...
fix(websearch_interception): surface a failed search as a web_search_tool_result_error block and end the turn
2026-09-18 18:34:11 -07:00
kerry
05cefb1480
fix(cost_calc): coerce string fireworks rates and drop the match fall-through
...
Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>
2026-09-19 01:33:41 +00:00
kerry-berri
fc0b37ff5d
Merge pull request #41904 from BerriAI/litellm_bedrock_batch_retrieve_sigv4_over_env_bearer
...
fix(bedrock): sign batch retrieve and cancel with deployment credentials when AWS_BEARER_TOKEN_BEDROCK is set
2026-09-18 18:33:10 -07:00
kerry
99d91d7205
fix(xai): reject non-success stt responses before parsing
...
Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>
2026-09-19 01:22:25 +00:00
mateo-berri
b1b7af884a
fix(websearch): forward the deployment api_base to agentic follow-up calls on /v1/messages
2026-09-18 18:19:48 -07:00
kerry
c9cd666b36
refactor(cost_calc): move the fireworks cache-read default under litellm/llms
...
Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>
2026-09-19 01:18:09 +00:00
mateo-berri
17c519c40a
test(custom_httpx): pass the token resolver and two-argument client factory in the realtime bridge test
2026-09-18 18:09:33 -07:00
kerry
1b305cd6b9
fix(cost_calc): default fireworks cached input to the documented 50% discount when the map has no cache-read rate
...
Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>
2026-09-19 01:07:07 +00:00