Commit graph

3503 commits

Author SHA1 Message Date
yuneng
d47008129c test(llms): migrate phase 7 provider unit tests to tests/unit
Move the wave 1 phase 7 batch (fireworks_ai, gemini, gigachat, github_copilot; 20 files) from tests/test_litellm to tests/unit after judging every test function under a behaviour mutation. Seven wiring or mock-echo tests that stayed green are deleted. The fireworks cost calculator tests get a local model_cost save/restore fixture since the tests/unit tree has no shared conftest for it

Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>
2026-09-20 08:08:12 +00:00
Mateo Wang
58065d46fd
Merge pull request #42071 from BerriAI/litellm_remove_dead_telemetry_flag 2026-09-19 21:48:02 -07:00
mateo-berri
19d77e2442 test(guardrails): type the recorder hook's request_data as a Mapping 2026-09-19 20:21:37 -07:00
mateo-berri
2e83871d54 test(guardrails): type the recorder hook's logging_obj as object 2026-09-19 18:52:28 -07:00
mateo
f820472488 chore: remove the dead telemetry flag from the SDK, proxy CLI and configs
Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>
2026-09-20 01:44:10 +00:00
mateo-berri
2a35dc5217 fix(guardrails): keep the undeliverable rewrite reason through copies and name the responses mismatch 2026-09-19 18:00:03 -07:00
mateo-berri
114fc16554 merge: origin/main into litellm_lit_7346_multi_choice_stream_guardrails 2026-09-19 17:34:43 -07:00
tin-berri
1944d40097
Merge pull request #41177 from BerriAI/litellm_autorouter_baseline_cache
fix(proxy): estimate auto-router baseline costs from durable cache history
2026-09-19 16:16:29 -07:00
Mateo Wang
56408b3915
Merge pull request #42031 from BerriAI/litellm_internal_copy_41781
fix(azure): drop tool_choice when the request has no tools (internal copy of #41781)
2026-09-19 16:03:32 -07:00
Mateo Wang
7651d6b550
Merge pull request #41942 from BerriAI/litellm_vllm_batch_runner
feat(batches): run hosted_vllm batches inside LiteLLM
2026-09-19 15:44:26 -07:00
mateo-berri
0ae5d7c2fa Merge remote-tracking branch 'origin/main' into litellm_pr41781_azure_tool_choice
# Conflicts:
#	tests/test_litellm/llms/azure/chat/test_azure_chat_gpt_transformation.py
2026-09-19 14:15:19 -07:00
Tin Chi Lo
ed40241d26 fix(proxy): estimate auto-router baseline costs from durable cache history 2026-09-19 12:44:47 -07:00
mateo-berri
9f0eb5082a fix(batches): authorize executed upload targets before the files api probe
A batch upload naming a model on a LiteLLM-executed provider now checks
that the key may call that model before the upstream server is probed for
a Files API, matching the order batch create already uses. Only targets on
an executed provider are checked here, so provider-model uploads keep
their existing behavior.

File content reads and writes move out of the storage backend into
ManagedFileContentRepository, so the backend no longer queries Prisma
directly.
2026-09-19 12:26:11 -07:00
yucheng-berri
4301ac4940
Merge pull request #41986 from BerriAI/litellm_revert_post_call_guardrail_context
revert(guardrails): drop the scoped request conversation and tools from post-call scans (#41220)
2026-09-19 12:06:30 -07:00
mateo-berri
1aca37e513 Merge remote-tracking branch 'origin/main' into litellm_vllm_batch_runner
# Conflicts:
#	tests/test_litellm/proxy/openai_files_endpoint/test_files_common_utils.py
2026-09-19 11:51:54 -07:00
yucheng
537cdaf487 Revert "Merge pull request #41220 from BerriAI/litellm_post_call_guardrail_context"
This reverts commit e40b90bbfa, reversing
changes made to d8d5437f55.
2026-09-19 17:45:15 +00:00
Joshua Valluru
fb56a14cd4 chore(mcp): merge main with unit test timeout safeguards 2026-09-19 09:42:08 -07:00
Mateo Wang
8c4c394ecc
Merge pull request #41960 from BerriAI/litellm_deepseek_off_peak_pricing
fix(cost): bill DeepSeek V4.1 Flash and V4 Pro at off-peak rates outside peak hours
2026-09-19 06:45:46 -07:00
Mateo Wang
b8d837b2ef
Merge pull request #40147 from abhirup7/fix/azure-image-generation-entra-id-auth
fix(azure): send the resolved Entra ID token on image generation requests
2026-09-19 04:44:07 -07:00
Mateo Wang
dad8c32d23
Merge pull request #41940 from BerriAI/litellm_rag_ingest_registry_store
fix(rag): resolve registry stores on /v1/rag/ingest and reject providers without ingestion
2026-09-19 04:42:54 -07:00
mateo-berri
e0db862781 fix(cost): bill DeepSeek V4.1 Flash and V4 Pro at their off-peak rates outside peak hours
DeepSeek charges half the listed rate outside 01:00-04:00 and 06:00-10:00 UTC
Monday to Friday, so every deepseek-flash, deepseek-v4-flash,
deepseek-v4-flash-vision-exp, and deepseek-v4-pro entry now carries an
off_peak_pricing block with those windows and the halved input, output, and
cache-hit rates. The generated cost map schema picks up the block, and the
regression tests pin the peak and off-peak cost of one call at fixed moments.
2026-09-19 04:31:55 -07:00
mateo-berri
608f8e2184 Merge remote-tracking branch 'origin/main' into litellm_vllm_batch_runner
# Conflicts:
#	tests/test_litellm/proxy/batches_endpoints/test_endpoints.py
2026-09-19 04:23:28 -07:00
mateo-berri
e4d01d1d78 fix(s3_vectors): reject a store id with an empty bucket or index part
A "bucket:" or ":index" id split into an empty name, so ingestion silently
generated a fresh index and search sent the empty name to AWS. Both sides now
raise the existing format error through the shared helper.
2026-09-19 03:52:48 -07:00
Mateo Wang
db04e7909e
Merge pull request #41934 from BerriAI/litellm_mistral_ocr_batches
feat(batches): support Mistral files/batches and per-page OCR batch cost tracking (internal copy of #40484)
2026-09-19 03:51:27 -07:00
mateo-berri
217ff78ae7 fix(mistral): read back files whose purpose Mistral never lets us upload as user_data 2026-09-19 03:04:45 -07:00
mateo-berri
fd45412c89 feat(batches): run hosted_vllm batches inside LiteLLM
vLLM serves no /v1/files or /v1/batches, so a hosted_vllm deployment can never
host a batch. Batch inputs for such a deployment now land in a LiteLLM-owned
storage backend, the batch is executed line by line through the deployment's
own chat, completion, embedding, or responses route, and the batch plus its
output and error files are served back from the database under the creating key
2026-09-19 01:47:43 -07:00
Mateo Wang
5fc510a6fd
Merge pull request #41938 from BerriAI/litellm_gemini_cache_control_messages
fix(anthropic): keep cache_control for Gemini targets on /v1/messages and normalize Anthropic ttl units
2026-09-19 01:20:48 -07:00
mateo-berri
5ad1847835 fix(gemini): drop ttl values whose expiry Google cannot store (past the year 9999) 2026-09-19 01:00:27 -07:00
mateo-berri
3684e5cbcb fix(gemini): drop ttl values outside the protobuf Duration range and the explanatory docstrings 2026-09-19 00:50:37 -07:00
mateo-berri
2303379c20 fix(gemini): keep the 2048 cache minimum on Gemini 2.5 Pro only, per Google's live cachedContents API 2026-09-19 00:07:22 -07:00
mateo-berri
3289e22834 fix(anthropic): keep cache_control for Gemini targets on /v1/messages and normalize Anthropic ttl units 2026-09-19 00:01:42 -07:00
mateo-berri
f89ca64481 fix(batches): honor deployment OCR page pricing in batch cost and answer 400 for unsupported Mistral file purposes 2026-09-18 23:57:41 -07:00
mateo-berri
6019e451ce Merge branch 'main' into feat/gemini-cache-control-pass-through 2026-09-18 23:42:58 -07:00
mateo-berri
cb80e8773e fix(policy_engine): fail open when an unended Messages stream has no text delta to carry the rewrite 2026-09-18 23:24:49 -07:00
mateo-berri
b3d9ba9e7b fix(policy_engine): deliver guardrail text rewrites on multi-choice, unfinished, and envelope-less streams
Post-call pipeline rewrites on buffered streams failed open on three shapes:
chat streams with n > 1 (the rebuilt response collapsed every choice into
index 0), streams that ended without a finish marker, and Responses streams
whose final event carried no response envelope.

The chat handler now rebuilds the ended stream one choice index at a time and
writes each choice's rewrite back to that choice's buffered deltas. The
Anthropic handler writes an unended stream's rewrite across its text deltas.
The Responses handler spreads an envelope-less rewrite over the buffered
output_text events, still failing open when a scanned event cannot be placed.

Tool-call rewrites on n > 1 chat streams keep failing open.
2026-09-18 22:59:52 -07:00
Joshua Valluru
f5ab563499 fix(mcp): preserve session expiry signals and scope dependency CI 2026-09-18 22:52:10 -07:00
mateo-berri
f3b198c1b7 fix(mistral): accept user_data as the OCR file purpose and keep OCR cost warnings single-line 2026-09-18 21:56:43 -07:00
mateo-berri
fb76b67e78 Merge remote-tracking branch 'origin/main' into litellm_mistral_ocr_batches
# Conflicts:
#	litellm/batches/batch_utils.py
2026-09-18 21:51:33 -07:00
Mateo Wang
faed57f92c
Merge pull request #41918 from BerriAI/litellm_websearch_followup_api_base
fix(websearch): forward the deployment api_base to agentic follow-up calls on /v1/messages
2026-09-18 21:41:41 -07:00
Mateo Wang
9486caf584
Merge pull request #41721 from BerriAI/litellm_vertex_chirp3_streaming_stt
feat(vertex_ai): stream Chirp speech-to-text over /v1/realtime
2026-09-18 20:34:56 -07:00
kerry-berri
f26afabe6d
Merge pull request #41914 from BerriAI/litellm_xai_audio_transcription
feat(xai): add speech-to-text (Grok Voice Transcribe) via /v1/audio/transcriptions
2026-09-18 20:03:00 -07:00
mubashir1osmani
8cf2606e2d fix(batches): mask pre-signed request auth headers before raw-request logging
A pre-signed batch/file request (Mistral, Bedrock) carries its auth header
inside the transformed request body, which pre_call logs verbatim into
raw_request_typed_dict and raw-request callbacks, leaking the provider key.
Mask the nested headers channel before handing the request to pre_call.

Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
2026-09-18 22:50:19 -04:00
kerry-berri
c4ddea7cfc
Merge pull request #41917 from BerriAI/litellm_fireworks_cache_read_default_discount
fix(cost_calc): default fireworks cached input to the documented 50% discount when the map has no cache-read rate
2026-09-18 18:48:58 -07:00
Mateo Wang
29bdd1ab47
Merge pull request #41905 from BerriAI/litellm_websearch_failed_search_error_block
fix(websearch_interception): surface a failed search as a web_search_tool_result_error block and end the turn
2026-09-18 18:34:11 -07:00
kerry
05cefb1480 fix(cost_calc): coerce string fireworks rates and drop the match fall-through
Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>
2026-09-19 01:33:41 +00:00
kerry-berri
fc0b37ff5d
Merge pull request #41904 from BerriAI/litellm_bedrock_batch_retrieve_sigv4_over_env_bearer
fix(bedrock): sign batch retrieve and cancel with deployment credentials when AWS_BEARER_TOKEN_BEDROCK is set
2026-09-18 18:33:10 -07:00
kerry
99d91d7205 fix(xai): reject non-success stt responses before parsing
Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>
2026-09-19 01:22:25 +00:00
mateo-berri
b1b7af884a fix(websearch): forward the deployment api_base to agentic follow-up calls on /v1/messages 2026-09-18 18:19:48 -07:00
kerry
c9cd666b36 refactor(cost_calc): move the fireworks cache-read default under litellm/llms
Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>
2026-09-19 01:18:09 +00:00
mateo-berri
17c519c40a test(custom_httpx): pass the token resolver and two-argument client factory in the realtime bridge test 2026-09-18 18:09:33 -07:00