Commit graph

18336 commits

Author SHA1 Message Date
Mateo Wang
1f6e5b60b5
Merge pull request #41952 from BerriAI/litellm_masker_memo_depth_fail_closed
fix(masker): memoize shared nodes and fail closed past the depth cap
2026-09-19 04:20:17 -07:00
Mateo Wang
43e835c3ca
Merge pull request #41950 from BerriAI/litellm_cost_callback_bounded_error_msg
fix(proxy): keep request metadata out of the cost tracking failure alert
2026-09-19 04:08:14 -07:00
mateo-berri
aceae8e566 test: drop the recursive detector allowlist entry for the removed _walk_payload 2026-09-19 04:07:07 -07:00
Mateo Wang
c39ec34553
Merge pull request #41234 from BerriAI/litellm_invalid_tool_choice_400
fix(utils): reject an untranslatable tool_choice with a 400 instead of a 500
2026-09-19 04:02:04 -07:00
mateo-berri
093fb78baf fix(masker): cut cycles at the first back-edge and walk pydantic dumps without self-recursion 2026-09-19 03:58:47 -07:00
Mateo Wang
db04e7909e
Merge pull request #41934 from BerriAI/litellm_mistral_ocr_batches
feat(batches): support Mistral files/batches and per-page OCR batch cost tracking (internal copy of #40484)
2026-09-19 03:51:27 -07:00
Mateo Wang
aa3e6df086
Merge pull request #41926 from BerriAI/litellm_fix_41515
fix(proxy): register transcribe as a known provider for model grants
2026-09-19 03:47:36 -07:00
Mateo Wang
2815d80fa4
Merge pull request #41948 from BerriAI/litellm_unit_shard_per_test_timeout
ci(unit): fail a hung test in 120s with a traceback instead of idling the shard to its step timeout
2026-09-19 03:36:37 -07:00
mateo-berri
7edafd1715 fix(masker): memoize shared nodes and fail closed past the depth cap 2026-09-19 03:28:50 -07:00
Mateo Wang
825e287f63
Merge pull request #41237 from BerriAI/litellm_internal_copy_36887
fix(cost): carry image and video input tokens through the Responses usage bridge (internal copy of #36887)
2026-09-19 03:27:42 -07:00
mateo-berri
9e8a847c5b fix(proxy): keep request metadata out of the cost tracking failure alert
The cost tracking callback f-stringed chosen_metadata, litellm_metadata,
and old_metadata into the failed_tracking_spend alert on every failure,
at every log level, so one 250-byte request produced a 23 KB alert
carrying the client's metadata, headers, and key-auth reprs four times
over. The alert now carries the exception, the traceback, the model, and
the call type; the metadata keys are logged once at debug level through
lazy formatting, so nothing is built at warning level
2026-09-19 03:21:28 -07:00
mateo-berri
4968e89f3c test(ci): drop the structure-only assertion on the shard script; the parametrized hang test covers both invocations 2026-09-19 03:20:48 -07:00
Mateo Wang
6b8db68fb8
Merge pull request #39102 from gaurav-pandey-zocdoc/litellm_budget_alert_wording
fix(alerting): clarify budget threshold messages
2026-09-19 03:15:16 -07:00
Mateo Wang
385932b4e3
Merge pull request #41941 from BerriAI/litellm_azure_content_safety_api_version_default
fix(guardrails): stop the Javelin api_version default leaking into Azure Content Safety
2026-09-19 03:09:38 -07:00
mateo-berri
217ff78ae7 fix(mistral): read back files whose purpose Mistral never lets us upload as user_data 2026-09-19 03:04:45 -07:00
mateo-berri
b8c2787bcf Merge remote-tracking branch 'origin/main' into litellm_internal_copy_36887 2026-09-19 02:54:13 -07:00
mateo-berri
343e1eeac8 ci(unit): fail a hung test in 120s with a traceback instead of idling the shard to its step timeout 2026-09-19 02:53:36 -07:00
mateo-berri
e327a6ae76 test(guardrails): drop regression docstrings from the api_version tests 2026-09-19 02:52:42 -07:00
mateo-berri
988bb65aa2 test: require a provider family's rows to reach its wildcard list 2026-09-19 02:50:31 -07:00
mateo-berri
2c2aa5df11 Merge remote-tracking branch 'origin/main' into litellm_invalid_tool_choice_400
# Conflicts:
#	tests/test_litellm/responses/litellm_completion_transformation/test_litellm_completion_responses.py
#	tests/test_litellm/test_main.py
2026-09-19 02:35:29 -07:00
mateo-berri
6e0356a7a0 test: fail a required shard when a cost map provider is unregistered 2026-09-19 02:34:27 -07:00
mateo-berri
24064e3b31 fix(guardrails): treat the stored Javelin api_version default as unset for Azure Content Safety
Guardrails created through POST /guardrails on older releases have api_version "v1" saved in the database, because the writer persists every default. Azure Content Safety never accepts that value, so those guardrails kept answering 404 after the default moved to None. The Azure base now resolves "v1" to 2024-09-01 the same way it resolves a missing value. Also restores the OpenAPI snapshot line that a Python 3.14 regeneration had dedented
2026-09-19 02:29:43 -07:00
mateo-berri
e0ebcb79fc Merge remote-tracking branch 'origin/main' into litellm_budget_alert_wording 2026-09-19 02:12:42 -07:00
mateo-berri
dd79c1f77d fix(cost): keep the deployment's OCR page rate when the model has no published price
When a deployment priced one OCR batch family and the other still needed a
published rate, a failed cost-map lookup returned zero for the whole line and
discarded the deployment rate that was already resolved. Those pages were
billed as free. The lookup failure now only logs, and the families the
deployment prices are billed at the configured rate
2026-09-19 01:59:39 -07:00
mateo-berri
6f4d1c5911 fix(guardrails): stop the Javelin api_version default leaking into Azure Content Safety
LitellmParams mixes every provider config model into one class, so the
Javelin api_version default of "v1" reached the Azure Content Safety
guardrails whenever config.yaml omitted api_version and Azure answered 404.
The shared field now defaults to None, Javelin keeps filling in "v1" itself,
and the Azure guardrails fall back to the documented 2024-09-01 at request
time so a DB update that omits api_version stays on the default too.
2026-09-19 01:41:35 -07:00
Mateo Wang
5fc510a6fd
Merge pull request #41938 from BerriAI/litellm_gemini_cache_control_messages
fix(anthropic): keep cache_control for Gemini targets on /v1/messages and normalize Anthropic ttl units
2026-09-19 01:20:48 -07:00
mateo-berri
00214ac371 test(router): cover the team-scoped credential deployment lookups
The router coverage gate in code-quality flags every router.py function
no router test calls by name, and the two helpers get_credential_deployment
gained (the team public-name lookup and the team-aware wildcard lookup)
were only reached through it. Each now has a test of its own: the
public-name lookup resolves only for the owning team, and the wildcard
lookup prefers the team's own pattern over the shared one and never hands
another team's wildcard deployment to a caller outside that team.
2026-09-19 01:05:01 -07:00
mateo-berri
5ad1847835 fix(gemini): drop ttl values whose expiry Google cannot store (past the year 9999) 2026-09-19 01:00:27 -07:00
mateo-berri
7133baa777 fix(vector_stores): run the full model grant check on caller-supplied model hints
The vector-store file routes accept a model hint through the ?model= query
param and the x-litellm-model header. That hint was authorized with a hand
rolled check that covered only the key allowlist and the team allowlist, so
a key restricted by its project's model grant, a team-member restriction, or a
key config still routed through the hinted deployment. Greptile flagged the
gap as a P1 on the replacement PR.

The hint now goes through the same authorize_model_for_key path the batches
and files routes use, which runs can_key_call_resolved_model with every rule
the proxy enforces elsewhere. Keys those extra rules deny now get a 403 on
these routes. The two remaining behavioral differences are edge cases the
old check tolerated: a key whose team_models is set without a team_id no
longer runs the team allowlist, and a key with a config set skips the key
allowlist, both matching the rest of the proxy.

The regression test caches a project whose grant excludes the hinted model
and asserts the request is refused before any deployment lookup. The two
patch() calls on litellm.proxy.proxy_server carry a test-quality-ok reason
because can_key_call_resolved_model reads prisma_client and
user_api_key_cache through a lazy module import with no injection seam.
2026-09-19 00:54:41 -07:00
mateo-berri
3684e5cbcb fix(gemini): drop ttl values outside the protobuf Duration range and the explanatory docstrings 2026-09-19 00:50:37 -07:00
mateo-berri
0feca8641f fix(batches): price model-encoded batch retrievals by their deployment
A batch retrieved by its model-encoded id takes the direct (non-router) path,
which resolved credentials without stamping the deployment's model_info, so a
completed batch on a deployment with its own per-page pricing was billed at
the published rate with an empty model_id on the spend row.

Extract the router's credential lookup into get_credential_deployment and
stamp the resolved deployment's model_info onto the retrieve call the way the
router does for routed calls.
2026-09-19 00:30:15 -07:00
Mateo Wang
5f1268c056
Merge pull request #41933 from BerriAI/litellm_deliver_multi_choice_stream_rewrites
fix(policy_engine): deliver guardrail text rewrites on multi-choice, unfinished, and envelope-less streams
2026-09-19 00:22:02 -07:00
mateo-berri
2303379c20 fix(gemini): keep the 2048 cache minimum on Gemini 2.5 Pro only, per Google's live cachedContents API 2026-09-19 00:07:22 -07:00
mateo-berri
3289e22834 fix(anthropic): keep cache_control for Gemini targets on /v1/messages and normalize Anthropic ttl units 2026-09-19 00:01:42 -07:00
mateo-berri
f89ca64481 fix(batches): honor deployment OCR page pricing in batch cost and answer 400 for unsupported Mistral file purposes 2026-09-18 23:57:41 -07:00
yucheng-berri
8afbc51cb7
Merge pull request #41685 from BerriAI/litellm_prompt_injection_llm_api_check_dispatch
* fix(proxy): dispatch llm_api_check moderation through during_call_hook

ProxyLogging.during_call_hook only ran async_moderation_hook for CustomGuardrail callbacks, so a
CustomLogger such as the prompt injection detector with llm_api_check enabled never called the
configured moderation model. Dispatch any CustomLogger that overrides async_moderation_hook and hand
the proxy router to every registered prompt injection detector at startup so that call can route

Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>

* fix(enterprise): resolve openai_moderations model at call time and default to omni-moderation-latest

Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>

* fix(proxy): keep queued moderation running past a V1 pre_call guardrail

A V1 CustomGuardrail with moderation_check pre_call returned out of
during_call_hook before asyncio.gather, abandoning already-queued
CustomLogger moderation coroutines and skipping every later callback.
Skip only that guardrail instead.

Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>

* fix(utils): skip null tool_calls when formatting prompts for moderation hooks

Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>

---------

Co-authored-by: yucheng <yucheng@berri.ai>
Co-authored-by: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>
2026-09-18 23:53:14 -07:00
mateo-berri
4fe1549432 fix(policy_engine): keep the per-choice rebuilt response's choices a list so legacy hook rewrites survive the model_dump round-trip 2026-09-18 23:53:02 -07:00
mateo-berri
6019e451ce Merge branch 'main' into feat/gemini-cache-control-pass-through 2026-09-18 23:42:58 -07:00
mateo-berri
cb80e8773e fix(policy_engine): fail open when an unended Messages stream has no text delta to carry the rewrite 2026-09-18 23:24:49 -07:00
Mateo Wang
56116079c8
Merge pull request #41930 from BerriAI/litellm_mat602_upstream_500_error_type
fix(exceptions): keep internal_server_error as the public type of an upstream 500
2026-09-18 23:19:52 -07:00
Mateo Wang
0a792c0f6b
Merge pull request #40399 from adssoccer1/feat/websearch-multi-query-schema
feat(websearch): let the model emit objective + multi-query search shapes
2026-09-18 23:03:48 -07:00
mateo-berri
b3d9ba9e7b fix(policy_engine): deliver guardrail text rewrites on multi-choice, unfinished, and envelope-less streams
Post-call pipeline rewrites on buffered streams failed open on three shapes:
chat streams with n > 1 (the rebuilt response collapsed every choice into
index 0), streams that ended without a finish marker, and Responses streams
whose final event carried no response envelope.

The chat handler now rebuilds the ended stream one choice index at a time and
writes each choice's rewrite back to that choice's buffered deltas. The
Anthropic handler writes an unended stream's rewrite across its text deltas.
The Responses handler spreads an envelope-less rewrite over the buffered
output_text events, still failing open when a scanned event cannot be placed.

Tool-call rewrites on n > 1 chat streams keep failing open.
2026-09-18 22:59:52 -07:00
Mateo Wang
4eb13a0b2a
Merge pull request #41843 from BerriAI/litellm_lit8064_unpin_derived_pricing
fix(proxy): unpin cost-map pricing copied into model_info and report pricing overrides
2026-09-18 22:58:04 -07:00
mateo-berri
5f6ffdc333 test: drop the docstring that restated the payload test's name 2026-09-18 22:48:42 -07:00
Mateo Wang
b7f07469bc
Merge pull request #41564 from BerriAI/litellm_responses_bridge_message_item_lit4622
fix(responses): announce message item before text events in the chat completions bridge
2026-09-18 22:36:28 -07:00
mateo-berri
5783a38e27 fix(proxy): enforce the unified batch model grant before the DB shortcut and skip it for registry-routed vector store models
retrieve_batch returned a terminal batch from the DB before checking that the key may use the model encoded in a unified batch id; the grant check now runs right after pre-call processing. The vector store file list helper authorized data["model"] through handle_model_based_routing even when the vector store registry set it server-side and even with no caller, which crashed on a None key; it now authorizes only a caller-supplied hint and resolves credentials directly.
2026-09-18 22:30:07 -07:00
mateo-berri
ddac683ec6 fix(exceptions): keep internal_server_error as the public type of an upstream 500
PR #40243 started carrying the upstream error body on InternalServerError so the Responses response.failed event can report the provider's code and message, and openai's APIError.__init__ took the body's type along with it. The proxy then answered an OpenAI-compatible upstream 500 with type server_error while a 502 and a 503 kept internal_server_error, and the integration contract in test_observed_routing.py went red. Pin the type the way RateLimitError pins throttling_error, keeping the body.
2026-09-18 22:28:58 -07:00
mateo-berri
12f831e863 Merge origin/main into feat/websearch-multi-query-schema
Resolves handler.py against main's SearchOutcome refactor: the file is
main's version plus this PR's substantive hunks only (the RichWebSearchInput
import, the rich= wiring at the three _execute_search call sites, the
_rich_search_input and _provider_supports_rich_search helpers, and the
_execute_search forwarding), so the 88-column re-wrap noise the PR carried
is gone and the diff against main is the feature alone. RichWebSearchInput
sits beside main's new SearchSucceeded/SearchFailed types, and main's two
_execute_search test stubs accept the new rich argument.
2026-09-18 22:15:35 -07:00
yuneng-jiang
12ddb35aad
Merge pull request #41924 from BerriAI/litellm_role_permissions_normalization
fix(proxy): parse role_permissions where it is read
2026-09-18 22:15:20 -07:00
mateo-berri
f3b198c1b7 fix(mistral): accept user_data as the OCR file purpose and keep OCR cost warnings single-line 2026-09-18 21:56:43 -07:00