Commit graph

34141 commits

Author SHA1 Message Date
Cesar Garcia
d384f7c320
Merge pull request #21233 from Chesars/feat/per-request-json-schema-validation
feat: support per-request enable_json_schema_validation for thread safety
2026-03-03 15:29:54 -03:00
Cesar Garcia
a8b5a876bf
Merge pull request #21491 from Chesars/fix/20998-remove-hardcoded-reasoning-summary
fix(anthropic): remove hardcoded reasoning summary in adapter
2026-03-03 15:29:11 -03:00
Chesars
833c1bc45d merge: resolve conflict with staging, remove hardcoded summary from reasoning test 2026-03-03 15:28:45 -03:00
Cesar Garcia
b2c7d2e049
Merge pull request #21577 from Chesars/fix/gemini-streaming-tool-calls-finish-reason
fix(gemini): correct streaming finish_reason for tool calls
2026-03-03 15:25:24 -03:00
Cesar Garcia
bca1964f70
Merge pull request #22603 from BerriAI/fix/helicone-vertex-gemini-provider-url
fix(helicone): correct provider URL for Vertex AI Gemini models
2026-03-03 15:23:51 -03:00
Chesars
4a88d85446 test: add provider_url routing test for vertex_ai/gemini models
Verifies that vertex_ai gemini models route to
aiplatform.googleapis.com instead of
generativelanguage.googleapis.com, preventing
regressions if the branch ordering changes.
2026-03-03 15:20:51 -03:00
Cesar Garcia
da941e4261
Merge pull request #22589 from Chesars/fix/vertex-preserve-any-type-schema
fix(vertex): preserve type schema semantics for JsonValuefields
2026-03-03 15:19:16 -03:00
Cesar Garcia
f77f28a5f8
Merge pull request #20920 from Chesars/refactor/files-main-credential-helpers
chore: code duplication in files/main.py using credential helpers
2026-03-03 15:18:17 -03:00
Cesar Garcia
fe8fa3abe0
Merge pull request #17308 from Chesars/fix/python-multipart-version-constraint
chore: update python-multipart constraint to >=0.0.18
2026-03-03 15:17:57 -03:00
Cesar Garcia
de415abd5a
Merge pull request #22653 from Chesars/fix/batch-encode-ids-x-litellm-model
fix(proxy): encode batch IDs when x-litellm-model header is used
2026-03-03 15:17:43 -03:00
Chesars
47f0390b9b fix: remove duplicate Pillow==11.0.0 pin (12.1.1 already on line 8) 2026-03-03 15:14:20 -03:00
Chesars
dc9f5a5cc4 fix(deps): update python-multipart to >=0.0.20 in CI and test configs 2026-03-03 15:10:39 -03:00
Chesars
dad7805b42 fix(deps): update python-multipart version to 0.0.22 in all files
Align requirements.txt, CI workflow, liccheck, and license cache
with the >=0.0.22 constraint already set in pyproject.toml.
2026-03-03 15:09:33 -03:00
Chesars
5005773909 fix(deps): relax python-multipart version constraint to >=0.0.22
The caret operator (^0.0.x) in zerover projects restricts to a single
patch version. Changed to >= to allow future patch updates.
2026-03-03 15:09:04 -03:00
Chesars
ead74dff11 refactor: move credential helpers to provider common_utils modules
Move get_openai_credentials() to litellm/llms/openai/common_utils.py
and get_azure_credentials() to litellm/llms/azure/common_utils.py
so they can be reused by batches/main.py and other modules.
Signatures now take individual params instead of GenericLiteLLMParams.
2026-03-03 15:02:24 -03:00
Chesars
fba19f089a refactor: reduce code duplication in files/main.py with credential helpers
Extract repeated OpenAI and Azure credential resolution logic into
_get_openai_credentials() and _get_azure_credentials() helper functions,
reducing ~270 lines of duplicated code across 5 file operations.
Also removes dead Vertex AI code path in create_file that was unreachable
since ProviderConfigManager.get_provider_files_config() handles it first.
2026-03-03 15:02:24 -03:00
Chesars
909e3ce6c9 test: create fresh ModelResponse per test to avoid shared mutable state 2026-03-03 14:54:11 -03:00
Chesars
1fe2e92d32 fix(main): forward enable_json_schema_validation to acompletion_with_mcp
The parameter was declared in completion() signature but not passed
to acompletion_with_mcp, causing per-request JSON schema validation
to silently fall back to the global default when MCP tools are present.
2026-03-03 14:34:32 -03:00
Chesars
59bde4a81a refactor(proxy): extract encode_batch_response_ids helper and fix list_batches encoding
Extract duplicated batch ID encoding logic into a shared helper
encode_batch_response_ids() in common_utils.py. Use it in create_batch,
retrieve_batch, and cancel_batch. Also add encoding to list_batches
when x-litellm-model is used.
2026-03-03 11:38:50 -03:00
Chesars
7506fd0426 fix(proxy): re-encode response IDs in cancel_batch for model-based routing 2026-03-03 11:22:43 -03:00
Chesars
9fd4c00b06 fix(proxy): re-encode response IDs in retrieve_batch for model-based routing
The provider returns raw IDs in the retrieve response (output_file_id,
error_file_id). These need to be encoded with model info so the client
can use them for subsequent file download calls through the proxy.
2026-03-03 11:06:38 -03:00
Chesars
9463de0c66 fix: correct indentation from commit suggestions and add missing Optional import 2026-03-03 10:51:48 -03:00
Cesar Garcia
096edface5
Update litellm/proxy/batches_endpoints/endpoints.py
Co-authored-by: greptile-apps[bot] <165735046+greptile-apps[bot]@users.noreply.github.com>
2026-03-03 10:46:14 -03:00
Cesar Garcia
7d664f0c09
Update tests/litellm/proxy/test_batch_x_litellm_model_encoding.py
Co-authored-by: greptile-apps[bot] <165735046+greptile-apps[bot]@users.noreply.github.com>
2026-03-03 10:45:00 -03:00
Cesar Garcia
3426b905ce
Update tests/litellm/proxy/test_batch_x_litellm_model_encoding.py
Co-authored-by: greptile-apps[bot] <165735046+greptile-apps[bot]@users.noreply.github.com>
2026-03-03 10:44:47 -03:00
Chesars
5ad0d03671 fix(proxy): encode batch IDs with model info when x-litellm-model header is used
When create_batch routes via x-litellm-model header, the response batch_id
was returned raw without model routing info. This meant retrieve_batch could
not determine which provider/credentials to use, defaulting to "openai"
instead of the correct provider (e.g., VLLM).

Now encodes batch_id, output_file_id, and error_file_id with model info
(same pattern as the model-embedded file_id flow in Scenario 1), so
retrieve_batch can decode and route back to the correct provider.
2026-03-03 10:26:03 -03:00
Sameer Kankute
9ffbd9e30e
Merge pull request #22464 from Point72/ephrimstanley/batch-fixes-feb27
Managed batches fixes for vertex
2026-03-03 18:53:53 +05:30
Krish Dholakia
67f90254ed
feat(guardrails): team-based guardrail registration and approval workflow (#22459)
* feat(guardrails): team-based guardrail registration and approval workflow

Add team-based guardrail submission system where teams can register
Generic Guardrail API guardrails for admin review. Includes:

- POST /guardrails/register endpoint for team-scoped submissions
- Admin review endpoints (list/get/approve/reject submissions)
- Team Guardrails tab in the UI dashboard
- extra_headers support for forwarding client headers to guardrail APIs
- Prisma schema migration for status, submitted_at, reviewed_at fields
- Documentation for team-based guardrails and static/dynamic headers

Co-Authored-By: Claude Opus 4.6 <noreply@anthropic.com>

* fix(guardrails): address review feedback - SSRF, silent failure, redundant query

- Validate api_base URL scheme (http/https only) and hostname in
  register_guardrail to prevent SSRF via team submissions
- Return warning field in approve response when in-memory initialization
  fails so admins know the guardrail won't work until next sync cycle
- Eliminate redundant DB query in list_guardrail_submissions by fetching
  all team guardrails once and deriving both filtered list and summary
  counts from the single result set

Co-Authored-By: Claude Opus 4.6 <noreply@anthropic.com>

* fix(guardrails): add pending_review status guard to reject endpoint

Prevent rejecting already-active or already-rejected guardrails, which
would create a DB/memory inconsistency (active in memory but rejected
in DB). Now mirrors the approve endpoint's status check.

Co-Authored-By: Claude Opus 4.6 <noreply@anthropic.com>

---------

Co-authored-by: Claude Opus 4.6 <noreply@anthropic.com>
2026-03-02 22:06:49 -08:00
Shivaang
213799282b
fix(openrouter): register OpenRouter as native Responses API provider (#22355)
OpenRouter supports the Responses API at /api/v1/responses with
encrypted_content for multi-turn stateless reasoning workflows.
Without native registration, requests fall through to the chat
completion bridge, which uses a different format (reasoning_details)
and drops encrypted_content entirely.

This adds OpenRouterResponsesAPIConfig to route requests directly to
OpenRouter's Responses API endpoint, preserving encrypted_content.

Fixes https://github.com/BerriAI/litellm/issues/22189

Co-authored-by: Krish Dholakia <krrishdholakia@gmail.com>
2026-03-02 22:02:59 -08:00
Jaeyeon Kim(김재연)
6bcba46dda
fix: set mock status_code in JWT OIDC discovery tests (#22361)
The _resolve_jwks_url method checks response.status_code != 200, but
MagicMock returns a MagicMock object for status_code which is always
truthy (!= 200). Explicitly set mock_response.status_code = 200 so the
tests exercise the intended code path.

Co-authored-by: Claude Opus 4.6 <noreply@anthropic.com>
2026-03-02 21:57:54 -08:00
ryan-crabbe
5b0238736c
Add incident report: cache eviction closes in-use httpx clients (#22309) 2026-03-02 21:49:48 -08:00
Ishaan Jaff
86d5b4c632
feat: add Nebius AI Studio models to model_prices_and_context_window.json (#22614)
Add 30 Nebius AI Studio models covering:
- Text-to-text: DeepSeek (R1, R1-0528, R1-Distill, V3, V3-0324), Meta Llama
  (3.1-8B/70B/405B, 3.3-70B), Qwen (3-235B/32B/30B/14B/4B, 2.5-72B/32B,
  2.5-Coder-7B, QwQ-32B), Mistral Nemo, NousResearch Hermes-3, NVIDIA
  Nemotron Ultra/Super, Google Gemma-3-27B, Llama-Guard-3
- Vision: Qwen2.5-VL-72B, Qwen2-VL-72B, Qwen2-VL-7B
- Embedding: BAAI/bge-en-icl, BAAI/bge-multilingual-gemma2, intfloat/e5-mistral-7b

Pricing sourced from https://nebius.com/prices-ai-studio (base flavor).
Context windows sourced from https://docs.nebius.com/studio/inference/models/

Co-authored-by: Cursor Agent <cursoragent@cursor.com>
Co-authored-by: Ishaan Jaff <ishaan-jaff@users.noreply.github.com>
2026-03-02 18:43:07 -08:00
Krish Dholakia
dfa2798169
Fix PR template: correct test directory path from tests/litellm/ to tests/test_litellm/ (#22612)
Co-authored-by: Cursor Agent <cursoragent@cursor.com>
2026-03-02 17:49:53 -08:00
Ishaan Jaff
bfceb7fc3f
feat(perplexity): add embedding support for pplx-embed-v1 models (#22610)
* feat: add Perplexity embedding support (pplx-embed-v1)

Add support for Perplexity AI's embedding models via the LLM HTTP handler:

Models:
- pplx-embed-v1-0.6b (1024 dims, 32K context, $0.004/1M tokens)
- pplx-embed-v1-4b (2560 dims, 32K context, $0.03/1M tokens)

Implementation:
- PerplexityEmbeddingConfig in litellm/llms/perplexity/embedding/
- Registered in ProviderConfigManager, __init__.py lazy imports, main.py dispatch
- Model pricing added to model_prices_and_context_window.json
- Supports dimensions and encoding_format parameters
- Uses base_llm_http_handler.embedding() pattern

Tests:
- 19 unit tests covering transformation, params, URLs, provider config, model info

Co-authored-by: Ishaan Jaff <ishaan-jaff@users.noreply.github.com>

* docs: add Perplexity AI embeddings documentation

- Create providers/perplexity_embedding.md with SDK and proxy usage examples
- Convert Perplexity from flat doc to category in sidebars.js
- Category includes existing chat/responses doc + new embeddings doc
- Covers pplx-embed-v1-0.6b and pplx-embed-v1-4b models
- Documents supported parameters (dimensions, encoding_format)
- Includes proxy config and curl examples

Co-authored-by: Ishaan Jaff <ishaan-jaff@users.noreply.github.com>

* fix: decode Perplexity base64_int8 embeddings to OpenAI-format float arrays

Perplexity returns embeddings as base64-encoded signed int8 values by default,
not float arrays like OpenAI. This commit adds decoding in
transform_embedding_response so the proxy returns standard OpenAI-compatible
float arrays (normalized to [-1, 1]).

- Added _decode_base64_embedding() static method
- Handles both base64 strings (decoded) and float lists (passthrough)
- Added 3 new tests for base64 decoding + passthrough

Co-authored-by: Ishaan Jaff <ishaan-jaff@users.noreply.github.com>

---------

Co-authored-by: Cursor Agent <cursoragent@cursor.com>
Co-authored-by: Ishaan Jaff <ishaan-jaff@users.noreply.github.com>
2026-03-02 17:37:50 -08:00
Kenan Yildirim
b8befb3403
Add CrowdStrike AIDR guardrail hook (#17876)
* Add CrowdStrike AIDR guardrail hook

* fixup! use apply_guardrail event hook

* fixup! update imports

* fix(guardrails): include AI response in CrowdStrike AIDR output events

Issue:
_build_guard_input_for_response() was:
- Sending only the original user input (messages).
- Not sending the AI provider response.

This fix will:
  - Extract response.choices from the ModelResponse object and include them in guard_input payload.
  - Thus, ensure AIDR output rules receive the AI-generated content for analysis.
  - Fix and update tests.

* fix(guardrails): prevent duplicate input events in CrowdStrike AIDR guardrail

Issue:
The CrowdStrike AIDR guardrail was running on during_call hooks wihtout event_hook configured.

This fix will:
- Set event_hook to ["pre_call", "post_call"] (AIDR admins will control what policy is applied)

This change will:
- Require default_on parameter
- Prevent duplicate API calls to AIDR for the same input
- Avoid unchecked AI provider API calls on during_call hook

* docs: add CrowdStrike AIDR to the list of Guardrails under Integrations

* docs: update CrowdStrike AIDR documentation page

---------

Co-authored-by: Konstantin Lapine <konstantin.lapine@crowdstrike.com>
2026-03-02 17:26:54 -08:00
Chesars
7d23106fcf fix(helicone): correct provider URL for Vertex AI Gemini models
Reorder elif branches so is_vertex_ai is checked before "gemini" in model.
Previously, Vertex AI Gemini models (e.g. vertex_ai/gemini-2.5-flash) matched
the "gemini" substring check first and were logged with the Google AI Studio
URL instead of the Vertex AI URL.
2026-03-02 19:15:37 -03:00
Cesar Garcia
2525d66dbe
Merge pull request #22584 from BerriAI/litellm_oss_staging_02_27_2026
Litellm oss staging 02 27 2026
2026-03-02 19:05:02 -03:00
Cesar Garcia
e559d4dd11
Merge pull request #22582 from BerriAI/litellm_oss_staging_02_26_2026
Litellm oss staging 02 26 2026
2026-03-02 18:51:37 -03:00
Chesars
6292c3dbdf merge: resolve conflicts with upstream/main
- anthropic.md: keep claude-opus-4-6 alias and claude-sonnet-4-6 entry
- transformation.py: take upstream's formatted effort_map with fallback
2026-03-02 18:49:24 -03:00
Cesar Garcia
835a2c3dc6
Merge pull request #22583 from Chesars/fix/add-bedrock-cache-token-pricing
fix(pricing): add missing cache token pricing for 24 Bedrock Claude models
2026-03-02 18:46:28 -03:00
Cesar Garcia
680b9ee9f2
Merge pull request #22586 from Chesars/fix/update-gemini-deprecation-dates
fix: update Gemini model deprecation dates
2026-03-02 18:45:30 -03:00
Cesar Garcia
a54a1d27d7
Merge pull request #22596 from Chesars/fix/add-dashscope-models-pricing
fix: add missing pricing for dashscope/qwen3.5-plus and dashscope/qwen3-vl-plus
2026-03-02 18:45:08 -03:00
Cesar Garcia
229eb5234d
Merge pull request #22601 from Chesars/fix/update-mistral-models-pricing
feat: add missing Mistral models and update pricing
2026-03-02 18:44:38 -03:00
Chesars
884f7c5e4e fix: update mistral-small-latest to match Small 3.2 specs
mistral-small-latest now points to Small 3.2 (since June 2025).
Updated pricing from $0.10/$0.30 to $0.06/$0.18 per 1M tokens,
context from 32k to 131k, and added vision support to match
mistral-small-3-2-2506.
2026-03-02 18:36:05 -03:00
Chesars
87fe521f46 fix: remove unused OpenAIImageGenerationOptionalParams import
Fixes ruff F401 in check_code_and_doc_quality CI check.
2026-03-02 18:24:29 -03:00
Chesars
abb7eb250a fix: remove retired Saba model from new entries
Saba was retired on 9/30/2025 per Mistral docs, replaced by Small 3.2.
2026-03-02 18:19:46 -03:00
Chesars
bd822a7a68 fix: add supports_response_schema to Ministral 3 models
Ministral 3 (3B, 8B, 14B) support structured outputs per Mistral docs.
2026-03-02 18:19:02 -03:00
Shivam Rawat
d5355602d5
added configurable env for mcp timeouts (#22287) 2026-03-02 13:13:41 -08:00
mubashir1osmani
ea8d22753d
docs: add fallback setup for virtual key with Loom video
docs: add fallback setup for virtual key with Loom video
2026-03-02 16:04:27 -05:00
mubashir1osmani
e96c4fed39
Update docs/my-website/docs/tutorials/fallbacks.md
Co-authored-by: greptile-apps[bot] <165735046+greptile-apps[bot]@users.noreply.github.com>
2026-03-02 16:03:55 -05:00