Commit graph

34410 commits

Author SHA1 Message Date
Sameer Kankute
9f627c67d8 Add tests for encrypted_content_affinity 2026-03-03 18:36:07 +05:30
Sameer Kankute
394c49d303 Add tests for encrypted_content_affinity 2026-03-03 18:36:07 +05:30
Sameer Kankute
92b2556282 Add encrypted_content_affinity in router 2026-03-03 18:34:46 +05:30
Sameer Kankute
058fac848e Add Encrypted-content-aware deployment affinity for the Router 2026-03-03 18:34:46 +05:30
Chesars
7977126181 fix(pricing): add 7 missing DashScope model pricing entries
Fixes #22646

Adds pricing for DashScope models that were missing from the cost map,
causing $0 spend tracking in the proxy dashboard:

- dashscope/qwen3-max-2026-01-23 (tiered, same as qwen3-max)
- dashscope/qwen3-next-80b-a3b-instruct ($0.15/$1.20 per 1M)
- dashscope/qwen3-next-80b-a3b-thinking ($0.15/$1.20 per 1M)
- dashscope/qwen3-vl-235b-a22b-instruct ($0.40/$1.60 per 1M)
- dashscope/qwen3-vl-235b-a22b-thinking ($0.40/$4.00 per 1M)
- dashscope/qwen3-vl-32b-instruct ($0.16/$0.64 per 1M)
- dashscope/qwen3-vl-32b-thinking ($0.16/$2.87 per 1M)
2026-03-03 09:52:17 -03:00
Chesars
409208771e fix(pricing): add 5 missing OpenRouter model pricing entries
Fixes #22609

Adds pricing for OpenRouter models that were routing correctly but
returning $0 for spend tracking due to missing cost map entries:

- openrouter/anthropic/claude-sonnet-4.6 ($3.00/$15.00 per 1M tokens)
- openrouter/google/gemini-3.1-pro-preview ($2.00/$12.00 per 1M tokens)
- openrouter/openai/gpt-5.1-codex-max ($1.25/$10.00 per 1M tokens)
- openrouter/qwen/qwen3-coder-plus ($1.00/$5.00 per 1M tokens)
- openrouter/z-ai/glm-5 ($0.80/$2.56 per 1M tokens)
2026-03-03 09:24:49 -03:00
Sameer Kankute
deb8fea6b1 Add blog post for gemini-3.1-flash-lite-preview 2026-03-03 15:19:16 +05:30
David Steele
565a52780b
Merge remote-tracking branch 'upstream/main' into pr-22553 2026-03-03 09:48:39 +00:00
Sameer Kankute
851be58775 Add day 0 support of gemini-3.1-flash-lite-preview 2026-03-03 15:07:47 +05:30
Sameer Kankute
213423cb45 Fix test case 2026-03-03 15:05:20 +05:30
Sameer Kankute
f8034f15ad Remove defualt hardcoded thinking levels for gemini 3 family 2026-03-03 14:51:12 +05:30
Harshit28j
d07689d2d7 bump: version 1.82.0 → 1.82.1 2026-03-03 11:59:58 +05:30
Krish Dholakia
67f90254ed
feat(guardrails): team-based guardrail registration and approval workflow (#22459)
* feat(guardrails): team-based guardrail registration and approval workflow

Add team-based guardrail submission system where teams can register
Generic Guardrail API guardrails for admin review. Includes:

- POST /guardrails/register endpoint for team-scoped submissions
- Admin review endpoints (list/get/approve/reject submissions)
- Team Guardrails tab in the UI dashboard
- extra_headers support for forwarding client headers to guardrail APIs
- Prisma schema migration for status, submitted_at, reviewed_at fields
- Documentation for team-based guardrails and static/dynamic headers

Co-Authored-By: Claude Opus 4.6 <noreply@anthropic.com>

* fix(guardrails): address review feedback - SSRF, silent failure, redundant query

- Validate api_base URL scheme (http/https only) and hostname in
  register_guardrail to prevent SSRF via team submissions
- Return warning field in approve response when in-memory initialization
  fails so admins know the guardrail won't work until next sync cycle
- Eliminate redundant DB query in list_guardrail_submissions by fetching
  all team guardrails once and deriving both filtered list and summary
  counts from the single result set

Co-Authored-By: Claude Opus 4.6 <noreply@anthropic.com>

* fix(guardrails): add pending_review status guard to reject endpoint

Prevent rejecting already-active or already-rejected guardrails, which
would create a DB/memory inconsistency (active in memory but rejected
in DB). Now mirrors the approve endpoint's status check.

Co-Authored-By: Claude Opus 4.6 <noreply@anthropic.com>

---------

Co-authored-by: Claude Opus 4.6 <noreply@anthropic.com>
2026-03-02 22:06:49 -08:00
Shivaang
213799282b
fix(openrouter): register OpenRouter as native Responses API provider (#22355)
OpenRouter supports the Responses API at /api/v1/responses with
encrypted_content for multi-turn stateless reasoning workflows.
Without native registration, requests fall through to the chat
completion bridge, which uses a different format (reasoning_details)
and drops encrypted_content entirely.

This adds OpenRouterResponsesAPIConfig to route requests directly to
OpenRouter's Responses API endpoint, preserving encrypted_content.

Fixes https://github.com/BerriAI/litellm/issues/22189

Co-authored-by: Krish Dholakia <krrishdholakia@gmail.com>
2026-03-02 22:02:59 -08:00
Jaeyeon Kim(김재연)
6bcba46dda
fix: set mock status_code in JWT OIDC discovery tests (#22361)
The _resolve_jwks_url method checks response.status_code != 200, but
MagicMock returns a MagicMock object for status_code which is always
truthy (!= 200). Explicitly set mock_response.status_code = 200 so the
tests exercise the intended code path.

Co-authored-by: Claude Opus 4.6 <noreply@anthropic.com>
2026-03-02 21:57:54 -08:00
ryan-crabbe
5b0238736c
Add incident report: cache eviction closes in-use httpx clients (#22309) 2026-03-02 21:49:48 -08:00
Sameer Kankute
18216ac07c Fix: Azure ai finetuning api 2026-03-03 10:48:00 +05:30
liweiguang
81ddf08494 fix: add missing supports_function_calling for deepinfra models
All 55 deepinfra models that had `supports_tool_choice: true` were
missing the `supports_function_calling` flag, causing
`litellm.supports_function_calling()` to incorrectly return False.

Fixes #22619

Co-Authored-By: Claude Opus 4.6 <noreply@anthropic.com>
2026-03-03 12:12:43 +08:00
Ishaan Jaff
86d5b4c632
feat: add Nebius AI Studio models to model_prices_and_context_window.json (#22614)
Add 30 Nebius AI Studio models covering:
- Text-to-text: DeepSeek (R1, R1-0528, R1-Distill, V3, V3-0324), Meta Llama
  (3.1-8B/70B/405B, 3.3-70B), Qwen (3-235B/32B/30B/14B/4B, 2.5-72B/32B,
  2.5-Coder-7B, QwQ-32B), Mistral Nemo, NousResearch Hermes-3, NVIDIA
  Nemotron Ultra/Super, Google Gemma-3-27B, Llama-Guard-3
- Vision: Qwen2.5-VL-72B, Qwen2-VL-72B, Qwen2-VL-7B
- Embedding: BAAI/bge-en-icl, BAAI/bge-multilingual-gemma2, intfloat/e5-mistral-7b

Pricing sourced from https://nebius.com/prices-ai-studio (base flavor).
Context windows sourced from https://docs.nebius.com/studio/inference/models/

Co-authored-by: Cursor Agent <cursoragent@cursor.com>
Co-authored-by: Ishaan Jaff <ishaan-jaff@users.noreply.github.com>
2026-03-02 18:43:07 -08:00
Krish Dholakia
dfa2798169
Fix PR template: correct test directory path from tests/litellm/ to tests/test_litellm/ (#22612)
Co-authored-by: Cursor Agent <cursoragent@cursor.com>
2026-03-02 17:49:53 -08:00
Ishaan Jaff
bfceb7fc3f
feat(perplexity): add embedding support for pplx-embed-v1 models (#22610)
* feat: add Perplexity embedding support (pplx-embed-v1)

Add support for Perplexity AI's embedding models via the LLM HTTP handler:

Models:
- pplx-embed-v1-0.6b (1024 dims, 32K context, $0.004/1M tokens)
- pplx-embed-v1-4b (2560 dims, 32K context, $0.03/1M tokens)

Implementation:
- PerplexityEmbeddingConfig in litellm/llms/perplexity/embedding/
- Registered in ProviderConfigManager, __init__.py lazy imports, main.py dispatch
- Model pricing added to model_prices_and_context_window.json
- Supports dimensions and encoding_format parameters
- Uses base_llm_http_handler.embedding() pattern

Tests:
- 19 unit tests covering transformation, params, URLs, provider config, model info

Co-authored-by: Ishaan Jaff <ishaan-jaff@users.noreply.github.com>

* docs: add Perplexity AI embeddings documentation

- Create providers/perplexity_embedding.md with SDK and proxy usage examples
- Convert Perplexity from flat doc to category in sidebars.js
- Category includes existing chat/responses doc + new embeddings doc
- Covers pplx-embed-v1-0.6b and pplx-embed-v1-4b models
- Documents supported parameters (dimensions, encoding_format)
- Includes proxy config and curl examples

Co-authored-by: Ishaan Jaff <ishaan-jaff@users.noreply.github.com>

* fix: decode Perplexity base64_int8 embeddings to OpenAI-format float arrays

Perplexity returns embeddings as base64-encoded signed int8 values by default,
not float arrays like OpenAI. This commit adds decoding in
transform_embedding_response so the proxy returns standard OpenAI-compatible
float arrays (normalized to [-1, 1]).

- Added _decode_base64_embedding() static method
- Handles both base64 strings (decoded) and float lists (passthrough)
- Added 3 new tests for base64 decoding + passthrough

Co-authored-by: Ishaan Jaff <ishaan-jaff@users.noreply.github.com>

---------

Co-authored-by: Cursor Agent <cursoragent@cursor.com>
Co-authored-by: Ishaan Jaff <ishaan-jaff@users.noreply.github.com>
2026-03-02 17:37:50 -08:00
Kenan Yildirim
b8befb3403
Add CrowdStrike AIDR guardrail hook (#17876)
* Add CrowdStrike AIDR guardrail hook

* fixup! use apply_guardrail event hook

* fixup! update imports

* fix(guardrails): include AI response in CrowdStrike AIDR output events

Issue:
_build_guard_input_for_response() was:
- Sending only the original user input (messages).
- Not sending the AI provider response.

This fix will:
  - Extract response.choices from the ModelResponse object and include them in guard_input payload.
  - Thus, ensure AIDR output rules receive the AI-generated content for analysis.
  - Fix and update tests.

* fix(guardrails): prevent duplicate input events in CrowdStrike AIDR guardrail

Issue:
The CrowdStrike AIDR guardrail was running on during_call hooks wihtout event_hook configured.

This fix will:
- Set event_hook to ["pre_call", "post_call"] (AIDR admins will control what policy is applied)

This change will:
- Require default_on parameter
- Prevent duplicate API calls to AIDR for the same input
- Avoid unchecked AI provider API calls on during_call hook

* docs: add CrowdStrike AIDR to the list of Guardrails under Integrations

* docs: update CrowdStrike AIDR documentation page

---------

Co-authored-by: Konstantin Lapine <konstantin.lapine@crowdstrike.com>
2026-03-02 17:26:54 -08:00
Chesars
7d23106fcf fix(helicone): correct provider URL for Vertex AI Gemini models
Reorder elif branches so is_vertex_ai is checked before "gemini" in model.
Previously, Vertex AI Gemini models (e.g. vertex_ai/gemini-2.5-flash) matched
the "gemini" substring check first and were logged with the Google AI Studio
URL instead of the Vertex AI URL.
2026-03-02 19:15:37 -03:00
Cesar Garcia
2525d66dbe
Merge pull request #22584 from BerriAI/litellm_oss_staging_02_27_2026
Litellm oss staging 02 27 2026
2026-03-02 19:05:02 -03:00
Cesar Garcia
e559d4dd11
Merge pull request #22582 from BerriAI/litellm_oss_staging_02_26_2026
Litellm oss staging 02 26 2026
2026-03-02 18:51:37 -03:00
Chesars
6292c3dbdf merge: resolve conflicts with upstream/main
- anthropic.md: keep claude-opus-4-6 alias and claude-sonnet-4-6 entry
- transformation.py: take upstream's formatted effort_map with fallback
2026-03-02 18:49:24 -03:00
Cesar Garcia
835a2c3dc6
Merge pull request #22583 from Chesars/fix/add-bedrock-cache-token-pricing
fix(pricing): add missing cache token pricing for 24 Bedrock Claude models
2026-03-02 18:46:28 -03:00
Cesar Garcia
680b9ee9f2
Merge pull request #22586 from Chesars/fix/update-gemini-deprecation-dates
fix: update Gemini model deprecation dates
2026-03-02 18:45:30 -03:00
Cesar Garcia
a54a1d27d7
Merge pull request #22596 from Chesars/fix/add-dashscope-models-pricing
fix: add missing pricing for dashscope/qwen3.5-plus and dashscope/qwen3-vl-plus
2026-03-02 18:45:08 -03:00
Cesar Garcia
229eb5234d
Merge pull request #22601 from Chesars/fix/update-mistral-models-pricing
feat: add missing Mistral models and update pricing
2026-03-02 18:44:38 -03:00
Chesars
884f7c5e4e fix: update mistral-small-latest to match Small 3.2 specs
mistral-small-latest now points to Small 3.2 (since June 2025).
Updated pricing from $0.10/$0.30 to $0.06/$0.18 per 1M tokens,
context from 32k to 131k, and added vision support to match
mistral-small-3-2-2506.
2026-03-02 18:36:05 -03:00
Chesars
87fe521f46 fix: remove unused OpenAIImageGenerationOptionalParams import
Fixes ruff F401 in check_code_and_doc_quality CI check.
2026-03-02 18:24:29 -03:00
Chesars
abb7eb250a fix: remove retired Saba model from new entries
Saba was retired on 9/30/2025 per Mistral docs, replaced by Small 3.2.
2026-03-02 18:19:46 -03:00
Chesars
bd822a7a68 fix: add supports_response_schema to Ministral 3 models
Ministral 3 (3B, 8B, 14B) support structured outputs per Mistral docs.
2026-03-02 18:19:02 -03:00
Shivam Rawat
d5355602d5
added configurable env for mcp timeouts (#22287) 2026-03-02 13:13:41 -08:00
mubashir1osmani
ea8d22753d
docs: add fallback setup for virtual key with Loom video
docs: add fallback setup for virtual key with Loom video
2026-03-02 16:04:27 -05:00
mubashir1osmani
e96c4fed39
Update docs/my-website/docs/tutorials/fallbacks.md
Co-authored-by: greptile-apps[bot] <165735046+greptile-apps[bot]@users.noreply.github.com>
2026-03-02 16:03:55 -05:00
Chesars
619f53d55a feat: add missing Mistral models and update outdated pricing
Add 9 new Mistral models (mistral-large-2512, mistral-medium-3-1-2508,
mistral-small-3-2-2506, ministral-3-3b/8b/14b-2512, saba-2502,
magistral-medium/small-1-2-2509) and update mistral-large-latest,
mistral-large-3, and mistral-medium-latest with correct pricing and
context windows.

Fixes #22585
2026-03-02 18:02:41 -03:00
mubashir1osmani
fac29f1963 docs: add fallback setup for virtual key with Loom video
Co-Authored-By: Claude Opus 4.6 <noreply@anthropic.com>
2026-03-02 15:57:36 -05:00
Chesars
bfa611cb45 docs: clarify why is_model_gpt_5_search_model uses substring matching
supports_web_search in model info flags models that can use web search
as a tool, not search-only models with restricted params.
2026-03-02 17:41:20 -03:00
Chesars
ec16bd3509 merge: resolve conflict with upstream/main in presidio.py
Take upstream's refactored PII handling with _unmask_pii_text and
_process_response_for_pii helpers. Add missing StreamingChoices import.
2026-03-02 17:40:22 -03:00
Chesars
f0e571413d fix: add missing pricing for dashscope/qwen3.5-plus and dashscope/qwen3-vl-plus
Fixes #22591 - These models were missing from the pricing JSON, causing
$0 cost tracking when routed via the dashscope/* wildcard.

Pricing sourced from official Alibaba Cloud Model Studio docs (international tier).
2026-03-02 17:24:19 -03:00
Chesars
5495003e60 fix: add missing Dict/Optional imports in ChatGPT streaming_utils
Fixes NameError at runtime when ChatGPTToolCallNormalizer is
instantiated. The imports were missed when type hints were changed
from Python 3.10+ syntax (dict[], str | None) to typing module
syntax (Dict[], Optional[str]).
2026-03-02 17:19:40 -03:00
Cesar Garcia
7ab46104e2
Merge pull request #22593 from BerriAI/revert-20516-fix/openrouter-native-model-double-strip
Revert "fix(adapter): double-stripping of model names with provider-matching prefixes"
2026-03-02 17:13:19 -03:00
Cesar Garcia
0da565f023
Revert "fix(adapter): double-stripping of model names with provider-matching prefixes" 2026-03-02 17:12:48 -03:00
Chesars
0e024221a8 refactor: remove no-op pass branch in process_items
Address Greptile review nit — the if + pass was a no-op since the
mutation was removed. Deleting the branch entirely.
2026-03-02 17:05:37 -03:00
Cesar Garcia
5d512f64fe
Merge pull request #22320 from tombii/fix/openrouter-native-model-double-stripping
fix(openrouter): pattern-based fix for native model double-stripping
2026-03-02 17:01:48 -03:00
Chesars
09ef5e67e5 refactor: move native OpenRouter check to get_llm_provider before strip
The previous check in _get_openai_compatible_provider_info() ran after
the model name was already split, so it never caught the second
get_llm_provider() call from the anthropic_messages bridge.

Moved the check to get_llm_provider() before the provider-list
stripping, using a pattern-based approach (custom_llm_provider ==
"openrouter" and model.startswith("openrouter/")) instead of a
hardcoded set. This covers all current and future native OpenRouter
models.

Updated tests to verify the bridge double-call scenario with
custom_llm_provider passed through.
2026-03-02 16:55:34 -03:00
Chesars
ee3475d187 fix: correct gemini/gemini-2.0-flash-lite-preview-02-05 deprecation_date
Update from 2025-12-02 to 2025-12-09 per
https://ai.google.dev/gemini-api/docs/deprecations
2026-03-02 15:54:24 -03:00
Chesars
53dc4ee7ef fix: revert gemini/gemini-2.0-flash-live-001 deprecation_date to 2025-12-09
The June 1 date is for Vertex AI, but this entry is for the Gemini API
where the shutdown date is December 9, 2025 per
https://ai.google.dev/gemini-api/docs/deprecations
2026-03-02 15:52:01 -03:00