Commit graph

34125 commits

Author SHA1 Message Date
Yuneng Jiang
895cd8b5cf
chore: fixes
Some checks failed
Unit Tests: Caching (Redis) / caching-redis (push) Has been cancelled
Unit Tests: Proxy DB Operations / proxy-db (auth-checks, tests/proxy_unit_tests/test_auth_checks.py tests/proxy_unit_tests/test_user_api_key_auth.py, 20, 8) (push) Has been cancelled
Unit Tests: Proxy DB Operations / proxy-db (key-generation, tests/proxy_unit_tests/test_key_generate_prisma.py, 30, 0) (push) Has been cancelled
Unit Tests: Proxy DB Operations / proxy-db (remaining, tests/proxy_unit_tests --ignore=tests/proxy_unit_tests/test_key_generate_prisma.py --ignore=tests/proxy_unit_tests/test_auth_checks.py --ignore=tests/proxy_unit_tests/test_user_api_key_auth.py, 20, 8) (push) Has been cancelled
Unit Tests: Security / security (push) Has been cancelled
2026-04-04 23:59:56 -07:00
Krrish Dholakia
69e4b4d485 feat: add some stubbed data - help show concept 2026-03-03 18:34:09 -08:00
Krrish Dholakia
f81152d5e1 Merge branch 'litellm_dev_02_25_2026_p2' into litellm_03_03_2026_tool_policies_demo
Resolve conflicts in schema files (formatting), _types.py (import style),
tool_registry_writer.py (keep raw SQL approach with agent_id enrichment),
proxy_server.py (keep InFlightRequestsMiddleware import), and
ToolPolicies.tsx (use extracted PolicySelect component).

Co-Authored-By: Claude Opus 4.6 <noreply@anthropic.com>
2026-03-03 13:30:18 -08:00
Krrish Dholakia
6feb9babc1 fix: working backend agent tracing with MCP tool calls 2026-03-03 11:28:36 -08:00
Krrish Dholakia
bde027f0d4 fix: fix connection switching and return more clearer errors 2026-03-03 10:25:13 -08:00
Krrish Dholakia
36ca05739c feat(ui/): cleanup 2026-03-03 10:01:23 -08:00
Krrish Dholakia
1dc2df14f2 feat(ui/): cleanup 2026-03-03 10:01:23 -08:00
Krish Dholakia
f2b3beac5d
Krrishdholakia/UI multi proxy url (#22684)
* feat(ui): Add multi-proxy URL switcher for control plane/data plane architecture

Allows users to manage and switch between multiple LiteLLM proxy instances from a single UI, enabling a control plane with multiple independent data planes.

Changes:
- Create ProxyConnectionContext to manage proxy connections in localStorage
- Add ProxySwitcher dropdown in navbar to switch between configured proxies
- Add ManageProxiesModal for adding/editing/removing proxy connections
- Update useAuthorized to use API key for remote proxies instead of cookie JWT
- Disable UI config fetch for remote proxies to prevent URL overwriting
- Add setProxyBaseUrl export to allow context to update proxy URL

When users switch proxies, the page reloads with the new proxy URL in the global, ensuring all 272 API functions use the correct endpoint. Remote proxies require CORS configured to allow the UI's origin.

Co-Authored-By: Claude Opus 4.6 <noreply@anthropic.com>

* docs: Add UI multi-proxy switcher section to control plane docs

Add documentation for the new UI proxy switcher feature with mermaid
diagrams showing the architecture flow, sequence diagram for the
switching workflow, and independent database topology.

Co-Authored-By: Claude Opus 4.6 <noreply@anthropic.com>

* fix(ui): Address Greptile review feedback on proxy switcher

- Fix removeConnection to use setProxyBaseUrl(null) instead of
  setProxyBaseUrl(defaultConn.url), consistent with switchConnection
- Replace hardcoded "Admin" role for remote proxies with actual role
  fetched from /user/info endpoint on the remote proxy

Co-Authored-By: Claude Opus 4.6 <noreply@anthropic.com>

---------

Co-authored-by: Claude Opus 4.6 <noreply@anthropic.com>
2026-03-03 10:00:45 -08:00
Krrish Dholakia
a3aef0d3ea feat(a2a_endpoints.py): fix tracing to avoid recreating logging objects for the same call
allows stable trace id usage
2026-03-02 21:38:39 -08:00
Krrish Dholakia
00e861b90b feat: set stable contextid for a2a calls - allows for easily passing to downstream llm/mcp calls 2026-03-02 21:00:33 -08:00
Krrish Dholakia
dcb4a4a5eb feat: initial grouping working 2026-03-02 19:45:32 -08:00
Krrish Dholakia
ce9b2965bc style(ui/): distinguish agent calls from llm calls on ui 2026-03-02 19:35:41 -08:00
Krish Dholakia
dfa2798169
Fix PR template: correct test directory path from tests/litellm/ to tests/test_litellm/ (#22612)
Co-authored-by: Cursor Agent <cursoragent@cursor.com>
2026-03-02 17:49:53 -08:00
Ishaan Jaff
bfceb7fc3f
feat(perplexity): add embedding support for pplx-embed-v1 models (#22610)
* feat: add Perplexity embedding support (pplx-embed-v1)

Add support for Perplexity AI's embedding models via the LLM HTTP handler:

Models:
- pplx-embed-v1-0.6b (1024 dims, 32K context, $0.004/1M tokens)
- pplx-embed-v1-4b (2560 dims, 32K context, $0.03/1M tokens)

Implementation:
- PerplexityEmbeddingConfig in litellm/llms/perplexity/embedding/
- Registered in ProviderConfigManager, __init__.py lazy imports, main.py dispatch
- Model pricing added to model_prices_and_context_window.json
- Supports dimensions and encoding_format parameters
- Uses base_llm_http_handler.embedding() pattern

Tests:
- 19 unit tests covering transformation, params, URLs, provider config, model info

Co-authored-by: Ishaan Jaff <ishaan-jaff@users.noreply.github.com>

* docs: add Perplexity AI embeddings documentation

- Create providers/perplexity_embedding.md with SDK and proxy usage examples
- Convert Perplexity from flat doc to category in sidebars.js
- Category includes existing chat/responses doc + new embeddings doc
- Covers pplx-embed-v1-0.6b and pplx-embed-v1-4b models
- Documents supported parameters (dimensions, encoding_format)
- Includes proxy config and curl examples

Co-authored-by: Ishaan Jaff <ishaan-jaff@users.noreply.github.com>

* fix: decode Perplexity base64_int8 embeddings to OpenAI-format float arrays

Perplexity returns embeddings as base64-encoded signed int8 values by default,
not float arrays like OpenAI. This commit adds decoding in
transform_embedding_response so the proxy returns standard OpenAI-compatible
float arrays (normalized to [-1, 1]).

- Added _decode_base64_embedding() static method
- Handles both base64 strings (decoded) and float lists (passthrough)
- Added 3 new tests for base64 decoding + passthrough

Co-authored-by: Ishaan Jaff <ishaan-jaff@users.noreply.github.com>

---------

Co-authored-by: Cursor Agent <cursoragent@cursor.com>
Co-authored-by: Ishaan Jaff <ishaan-jaff@users.noreply.github.com>
2026-03-02 17:37:50 -08:00
Kenan Yildirim
b8befb3403
Add CrowdStrike AIDR guardrail hook (#17876)
* Add CrowdStrike AIDR guardrail hook

* fixup! use apply_guardrail event hook

* fixup! update imports

* fix(guardrails): include AI response in CrowdStrike AIDR output events

Issue:
_build_guard_input_for_response() was:
- Sending only the original user input (messages).
- Not sending the AI provider response.

This fix will:
  - Extract response.choices from the ModelResponse object and include them in guard_input payload.
  - Thus, ensure AIDR output rules receive the AI-generated content for analysis.
  - Fix and update tests.

* fix(guardrails): prevent duplicate input events in CrowdStrike AIDR guardrail

Issue:
The CrowdStrike AIDR guardrail was running on during_call hooks wihtout event_hook configured.

This fix will:
- Set event_hook to ["pre_call", "post_call"] (AIDR admins will control what policy is applied)

This change will:
- Require default_on parameter
- Prevent duplicate API calls to AIDR for the same input
- Avoid unchecked AI provider API calls on during_call hook

* docs: add CrowdStrike AIDR to the list of Guardrails under Integrations

* docs: update CrowdStrike AIDR documentation page

---------

Co-authored-by: Konstantin Lapine <konstantin.lapine@crowdstrike.com>
2026-03-02 17:26:54 -08:00
Cesar Garcia
2525d66dbe
Merge pull request #22584 from BerriAI/litellm_oss_staging_02_27_2026
Litellm oss staging 02 27 2026
2026-03-02 19:05:02 -03:00
Cesar Garcia
e559d4dd11
Merge pull request #22582 from BerriAI/litellm_oss_staging_02_26_2026
Litellm oss staging 02 26 2026
2026-03-02 18:51:37 -03:00
Chesars
6292c3dbdf merge: resolve conflicts with upstream/main
- anthropic.md: keep claude-opus-4-6 alias and claude-sonnet-4-6 entry
- transformation.py: take upstream's formatted effort_map with fallback
2026-03-02 18:49:24 -03:00
Cesar Garcia
835a2c3dc6
Merge pull request #22583 from Chesars/fix/add-bedrock-cache-token-pricing
fix(pricing): add missing cache token pricing for 24 Bedrock Claude models
2026-03-02 18:46:28 -03:00
Cesar Garcia
680b9ee9f2
Merge pull request #22586 from Chesars/fix/update-gemini-deprecation-dates
fix: update Gemini model deprecation dates
2026-03-02 18:45:30 -03:00
Cesar Garcia
a54a1d27d7
Merge pull request #22596 from Chesars/fix/add-dashscope-models-pricing
fix: add missing pricing for dashscope/qwen3.5-plus and dashscope/qwen3-vl-plus
2026-03-02 18:45:08 -03:00
Cesar Garcia
229eb5234d
Merge pull request #22601 from Chesars/fix/update-mistral-models-pricing
feat: add missing Mistral models and update pricing
2026-03-02 18:44:38 -03:00
Chesars
884f7c5e4e fix: update mistral-small-latest to match Small 3.2 specs
mistral-small-latest now points to Small 3.2 (since June 2025).
Updated pricing from $0.10/$0.30 to $0.06/$0.18 per 1M tokens,
context from 32k to 131k, and added vision support to match
mistral-small-3-2-2506.
2026-03-02 18:36:05 -03:00
Chesars
87fe521f46 fix: remove unused OpenAIImageGenerationOptionalParams import
Fixes ruff F401 in check_code_and_doc_quality CI check.
2026-03-02 18:24:29 -03:00
Chesars
abb7eb250a fix: remove retired Saba model from new entries
Saba was retired on 9/30/2025 per Mistral docs, replaced by Small 3.2.
2026-03-02 18:19:46 -03:00
Chesars
bd822a7a68 fix: add supports_response_schema to Ministral 3 models
Ministral 3 (3B, 8B, 14B) support structured outputs per Mistral docs.
2026-03-02 18:19:02 -03:00
Shivam Rawat
d5355602d5
added configurable env for mcp timeouts (#22287) 2026-03-02 13:13:41 -08:00
mubashir1osmani
ea8d22753d
docs: add fallback setup for virtual key with Loom video
docs: add fallback setup for virtual key with Loom video
2026-03-02 16:04:27 -05:00
mubashir1osmani
e96c4fed39
Update docs/my-website/docs/tutorials/fallbacks.md
Co-authored-by: greptile-apps[bot] <165735046+greptile-apps[bot]@users.noreply.github.com>
2026-03-02 16:03:55 -05:00
Chesars
619f53d55a feat: add missing Mistral models and update outdated pricing
Add 9 new Mistral models (mistral-large-2512, mistral-medium-3-1-2508,
mistral-small-3-2-2506, ministral-3-3b/8b/14b-2512, saba-2502,
magistral-medium/small-1-2-2509) and update mistral-large-latest,
mistral-large-3, and mistral-medium-latest with correct pricing and
context windows.

Fixes #22585
2026-03-02 18:02:41 -03:00
mubashir1osmani
fac29f1963 docs: add fallback setup for virtual key with Loom video
Co-Authored-By: Claude Opus 4.6 <noreply@anthropic.com>
2026-03-02 15:57:36 -05:00
Chesars
bfa611cb45 docs: clarify why is_model_gpt_5_search_model uses substring matching
supports_web_search in model info flags models that can use web search
as a tool, not search-only models with restricted params.
2026-03-02 17:41:20 -03:00
Chesars
ec16bd3509 merge: resolve conflict with upstream/main in presidio.py
Take upstream's refactored PII handling with _unmask_pii_text and
_process_response_for_pii helpers. Add missing StreamingChoices import.
2026-03-02 17:40:22 -03:00
Chesars
f0e571413d fix: add missing pricing for dashscope/qwen3.5-plus and dashscope/qwen3-vl-plus
Fixes #22591 - These models were missing from the pricing JSON, causing
$0 cost tracking when routed via the dashscope/* wildcard.

Pricing sourced from official Alibaba Cloud Model Studio docs (international tier).
2026-03-02 17:24:19 -03:00
Chesars
5495003e60 fix: add missing Dict/Optional imports in ChatGPT streaming_utils
Fixes NameError at runtime when ChatGPTToolCallNormalizer is
instantiated. The imports were missed when type hints were changed
from Python 3.10+ syntax (dict[], str | None) to typing module
syntax (Dict[], Optional[str]).
2026-03-02 17:19:40 -03:00
Cesar Garcia
7ab46104e2
Merge pull request #22593 from BerriAI/revert-20516-fix/openrouter-native-model-double-strip
Revert "fix(adapter): double-stripping of model names with provider-matching prefixes"
2026-03-02 17:13:19 -03:00
Cesar Garcia
0da565f023
Revert "fix(adapter): double-stripping of model names with provider-matching prefixes" 2026-03-02 17:12:48 -03:00
Cesar Garcia
5d512f64fe
Merge pull request #22320 from tombii/fix/openrouter-native-model-double-stripping
fix(openrouter): pattern-based fix for native model double-stripping
2026-03-02 17:01:48 -03:00
Chesars
09ef5e67e5 refactor: move native OpenRouter check to get_llm_provider before strip
The previous check in _get_openai_compatible_provider_info() ran after
the model name was already split, so it never caught the second
get_llm_provider() call from the anthropic_messages bridge.

Moved the check to get_llm_provider() before the provider-list
stripping, using a pattern-based approach (custom_llm_provider ==
"openrouter" and model.startswith("openrouter/")) instead of a
hardcoded set. This covers all current and future native OpenRouter
models.

Updated tests to verify the bridge double-call scenario with
custom_llm_provider passed through.
2026-03-02 16:55:34 -03:00
Chesars
ee3475d187 fix: correct gemini/gemini-2.0-flash-lite-preview-02-05 deprecation_date
Update from 2025-12-02 to 2025-12-09 per
https://ai.google.dev/gemini-api/docs/deprecations
2026-03-02 15:54:24 -03:00
Chesars
53dc4ee7ef fix: revert gemini/gemini-2.0-flash-live-001 deprecation_date to 2025-12-09
The June 1 date is for Vertex AI, but this entry is for the Gemini API
where the shutdown date is December 9, 2025 per
https://ai.google.dev/gemini-api/docs/deprecations
2026-03-02 15:52:01 -03:00
Chesars
ad1ab9e874 fix: add deprecation_date for gemini/gemini-3-pro-preview (Gemini API)
Gemini API shuts down gemini-3-pro-preview on 2026-03-09, per
https://ai.google.dev/gemini-api/docs/deprecations
2026-03-02 15:48:00 -03:00
Chesars
9c8620db00 fix: update Gemini model deprecation dates per Google notifications
- gemini-3-pro-preview: add deprecation_date 2026-03-26 (Vertex AI)
- gemini-2.0-flash / flash-001: update to 2026-06-01
- gemini-2.0-flash-lite / lite-001: update to 2026-06-01
- gemini/gemini-2.0-flash-live-001: update to 2026-06-01
- Also updated deepinfra, openrouter, vercel_ai_gateway variants
2026-03-02 15:40:40 -03:00
Chesars
5c4d3d85e5 fix(pricing): add missing cache token pricing for 24 Bedrock Claude models
Bedrock Claude models were missing cache_read_input_token_cost and
cache_creation_input_token_cost fields, causing cache tokens to be
billed at the full input rate instead of the discounted cache rate.

Added pricing using Bedrock's documented multipliers (0.1x for cache
read, 1.25x for cache write) consistent with all existing entries.
2026-03-02 15:02:21 -03:00
Sameer Kankute
92407ec0d4
Merge pull request #22340 from BerriAI/litellm_oss_staging_02_28_2026
Litellm oss staging 02 28 2026
2026-03-02 19:48:03 +05:30
Sameer Kankute
fc41f46f0f Fix vertex ai function calls 2026-03-02 19:43:24 +05:30
Sameer Kankute
e40e913622 Fix vertex ai function calls 2026-03-02 19:42:18 +05:30
Sameer Kankute
7e2f2a8ffa Fix inflight mypy 2026-03-02 19:41:32 +05:30
Chesars
3007010f21 style: add trailing newline to test file 2026-03-02 19:21:27 +05:30
Chesars
273cf12afa fix(gemini): add missing role="user" to function response content blocks
Gemini API only accepts "user" and "model" roles. Function responses were
being sent without a role field, causing 400 errors on multi-turn tool
calling conversations.

Fixes #22003
Fixes #20690
2026-03-02 19:21:27 +05:30