Commit graph

22340 commits

Author SHA1 Message Date
Krrish Dholakia
36f18411d0 fix(page.tsx): create pattern for loading in ui config before making network requests
ensures requests are formatted correctly
2025-06-02 15:13:49 -07:00
Krrish Dholakia
f3784b7b02 fix(/_types.py): add .well-known config to as public route 2025-06-02 14:23:03 -07:00
Krrish Dholakia
93a98576ac feat(_types.py): add litellm well known config as public route
allows ui to query it
2025-06-02 14:22:36 -07:00
Krrish Dholakia
c4a15dcdcb feat(ui_discovery_endpoints.py): add new public .well-known/ route for litellm ui config
returns the server root path and proxy base url for constructing api calls
2025-06-02 14:21:41 -07:00
Krrish Dholakia
66ce04d9e9 fix(onboarding_link.tsx): fix onboarding link when custom server path is set 2025-06-02 14:05:13 -07:00
Krrish Dholakia
53aae8f12a fix: fix linting error 2025-06-02 13:53:20 -07:00
Krrish Dholakia
330b6ab437 refactor: remove uneccessary references to proxybaseurl in ui code - reduce potential for errors 2025-06-02 13:51:55 -07:00
Krrish Dholakia
a10844f358 fix(networking.tsx): handle updating proxy base url for non-local instances 2025-06-02 13:42:31 -07:00
Krrish Dholakia
acf0fa803c fix(networking.tsx): update proxy base url with custom root path 2025-06-02 13:29:17 -07:00
Krrish Dholakia
c8240c41df feat(ui_sso.py): allows ui to call correct endpoint 2025-06-02 12:11:50 -07:00
Krrish Dholakia
b4393c0fdf feat(ui_sso.py): add server root path to ui token 2025-06-02 11:39:57 -07:00
Krrish Dholakia
ff81c02859 refactor(proxy_server.py): refactor all ui login endpoints to use same returned ui token object 2025-06-02 10:59:12 -07:00
Krrish Dholakia
189ec6e476 fix(proxy_server.py): create typed dict for ui returned token
allows better documentation of expected params
2025-06-02 10:49:39 -07:00
Krrish Dholakia
e9d33c4d02 fix(ui/): working custom server root path for login 2025-06-02 10:41:59 -07:00
Krrish Dholakia
526f5dd907 fix(ui/): working custom auth uptil login success event 2025-06-02 10:20:26 -07:00
Krrish Dholakia
755ef77259 fix(proxy_server.py): working swagger on custom base
removes the swagger monkey patch - this seems to render the swagger on custom base paths
2025-06-02 09:12:56 -07:00
Krrish Dholakia
9630386f2b docs: add release candidate notice 2025-06-01 22:39:57 -07:00
Krish Dholakia
83becdbc11
Litellm doc fixes 05 31 2025 (#11305)
* docs: cleanup

* docs: add anthropic file tutorial

* docs: add to sidebar
2025-06-01 00:53:56 -07:00
Ishaan Jaff
bdfa24be23 update doc v1.72.0.rc 2025-05-31 20:57:48 -07:00
Krish Dholakia
a40b81cd6b
Rate Limiting: Check all slots on redis, Reduce number of cache writes (#11299)
* fix(base_routing_strategy.py): compress increments to redis - reduces write ops

* fix(base_routing_strategy.py): make get and reset in memory keys atomic

* fix(base_routing_strategy.py): don't reset keys - causes discrepency on subsequent requests to instance

* fix(parallel_request_limiter.py): retrieve values of previous slots from cache

more accurate rate limiting with sliding window

* fix: fix test

* fix: fix linting error
2025-05-31 18:32:13 -07:00
Ishaan Jaff
10fa45d987 docs fix 2025-05-31 16:29:19 -07:00
AyrennC
8ae79178ae
feat: Add audio parameter support to gemini tts models (#11287)
* feat: Add Gemini TTS audio parameter support

- Add is_model_gemini_audio_model() method to detect TTS models
- Include 'audio' parameter in supported params for TTS models
- Map OpenAI audio parameter to Gemini speechConfig format
- Add _extract_audio_response_from_parts() method to transform audio
  output to openai format

* updated unit-test to use pcm16

* - created typedict for speechconfig
- simplified gemini tts model detection
- moved gemini_tts test to test_litellm

* simplified is_model_gemini_audio_model more
2025-05-31 16:20:19 -07:00
Ishaan Jaff
13dc757873 bump: version 1.71.3 → 1.72.0 2025-05-31 15:54:01 -07:00
Ishaan Jaff
3f616423a4 docs fixes 2025-05-31 15:30:53 -07:00
Ishaan Jaff
ab2f066df8 docs prometheus 2025-05-31 14:26:42 -07:00
Krish Dholakia
06484f6e5a
Xai, VertexAI, Google AI Studio - live web search support in OpenAI format (#11251)
* build(model_prices_and_context_window.json): fix 'supports_web_search' flag - openai only supports it on 2 models - gpt-4o-search-preview and gpt-4o-mini-search-preview

* feat(xai/chat): add xai web search options param support

* test: add max tokens to test

xai output very verbose

* build(xai/): add web search support for all xai models

* build(model_prices_and_cost.json): add gemini-2.0 supports web search

* feat(gemini/): map openai 'web_search_options' to google's 'googlesearch' tool

* build(model_prices_and_context_window.json): add supports_web_search for vertex_ai/gemini-2 models

* fix: fix circular reference error

* fix(convert_dict_to_response.py): handle scenario where xai returns finish reason as 'stop' for tool calls

* fix: reduce function size

* fix: import session handling

* Revert "fix: import session handling"

This reverts commit deb257dc10.

* fix: linting pin mypy

* [Feat]: Guardrails - Add streaming for bedrock post guard (#11247)

* feat: add streaming for bedrock post guard

* fix: bedrock guardrails

* fix: add clear comments

* Update litellm/proxy/guardrails/guardrail_hooks/bedrock_guardrails.py

Co-authored-by: Copilot <175728472+Copilot@users.noreply.github.com>

* Update litellm/proxy/guardrails/guardrail_hooks/bedrock_guardrails.py

Co-authored-by: Copilot <175728472+Copilot@users.noreply.github.com>

* fix: clean up bedrock guardrails

---------

Co-authored-by: Copilot <175728472+Copilot@users.noreply.github.com>

* [Fix] Responses API - Session management  (#11254)

* fix: import session handling

* fix: imports for session handler

* tests: tests for session handler

* Update enterprise/litellm_enterprise/enterprise_callbacks/session_handler.py

Co-authored-by: Copilot <175728472+Copilot@users.noreply.github.com>

---------

Co-authored-by: Copilot <175728472+Copilot@users.noreply.github.com>

* bump: bump litellm enterprise

* fixes: test_create_user_default_budget

* fix(xai/): filter 'strict' on tool call

* test: update test for new error string

* fix(utils.py): default to None if not set in  model cost map

ensures consistent usage of 'supports_[x]' flags

* fix(fireworks_ai/): support fireworks ai document inlining on pdf's sent via openai 'file' message type

* test: update test

* test: name filter_value_from_dict

* fix(fireworks_ai/): handle cache control flag in messages

* fix(xai/chat): fix check

---------

Co-authored-by: Ishaan Jaff <ishaanjaffer0324@gmail.com>
Co-authored-by: Copilot <175728472+Copilot@users.noreply.github.com>
2025-05-31 14:26:16 -07:00
Ishaan Jaff
cdfb6b8c37 docs prometheus end user tracking 2025-05-31 14:21:37 -07:00
Ishaan Jaff
170af8f2c8
[Docs] 1.72.0-stable release note (#11295)
* draft 1.72.0 stable

* docs - note on aiohttp transport

* docs - section for guardrails

* clean up key highlights

* docs aiohttp transport

* docs cleanup

* docs organize logging/guardrail section

* docs logging+guardrails

* docs add prometheus note

* docs fixes release note

* docs 1.72.0-stable

* docs vector store permissions
2025-05-31 14:15:16 -07:00
Ishaan Jaff
16efb8db67 Revert "Make gemini stream thinking as reasoning_content (#11290)"
This reverts commit e0daa3da68.
2025-05-31 13:29:51 -07:00
Ishaan Jaff
6863073aa4 fix: tests 2025-05-31 13:14:37 -07:00
Ishaan Jaff
3b930f6736 fix: aiohttp handle transfer encoding errors gracefully 2025-05-31 13:02:07 -07:00
Ishaan Jaff
7d47417906 test: fixes 2025-05-31 12:42:56 -07:00
Ishaan Jaff
236975a742 fix: allow users to disable aiohttp transport 2025-05-31 12:32:24 -07:00
Ishaan Jaff
e011167317 docs DISABLE_AIOHTTP_TRANSPORT 2025-05-31 12:30:52 -07:00
Ishaan Jaff
f95754c67f (UI) new build 2025-05-31 12:25:29 -07:00
Ishaan Jaff
ebf05c10a9 (ui) fix view 2025-05-31 12:21:40 -07:00
Ishaan Jaff
0dca4780c5 ui - fix permission checks 2025-05-31 12:19:37 -07:00
Ishaan Jaff
b0f2d969e7 (ui) fix passing premium user 2025-05-31 12:15:56 -07:00
Ishaan Jaff
7b4fb48bd1 ui new build 2025-05-31 12:08:48 -07:00
Ishaan Jaff
3be42fd744 ui fixes 2025-05-31 12:08:09 -07:00
Ishaan Jaff
75f87724bd chore - vector store permissions enterprise 2025-05-31 12:01:51 -07:00
Ishaan Jaff
372de1476b (chore): mark object permissions as enterprise 2025-05-31 11:52:59 -07:00
Krish Dholakia
39849627f7
feat(parallel_request_limiter_v2.py): add sliding window logic (#11283)
* feat(parallel_request_limiter_v2.py): add sliding window logic

allows rate limiting to work across minutes

* fix(parallel_request_limiter_v2.py): decrement usage on rate limit error

* fix(base_routing_strategy.py): fix merge from redis - preserve values in in-memory cache during gap b/w push to redis and read from redis

* fix(base_routing_strategy.py): catch the delta change during redis sync

ensures values are kept in sync

* fix(parallel_request_limiter_v2.py): update tpm tracking to use slot key logic

* fix: fix linting error

* test: update testing

* test: update tests

* test: skip on rate limit or internal server errors

* test: use pytest fixture instead

* test: bump mistral model
2025-05-31 10:06:42 -07:00
Ishaan Jaff
1a05f8d9e2 UI QA fixes 2025-05-31 09:41:22 -07:00
Ishaan Jaff
68fd17d15e
[Fix] QA Fixes - Vector Store Object Permissions (#11291)
* fix: QA for key,team,org permissions

* fix: add_vector_store_to_registry

* fix: refactor bedrock guard

* fix: refactor using us east 1 with vector stores

* fix: code QA checks

* fix: testing for mgmt endpoints
2025-05-31 09:41:05 -07:00
Adam Holmberg
e0daa3da68
Make gemini stream thinking as reasoning_content (#11290)
When "Thought": True, return text as reasoning_content instead of
content.

fixes #10563
fixes #11000
2025-05-31 09:13:00 -07:00
Krrish Dholakia
51f716c762 build(VLLM-Passthrough-with-loadbalancing-support-(enables-using-model-list-for-VLLM-/classify-endpoint)): Closes #11205 2025-05-31 09:00:04 -07:00
மனோஜ்குமார் பழனிச்சாமி
0fd4ee2f94
Increase timeout (#11288) 2025-05-31 07:31:14 -07:00
Bryan Low
d77b825814
Swap Cohere and Cohere Chat provider (#11173)
* fix cohere rerank provider

* swap cohere and cohere chat
2025-05-31 01:20:37 -07:00
Shuai Zhang
712e042aa4
fix(secret-managers): Break AzureCredentialType restriction on AZURE_CREDENTIAL (#11272) 2025-05-31 01:03:08 -07:00