Krrish Dholakia
a10844f358
fix(networking.tsx): handle updating proxy base url for non-local instances
2025-06-02 13:42:31 -07:00
Krrish Dholakia
acf0fa803c
fix(networking.tsx): update proxy base url with custom root path
2025-06-02 13:29:17 -07:00
Krrish Dholakia
c8240c41df
feat(ui_sso.py): allows ui to call correct endpoint
2025-06-02 12:11:50 -07:00
Krrish Dholakia
b4393c0fdf
feat(ui_sso.py): add server root path to ui token
2025-06-02 11:39:57 -07:00
Krrish Dholakia
ff81c02859
refactor(proxy_server.py): refactor all ui login endpoints to use same returned ui token object
2025-06-02 10:59:12 -07:00
Krrish Dholakia
189ec6e476
fix(proxy_server.py): create typed dict for ui returned token
...
allows better documentation of expected params
2025-06-02 10:49:39 -07:00
Krrish Dholakia
e9d33c4d02
fix(ui/): working custom server root path for login
2025-06-02 10:41:59 -07:00
Krrish Dholakia
526f5dd907
fix(ui/): working custom auth uptil login success event
2025-06-02 10:20:26 -07:00
Krrish Dholakia
755ef77259
fix(proxy_server.py): working swagger on custom base
...
removes the swagger monkey patch - this seems to render the swagger on custom base paths
2025-06-02 09:12:56 -07:00
Krrish Dholakia
9630386f2b
docs: add release candidate notice
2025-06-01 22:39:57 -07:00
Krish Dholakia
83becdbc11
Litellm doc fixes 05 31 2025 ( #11305 )
...
* docs: cleanup
* docs: add anthropic file tutorial
* docs: add to sidebar
2025-06-01 00:53:56 -07:00
Ishaan Jaff
bdfa24be23
update doc v1.72.0.rc
2025-05-31 20:57:48 -07:00
Krish Dholakia
a40b81cd6b
Rate Limiting: Check all slots on redis, Reduce number of cache writes ( #11299 )
...
* fix(base_routing_strategy.py): compress increments to redis - reduces write ops
* fix(base_routing_strategy.py): make get and reset in memory keys atomic
* fix(base_routing_strategy.py): don't reset keys - causes discrepency on subsequent requests to instance
* fix(parallel_request_limiter.py): retrieve values of previous slots from cache
more accurate rate limiting with sliding window
* fix: fix test
* fix: fix linting error
2025-05-31 18:32:13 -07:00
Ishaan Jaff
10fa45d987
docs fix
2025-05-31 16:29:19 -07:00
AyrennC
8ae79178ae
feat: Add audio parameter support to gemini tts models ( #11287 )
...
* feat: Add Gemini TTS audio parameter support
- Add is_model_gemini_audio_model() method to detect TTS models
- Include 'audio' parameter in supported params for TTS models
- Map OpenAI audio parameter to Gemini speechConfig format
- Add _extract_audio_response_from_parts() method to transform audio
output to openai format
* updated unit-test to use pcm16
* - created typedict for speechconfig
- simplified gemini tts model detection
- moved gemini_tts test to test_litellm
* simplified is_model_gemini_audio_model more
2025-05-31 16:20:19 -07:00
Ishaan Jaff
13dc757873
bump: version 1.71.3 → 1.72.0
2025-05-31 15:54:01 -07:00
Ishaan Jaff
3f616423a4
docs fixes
2025-05-31 15:30:53 -07:00
Ishaan Jaff
ab2f066df8
docs prometheus
2025-05-31 14:26:42 -07:00
Krish Dholakia
06484f6e5a
Xai, VertexAI, Google AI Studio - live web search support in OpenAI format ( #11251 )
...
* build(model_prices_and_context_window.json): fix 'supports_web_search' flag - openai only supports it on 2 models - gpt-4o-search-preview and gpt-4o-mini-search-preview
* feat(xai/chat): add xai web search options param support
* test: add max tokens to test
xai output very verbose
* build(xai/): add web search support for all xai models
* build(model_prices_and_cost.json): add gemini-2.0 supports web search
* feat(gemini/): map openai 'web_search_options' to google's 'googlesearch' tool
* build(model_prices_and_context_window.json): add supports_web_search for vertex_ai/gemini-2 models
* fix: fix circular reference error
* fix(convert_dict_to_response.py): handle scenario where xai returns finish reason as 'stop' for tool calls
* fix: reduce function size
* fix: import session handling
* Revert "fix: import session handling"
This reverts commit deb257dc10 .
* fix: linting pin mypy
* [Feat]: Guardrails - Add streaming for bedrock post guard (#11247 )
* feat: add streaming for bedrock post guard
* fix: bedrock guardrails
* fix: add clear comments
* Update litellm/proxy/guardrails/guardrail_hooks/bedrock_guardrails.py
Co-authored-by: Copilot <175728472+Copilot@users.noreply.github.com>
* Update litellm/proxy/guardrails/guardrail_hooks/bedrock_guardrails.py
Co-authored-by: Copilot <175728472+Copilot@users.noreply.github.com>
* fix: clean up bedrock guardrails
---------
Co-authored-by: Copilot <175728472+Copilot@users.noreply.github.com>
* [Fix] Responses API - Session management (#11254 )
* fix: import session handling
* fix: imports for session handler
* tests: tests for session handler
* Update enterprise/litellm_enterprise/enterprise_callbacks/session_handler.py
Co-authored-by: Copilot <175728472+Copilot@users.noreply.github.com>
---------
Co-authored-by: Copilot <175728472+Copilot@users.noreply.github.com>
* bump: bump litellm enterprise
* fixes: test_create_user_default_budget
* fix(xai/): filter 'strict' on tool call
* test: update test for new error string
* fix(utils.py): default to None if not set in model cost map
ensures consistent usage of 'supports_[x]' flags
* fix(fireworks_ai/): support fireworks ai document inlining on pdf's sent via openai 'file' message type
* test: update test
* test: name filter_value_from_dict
* fix(fireworks_ai/): handle cache control flag in messages
* fix(xai/chat): fix check
---------
Co-authored-by: Ishaan Jaff <ishaanjaffer0324@gmail.com>
Co-authored-by: Copilot <175728472+Copilot@users.noreply.github.com>
2025-05-31 14:26:16 -07:00
Ishaan Jaff
cdfb6b8c37
docs prometheus end user tracking
2025-05-31 14:21:37 -07:00
Ishaan Jaff
170af8f2c8
[Docs] 1.72.0-stable release note ( #11295 )
...
* draft 1.72.0 stable
* docs - note on aiohttp transport
* docs - section for guardrails
* clean up key highlights
* docs aiohttp transport
* docs cleanup
* docs organize logging/guardrail section
* docs logging+guardrails
* docs add prometheus note
* docs fixes release note
* docs 1.72.0-stable
* docs vector store permissions
2025-05-31 14:15:16 -07:00
Ishaan Jaff
16efb8db67
Revert "Make gemini stream thinking as reasoning_content ( #11290 )"
...
This reverts commit e0daa3da68 .
2025-05-31 13:29:51 -07:00
Ishaan Jaff
6863073aa4
fix: tests
2025-05-31 13:14:37 -07:00
Ishaan Jaff
3b930f6736
fix: aiohttp handle transfer encoding errors gracefully
2025-05-31 13:02:07 -07:00
Ishaan Jaff
7d47417906
test: fixes
2025-05-31 12:42:56 -07:00
Ishaan Jaff
236975a742
fix: allow users to disable aiohttp transport
2025-05-31 12:32:24 -07:00
Ishaan Jaff
e011167317
docs DISABLE_AIOHTTP_TRANSPORT
2025-05-31 12:30:52 -07:00
Ishaan Jaff
f95754c67f
(UI) new build
2025-05-31 12:25:29 -07:00
Ishaan Jaff
ebf05c10a9
(ui) fix view
2025-05-31 12:21:40 -07:00
Ishaan Jaff
0dca4780c5
ui - fix permission checks
2025-05-31 12:19:37 -07:00
Ishaan Jaff
b0f2d969e7
(ui) fix passing premium user
2025-05-31 12:15:56 -07:00
Ishaan Jaff
7b4fb48bd1
ui new build
2025-05-31 12:08:48 -07:00
Ishaan Jaff
3be42fd744
ui fixes
2025-05-31 12:08:09 -07:00
Ishaan Jaff
75f87724bd
chore - vector store permissions enterprise
2025-05-31 12:01:51 -07:00
Ishaan Jaff
372de1476b
(chore): mark object permissions as enterprise
2025-05-31 11:52:59 -07:00
Krish Dholakia
39849627f7
feat(parallel_request_limiter_v2.py): add sliding window logic ( #11283 )
...
* feat(parallel_request_limiter_v2.py): add sliding window logic
allows rate limiting to work across minutes
* fix(parallel_request_limiter_v2.py): decrement usage on rate limit error
* fix(base_routing_strategy.py): fix merge from redis - preserve values in in-memory cache during gap b/w push to redis and read from redis
* fix(base_routing_strategy.py): catch the delta change during redis sync
ensures values are kept in sync
* fix(parallel_request_limiter_v2.py): update tpm tracking to use slot key logic
* fix: fix linting error
* test: update testing
* test: update tests
* test: skip on rate limit or internal server errors
* test: use pytest fixture instead
* test: bump mistral model
2025-05-31 10:06:42 -07:00
Ishaan Jaff
1a05f8d9e2
UI QA fixes
2025-05-31 09:41:22 -07:00
Ishaan Jaff
68fd17d15e
[Fix] QA Fixes - Vector Store Object Permissions ( #11291 )
...
* fix: QA for key,team,org permissions
* fix: add_vector_store_to_registry
* fix: refactor bedrock guard
* fix: refactor using us east 1 with vector stores
* fix: code QA checks
* fix: testing for mgmt endpoints
2025-05-31 09:41:05 -07:00
Adam Holmberg
e0daa3da68
Make gemini stream thinking as reasoning_content ( #11290 )
...
When "Thought": True, return text as reasoning_content instead of
content.
fixes #10563
fixes #11000
2025-05-31 09:13:00 -07:00
Krrish Dholakia
51f716c762
build(VLLM-Passthrough-with-loadbalancing-support-(enables-using-model-list-for-VLLM-/classify-endpoint)): Closes #11205
2025-05-31 09:00:04 -07:00
மனோஜ்குமார் பழனிச்சாமி
0fd4ee2f94
Increase timeout ( #11288 )
2025-05-31 07:31:14 -07:00
Bryan Low
d77b825814
Swap Cohere and Cohere Chat provider ( #11173 )
...
* fix cohere rerank provider
* swap cohere and cohere chat
2025-05-31 01:20:37 -07:00
Shuai Zhang
712e042aa4
fix(secret-managers): Break AzureCredentialType restriction on AZURE_CREDENTIAL ( #11272 )
2025-05-31 01:03:08 -07:00
VigneshwarRajasekaran
9df61ef08b
Wrong parameter mapping for "frequency_penalty" to "repeat_penalty" in Ollama/completion/transformation.py ( #11284 )
2025-05-31 00:58:19 -07:00
Ishaan Jaff
15ea80d2cf
ui new build
2025-05-30 23:08:07 -07:00
Ishaan Jaff
c5873c6f1f
UI QA fixes/cleanup
2025-05-30 23:06:45 -07:00
Ishaan Jaff
310d97c982
fix: fix linting error
2025-05-30 22:51:10 -07:00
Ishaan Jaff
d7f19bbfe3
[Bug]: Performance Fix Max langfuse clients reached: 20 is greater than 20 ( #11285 )
...
* fix: initializing langfuse clients
* fix: initializing langfuse clients
* tests: tests for langfuse cache
2025-05-30 22:34:39 -07:00
Ishaan Jaff
7e49b4e2a0
[Feat] Enforce Vector Store Access Controls on LiteLLM Auth ( #11281 )
...
* fix LiteLLM_ObjectPermissionTable
* fix include object_permission for list key
* fix key list to inclue obj permissions
* fix object permissions for vector stores on key info
* add key edit view with vector stores
* allow editing vector stores permissions
* fixes obj permissions
* feat: add obj permission on UI
* fix: add object_permission:true
* ui show org vector stores on org info
* fix: show object permissions on /org/info
* feat: allow updating obj permissions for keys
* fixes: key object permissions
* fixes: team object permissions
* fixes: org object permissions
* fix vector store selector for Orgs
* feat: add auth checks for vector store permissions
* feat: working auth checks for vector store permissions
* test vector stores auth checks
* Update litellm/proxy/_types.py
Co-authored-by: Copilot <175728472+Copilot@users.noreply.github.com>
* fix: linting
---------
Co-authored-by: Copilot <175728472+Copilot@users.noreply.github.com>
2025-05-30 22:20:11 -07:00
Ishaan Jaff
2f8eb1dcc3
fix: dont require mcp pip for proxy ( #11282 )
2025-05-30 22:19:53 -07:00