Krish Dholakia
06484f6e5a
Xai, VertexAI, Google AI Studio - live web search support in OpenAI format ( #11251 )
...
* build(model_prices_and_context_window.json): fix 'supports_web_search' flag - openai only supports it on 2 models - gpt-4o-search-preview and gpt-4o-mini-search-preview
* feat(xai/chat): add xai web search options param support
* test: add max tokens to test
xai output very verbose
* build(xai/): add web search support for all xai models
* build(model_prices_and_cost.json): add gemini-2.0 supports web search
* feat(gemini/): map openai 'web_search_options' to google's 'googlesearch' tool
* build(model_prices_and_context_window.json): add supports_web_search for vertex_ai/gemini-2 models
* fix: fix circular reference error
* fix(convert_dict_to_response.py): handle scenario where xai returns finish reason as 'stop' for tool calls
* fix: reduce function size
* fix: import session handling
* Revert "fix: import session handling"
This reverts commit deb257dc10 .
* fix: linting pin mypy
* [Feat]: Guardrails - Add streaming for bedrock post guard (#11247 )
* feat: add streaming for bedrock post guard
* fix: bedrock guardrails
* fix: add clear comments
* Update litellm/proxy/guardrails/guardrail_hooks/bedrock_guardrails.py
Co-authored-by: Copilot <175728472+Copilot@users.noreply.github.com>
* Update litellm/proxy/guardrails/guardrail_hooks/bedrock_guardrails.py
Co-authored-by: Copilot <175728472+Copilot@users.noreply.github.com>
* fix: clean up bedrock guardrails
---------
Co-authored-by: Copilot <175728472+Copilot@users.noreply.github.com>
* [Fix] Responses API - Session management (#11254 )
* fix: import session handling
* fix: imports for session handler
* tests: tests for session handler
* Update enterprise/litellm_enterprise/enterprise_callbacks/session_handler.py
Co-authored-by: Copilot <175728472+Copilot@users.noreply.github.com>
---------
Co-authored-by: Copilot <175728472+Copilot@users.noreply.github.com>
* bump: bump litellm enterprise
* fixes: test_create_user_default_budget
* fix(xai/): filter 'strict' on tool call
* test: update test for new error string
* fix(utils.py): default to None if not set in model cost map
ensures consistent usage of 'supports_[x]' flags
* fix(fireworks_ai/): support fireworks ai document inlining on pdf's sent via openai 'file' message type
* test: update test
* test: name filter_value_from_dict
* fix(fireworks_ai/): handle cache control flag in messages
* fix(xai/chat): fix check
---------
Co-authored-by: Ishaan Jaff <ishaanjaffer0324@gmail.com>
Co-authored-by: Copilot <175728472+Copilot@users.noreply.github.com>
2025-05-31 14:26:16 -07:00
Ishaan Jaff
16efb8db67
Revert "Make gemini stream thinking as reasoning_content ( #11290 )"
...
This reverts commit e0daa3da68 .
2025-05-31 13:29:51 -07:00
Ishaan Jaff
6863073aa4
fix: tests
2025-05-31 13:14:37 -07:00
Ishaan Jaff
7d47417906
test: fixes
2025-05-31 12:42:56 -07:00
Krish Dholakia
39849627f7
feat(parallel_request_limiter_v2.py): add sliding window logic ( #11283 )
...
* feat(parallel_request_limiter_v2.py): add sliding window logic
allows rate limiting to work across minutes
* fix(parallel_request_limiter_v2.py): decrement usage on rate limit error
* fix(base_routing_strategy.py): fix merge from redis - preserve values in in-memory cache during gap b/w push to redis and read from redis
* fix(base_routing_strategy.py): catch the delta change during redis sync
ensures values are kept in sync
* fix(parallel_request_limiter_v2.py): update tpm tracking to use slot key logic
* fix: fix linting error
* test: update testing
* test: update tests
* test: skip on rate limit or internal server errors
* test: use pytest fixture instead
* test: bump mistral model
2025-05-31 10:06:42 -07:00
Ishaan Jaff
68fd17d15e
[Fix] QA Fixes - Vector Store Object Permissions ( #11291 )
...
* fix: QA for key,team,org permissions
* fix: add_vector_store_to_registry
* fix: refactor bedrock guard
* fix: refactor using us east 1 with vector stores
* fix: code QA checks
* fix: testing for mgmt endpoints
2025-05-31 09:41:05 -07:00
Adam Holmberg
e0daa3da68
Make gemini stream thinking as reasoning_content ( #11290 )
...
When "Thought": True, return text as reasoning_content instead of
content.
fixes #10563
fixes #11000
2025-05-31 09:13:00 -07:00
Krrish Dholakia
51f716c762
build(VLLM-Passthrough-with-loadbalancing-support-(enables-using-model-list-for-VLLM-/classify-endpoint)): Closes #11205
2025-05-31 09:00:04 -07:00
Shuai Zhang
712e042aa4
fix(secret-managers): Break AzureCredentialType restriction on AZURE_CREDENTIAL ( #11272 )
2025-05-31 01:03:08 -07:00
Ishaan Jaff
d7f19bbfe3
[Bug]: Performance Fix Max langfuse clients reached: 20 is greater than 20 ( #11285 )
...
* fix: initializing langfuse clients
* fix: initializing langfuse clients
* tests: tests for langfuse cache
2025-05-30 22:34:39 -07:00
Ishaan Jaff
7e49b4e2a0
[Feat] Enforce Vector Store Access Controls on LiteLLM Auth ( #11281 )
...
* fix LiteLLM_ObjectPermissionTable
* fix include object_permission for list key
* fix key list to inclue obj permissions
* fix object permissions for vector stores on key info
* add key edit view with vector stores
* allow editing vector stores permissions
* fixes obj permissions
* feat: add obj permission on UI
* fix: add object_permission:true
* ui show org vector stores on org info
* fix: show object permissions on /org/info
* feat: allow updating obj permissions for keys
* fixes: key object permissions
* fixes: team object permissions
* fixes: org object permissions
* fix vector store selector for Orgs
* feat: add auth checks for vector store permissions
* feat: working auth checks for vector store permissions
* test vector stores auth checks
* Update litellm/proxy/_types.py
Co-authored-by: Copilot <175728472+Copilot@users.noreply.github.com>
* fix: linting
---------
Co-authored-by: Copilot <175728472+Copilot@users.noreply.github.com>
2025-05-30 22:20:11 -07:00
Ishaan Jaff
ea841eeb9b
[Feat] UI - show vector store permissions for Key, Team, Org ( #11277 )
...
* fix LiteLLM_ObjectPermissionTable
* fix include object_permission for list key
* fix key list to inclue obj permissions
* fix object permissions for vector stores on key info
* add key edit view with vector stores
* allow editing vector stores permissions
* fixes obj permissions
* feat: add obj permission on UI
* fix: add object_permission:true
* ui show org vector stores on org info
* fix: show object permissions on /org/info
* feat: allow updating obj permissions for keys
* fixes: key object permissions
* fixes: team object permissions
* fixes: org object permissions
* fix vector store selector for Orgs
2025-05-30 17:23:50 -07:00
Krish Dholakia
44a69421ea
Anthropic - Files API w/ form-data support on passthrough + File ID support on /chat/completion ( #11256 )
...
* fix(anthropic/chat): support passing 'file_id' param to anthropic
Partial fix for LIT-200
* feat(anthropic/chat): use correct anthropic content block based on file object
* fix(anthropic/chat): fix file id for container_upload message type
* fix(anthropic/chat/transformation.py): fix check for adding code execution to tool calls - needed for 'container_upload' message type
* fix(llm_passthrough_endpoints.py): support reading form data for anthropic passthrough
* refactor(llm_passthrough_endpoints.py): refactor block into function for easier testing
* test: add unit test
* fix: don't pass in empty tools list
* [Fix] Responses API - Session management (#11254 )
* fix: import session handling
* fix: imports for session handler
* tests: tests for session handler
* Update enterprise/litellm_enterprise/enterprise_callbacks/session_handler.py
Co-authored-by: Copilot <175728472+Copilot@users.noreply.github.com>
---------
Co-authored-by: Copilot <175728472+Copilot@users.noreply.github.com>
* bump: bump litellm enterprise
* fixes: test_create_user_default_budget
* fix: fix linting error
---------
Co-authored-by: Ishaan Jaff <ishaanjaffer0324@gmail.com>
Co-authored-by: Copilot <175728472+Copilot@users.noreply.github.com>
2025-05-29 23:54:24 -07:00
Ishaan Jaff
5e6f6ddc52
[Feat]: Add Bedrock InvokeAgents as a /chat/completions route on LiteLLM ( #11239 )
...
* feat: init structure for bedrock AGENTs
* feat: add basic routing for bedrock AGENTs
* feat: add basic transforms for bedrock AGENTs
* fix: url for bedrock agent runtime
* fix: working agents request
* feat: working agents non-streaming request
* feat: bedrock agents
* feat: add streaming for bedrock agents
* feat: add cost tracking for bedrock agents
* docs litellm with bedrock agents
* fix: linting errors
* test: invoke agents tests
2025-05-29 16:48:55 -07:00
Krish Dholakia
1995c7aad5
fix(utils.py): support non default params for audio transcription ( #11212 )
...
* fix(utils.py): support non default params for audio transcription
allows passing provider specific params straight through on transcription calls
* fix(gpt_transformation.py): fix o_series model routing
call _transform_request on async event
* refactor: refactor tests
* test(test_azure_chat_o_series_transformation.py): add unit test for azure o series error
* test: update test
* test: update json
* fix: fix mutiple keyword error
2025-05-28 22:24:02 -07:00
Krish Dholakia
ba39f9e360
Helicone base url support + fix for embedding cache hits on str input ( #11211 )
...
* fix(helicone.py): add helicone api base support
Fixes https://github.com/BerriAI/litellm/issues/10825
* test: add unit test for cache hit response on embedding calls
* fix(caching_handler.py): fix handling cache hit on embedding when input is string
Fixes LIT-197
* docs(helicone_integration.md): document new helicone api base param
2025-05-28 22:02:55 -07:00
Ishaan Jaff
711c931c71
test: fix test_key_generation_with_object_permission
2025-05-28 21:24:35 -07:00
Ishaan Jaff
4e6c4beef8
[Feat] Permission management vector stores on LiteLLM Key, Team, Orgs ( #11213 )
...
* fix: init commit for object permissions
* fix: init commit for object permissions
* fix: add vector_store_id to permissions
* fix vector store selector
* feat:add vector store permission mgmt
* feat: ui add allowed vector stores dropdown
* feat: add new vector store object permissions
* testing: key mgmt
* fix: stor vector store permissions on team
* ui select vector store for teams
* ui add vector store settings for orgs
* feat: allow setting org vector store permissions
* test: adding team permissions for vector stores
2025-05-28 16:58:53 -07:00
Vinnie-Singleton-NN
178a614d4a
Add sentry sample rate ( #10283 )
...
* Add SENTRY_API_SAMPLE_RATE configuration option for Sentry SDK
* removed print line
* Update Sentry documentation with sample rate information
---------
Co-authored-by: Vinnie <vinnie@Vinnies-MacBook-Pro.local>
2025-05-28 16:44:10 -07:00
Krish Dholakia
05e0a6d8d5
Return anthropic thinking blocks on streaming + VertexAI Minor Fixes & Improvements (Thinking, Global regions, Parallel tool calling) ( #11194 )
...
* fix(anthropic/chat/handler.py): Fixes https://github.com/BerriAI/litellm/issues/10328
Adopts changes from https://github.com/BerriAI/litellm/pull/10329
* fix(vertex_and_google_ai_studio.py): don't set 'include thoughts' if thinking budget = 0
VertexAI raises errors
* fix(vertex_llm_base.py): new function for deciding the api base, handles 'global' api base
Fixes https://github.com/BerriAI/litellm/issues/11190
* fix(vertex_ai/partner_models): fix instrumentation for custom api base check
* refactor(vertex_ai/partner): refactor function to keep below 50 LOC
* fix(vertex_ai/gemini): remove parallel tool calls error for >1 tool - just ignore (prevent call from failing)
* fix: fix linting error
2025-05-27 23:07:13 -07:00
Krish Dholakia
7072466775
VertexAI - codeExecution tool support + anyOf handling ( #11195 )
...
* fix(vertex_and_google_ai_studio_gemini.py): handle both camel case and underscores in the tool for vertex ai code execution
support vertex ai code execution
* docs(vertex.md): add code execution example to vertex ai
* fix(vertex_ai/common_utils.py): when anyof in field, just select anyof - don't include other k,v pairs - vertex throws error
Fixes https://github.com/BerriAI/litellm/issues/11164
* fix(common_utils.py): add title field inside anyof - to retain some description
Addresses https://github.com/BerriAI/litellm/issues/11164#issuecomment-2914728385
2025-05-27 21:23:14 -07:00
Ishaan Jaff
1a17755c60
test: fix test_ensure_initialize_azure_sdk_client_always_used
2025-05-27 19:02:11 -07:00
Ishaan Jaff
0590b1eb3a
[Fix] Prometheus Metrics - Do not track end_user by default + expose flag to enable tracking end_user on prometheus ( #11192 )
...
* fix: testing for disabling end user on metrics
* fix: fixes for test_prometheus_factory
* Delete litellm/model_prices_and_context_window_backup.json
* fix: issues with merge conflicts
* fix: test_get_end_user_id_for_cost_tracking_prometheus_only
* Update tests/test_litellm/integrations/test_prometheus.py
Co-authored-by: Copilot <175728472+Copilot@users.noreply.github.com>
---------
Co-authored-by: Copilot <175728472+Copilot@users.noreply.github.com>
2025-05-27 17:06:58 -07:00
Krish Dholakia
4c82dd9b27
Ollama Chat - parse tool calls on streaming ( #11171 )
...
* fix(user_api_key_auth.py): fix else block
Fixes https://github.com/BerriAI/litellm/issues/11170
* refactor(ollama/chat): refactor to base config pattern
easier to maintain fixes
* fix(ollama/chat): support tool call parsing on streaming
Closes https://github.com/BerriAI/litellm/issues/11104
* test: update import location
* fix: cleanup unused import
* fix: fix ruff check error
* test: update import
* test: update test on ci
* ci: cleanup
* fix: fix chekc
* fix: fix api key check order
* test: fix import
* ci: fix script
* test: fix imports
* fix: fix tests
2025-05-27 16:14:49 -07:00
Krrish Dholakia
0017d5f1db
build(ui/): Allow empty values in daily agg table + reintroduce 'unassigned' teams in spend tracking
2025-05-26 22:03:34 -07:00
Ishaan Jaff
4d2edc4e7a
[Fixes] Aiohttp transport fixes - add handling for aiohttp.ClientPayloadError and ssl_verification settings ( #11162 )
...
* fix: AiohttpResponseStream transport
* fix: use AiohttpResponseStream transport by default
* fix: AiohttpResponseStream transport
* fixes: mapping aiohttp exceptions
* fixes: aiohttp rollout
* fixes: add support ssl_verify for aiohttp
* fixes: add support ssl_verify for aiohttp
* fixes: remove duplicates
2025-05-26 21:14:35 -07:00
Ishaan Jaff
e606bfe31d
[Feat - Contributor PR] Add Video support for Bedrock Converse ( #11166 )
...
* feat: add video support for bedrock converse api (#11043 )
* fixes: bedrock add video support
* fixes: bedrock add video support
---------
Co-authored-by: yytdfc <fuchen@foxmail.com>
2025-05-26 20:17:07 -07:00
Krish Dholakia
ef42461c1e
Litellm fix GitHub action testing ( #11163 )
...
* test: add __init__.py files
* refactor: rename test folder to avoid naming conflict
* test: update workflows
* test: update tests
* test: update imports
* test: update tests
* test: remove unused import
* ci(test-litellm.yml): add pytest retry to github workflow
* test: fix test
2025-05-26 14:41:42 -07:00