Commit graph

22443 commits

Author SHA1 Message Date
Krish Dholakia
c42740a4b9
Simplify experimental multi-instance rate limiter - more accurate (#11424)
* refactor: comment out circuit breaker

causes incorrect rate limiting in high traffic

* fix(base_routing_strategy.py): don't reset value if redis val is lower than current in-memory value

Fixes issue where redis might be trailing in-memory value

* fix(parallel_request_limiter_v2.py): if in-memory higher than redis, don't reset value; add previous slot keys to redis increment to correctly 'get' them

* fix(parallel_request_limiter_v3.py): v3 implementation of parallel request limiter

does not use background redis syncing - increments redis in call

 simplify rate limiting logic, to improve accuracy

* fix: fix ruff errors

* fix(parallel_request_limiter_v3.py): don't decrement limit on post call success - causes double decrements

* fix(parallel_request_limiter_v3.py): working accurate multi-instance logic

ensured just 100 requests allowed on 100 users, 10 ramp up, 100 rpm limit key, 2 instances

* fix(parallel_request_limiter_v3.py): working accurate rate limiting with time window resets

allows rate limiting to work across multiple windows

* test: add unit tests for v3 rate limiter

* fix(parallel_request_limiter_v3.py): return window value into in-memory cache

allows in-memory cache checks to be used correctly

* refactor(parallel_request_limiter_v3.py): refactor rate limiting to work for multiple window/counter key pairs

enables using for user/team/model rate limiting

* feat(parallel_request_limiter_v3.py): working rate limiting, across key/user/team/end-user

* fix(parallel_request_limiter_v3.py): add model specific rate limiting

* fix(parallel_request_limiter_v3.py): ignore if no rate limits set

skip unecessary rate limit checks - if no limits set

* fix(parallel_request_limiter_v3.py): initial commit bringing token rate limits back

* fix(parallel_request_limiter_v3.py): increment by value in list + update assertions to handle tokens + max parallel requests

* test(parallel_request_limiter_v3.py): more testing

* fix(parallel_request_limiter.py): working in-memory cache limiter

* fix(redis_cache.py): ignore linting error - use safe hasattr

* fix(parallel_request_limiter_v3.py): fix linting error

* refactor: remove redundant parallel_Request_limiter_v2.py

old / inaccurate implementation

* test: update tests

* style: cleanup

* test: update test

* docs(config_settings.md): document new env var

* test(test_base_routing_strategy.py): update test
2025-06-07 11:10:55 -07:00
Cole McIntosh
1d86fc84fe
Update web search documentation for new provider support (xAI, VertexAI, Google AI Studio) (#11515)
* Update web_search.md to include new supported providers and models, enhance web search options, and improve documentation for using web search with various AI models.

* Update LiteLLM version in web_search.md to reflect the latest stable release.

* Fix formatting in web_search.md for model declaration consistency.
2025-06-07 09:13:03 -07:00
Krish Dholakia
bc7dd9fa6d
Litellm dev 06 06 2025 p1 (#11496)
* fix(proxy/_types.py): add key masking to audit logs - prevent leaking sk- keys

* fix(navbar.tsx): fix getting image url when proxy base url is null
2025-06-07 09:12:16 -07:00
Tu Vu
3b7746a13b
Update the correct test directory in contributing_code.md (#11511) 2025-06-07 07:35:01 -07:00
Ishaan Jaff
5299c4bb6e docs - stable release v1.72.0 2025-06-06 20:58:02 -07:00
Ishaan Jaff
18081cf250 test fix test_aaauser_personal_budgets 2025-06-06 20:55:27 -07:00
Ishaan Jaff
bc835c6044 test_lm_studio_completion 2025-06-06 20:41:00 -07:00
Ishaan Jaff
362e358a77
[Feat] Allow using litellm.completion with /v1/messages API Spec (use gpt-4, gemini etc with claude code) (#11502)
* feat: add anthropic stream wrapper

* feat: add AnthropicExperimentalPassThroughConfig

* feat: working non streaming anthropic

* feat: working streaming anthropic-litellm bridge

* test - anthropic OpenAI bridge tests

* fix: add sync support for anthropic_messages

* fix: using is async check

* fix: ensure streams are SSE

* fix: imports

* fix code qa check

* fix: linting errors

* test_sync_openai_messages

* cleanup remove stash file
2025-06-06 20:35:53 -07:00
Tu Vu
bb45844ad8
Update model version in deploy.md (#11506) 2025-06-06 20:35:14 -07:00
Tu Vu
bc0e93e8a1
Remove retired version gpt-3.5 from configs.md (#11508) 2025-06-06 20:34:45 -07:00
Ishaan Jaff
08d6f3e142 Revert "Enhance proxy CLI with Rich formatting and improved user experience (#11420)"
This reverts commit 3b911ba1b2.
2025-06-06 17:55:45 -07:00
Ishaan Jaff
3be1dab1e1 Revert "fix don't require rich for litellm python SDK"
This reverts commit b76e7acbb8.
2025-06-06 17:55:36 -07:00
Ishaan Jaff
b76e7acbb8 fix don't require rich for litellm python SDK 2025-06-06 17:47:17 -07:00
Cole McIntosh
3b911ba1b2
Enhance proxy CLI with Rich formatting and improved user experience (#11420)
* Enhance proxy CLI with Rich formatting and improved user experience

- Integrated Rich library for better console output in `proxy_cli.py`, including version display, health check results, and test completion responses.
- Updated health check and test completion methods to provide progress indicators and formatted tables.
- Refactored feedback display in `proxy_server.py` to use Rich for a more visually appealing user interface.
- Adjusted tests in `test_proxy_cli.py` to mock console output instead of using print statements, ensuring compatibility with Rich formatting.

* fix linting error

* refactor(proxy_cli.py): simplify DB setup logging

- Removed progress indicators for IAM token generation and environment variable decryption to simplify the code.
- Consolidated the logic for generating the database URL and setting environment variables.
- Enhanced error handling for configuration loading and database setup, ensuring clearer feedback

* Update test-linting workflow to include proxy-dev dependencies in Poetry installation

* Enhance proxy server initialization with Rich console for improved model display. Added support for loading model parameters from environment variables and refined provider identification logic. Fallback to original print formatting if Rich is not available.

* Refactor feedback handling: Moved feedback message generation and custom warning display to utils.py. Enhanced feedback box with rich formatting and fallback to ASCII for environments without rich. Cleaned up proxy_server.py by removing obsolete code.

* fix linting error

* Refactor model initialization display: Moved model initialization logic to a new utility function `display_model_initialization` for improved readability and maintainability. Enhanced model provider extraction with a dedicated function. Fallback to basic logging if Rich console is unavailable.

* Refactor model provider extraction: Replace the `_extract_provider_from_model` function with a more robust approach using `get_llm_provider`. Implement fallback logic for provider identification and improve error handling. Ensure compatibility with Rich console for model initialization display.
2025-06-06 17:16:53 -07:00
Krrish Dholakia
079d397c6d fix(vertex_and_google_ai_studio.py): preserve vertex response id 2025-06-06 14:42:49 -07:00
Ishaan Jaff
4749008fbe
docs: add redis version requirement (#11499) 2025-06-06 14:22:47 -07:00
Cole McIntosh
e191e72746
Fix: Respect user_header_name property for budget selection and user identification (#11419)
* Refactor get_end_user_id_from_request_body to support user ID retrieval from custom headers and multiple request body formats. Enhance tests to cover various scenarios including header precedence and fallback mechanisms.

* Refactor get_end_user_id_from_request_body function to accept request_body as the first parameter, improving clarity and flexibility. Update tests for compatibility and add new cases to ensure correct functionality across various request body formats.

* Update _user_api_key_auth_builder and user_api_key_auth to pass request object to get_end_user_id_from_request_body, enhancing user ID retrieval from request data.

* refactor(auth_utils.py): update get_end_user_id_from_request_body to accept request_headers instead of request, and adjust related function calls in user_api_key_auth and tests

* refactor(tests): update mock request handling in LLM pass-through endpoint tests

- Replaced the Request object with a Mock for better flexibility in testing.
- Enhanced mock setup to include user API key handling and virtual key retrieval.
- Updated test calls to reflect changes in mock request structure and added necessary patches for new dependencies.

* refactor(vertex_and_google_ai_studio_gemini.py): remove redundant variable declaration for url_context_metadata, linting error
2025-06-06 14:21:02 -07:00
Cole McIntosh
f99e450d38
Update Makefile and add CONTRIBUTING.md to guide contributors on best practices and submission process (#11485)
- Introduced a comprehensive contributing guide outlining the checklist for PR submissions, including signing the Contributor License Agreement, adding tests, and ensuring code quality.
- Updated README.md to link to the new CONTRIBUTING.md and provide a quick start for contributors.
- Enhanced Makefile with additional commands for installation and testing to streamline the development workflow.
2025-06-06 14:19:28 -07:00
Cole McIntosh
1ceb9f9621
Merge pull request #11455 from colesmcintosh/429-fireworks-mapping
Fix Fireworks AI rate limit exception mapping - detect "rate limit" text in error messages
2025-06-06 15:06:15 -06:00
Krrish Dholakia
0c9f992af0 test: update to handle gemini-flash empty responses 2025-06-06 13:37:29 -07:00
Ishaan Jaff
96cba0148b
[Bug Fix] Fix: _transform_responses_api_content_to_chat_completion_content` doesn't support file content type (#11494)
* Handle file content type transformation in responses api (#11310)

* Handle file content type transformation in responses api

* change to use input_file

* -

* TestLiteLLMCompletionResponsesConfig

* test: TestLiteLLMCompletionResponsesConfig

* fix: fix linting

---------

Co-authored-by: Jayme Gordon <jayme_gordon@icloud.com>
2025-06-06 13:20:46 -07:00
Krrish Dholakia
196f05016e build(config.yml): weaken grype check for nightly - python3.13 known cve - blocked on chainguard to release a fix
Will call this out on release note, but this an upstream python3.13 issue  and shouldn't block nightly releases
2025-06-06 12:06:23 -07:00
Ishaan Jaff
eb02cf1a2d
Revert "Nebius model pricing info updted (#11445)" (#11493)
This reverts commit 32281de91f.
2025-06-06 11:04:21 -07:00
Fadil Rahman
fb5f2c5441
Add batch polling to python code in batches docs (#11286) 2025-06-06 10:49:23 -07:00
AyrennC
900c6e5d7b
[Docs] Add audio / tts section for gemini and vertex (#11306)
* added audio and tts doc for gemini

* updated gemini and vertex audio gen doc to be more concise
2025-06-06 10:48:50 -07:00
Krrish Dholakia
05125a9691 fix: fix linting 2025-06-06 10:43:57 -07:00
Akim Tsvigun
32281de91f
Nebius model pricing info updted (#11445) 2025-06-06 10:43:04 -07:00
Ishaan Jaff
2aa75e1403
add codex-mini-latest (#11492) 2025-06-06 10:39:09 -07:00
Ishaan Jaff
fdaad51015
Feat: add add azure endpoint for image endpoints (#11482)
* feat: add add azure endpoint for image endpoints

* test: azure image routes working as expected

* test azure routes
2025-06-06 10:38:37 -07:00
Krrish Dholakia
1af0b62578 refactor: cleanup huggingface rerank transformation 2025-06-06 10:30:44 -07:00
Krrish Dholakia
b70574017c docs: document new env vars 2025-06-06 09:59:13 -07:00
Peter Dave Hello
b452f82045
Add Google Gemini 2.5 Pro Preview 06-05 (#11447) 2025-06-06 09:28:53 -07:00
Krrish Dholakia
e5f228abd5 fix(utils.py): handle litellm proxy case for checking model info 2025-06-06 09:24:41 -07:00
Krrish Dholakia
e4d1d88d15 fix: remove redundant f-string 2025-06-06 09:18:20 -07:00
Krrish Dholakia
e2da29c54d test: update test 2025-06-06 09:15:08 -07:00
Krrish Dholakia
3608db5ffe fix(prometheus.py): update tests 2025-06-06 09:12:54 -07:00
Cole McIntosh
1b7056f281
fix(vertex_and_google_ai_studio_gemini.py): remove redundant initialization of url_context_metadata, linting error (#11486) 2025-06-06 09:01:21 -07:00
Krrish Dholakia
398fef8391 fix(bedrock/): add generic support for tool calling on bedrock models
Closes https://github.com/BerriAI/litellm/issues/11430
2025-06-05 23:37:02 -07:00
Krish Dholakia
603bd73a17
Gemini - web search cost tracking + Update max output tokens for nova models
* fix(vertex_and_google_ai_studio_gemini.py): add web search request tracking

Enables cost calculation for google web search

* fix(vertex_and_gemini): use common processing logic across stream / non-stream calls

* fix(vertex_And_google_ai_studio_Gemini.py): fix initial choice

* fix: fix linting error

* fix: add initial support for google search cost tracking

* fix(tool_call_cost_tracking.py): working tool cost tracking for gemini

* fix(vertex_ai/gemini/cost_calculator.py): add google web search tool cost tracking for vertex ai

Closes LIT-210

* fix: fix check

* build(model_prices_and_context_window.json): fix amazon nova max output tokens

Closes https://github.com/BerriAI/litellm/issues/11441

* fix: fix ruff check
2025-06-05 23:25:18 -07:00
cainiaoit
be12416863
feat: add HuggingFace rerank provider support (#11438)
++ feat: add HuggingFace rerank provider support

feat: add HuggingFace rerank provider support

feat: add HuggingFace rerank provider support

feat: add HuggingFace rerank provider support

feat: add HuggingFace rerank provider support
2025-06-05 23:23:01 -07:00
Pedro Azevedo
5d516aace1
fix: supports_function_calling works with llm_proxy models (#11381)
* Add tests for function calling support in LiteLLM proxy models

- Introduced a new test script `test_proxy_function_calling.py` to validate function calling capabilities for both direct and proxied models.
- Created a comprehensive test suite in `tests/litellm_utils_tests/test_proxy_function_calling.py` using pytest, covering various model configurations and edge cases.
- Implemented parameterized tests to ensure consistency between direct and proxied model function calling support.
- Added tests for specific proxy models, edge cases, and import verification for the `supports_function_calling` function.
- Included a demonstration test to highlight the current issue with proxy model resolution.

* feat: add fallback handling for litellm_proxy models in model info retrieval

* feat: enhance proxy function calling tests with custom model name handling and documentation

* fix: add type ignore comments for custom logger callback initialization

* fix: remove styling diff

* fix: style

* fix(utils.py): remove outdated comment regarding litellm_proxy models

* feat(utils.py): add proxy model handling for underlying model extraction

* feat(utils.py): enhance model name handling for litellm_proxy integration

* refactor(utils.py): remove unused _handle_proxy_model_names function
2025-06-05 23:15:33 -07:00
Krrish Dholakia
1ca85161a5 docs(users.md): clarify how budgets are applied 2025-06-05 23:12:24 -07:00
Ishaan Jaff
c99daef689
[Fix]: /v1/messages - return streaming usage statistics when using litellm with bedrock models (#11469)
* fix: using litellm with claude code bedrock

* fix: usage for bedrock with /messages

* fix: bedrock_sse_wrapper

* tests: test for test_chunk_parser_usage_transformation

* test fix
2025-06-05 21:18:19 -07:00
Ishaan Jaff
f0cb80ec50
[Feat] Return response_id == upstream response ID for VertexAI + Google AI studio (Stream+Non stream) (#11456)
* fix: vertexAI return responseID

* fix: vertexAI return responseID

* test_vertex_ai_response_id

* test: test_vertex_ai_streaming_response_id

* test_vertex_ai_streaming_response_id
2025-06-05 20:18:55 -07:00
Ishaan Jaff
23627d6a26
[Fix] [Bug]: Knowledge Base Call returning error (#11467)
* fix:get_and_pop_recognised_vector_store_tools

* test: tools wwith vector stores

* test - bedrock kb tools

* fix: add clear comment

* fix: vector store tools
2025-06-05 18:24:36 -07:00
RMeans
742405f6cf
Add pangea to guardrails sidebar (#11464) 2025-06-05 18:11:52 -07:00
Ishaan Jaff
18ea65218b
[Feat] Make batch size for maximum retention in spend logs a controllable parameter (#11459)
* feat: add SPEND_LOG_CLEANUP_BATCH_SIZE

* docs update

* test: test_cleanup_batch_size_env_var
2025-06-05 17:11:51 -07:00
Krish Dholakia
d05eda0311
Custom Root Path Improvements: don't require reserving /litellm route (#11460)
* fix(proxy_server.py): initial commit with asset prefix rewriting for custom base path

Closes https://github.com/BerriAI/litellm/issues/11451

* docs(litellm_proxy.md): clarify version requirement

* fix(proxy_server.py): replace litellm well known route with custom server root path

Ensures UI calls correct endpoint

* build(ui/): update ui build
2025-06-05 16:36:47 -07:00
Cole McIntosh
a3da7f1876
Add AGENTS.md (#11461) 2025-06-05 16:29:28 -07:00
Cole McIntosh
08239357cf Add ExceptionCheckers class for improved error string detection
Introduce the ExceptionCheckers class to encapsulate methods for checking error conditions in exception strings, specifically for identifying rate limit errors. Update the Fireworks AI exception mapping tests to cover various scenarios, including standard 429 errors and text-based detection, ensuring accurate mapping to RateLimitError. Enhance test coverage for both positive and negative cases of rate limit detection.
2025-06-05 17:15:53 -06:00