* Vllm rerank (#12737)
* Add Hosted VLLM rerank provider integration
This commit implements the Hosted VLLM rerank provider integration for LiteLLM. The integration includes:
Adding Hosted VLLM as a supported rerank provider in the main rerank function
Implementing the HostedVLLMRerank handler class for making API requests
Creating a transformation class to convert Hosted VLLM responses to LiteLLM's standardized format
The integration supports both synchronous and asynchronous rerank operations. API credentials can be provided directly or through environment variables (HOSTED_VLLM_API_KEY and HOSTED_VLLM_API_BASE).
Notable features:
Proper error handling for missing credentials
Standard response transformation
Support for common rerank parameters (top_n, return_documents, etc.)
Proper token usage tracking
This expands LiteLLM's rerank provider ecosystem to include Hosted VLLM alongside existing providers like Cohere, Together AI, Azure AI, and Bedrock.
* refactor(rerank): use base_llm_http_handler for hosted_vllm rerank
- Replace custom HostedVLLMRerank handler with base_llm_http_handler
- Implement proper HostedVLLMRerankConfig inheriting from BaseRerankConfig
- Follow Cohere-compatible implementation pattern
- Clean up unnecessary comments
* Fix lint errors in hosted_vllm rerank transformer: remove unused imports
* Fix linting errors in rerank transformation modules
* fix: resolve type errors in Hosted VLLM rerank module
---------
Co-authored-by: Philip D'Souza <philip.dsouza@macro4.com>
Co-authored-by: Philip D'Souza <philip.a.dsouza@gmail.com>
* added a few tests
---------
Co-authored-by: Philip D'Souza <philip.dsouza@macro4.com>
Co-authored-by: Philip D'Souza <philip.a.dsouza@gmail.com>
Replace MagicMock with AsyncMock for litellm_teamtable.update to fix:
TypeError: object MagicMock can't be used in 'await' expression
The test was failing because it tried to await a MagicMock object.
Added AsyncMock for the update method to properly handle async operations.
* fix bug
When Max Budget, TPM, RPM, Expire Key are set on Regenerate Key, these values are not reflected on the settings page. Refreshing the page is required.
* fix editing settings after key generation
* fix(team_info.tsx): allow setting custom key duration
more flexible than previous pre-set options
* feat(team_info.tsx): show how many user + service account keys have been created within a team
* fix(team_endpoints.py): ensure user id correctly added when new team created with user email as member
Fixes issue where user not correctly added to team on /team/new
* fix(internal_user_endpoints.py): make user email validation check case insensitive
Fixes issue where uppercase email was added even when lowercase email existed
* test: update test
* build: move build_and_test to use prisma migrate
* feat(proxy_setting_endpoints.py): encrypt env var before storing in db
Ensures env var can be read when loaded in from DB
Fixes issue when trying to add SSO from admin UI
* test: update tests
* Check content and order of trimmed messages
* Assert tool calls are preserved if below max_tokens
* Unreverse order of tool calls
* Return tool calls alongside other messages
* Write test for trimming untokenizable field
* Return original messages in case of exception
* Add concise Claude Code + LiteLLM Gateway tutorial
- Create focused tutorial matching existing tutorial style
- Step-by-step guide from installation to advanced configurations
- Multi-provider configuration examples (AWS Bedrock, Azure OpenAI, Load Balancing)
- Based on Anthropic's official LiteLLM configuration documentation
- Added to sidebar with clean title 'Use LiteLLM with Claude Code'
- Fixed sidebar reference from 'secret' to 'set_keys' for proper document resolution
* Update config_settings.md to correct documentation links for key management and Hashicorp Vault settings. Changed references from 'secret.md' to 'set_keys.md' for improved clarity and accuracy.
* Update sidebar and config_settings.md to reflect changes in key management documentation. Changed sidebar reference from 'set_keys' to 'secret' and updated links in config_settings.md for Hashicorp Vault settings to point to 'secret.md' for improved accuracy.
* Remove extra tutorial and update sidebar accordingly
* Update tutorial title from 'WebUI' to 'Open WebUI' for clarity and consistency in documentation.
* Remove Python version requirement from Claude Responses API tutorial for clarity and to align with updated prerequisites.
* feat: add input_fidelity parameter for OpenAI image generation
- Add input_fidelity to OpenAIImageGenerationOptionalParams type
- Update image_generation function signature to accept input_fidelity
- Add input_fidelity to default_params in get_optional_params_image_gen
- Include input_fidelity in openai_params list for proper handling
- Update documentation with input_fidelity parameter description
- Add test for input_fidelity parameter functionality
This enables control over how closely the model follows the input prompt
for gpt-image-1 model, improving prompt adherence and image quality.
* feat: add input_fidelity to optional parameters for image generation
- Include input_fidelity in the list of OpenAIImageGenerationOptionalParams
- This addition enhances the flexibility of image generation by allowing control over input fidelity.
* test: enhance test for gpt-image-1 with input_fidelity parameter
- Update test_gpt_image_1_with_input_fidelity to include mocking of OpenAI response
- Validate that the OpenAI client is called with correct parameters, including input_fidelity
- Improve response validation to ensure expected output structure and values
* Add comprehensive GitHub Copilot + LiteLLM integration tutorial
- Complete setup guide from installation to production deployment
- Multiple configuration examples including authentication, load balancing, and cost tracking
- Docker and Kubernetes deployment configurations
- Troubleshooting section with common issues and solutions
- Best practices for security, monitoring, and reliability
- Usage examples for code completion, chat interface, and direct API integration
* Add concise GitHub Copilot + LiteLLM tutorial
- Create focused tutorial matching Gemini CLI style
- Step-by-step guide from installation to production deployment
- Multi-provider configuration examples (OpenAI, Anthropic, Bedrock)
- Load balancing and fallback configuration
- Docker deployment instructions
- Troubleshooting section with common issues
- Updated sidebar with clean title 'Use LiteLLM with GitHub Copilot'
* Refactor GitHub Copilot integration tutorial
- Removed outdated production deployment and direct API usage sections
- Streamlined troubleshooting steps for clarity
- Ensured documentation aligns with current best practices and configurations
* Add proper credit to Sergio Pino for GitHub Copilot tutorial
- Reference original DEV.to article in info box
- Add credits section acknowledging foundational work
- Maintain attribution to original author's guide