* feat: integrate Google Cloud Model Armor guardrails (LIT-298)
- Add ModelArmorGuardrail class that extends CustomGuardrail and VertexBase
- Support for both pre-call (sanitizeUserPrompt) and post-call (sanitizeModelResponse) sanitization
- Integrate with existing Vertex AI authentication using VertexBase
- Add configuration model for Model Armor in guardrail types
- Register Model Armor in guardrail initializers and registry
- Include comprehensive test suite for Model Armor functionality
- Support content masking for both requests and responses
- Handle streaming responses with content sanitization
This integration allows LiteLLM to use Google Cloud Model Armor API for
content moderation and sanitization, providing similar functionality to
Bedrock Guards but using Google Cloud's security infrastructure.
* fix: remove unused imports flagged by ruff linter
- Remove unused asyncio import at top level (moved to local import where needed)
- Remove unused TextCompletionResponse import
* fix: remove additional unused imports
- Remove unused TextChoices import from line 34
- Remove duplicate asyncio import from line 384
- Replace asyncio.iscoroutine() with hasattr check for __await__
* fix: remove final unused imports
- Remove unused 'import sys' from line 9
- Remove unused 'StreamingChoices' from imports
* fix: remove unused import from model_armor.py
- Remove unused 'import os' from the top of the file
* fix: remove commented-out header from model_armor.py
- Eliminate unnecessary comments at the top of the file to improve code clarity.
* feat(guardrails): Add Model Armor UI support
- Convert model_armor.py to a directory structure with __init__.py for dynamic discovery
- Add get_config_model() method to ModelArmorGuardrail class for UI integration
- Add ui_friendly_name() to ModelArmorConfigModel returning "Google Cloud Model Armor"
- Remove manual registration from guardrail_registry.py to use dynamic discovery
- Model Armor now appears in the UI guardrails dropdown with proper configuration fields
This enables users to configure Model Armor guardrails through the LiteLLM UI interface.
* fix(guardrails): Fix undefined name 'GuardrailConfigModel' in Model Armor
- Import TYPE_CHECKING and GuardrailConfigModel from base module
- Fixes F821 linting error for undefined name in type annotation
- Follows same pattern as other guardrail implementations
* fix(guardrails): Fix Model Armor type errors and config model inheritance
- Create ModelArmorGuardrailConfigModel that properly inherits from GuardrailConfigModel base class
- Move config model to litellm/types/proxy/guardrails/guardrail_hooks/model_armor.py following convention
- Update get_config_model() to return the properly typed config model
- Remove ModelArmorConfigModel from LitellmParams inheritance chain
- Add template_id field to BaseLitellmParams instead
This fixes the mypy type errors and follows the same pattern as other guardrail implementations.
* fix(guardrails): Add missing Model Armor fields to BaseLitellmParams
- Add location, credentials, api_endpoint, and fail_on_error fields
- Fixes mypy errors about missing attributes in LitellmParams
- All Model Armor configuration parameters are now properly defined
* fix: Apply PR review feedback for Model Armor guardrail
- Move test file to tests/test_litellm/ for GitHub Actions
- Extract only last consecutive user messages to avoid context limits
- Use get_content_from_model_response helper for response extraction
- Handle non-ModelResponse types (e.g., TTS) gracefully
- Maintain newline separation for multi-part content
* refactor: Simplify message content extraction in ModelArmorGuardrail
- Removed the custom _extract_content_from_messages method.
- Integrated get_last_user_message helper for improved content extraction.
- Updated test to reflect changes in content formatting.
* refactor: Remove unused import in model_armor.py
- Deleted the unused import of AllMessageValues to clean up the codebase.
* Add unit tests for Model Armor guardrail functionality
- Implement tests for pre-call and post-call hooks, including content sanitization and blocking behavior.
- Validate error handling for API responses and credential management.
- Test streaming responses and handling of list content in user messages.
- Ensure proper assertions for API interactions and response sanitization.
* Add comprehensive test coverage for Model Armor guardrail
- Add test for requests with no messages field
- Add test for empty message content handling
- Add test for system/assistant-only messages
- Add test for fail_on_error=False behavior
- Add test for custom API endpoint configuration
- Add test for dictionary credentials (non-file path)
- Add test for action=NONE response handling
- Add test for missing sanitized_text field fallback
- Add test for non-text response types (TTS/image)
- Add test for auth token refresh behavior
All tests ensure robust edge case handling and proper error management.
* feat: improve Model Armor handling of non-ModelResponse types
- Add debug logging when skipping non-text responses (TTS, images, etc.)
- Improve docstring to clarify behavior for non-text responses
- Add test coverage for non-ModelResponse handling
- Ensure guardrail gracefully skips processing for response types it cannot handle
* build: move build_and_test to use prisma migrate
* fix(guardrails_ai.py): default to guardrail accepting 'llmOutput' as the input param
enables same guardrail to work for pre call and post call
* fix(__init__.py): set default value
* fix(guardrails_ai.py): updates
* fix: fix linting error
* fix(prompt_templates/factory.py): handle anthropic cache control on individual tool results
Fixes issue where cache control on individual tool result was being ignored
* test(test_vertex_And_google_ai_studio_gemini.py): initial unit test covering translation for grounding metadata on streaming chunk
* fix(google_genai/adapters/transformation.py): enable calling non-googlegenai models via streaming
Fixes https://github.com/BerriAI/litellm/issues/12562
* test(test_openai.py): add unit test asserting streaming works as expected
* Vllm rerank (#12737)
* Add Hosted VLLM rerank provider integration
This commit implements the Hosted VLLM rerank provider integration for LiteLLM. The integration includes:
Adding Hosted VLLM as a supported rerank provider in the main rerank function
Implementing the HostedVLLMRerank handler class for making API requests
Creating a transformation class to convert Hosted VLLM responses to LiteLLM's standardized format
The integration supports both synchronous and asynchronous rerank operations. API credentials can be provided directly or through environment variables (HOSTED_VLLM_API_KEY and HOSTED_VLLM_API_BASE).
Notable features:
Proper error handling for missing credentials
Standard response transformation
Support for common rerank parameters (top_n, return_documents, etc.)
Proper token usage tracking
This expands LiteLLM's rerank provider ecosystem to include Hosted VLLM alongside existing providers like Cohere, Together AI, Azure AI, and Bedrock.
* refactor(rerank): use base_llm_http_handler for hosted_vllm rerank
- Replace custom HostedVLLMRerank handler with base_llm_http_handler
- Implement proper HostedVLLMRerankConfig inheriting from BaseRerankConfig
- Follow Cohere-compatible implementation pattern
- Clean up unnecessary comments
* Fix lint errors in hosted_vllm rerank transformer: remove unused imports
* Fix linting errors in rerank transformation modules
* fix: resolve type errors in Hosted VLLM rerank module
---------
Co-authored-by: Philip D'Souza <philip.dsouza@macro4.com>
Co-authored-by: Philip D'Souza <philip.a.dsouza@gmail.com>
* added a few tests
---------
Co-authored-by: Philip D'Souza <philip.dsouza@macro4.com>
Co-authored-by: Philip D'Souza <philip.a.dsouza@gmail.com>
Replace MagicMock with AsyncMock for litellm_teamtable.update to fix:
TypeError: object MagicMock can't be used in 'await' expression
The test was failing because it tried to await a MagicMock object.
Added AsyncMock for the update method to properly handle async operations.
* fix bug
When Max Budget, TPM, RPM, Expire Key are set on Regenerate Key, these values are not reflected on the settings page. Refreshing the page is required.
* fix editing settings after key generation
* fix(team_info.tsx): allow setting custom key duration
more flexible than previous pre-set options
* feat(team_info.tsx): show how many user + service account keys have been created within a team
* fix(team_endpoints.py): ensure user id correctly added when new team created with user email as member
Fixes issue where user not correctly added to team on /team/new
* fix(internal_user_endpoints.py): make user email validation check case insensitive
Fixes issue where uppercase email was added even when lowercase email existed
* test: update test
* build: move build_and_test to use prisma migrate
* feat(proxy_setting_endpoints.py): encrypt env var before storing in db
Ensures env var can be read when loaded in from DB
Fixes issue when trying to add SSO from admin UI
* test: update tests
* Check content and order of trimmed messages
* Assert tool calls are preserved if below max_tokens
* Unreverse order of tool calls
* Return tool calls alongside other messages
* Write test for trimming untokenizable field
* Return original messages in case of exception
* Add concise Claude Code + LiteLLM Gateway tutorial
- Create focused tutorial matching existing tutorial style
- Step-by-step guide from installation to advanced configurations
- Multi-provider configuration examples (AWS Bedrock, Azure OpenAI, Load Balancing)
- Based on Anthropic's official LiteLLM configuration documentation
- Added to sidebar with clean title 'Use LiteLLM with Claude Code'
- Fixed sidebar reference from 'secret' to 'set_keys' for proper document resolution
* Update config_settings.md to correct documentation links for key management and Hashicorp Vault settings. Changed references from 'secret.md' to 'set_keys.md' for improved clarity and accuracy.
* Update sidebar and config_settings.md to reflect changes in key management documentation. Changed sidebar reference from 'set_keys' to 'secret' and updated links in config_settings.md for Hashicorp Vault settings to point to 'secret.md' for improved accuracy.
* Remove extra tutorial and update sidebar accordingly
* Update tutorial title from 'WebUI' to 'Open WebUI' for clarity and consistency in documentation.
* Remove Python version requirement from Claude Responses API tutorial for clarity and to align with updated prerequisites.
* feat: add input_fidelity parameter for OpenAI image generation
- Add input_fidelity to OpenAIImageGenerationOptionalParams type
- Update image_generation function signature to accept input_fidelity
- Add input_fidelity to default_params in get_optional_params_image_gen
- Include input_fidelity in openai_params list for proper handling
- Update documentation with input_fidelity parameter description
- Add test for input_fidelity parameter functionality
This enables control over how closely the model follows the input prompt
for gpt-image-1 model, improving prompt adherence and image quality.
* feat: add input_fidelity to optional parameters for image generation
- Include input_fidelity in the list of OpenAIImageGenerationOptionalParams
- This addition enhances the flexibility of image generation by allowing control over input fidelity.
* test: enhance test for gpt-image-1 with input_fidelity parameter
- Update test_gpt_image_1_with_input_fidelity to include mocking of OpenAI response
- Validate that the OpenAI client is called with correct parameters, including input_fidelity
- Improve response validation to ensure expected output structure and values