When using Model Armor guardrail with explicit project_id in config,
the project_id was being overwritten to None due to incorrect
initialization order between ModelArmorGuardrail and VertexBase parent class.
This fix ensures that user-provided project_id is preserved by initializing
parent classes before setting instance attributes.
Fixes#12757
* feat: initial commit for forwarding client headers by model group
* fix(router.py): support new forwarclientsideheadersbymodelgroup class
enables headers to be forwarded to backend model, by model group
* fix(proxy_server.py): load in model group settings from config correctly
* refactor(litellm_pre_call_utils.py): litellm_pre_call_utils.py
introduce new 'secret_fields' field
includes raw request headers (not the sanitized ones used for logging) - needed to support forwarding clientside headers to llm api
* feat(router.py): log the deployment model name as well
allows wildcard models to support forward_client_headers_to_llm_api
* test(test_router.py): add more unit testing
* feat(router.py): specify the model group alias in metadata kwargs
allows usage for internal routing logic
* fix: fix ruff check errors
* fix(router.py): refactor to cleanup optional pre-call checks
* fix: fix ruff check
* test: add missing unit test
* fix(add_model_modes.tsx): add 'batch' mode to ui
* fix(main.py): support health checks on batches + support litellm_credentials on batches
* fix(add_model_tab.tsx): clarify what 'team' on add model means
- Changed groq/moonshotai-kimi-k2-instruct to groq/moonshotai/kimi-k2-instruct in model_prices_and_context_window.json
- Added groq/moonshotai/kimi-k2-instruct and groq/qwen-qwq-32b to the supported models table in Groq documentation
* fix(prompt_templates/factory.py): handle anthropic cache control on individual tool results
Fixes issue where cache control on individual tool result was being ignored
* test(test_vertex_And_google_ai_studio_gemini.py): initial unit test covering translation for grounding metadata on streaming chunk
* fix(vertex_and_google_ai_studio.py): ensure grounding metadata is preserved on streaming
Closes https://github.com/BerriAI/litellm/issues/10237
* fix(core_helpers.py): include usage in expected openai keys
* feat(bulk_edit_user.tsx): initial working ui for editing users in bulk on the ui
easier to give access / assign to a default team
* feat(team_endpoints-+-bulk_edit_users.tsx): add bulk adding users to teams
make it easier to add existing users to a default team
* fix(bulk_edit_user.tsx): fix ui linting error
* fix: fix linting error
* feat: add v0 provider support to LiteLLM
- Add v0 as a new OpenAI-compatible provider
- Support all three v0 models: v0-1.0-md, v0-1.5-md, v0-1.5-lg
- Configure correct token limits and pricing for each model
- Enable vision support for all v0 models (multimodal)
- Add provider detection for v0/ prefix and api.v0.dev endpoint
- Include comprehensive unit tests for the provider
The v0 provider uses the standard OpenAI-compatible implementation
and supports all standard features including streaming, function
calling, and system messages.
* fix: add v0 provider to ProviderConfigManager
Add V0ChatConfig to the get_provider_chat_config method to fix
test_supports_tool_choice test failure. The v0 provider needs to
be included in the provider config manager to return the correct
configuration for tool choice support detection.
* docs: add documentation for v0 provider
- Add comprehensive v0 provider documentation
- Cover all supported models and their capabilities
- Include examples for SDK usage, proxy configuration, and all features
- Document supported OpenAI parameters based on v0 API docs
- Add v0 to the providers sidebar navigation
* fix: correct v0 supported OpenAI parameters
Based on review feedback and v0 API documentation:
- v0 only supports: messages, model, stream, tools, tool_choice
- Remove unsupported parameters like temperature, max_tokens, etc.
- Update tests to verify correct parameter set
- Update documentation to reflect actual API capabilities
- Remove JSON mode example as response_format is not supported
Reference: https://v0.dev/docs/v0-model-api#request-body
* fix: remove supports_response_schema from v0 models
Remove the supports_response_schema property from all v0 models in the model configuration files as v0 does not support this feature.
Models updated:
- v0/v0-1.0-md
- v0/v0-1.5-md
- v0/v0-1.5-lg
* fix(lowest_latency.py): Handle ZeroDivisionError with zero completion tokens (#12641)
Fixes ZeroDivisionError when LLM responses have zero completion tokens, which can
occur with Gemini models on very long contexts that only use tool calls.
Changes:
- Add check for completion_tokens > 0 before division in log_success_event
- Handle both timedelta and float types for response times (supporting both time.time() and datetime)
- Apply fix to both occurrences of the division operation in the file
- Add comprehensive tests for zero completion token scenarios
This ensures the lowest latency routing continues to work properly even when
models return responses with no completion tokens.
* test: Move lowest_latency zero tokens test to test_litellm for CI execution
* refactor: Use safe_divide helper to eliminate code duplication
- Added safe_divide utility function to litellm_core_utils.core_helpers
- Handles both timedelta and float types for numerator
- Prevents ZeroDivisionError and negative denominator issues
- Replaced duplicated division logic in lowest_latency.py
- Added comprehensive unit tests for safe_divide function
This improves code quality by reducing duplication and centralizing the
division safety logic in a reusable helper function.
* refactor: Rename to safe_divide_seconds for clarity
- Renamed safe_divide to safe_divide_seconds to better indicate it handles time durations
- Simplified implementation with single-line ternary for seconds conversion
- Updated all references and tests accordingly
The function name now clearly indicates it's specifically for dividing
time durations (in seconds) by a denominator.
* refactor: Simplify safe_divide_seconds to only accept float arguments
- Changed safe_divide_seconds to accept only float parameters for consistency
- Callers now handle timedelta to seconds conversion explicitly
- Removed timedelta-specific test cases
* fix: Remove unused timedelta import
* fix: Remove Union[timedelta, float] type annotations
Since safe_divide_seconds now only accepts floats, we handle the
timedelta conversion explicitly in the code. The type annotations
are no longer needed and can be simplified.
* Revert "fix: Remove Union[timedelta, float] type annotations"
This reverts commit 19c6c62154.
* fix: Clean up implementation
- Remove unnecessary type annotations
- Use safe_divide_seconds utility for zero-division protection
- Handle both timedelta and float types explicitly
- Maintain compatibility with both datetime and time.time() usage