* fix(proxy/_types.py): add missing comma for `/v2/rerank`
Enables non admins to access `/v2/rerank` endpoint
* fix(proxy_track_cost_callback.py): add patch to handle scenario where both 'litellm_metadata' and 'metadata' exist
* ui fix bedrock guard
* polish: logo should appear after selecting provider
* fix ui config bedrock
* fix: refactor - use specific configs per provider
* fix: refactor - use specific configs per provider
* feat: ui, show provider specific params for guardrails
* fix: updated type of LiteLLM params for guardrails
* fix: updated type of LiteLLM params for guardrails
* ui, use endpoint for adding presidio, bedrock guardrails
* fix: linting error
* add llama guard and secret detector on UI
* add aim on ui
* allow adding lakera AI on litellm ui
* fix: fixes for params to init guardrails
* test: test_guardrail_info_response
* test: test_initialize_presidio_guardrail
* fix: init guardrails
* fix: init guardrails
* add showSearch
* working bedrock guard
* Add --only-models-matching-regex option
to `models import` which only processes models where
`litelllm_params.model` matches the regex
* Add test_models_import_only_models_matching_regex
* Print each model we're importing
* Add --only-access-groups-matching-regex option
to `models import` which only processes models where at least one item
in `model_info.access_groups` matches the regex. Add a unit test.
* Add `models import` examples to README.md
Add `models import` examples to proxy/client/cli/README.md
* ruff format litellm/proxy/client/cli/commands/models.py
* Make `models import` display tabular output
* models import refactoring
* Fix failing tests in test_models_commands.py
* Refactor import_models to make it shorter and more readable
* Extract from `import_models` a function called `get_model_list_from_yaml_file`
* Fix mypy error
* Add more specific typing
for better understandability and Intellisense
* More import_models refactoring
* More refactoring
* More refactoring
* Write unit tests for format_iso_datetime_str
* Add more unit tests
* ruff format tests/litellm/proxy/client/cli/test_models_commands.py
* ruff format litellm/proxy/client/cli/commands/models.py
* Make test_format_timestamp use UTC time
* fix(embeddings): use non default tokenizer when passing list of lists of tokens (int)
* feat(embeddings): allow for passthrough of list of lists of tokens to hosted_vllm models
* Revert "fix(embeddings): use non default tokenizer when passing list of lists of tokens (int)"
This reverts commit a48acd95f8.
* refactor(embeddings): use a list to verify if provider accept as input a list of tokens
* fix(embeddings): verify the model name before validating if provider accept a arrays of tokens as input
When passing a list of tokens as input, verify the provider of the model by going through the list of models (`llm_model_list`). First, it check for model name then get the provider and verify if it accept or not arrays of tokens. If yes, then pass, else decode.
Previously, it was verifying provider and model name at the same time resulting in decoding even if the current model checked was not the target one (looping onto `llm_model_list`)
* test(embedding): add unit test to bypass decode for some providers with input as array of tokens
Ref: https://github.com/BerriAI/litellm/issues/10113
* fix(duration_parser.py): support `mo` unit
* test(test_key_management_endpoints.py): add test confirming generate_key_helper_fn uses predictable budgets
Closes https://github.com/BerriAI/litellm/issues/10800
* fix(anthropic/chat/transformation.py): add tool use cost tracking
* fix(anthropic/): refactor how hosted tool usage tracking is done
keep it separate from prompt / completion token details
* fix(anthropic/): add web search tool cost tracking
accurate cost tracking
* feat(anthropic/chat/transformation.py): map openai 'web_search_options' param to anthropic hosted tool
Allows calling anthropic web search in same format as openai
* feat(anthropic/chat/transformation.py): support unified anthropic 'web_search_options' param
Allows calling anthropic's web search tool in the openai format
* feat(anthropic/chat/transformation.py): map openai 'search_context_size' to anthropic 'max_uses' param
Translate search effort across both providers
* fix: mark web_search_options param as supported by openai + azure
* fix: fix linting error
* fix: fix linting errors
* fix: fix linting error
* fix: check if usage hasattr
* fix: pass web search options param
* add function to check config flag
* added unit tests
* convert to seconds support
* added in settings.md
* Updated config_settings.md
* remove extra point
* change config var
* resolve conflict
* feat(cohere/embed): v2 embed api support
adds output_dimensions param support
* fix(cohere/embed): migrate to v2 embedding
Adds output dimension support
* fix: maintain /v1/embedding compatibility for bedrock cohere
Bedrock cohere is still using /v1 endpoints
* fix: fix linting error
* fix: fix passing extra headers
* test: update tests
* fix(litellm_logging.py): log custom headers in requester metadata
allows passing along custom headers from client to logging integration - e.g. `x-correlation-id`
* refactor: move enterprise code out of OSS package
work towards simplified CE version of docker image
* test: update test
* fix: fix linting error
* add test_function_calling_with_tool_response to base llm tests
* run test suite for nova
* update test_function_calling_with_tool_response
* allowed ToolJsonSchemaBlock keys
* fix ToolJsonSchemaBlock
* add back pytest fixture
* test: test_prompt_caching
* fix: bump: DEFAULT_MAX_RECURSE_DEPTH
* fix: bump: DEFAULT_MAX_RECURSE_DEPTH
* test: test_vertex_ai_complex_response_schema
* fix: allow all constants to be overriden
* fix: allow all numeric constants to be overriden with env vars
* fix: remove dup DEFAULT_MAX_TOKENS in constants.py
* document all constants env vars
* docs - DEFAULT_PROMPT_INJECTION_SIMILARITY_THRESHOLD
* fix(main.py): use base model instead of user model if given
Fixes https://github.com/BerriAI/litellm/issues/10760
* feat(azure/image_generation/__init__.py): make azure image gen check more robust
Fixes https://github.com/BerriAI/litellm/issues/10760
* fix(user_api_key_auth.py): support bearer token auth for `x-litellm-api-key` header
Fixes earlier regression on vertex ai passthrough auth
* fix(user_api_key_auth.py): refactor get api key into separate function
enables easier testing
* fix: cleanup
* fix: fix linting error
* fix: cleanup
* test: update tests
* Add new model provider Novita AI (#7582)
* feat: add new model provider Novita AI
* feat: use deepseek r1 model for examples in Novita AI docs
* fix: fix tests
* fix: fix tests for novita
* fix: fix novita transformation
* ci: fix ci yaml
* fix: fix novita transformation and test (#10056)
---------
Co-authored-by: Jason <ggbbddjm@gmail.com>
* fix(factory.py): Add reasoning content handling for missing assistant content
* fix(factory.py): Improve handling of thinking blocks for assistant content
* test(factory.py): Add test for Bedrock processing of thinking blocks with None content
* Fixed Json.dumps in JSON Schema Validation Error
* Added Response Schema to Ollama chat for structured response
* Added Test cases
* refactor(ollama): remove redundant response_format check
The response_format parameter conversion is already handled in utils.py's
get_optional_params function, making the duplicate check in ollama_chat.py
unnecessary. This change removes the redundant code while maintaining the
same functionality.
* Support pdf url's to openai (#10640)
* fix(gpt_transformation.py): support pdf url input to openai
pass as base64 as openai doesn't support image url's
* fix(openai.py): support async message transformation
allows async get request to convert url to base64
* fix(gpt_transformation.py): fix linting errrors and use common components across sync + async flows
* fix: fix linting errors
* fix(openai.py): pop correct var
* Fix sagemaker chat calls - content length error (#10607)
* fix(sagemaker_chat/): support passing dynamic aws params
previously being ignored
* refactor(sagemaker/chat): more refactoring
* fix(sagemaker_chat/): make sure streaming is correctly handled post-refactor
* refactor: more refactoring to support using signed json str
* fix(sagemaker/chat): working sync streaming post refactor
* fix(sagemaker/chat): support async streaming post refactor
* fix(llm_http_handler.py): await async function
* fix: remove print statements
* test: update test
* test: update test
* fix(llm_http_handler.py): retain passing in data as json str
* test: update test
* fix(base_model_iterator.py): fix linting error
* test: test auth
* fix: fix linting error
* test: update test
* test: update translation test
* fix(gpt_transformation.py): handle awaitable/non-awaitable object
* fix: handle async flow for message transformation on openai compatible api's
* test: cleanup testing
* test: update test
* test(test_router.py): use model with higher quota
* test: simplify test
* test: update test