Commit graph

14665 commits

Author SHA1 Message Date
Cesar Garcia
63a97db663
feat(voyage): add rerank API support (#17744)
* feat(voyage): add rerank API support

Add support for Voyage AI rerank models (rerank-2.5, rerank-2.5-lite,
rerank-2, rerank-2-lite) to the LiteLLM rerank API.

Changes:
- Add VoyageRerankConfig transformation class
- Register voyage provider in rerank_api/main.py
- Add voyage case in utils.py get_provider_rerank_config
- Add rerank-2.5 and rerank-2.5-lite models to pricing JSON
- Add unit tests for transformation logic
- Update documentation for voyage.md and rerank.md

Usage:
```python
from litellm import rerank

response = rerank(
    model="voyage/rerank-2.5",
    query="What is the capital of France?",
    documents=["Paris is...", "London is..."],
    top_n=3,
)
```

* refactor(voyage): simplify rerank transformation code

Remove verbose docstrings to align with other providers (jina_ai pattern).
No functional changes - 168 lines vs 169 for jina_ai.

* fix(voyage): remove incorrect input_cost_per_query from rerank models

Voyage AI charges per token, not per query. The input_cost_per_query
field was incorrectly set to the same value as input_cost_per_token
in the existing rerank-2 and rerank-2-lite models.

Removes input_cost_per_query from all Voyage rerank models:
- voyage/rerank-2
- voyage/rerank-2-lite
- voyage/rerank-2.5
- voyage/rerank-2.5-lite

Pricing source: https://docs.voyageai.com/docs/pricing
2025-12-09 17:34:09 -08:00
Ishaan Jaff
3631e8fa1d
[Feat] Containers API - add new container API file management + UI Interface (#17745)
* test_router_acreate_container_without_model

* _init_containers_api_endpoints

* test_init_containers_api_endpoints

* init container files endpoints

* init files api

* init container files API

* add containers api file content

* add code interpreter output UI

* add code interpreter input ui

* refactor code interpreter ui

* fix: require model selection

* cleaner container provision

* fix ContainerFileObject

* def container_file_content_handler(
add

* add retrieve_container_file_content

* aretrieve_container_file_content

* UI fix model

* fix linting errors
2025-12-09 17:33:26 -08:00
Ishaan Jaff
142567e143
[Fix] Containers API - Allow using LIST, Create Containers using custom-llm-provider (#17740)
* test_router_acreate_container_without_model

* _init_containers_api_endpoints

* test_init_containers_api_endpoints
2025-12-09 17:00:35 -08:00
yuneng-jiang
d99cf81386 Fixing test 2025-12-09 16:11:17 -08:00
ephrimstanley
a91dda1194
Return 403 instead of 503 for unauthorized routes (#17723) 2025-12-09 15:16:11 -08:00
yuneng-jiang
bdfc3308b1
Merge pull request #17741 from BerriAI/litellm_ui_cred_fix_2
[Fix] Change credential encryption to only affect db credentials
2025-12-09 14:04:20 -08:00
yuneng-jiang
879ae45421 Change credential encryption to only affect db credentials 2025-12-09 13:36:40 -08:00
YutaSaito
80a18f989a
feat: propagate Langfuse trace_id (#17669) 2025-12-09 12:25:52 -08:00
yuneng-jiang
68419cfe4e Merge remote-tracking branch 'origin' into litellm_tag_spend_dedupe 2025-12-09 11:59:47 -08:00
yuneng-jiang
39bf7a9f7c Merge remote-tracking branch 'origin' into litellm_allow_custom_mount_paths 2025-12-09 11:58:05 -08:00
yuneng-jiang
4bffbe0bff Merge remote-tracking branch 'origin' into litellm_sso_config_2 2025-12-09 11:56:30 -08:00
yuneng-jiang
9253a9c365
Merge pull request #17689 from BerriAI/litellm_ui_settings_backend
[Feature] Get and Update Backend Routes for UI Settings
2025-12-09 11:55:28 -08:00
yuneng-jiang
e4442c2946 Change UI Settings to a dedicated table 2025-12-09 11:19:53 -08:00
Derek Duenas
3322523e07
Passthrough in response (#17102)
* attempt to implement the passthrough feature

* Formatting and small change

* Fix formatting

* feat: grayswan guardrail overwrite ModelResponse in passthrough mode

* fix missing exception error catching on certain
endpoints

* fix wrong call site

* fix: patch anthropic endpoint internal error on streaming obj

* fix grayswan testcase

* feat: update the violation response to more natural

* Formatting

* move passthrough exception definition to custom_guardrail.

* Enhancement: show whether the blocked at input or output

* update exception name

* fix a typo in testing unit.

---------

Co-authored-by: Xiaohan Fu <xiaohan@grayswan.ai>
2025-12-09 10:45:45 -08:00
yuneng-jiang
305a7c6bd5 Merge remote-tracking branch 'origin' into litellm_ui_settings_backend 2025-12-09 10:21:52 -08:00
Sameer Kankute
baed9fcea2
Merge branch 'main' into litellm_videos_bugs_2 2025-12-09 23:30:13 +05:30
Sameer Kankute
878c86e632 Add tests for handling video with litellm param 2025-12-09 23:27:20 +05:30
Sameer Kankute
237f6b991f
Merge pull request #17708 from BerriAI/litellm_videos_bug_fixes
Fix error about encoding video id for azure
2025-12-09 23:11:17 +05:30
Raghav Jhavar
face8173b0 fix failing test 2025-12-09 19:42:26 +07:00
Sameer Kankute
b39c21d90c fix: Add _delete_nested_value_custom to recursive function ignore list
The _delete_nested_value_custom function is recursive but has bounded depth
(limited by the number of path segments), preventing infinite recursion.
This is necessary for nested field removal in additional_drop_params.
2025-12-09 17:56:13 +05:30
mzagar
9bb0e7dd75 feat: Replace jsonpath-ng with custom minimal parser for additional_drop_params 2025-12-09 17:30:30 +05:30
mzagar
da2aa2ba8d feat: Add nested field removal support to additional_drop_params using JSONPath 2025-12-09 17:29:56 +05:30
Raghav Jhavar
9e85dcbd60 read responses api usage 2025-12-09 18:14:00 +07:00
Sameer Kankute
a7e5be8e03 Fix error about encoding video id for azure 2025-12-09 16:36:11 +05:30
yuneng-jiang
dab4c9d8ab Merge remote-tracking branch 'origin' into litellm_sso_config_2 2025-12-08 21:19:50 -08:00
yuneng-jiang
8039c5052f
Merge pull request #17681 from BerriAI/litellm_ui_fallback_login_alert
[Fix] Change deprecation banner to only show on /sso/key/generate
2025-12-08 21:18:38 -08:00
YutaSaito
5b926ae2c6
fix: resolve UI session MCP permissions across real teams (#17620)
* fix: resolve UI session MCP permissions across real teams

* fix: remove user from logs
2025-12-08 19:03:01 -08:00
Earl St Sauver
ad9c69860e
Fix Cerebras context window errors not recognized (#17587)
Add detection for Cerebras's context window exceeded error format:
"Current length is X while limit is Y"

This ensures LiteLLM raises ContextWindowExceededError instead of
generic BadRequestError when Cerebras API calls exceed the model's
context limit, enabling downstream libraries like DSPy to properly
catch and handle these errors for automatic context management.
2025-12-08 19:02:06 -08:00
Cesar Garcia
a7ad8a36a4
chore: cleanup unused scripts and fix misplaced test file (#17611)
Remove scripts/ directory containing unused development/debug scripts:
- mock_ibm_guardrails_server.py
- test_groq_streaming_issue.py (debug for #12660)
- test_mock_ibm_guardrails.py
- update_readme_providers_table.py

Move misplaced test file to correct location:
- test_litellm/ -> tests/test_litellm/ (from PR #17221)
2025-12-08 19:00:55 -08:00
Cesar Garcia
0295f912be
fix(openai): include 'user' param for responses API models (#17648)
The 'user' parameter was being ignored when using responses API models
(e.g., model="openai/responses/gpt-4.1") because the model name check
in get_supported_openai_params() didn't account for the "responses/" prefix.

Fix: Normalize the model name by stripping "responses/" prefix before
checking if the model is in the list of supported OpenAI models.

This is a minimal, non-breaking change that:
- Adds 2 lines of code in gpt_transformation.py
- Only affects the parameter support check, not the model variable itself
- Includes unit and integration tests
2025-12-08 18:52:47 -08:00
Cesar Garcia
fd9ff90307
fix(responses): prevent streaming tool_calls from being dropped when text + tool_calls (#17652)
When OpenAI Responses API returns both text AND tool_calls, the bridge
transformation was emitting is_finished=True after the text message completed,
causing subsequent tool_call chunks to be dropped.

The fix:
- response.output_item.done for messages no longer emits is_finished=True
- Added handler for response.completed to properly signal stream end
2025-12-08 18:51:59 -08:00
Emil Svensson
61e737e361
fix Azure AI Anthropic api-key header and passthrough cost calculation (#17656)
* refactor: remove api-key conversion logic for Azure Anthropic

Co-authored-by: Erdem Halil <erdemhalil@users.noreply.github.com>

* fix(passthrough): pass custom_llm_provider to completion_cost for Azure AI Anthropic

The passthrough logging for Anthropic was failing when using Azure AI Anthropic
because the completion_cost function was not receiving the custom_llm_provider
parameter, causing it to fail with "LLM Provider NOT provided" error.

This fix:
- Retrieves custom_llm_provider from logging_obj.model_call_details
- Prepends provider prefix to model name for cost calculation
- Passes both formatted model and custom_llm_provider to completion_cost
- Centralizes provider prefix logic in _create_anthropic_response_logging_payload

This ensures cost calculation works correctly for Azure AI Anthropic requests
with models like azure_ai/claude-sonnet-4-5_gb_20250929.

Co-authored-by: Erdem Halil <erdemhalil@users.noreply.github.com>

* test: add unit tests for Azure AI Anthropic fixes

- Add tests for custom_llm_provider cost calculation in passthrough logging
- Add tests for ProviderConfigManager returning AzureAnthropicMessagesConfig
- Update existing tests to reflect removal of api-key to x-api-key conversion

Co-authored-by: Erdem Halil <erdemhalil@users.noreply.github.com>

---------

Co-authored-by: Erdem Halil <erdemhalil@users.noreply.github.com>
2025-12-08 18:50:26 -08:00
Cesar Garcia
7c2e2111c0
fix(router): handle tools=None in filter_web_search_deployments (#17684)
Fixes #17672

Changed `request_kwargs.get("tools", [])` to `request_kwargs.get("tools") or []`
to handle the case where tools is explicitly set to None.
2025-12-08 18:36:46 -08:00
yuneng-jiang
3683614a6b Get and Update route for UI Settings 2025-12-08 18:07:00 -08:00
Ishaan Jaff
a904067d38
[Feat] New model - add bedrock writer models (#17685)
* add new bedrock models

* test bedrock writer models

* docs bedrock writer palmyra

* add palymra models

* add bedrock writer models

* docs fix
2025-12-08 17:49:06 -08:00
Ishaan Jaff
074445edb1
[Fix] AI Gateway Auth - allow using wildcard patterns for public routes (#17686)
* edit auth utils to allow wildcard patterns

* docs fix private / public routes

* test_route_in_additional_public_routes_wildcard_match
2025-12-08 17:39:53 -08:00
Ishaan Jaff
2f335ac5a6
[Feat] Dynamic Rate Limiter - allow specifying ttl for in memory cache (#17679)
* fix _get_saturation_value_from_cache

* fix _get_saturation_check_cache_ttl

* fix test_saturation_check_cache_ttl_configuration

* docs saturation_check_cache_ttl
2025-12-08 17:20:52 -08:00
yuneng-jiang
8338bd9c53 Change deprecation banner to only show on /sso/key/generate 2025-12-08 16:30:37 -08:00
Ishaan Jaff
601da4a3d1
[Feat] New model - add nvidia nim llama-3.2-nv-rerankqa-1b-v2 (#17670)
* fix get_nvidia_nim_rerank_config

* add NvidiaNimRankingConfig

* add get_nvidia_nim_rerank_config

* add test_nvidia_nim_rerank_ranking_endpoint

* add /ranking model provider support

* feat: add nvidia/llama-3.2-nv-rerankqa-1b-v2
2025-12-08 15:25:23 -08:00
yuneng-jiang
3394dbb363 Remove SSO config values from old config table on update 2025-12-08 13:06:39 -08:00
_juliettech
ee0812a297
Add Helicone as a provider and update observability documentation (#17663)
* Add Helicone as a provider to liteLLM

* Add Helicone provider integration
2025-12-08 12:34:11 -08:00
vasilisazayka
c87874c29e
[New provider] Sap gen ai hub (#16053)
* add sap gen ai hub

* add async tests

* add async and streaming support

* add embedding model support

* add embedding support

* remove unused import

* fix structured output

* clean-up

* remove timeout and add tool support

* remove unused code

* fix(sap): improve streaming robustness; restore embed URL builder compatibility
- sap/embed/transformation: add api_key and litellm_params to get_complete_url to align with core flow and prevent failures
- sap/chat/handler: wrap async/sync streaming iterators to safely handle Stop(Async)Iteration and errors
- sap/chat/transformation: remove unused imports and dead code

* fix(sap): linter fix

* fix(sap): made gen_ai_hub optional: import check + OptionalDependencyError with install hint if missing.

* test(sap): add chat/stream/async tests and OptionalDependencyError check

* Fix tool call handling in SAP GenAI Hub transformation
Add sap models to model_prices_and_context_window.json and model_prices_and_context_window_backup.json

* fix(sap): delete unnecessary code, linter fix

* fix(sap): - refactor chat transformation
- add support of list and dict content

* fix(sap): - fix tests

* fix(sap): - fix lint

* Update transformation.py

* fix(sap): fix model description and fix after rebase

* change(sap): - http calls in chat handler, response transformation and auth handling without sap sdk.

* change(sap): switching to v2 (chat handler, chat transformation), code clean up

* add deployment discovery and improved crendentials handling

* add deployment discovery and improved crendentials handling

* change(sap): - fix sync stream

* change(sap): - fix sync stream

* fix(sap): - fix response format

* fix(sap): - switch embedding to v2 and http request
- reimplement stream creator
- improve request transformation

* fix async streaming

* fix(sap): linters, transformation models, remove sap dependency test

* fix(sap): code clean up

* add unit test for sap chat completion

* linters fix

* move token, rg and base_url to properties

* (sap): add embedding unit test

Signed-off-by: Vasilisa Parshikova <vasilisa.parshikova@sap.com>

* fix(sap): bypass response format for some models

Signed-off-by: Vasilisa Parshikova <vasilisa.parshikova@sap.com>

* fix(sap): fix chat transformation and list of supported params

Signed-off-by: Vasilisa Parshikova <vasilisa.parshikova@sap.com>

* fix(sap): fix lint

* add sap service key module parameter

* fix(sap): remove unused code

* fix(sap): remove prices

* add service key support

* fix(sap): - add message content validations
- change get_supported_openai_params in chat transformation

* typo in mock

* fix(sap): - fix in supported params map

* fix(sap): - fix in message content validation

* fix(sap): - fix in message content validation

* fix(sap): - use litellm client for credentials

* fix(sap): - linter fix

* fix(sap): - use build in custom_http_client
- move credentials handling to transformation

* fix(sap): - handle stream_options

* fix(sap): - fix tests

* fix(sap): - code clean up, linter fix

* skip other authentication options when creds are provided

* fix local variable

---------

Signed-off-by: Vasilisa Parshikova <vasilisa.parshikova@sap.com>
Co-authored-by: Mathis Boerner <mathis.boerner@sap.com>
Co-authored-by: karimmohraz <37623804+karimmohraz@users.noreply.github.com>
Co-authored-by: Karim <karim.mohraz@sap.com>
2025-12-08 12:31:06 -08:00
Alexsander Hamir
958c190134
Fix flanky tests (#17665)
* Fix test_delete_polling_removes_from_cache mock setup

- Mock async_delete_cache to properly execute the real implementation path
- Ensures init_async_client() is called and delete() is invoked on the returned client
- Fixes AssertionError: Expected 'delete' to be called once. Called 0 times.

* fix: resolve timeout in add_model_tab test by mocking useProviderFields hook

- Mock useProviderFields hook to prevent network calls and React Query delays
- Use waitFor to properly handle async operations
- Test now passes reliably without 10s timeout

* fix: add test timeout to prevent CI timeout failure

- Add 15 second timeout to 'should display Test Connect and Add Model buttons' test
- Test takes ~6 seconds locally, but CI was timing out at default 5 second limit
- Ensures test has sufficient time to complete in CI environment

* test: quarantine flaky test_oidc_circleci_with_azure

Quarantine test that fails with 401 Unauthorized from Azure OAuth.
The test is flaky and blocks CI builds. Marked with @pytest.mark.skip
until Azure authentication can be fixed or migrated to our own account.
2025-12-08 12:21:26 -08:00
Sameer Kankute
05f800fe7d
Merge pull request #17653 from BerriAI/litellm_fireworks_rerank_model
(Feat) Add fireworks rerank support
2025-12-08 21:33:08 +05:30
Sameer Kankute
7aaab32313
Merge pull request #17651 from BerriAI/litellm_audio_caching_fix
Use audio content for caching
2025-12-08 20:57:09 +05:30
Sameer Kankute
76469182bc
Merge pull request #17641 from BerriAI/litellm_responses_api_usage_populated
Add usage details in responses usage object
2025-12-08 20:38:31 +05:30
Sameer Kankute
87cf6f3ffe Add fireworks rerank support 2025-12-08 20:29:50 +05:30
Sameer Kankute
f486fb2283 Use audio content for caching 2025-12-08 19:31:22 +05:30
Cesar Garcia
b6b155d67b
fix(anthropic): handle partial JSON chunks in streaming responses (#17493)
Fixes #17473 - Anthropic streaming fails with JSONDecodeError when
network fragmentation causes SSE data to arrive in partial chunks.

Changes:
- Add accumulated_json buffer and chunk_type to ModelResponseIterator
- Add _handle_accumulated_json_chunk() to accumulate partial JSON
- Add _parse_sse_data() to handle both complete and partial chunks
- Modify __next__ and __anext__ to use accumulation logic
- Add unit tests for partial chunk handling
2025-12-07 23:34:42 -08:00
Tamir Kiviti
0f5694c8eb
add onyx guardrail hooks integration (#16591)
* add onyx guardrail hooks integration

* fix lint issue

* fix lint issue

* update PR to use the new custom guardrail interface

* lint fix
2025-12-07 23:33:28 -08:00