Commit graph

28530 commits

Author SHA1 Message Date
Sameer Kankute
bcac9e41f6 Add support for computer use for gemini 2025-12-10 10:34:08 +05:30
Krish Dholakia
254c1155a2
Remove streaming_logging.md documentation (#17739)
Co-authored-by: Cursor Agent <cursoragent@cursor.com>
2025-12-09 18:26:08 -08:00
Cesar Garcia
01dec55c2f
fix(anthropic): preserve server_tool_use and web_search_tool_result in multi-turn conversations (#17746)
- Extract web_search_tool_result blocks in extract_response_content()
- Store web_search_results in provider_specific_fields for round-trip
- Detect srvtoolu_ prefix to reconstruct as server_tool_use (not tool_use)
- Add corresponding web_search_tool_result after server_tool_use blocks

This ensures multi-turn conversations with Anthropic web search + custom
tools work correctly without Anthropic expecting tool_result for server-
side tool executions.
2025-12-09 18:25:23 -08:00
Krish Dholakia
9fa6c51678
Fix: Add Gemini context window exception mapping (#17751)
Co-authored-by: Cursor Agent <cursoragent@cursor.com>
2025-12-09 18:24:13 -08:00
Krrish Dholakia
f23a78fddc test: remove expensive test 2025-12-09 18:23:03 -08:00
Krrish Dholakia
c276a87ab0 fix(anthropic/chat/transformation.py): pass output_config + thinking to claude opus 4.5 2025-12-09 18:21:41 -08:00
Sameer Kankute
8fafd81f9d
Merge pull request #17732 from BerriAI/litellm_videos_bugs_2
Fix : use litellm params for all videos apis
2025-12-10 07:49:55 +05:30
Krrish Dholakia
d3531be9a0 docs(community.md): add new integration partner doc 2025-12-09 18:17:14 -08:00
Ishaan Jaff
42f5770cfa
[docs] add docs for containers files api + code interpreter on LiteLLM (#17749)
* add new container api on OpenAI

* add related

* docs fix

* docs code interpreter

* code interp

* docs code interptert

* docs code int

* docs code interp

* docs code interp
2025-12-09 18:11:28 -08:00
Shivam Rawat
4ada6bee49
fixed flex tier pricing (#17748) 2025-12-09 18:10:51 -08:00
Krish Dholakia
8d5e6cc62d
Add community doc link (#17734)
* Add community contribution guide for integration partners

Co-authored-by: krrishdholakia <krrishdholakia@gmail.com>

* Update community docs to direct users to #integration-partners

Co-authored-by: krrishdholakia <krrishdholakia@gmail.com>

---------

Co-authored-by: Cursor Agent <cursoragent@cursor.com>
2025-12-09 18:10:00 -08:00
Cesar Garcia
63a97db663
feat(voyage): add rerank API support (#17744)
* feat(voyage): add rerank API support

Add support for Voyage AI rerank models (rerank-2.5, rerank-2.5-lite,
rerank-2, rerank-2-lite) to the LiteLLM rerank API.

Changes:
- Add VoyageRerankConfig transformation class
- Register voyage provider in rerank_api/main.py
- Add voyage case in utils.py get_provider_rerank_config
- Add rerank-2.5 and rerank-2.5-lite models to pricing JSON
- Add unit tests for transformation logic
- Update documentation for voyage.md and rerank.md

Usage:
```python
from litellm import rerank

response = rerank(
    model="voyage/rerank-2.5",
    query="What is the capital of France?",
    documents=["Paris is...", "London is..."],
    top_n=3,
)
```

* refactor(voyage): simplify rerank transformation code

Remove verbose docstrings to align with other providers (jina_ai pattern).
No functional changes - 168 lines vs 169 for jina_ai.

* fix(voyage): remove incorrect input_cost_per_query from rerank models

Voyage AI charges per token, not per query. The input_cost_per_query
field was incorrectly set to the same value as input_cost_per_token
in the existing rerank-2 and rerank-2-lite models.

Removes input_cost_per_query from all Voyage rerank models:
- voyage/rerank-2
- voyage/rerank-2-lite
- voyage/rerank-2.5
- voyage/rerank-2.5-lite

Pricing source: https://docs.voyageai.com/docs/pricing
2025-12-09 17:34:09 -08:00
Ishaan Jaff
3631e8fa1d
[Feat] Containers API - add new container API file management + UI Interface (#17745)
* test_router_acreate_container_without_model

* _init_containers_api_endpoints

* test_init_containers_api_endpoints

* init container files endpoints

* init files api

* init container files API

* add containers api file content

* add code interpreter output UI

* add code interpreter input ui

* refactor code interpreter ui

* fix: require model selection

* cleaner container provision

* fix ContainerFileObject

* def container_file_content_handler(
add

* add retrieve_container_file_content

* aretrieve_container_file_content

* UI fix model

* fix linting errors
2025-12-09 17:33:26 -08:00
Ishaan Jaff
142567e143
[Fix] Containers API - Allow using LIST, Create Containers using custom-llm-provider (#17740)
* test_router_acreate_container_without_model

* _init_containers_api_endpoints

* test_init_containers_api_endpoints
2025-12-09 17:00:35 -08:00
yuneng-jiang
9a7831dd33
Merge pull request #17697 from BerriAI/litellm_ui_settings_ui
[Feature] UI - UI Settings
2025-12-09 15:58:08 -08:00
ephrimstanley
a91dda1194
Return 403 instead of 503 for unauthorized routes (#17723) 2025-12-09 15:16:11 -08:00
Krrish Dholakia
cc8a46b6b2 bump: version 1.80.9 → 1.80.10 2025-12-09 14:55:43 -08:00
yuneng-jiang
bdfc3308b1
Merge pull request #17741 from BerriAI/litellm_ui_cred_fix_2
[Fix] Change credential encryption to only affect db credentials
2025-12-09 14:04:20 -08:00
Alexsander Hamir
9a0432d013
[Fix] Perf - Reduce memory accumulation of spend_logs (#17742)
Replaces time-based spend log processing with a queue-size-based approach to improve performance and reduce memory usage. The new implementation:

- Adds background task that monitors spend_log_transactions queue size
- Triggers processing when queue reaches configurable threshold (default: 100)
- Implements exponential backoff when queue is idle to reduce CPU usage
- Increases batch processing limits (BATCH_SIZE: 100→1000, MAX_LOGS_PER_INTERVAL: 1000→10000)
- Adds configurable environment variables: SPEND_LOG_QUEUE_SIZE_THRESHOLD, SPEND_LOG_QUEUE_POLL_INTERVAL

This approach prevents unbounded queue growth while reducing unnecessary polling overhead.
2025-12-09 14:00:34 -08:00
yuneng-jiang
879ae45421 Change credential encryption to only affect db credentials 2025-12-09 13:36:40 -08:00
YutaSaito
80a18f989a
feat: propagate Langfuse trace_id (#17669) 2025-12-09 12:25:52 -08:00
yuneng-jiang
9253a9c365
Merge pull request #17689 from BerriAI/litellm_ui_settings_backend
[Feature] Get and Update Backend Routes for UI Settings
2025-12-09 11:55:28 -08:00
Shivam Rawat
43a7bbeeaf
added note for using Azure Active Directory Tokens with all the other endpoints (#17733) 2025-12-09 11:51:28 -08:00
yuneng-jiang
aa450e7ebe
Merge pull request #17738 from BerriAI/litellm_doc_update_1805
[Docs] Adding known issues to 1.80.5-stable docs
2025-12-09 11:46:08 -08:00
yuneng-jiang
431884f591 Adding known issues to 1.80.5-stable docs 2025-12-09 11:45:16 -08:00
yuneng-jiang
1b86269584 Bump version to include new schema 2025-12-09 11:25:24 -08:00
yuneng-jiang
7957244367 bump: version 0.4.11 → 0.4.12 2025-12-09 11:24:18 -08:00
yuneng-jiang
36183c3a9b Adding migration 2025-12-09 11:23:26 -08:00
yuneng-jiang
e4442c2946 Change UI Settings to a dedicated table 2025-12-09 11:19:53 -08:00
Derek Duenas
3322523e07
Passthrough in response (#17102)
* attempt to implement the passthrough feature

* Formatting and small change

* Fix formatting

* feat: grayswan guardrail overwrite ModelResponse in passthrough mode

* fix missing exception error catching on certain
endpoints

* fix wrong call site

* fix: patch anthropic endpoint internal error on streaming obj

* fix grayswan testcase

* feat: update the violation response to more natural

* Formatting

* move passthrough exception definition to custom_guardrail.

* Enhancement: show whether the blocked at input or output

* update exception name

* fix a typo in testing unit.

---------

Co-authored-by: Xiaohan Fu <xiaohan@grayswan.ai>
2025-12-09 10:45:45 -08:00
yuneng-jiang
305a7c6bd5 Merge remote-tracking branch 'origin' into litellm_ui_settings_backend 2025-12-09 10:21:52 -08:00
yuneng-jiang
486bcbe4e8 Revert "UI Settings Frontend"
This reverts commit 51275e7014.
2025-12-09 10:21:30 -08:00
Sameer Kankute
baed9fcea2
Merge branch 'main' into litellm_videos_bugs_2 2025-12-09 23:30:13 +05:30
Sameer Kankute
878c86e632 Add tests for handling video with litellm param 2025-12-09 23:27:20 +05:30
Sameer Kankute
237f6b991f
Merge pull request #17708 from BerriAI/litellm_videos_bug_fixes
Fix error about encoding video id for azure
2025-12-09 23:11:17 +05:30
Sameer Kankute
320f861916 Fix : use litellm params for other video apis 2025-12-09 23:03:29 +05:30
Krish Dholakia
81f0bbad73
Add Azure AI Search to supported vector stores (#17726)
Co-authored-by: Cursor Agent <cursoragent@cursor.com>
2025-12-09 09:04:04 -08:00
Sameer Kankute
a7e5be8e03 Fix error about encoding video id for azure 2025-12-09 16:36:11 +05:30
yuneng-jiang
8039c5052f
Merge pull request #17681 from BerriAI/litellm_ui_fallback_login_alert
[Fix] Change deprecation banner to only show on /sso/key/generate
2025-12-08 21:18:38 -08:00
yuneng-jiang
0a6038d4c5 UI Settings Frontend 2025-12-08 21:12:43 -08:00
yuneng-jiang
51275e7014 UI Settings Frontend 2025-12-08 21:12:06 -08:00
YutaSaito
5b926ae2c6
fix: resolve UI session MCP permissions across real teams (#17620)
* fix: resolve UI session MCP permissions across real teams

* fix: remove user from logs
2025-12-08 19:03:01 -08:00
chenzhaofei01
458032b083
fix dashscope default api_base error (#17584) 2025-12-08 19:02:55 -08:00
Earl St Sauver
ad9c69860e
Fix Cerebras context window errors not recognized (#17587)
Add detection for Cerebras's context window exceeded error format:
"Current length is X while limit is Y"

This ensures LiteLLM raises ContextWindowExceededError instead of
generic BadRequestError when Cerebras API calls exceed the model's
context limit, enabling downstream libraries like DSPy to properly
catch and handle these errors for automatic context management.
2025-12-08 19:02:06 -08:00
Cesar Garcia
a7ad8a36a4
chore: cleanup unused scripts and fix misplaced test file (#17611)
Remove scripts/ directory containing unused development/debug scripts:
- mock_ibm_guardrails_server.py
- test_groq_streaming_issue.py (debug for #12660)
- test_mock_ibm_guardrails.py
- update_readme_providers_table.py

Move misplaced test file to correct location:
- test_litellm/ -> tests/test_litellm/ (from PR #17221)
2025-12-08 19:00:55 -08:00
Ishaan Jaff
b673177b22
Add 227 new Fireworks AI models (#17692) 2025-12-08 18:57:43 -08:00
Sergei Silnov
f5dc8b7c38
fix: Use python instead of wget for healthcheck in docker-compose.yml (#17646)
Fixes #17645
2025-12-08 18:54:54 -08:00
Chetan Choudhary
38eda3409a
docs: Add SumoLogic integration documentation (#17647)
* docs: Add SumoLogic integration documentation

* minor update
2025-12-08 18:54:07 -08:00
Cesar Garcia
0295f912be
fix(openai): include 'user' param for responses API models (#17648)
The 'user' parameter was being ignored when using responses API models
(e.g., model="openai/responses/gpt-4.1") because the model name check
in get_supported_openai_params() didn't account for the "responses/" prefix.

Fix: Normalize the model name by stripping "responses/" prefix before
checking if the model is in the list of supported OpenAI models.

This is a minimal, non-breaking change that:
- Adds 2 lines of code in gpt_transformation.py
- Only affects the parameter support check, not the model variable itself
- Includes unit and integration tests
2025-12-08 18:52:47 -08:00
Cesar Garcia
fd9ff90307
fix(responses): prevent streaming tool_calls from being dropped when text + tool_calls (#17652)
When OpenAI Responses API returns both text AND tool_calls, the bridge
transformation was emitting is_finished=True after the text message completed,
causing subsequent tool_call chunks to be dropped.

The fix:
- response.output_item.done for messages no longer emits is_finished=True
- Added handler for response.completed to properly signal stream end
2025-12-08 18:51:59 -08:00