Commit graph

31779 commits

Author SHA1 Message Date
Aaron Yim
d4031c8ba6
Add OpenRouter Kimi K2.5 (#19872)
Co-authored-by: Claude Opus 4.5 <noreply@anthropic.com>
2026-01-28 22:34:48 -08:00
Christopher Chase
87bdfb0253
fix(hosted_vllm): route through base_llm_http_handler to support ssl_verify (#19893)
* fix(hosted_vllm): route through base_llm_http_handler to support ssl_verify

The hosted_vllm provider was falling through to the OpenAI catch-all path
which doesn't pass ssl_verify to the HTTP client. This adds an explicit
elif branch that routes hosted_vllm through base_llm_http_handler.completion()
which properly passes ssl_verify to the httpx client.

- Add explicit hosted_vllm branch in main.py completion()
- Add ssl_verify tests for sync and async completion
- Update existing audio_url test to mock httpx instead of OpenAI client

* feat(hosted_vllm): add embedding support with ssl_verify

- Add HostedVLLMEmbeddingConfig for embedding transformations
- Register hosted_vllm embedding config in utils.py
- Add lazy import for embedding transformation module
- Add unit test for ssl_verify parameter handling
2026-01-28 22:33:07 -08:00
Bernardo Donadio
ba17f51812
fix(proxy): prevent provider-prefixed model leaks (#19943)
* fix(proxy): prevent provider-prefixed model leaks

Proxy clients should not see LiteLLM internal provider prefixes (e.g. hosted_vllm/...) in the OpenAI-compatible response model field.

This patch sanitizes the client-facing model name for both:
- Non-streaming responses returned from base_process_llm_request
- Streaming SSE chunks emitted by async_data_generator

Adds regression tests covering vLLM-style hosted_vllm routing for both streaming and non-streaming paths.

* chore(lint): suppress PLR0915 in proxy handler

Ruff started flagging ProxyBaseLLMRequestProcessing.base_process_llm_request() for too many statements after the hotpatch changes.

Add an explicit '# noqa: PLR0915' on the function definition to avoid a large refactor in a hotpatch.

* refactor(proxy): make model restamp explicit

Replace silent try/except/pass and type ignores with explicit model restamping.

- Logs an error when the downstream response model differs from the client-requested model
- Overwrites the OpenAI `model` field to the client-requested value to avoid leaking internal provider-prefixed identifiers
- Applies the same behavior to streaming chunks, logging the mismatch only once per stream

* chore(lint): drop PLR0915 suppression

The model restamping bugfix made `base_process_llm_request()` slightly exceed Ruff's
PLR0915 (too-many-statements) threshold, requiring a `# noqa` suppression.

Collapse consecutive `hidden_params` extractions into tuple unpacking so the
function falls back under the lint limit and remove the suppression.

No functional change intended; this keeps the proxy model-field bugfix intact
while aligning with project linting rules.

* chore(proxy): log model mismatches as warnings

These model-restamping logs are intentionally verbose: a mismatch is a useful signal
that an internal provider/deployment identifier may be leaking into the public
OpenAI response `model` field.

- Downgrade model mismatch logs from error -> warning
- Keep error logs only for cases where the proxy cannot read/override the model

* fix(proxy): preserve client model for streaming aliasing

Pre-call processing can rewrite request_data['model'] via model alias maps.\n\nOur streaming SSE generator was using the rewritten value when restamping chunk.model, which caused the public 'model' field to differ between streaming and non-streaming responses for alias-based requests.\n\nStash the original client model in request_data as _litellm_client_requested_model after the model has been routed, and prefer it when overriding the outgoing chunk model. Add a regression test for the alias-mapping case.

* chore(lint): satisfy PLR0915 in streaming generator

Ruff started flagging async_data_generator() for too many statements after adding model restamping logic.\n\nExtract the client-model selection + chunk restamping into small helpers to keep behavior unchanged while meeting the project's PLR0915 threshold.
2026-01-28 22:26:38 -08:00
michelligabriele
dcf5f07e5e
fix(proxy): add datadog_llm_observability to /health/services allowed list (#19952)
The /health/services endpoint rejected datadog_llm_observability as an
unknown service, even though it was registered in the core callback
registry and __init__.py. Added it to both the Literal type hint and
the hardcoded validation list in the health endpoint.
2026-01-28 22:16:27 -08:00
Sameer Kankute
654edbd15c Fix Only allowed to call routes: ['llm_api_routes']. Tried to call route: /batches/bGl0ZWxsbV9wcm/cancel 2026-01-29 11:42:49 +05:30
Sameer Kankute
70684ca86f Fix File access permissions for .retreive and .delete 2026-01-29 11:19:24 +05:30
Cesar Garcia
2a48d12507
fix(docker): add libsndfile to main Dockerfile for ARM64 audio processing (#19776)
Fixes #16920 for users of the stable release images.

The previous fix (PR #18092) added libsndfile to docker/Dockerfile.alpine,
but stable releases are built from the main Dockerfile (Wolfi-based),
not the Alpine variant.
2026-01-28 21:33:41 -08:00
Cesar Garcia
c7453c01f9
Fix stream_chunk_builder to preserve images from streaming chunks (#19654)
Fixes #19478

The stream_chunk_builder function was not handling image chunks from
models like gemini-2.5-flash-image. When streaming responses were
reconstructed (e.g., for caching), images in delta.images were lost.

This adds handling for image_chunks similar to how audio, annotations,
and other delta fields are handled.
2026-01-28 21:31:06 -08:00
Sameer Kankute
833cf6a2cf Fix: Batch cancellation ownership bug 2026-01-29 10:54:42 +05:30
yuneng-jiang
f2d2ed5a0d
Merge pull request #19953 from BerriAI/litellm_key_alias_spend_usage_report
[Feature] UI - Usage Export: Breakdown by Teams and Keys
2026-01-28 20:29:22 -08:00
yuneng-jiang
507f4c45a0
Merge pull request #19976 from BerriAI/ui_build_yj_2
[Infra] Remove _experimental/out routes from gitignore + UI Build
2026-01-28 20:15:06 -08:00
yuneng-jiang
12a4d14980 chore: update Next.js build artifacts (2026-01-29 04:12 UTC, node v22.16.0) 2026-01-28 20:12:20 -08:00
yuneng-jiang
20bab33e36 removing _experimental out routes from gitignore 2026-01-28 20:11:35 -08:00
Harshit Jain
8e2fa7969c
Fix/router search tools v2 (#19840)
* fix(proxy_server): pass search_tools to Router during DB-triggered initialization

* fix search tools from db

* add missing statement to handle from db

* fix import issues to pass lint errors
2026-01-28 19:45:35 -08:00
Cesar Garcia
8a26033a4b
fix(vertex_ai): convert image URLs to base64 in tool messages for Anthropic (#19896)
* fix(vertex_ai): convert image URLs to base64 in tool messages for Anthropic

Fixes #19891

Vertex AI Anthropic models don't support URL sources for images. LiteLLM
already converted image URLs to base64 for user messages, but not for tool
messages (role='tool'). This caused errors when using ToolOutputImage with
image_url in tool outputs.

Changes:
- Add force_base64 parameter to convert_to_anthropic_tool_result()
- Pass force_base64 to create_anthropic_image_param() for tool message images
- Calculate force_base64 in anthropic_messages_pt() based on llm_provider
- Add unit tests for tool message image handling

* chore: remove extra comment from test file header
2026-01-28 19:42:51 -08:00
Sameer Kankute
2a1bfd39aa
Merge pull request #19974 from BerriAI/litellm_model_map_fix_jan_29
fix gemini gemini-robotics-er-1.5-preview entry
2026-01-29 09:07:54 +05:30
Sameer Kankute
be8a76f270 fix gemini gemini-robotics-er-1.5-preview entry 2026-01-29 09:06:44 +05:30
rushilchugh01
562f0a0282
feat: Add new OpenRouter models: xiaomi/mimo-v2-flash, z-ai/glm-4.7, z-ai/glm-4.7-flash, and minimax/minimax-m2.1. to model prices and context window (#19938)
Co-authored-by: Rushil Chugh <Rushil>
2026-01-28 18:56:20 -08:00
Ishaan Jaff
9c5fed4f52
[Feat] LiteLLM Vector Stores - Add permission management for users, teams (#19972)
* fix: create_vector_store_in_db

* add team/user to LiteLLM_ManagedVectorStore

* add _check_vector_store_access

* add new fields

* test_check_vector_store_access

* add vector_store/list endpoints

* fix code QA checks
2026-01-28 18:55:40 -08:00
yuneng-jiang
e796b9eb22
Merge pull request #19963 from BerriAI/litellm_ui_spend_logs_em_search
[Feature] UI - Logs: Adding Error message search to ui spend logs
2026-01-28 18:15:35 -08:00
yuneng-jiang
58dd3bd134 fixing sorting for v2/model/info 2026-01-28 18:07:22 -08:00
yuneng-jiang
632e8cf2f6
Merge pull request #19970 from BerriAI/litellm_ui_column_sort_component
[Feature] UI - Tables: Reusable Table Sort Component
2026-01-28 18:02:34 -08:00
Ishaan Jaff
dcca8c7350
[Feat] - Search API add /list endpoint to list what search tools exist in router (#19969)
* feat: List all available search tools configured in the router.

* add debugging search API

* add debugging search API
2026-01-28 17:58:17 -08:00
Alexsander Hamir
69bd4426e8
[Release Day] - Fixed CI/CD issues & changed processes (#19902) 2026-01-28 17:57:24 -08:00
yuneng-jiang
e9056671f9 Fixing sorting API calls 2026-01-28 17:53:38 -08:00
yuneng-jiang
92f9d8f86e Reusable Table Sort Component 2026-01-28 17:42:57 -08:00
Neha Prasad
a785eecf7a
fix Prompt Studio history to load tools and system messages (#19920) 2026-01-28 17:19:59 -08:00
Ishaan Jaff
d12ce3cd5d
[Fix] VertexAI Pass through - fix regression that caused vertex ai passthroughs to stop working for router models (#19967)
* fix(vertex_ai): replace custom model names with actual Vertex AI model names in passthrough URLs (#19948)

When the passthrough URL already contains project and location, the code
was skipping the deployment lookup and forwarding the URL as-is to Vertex AI.
For custom model names like gcp/google/gemini-2.5-flash, Vertex AI returned
404 because it only knows the actual model name (gemini-2.5-flash).

The fix makes the deployment lookup always run, so the custom model name
gets replaced with the actual Vertex AI model name before forwarding.

* add _resolve_vertex_model_from_router

* fix: get_llm_provider

* Potential fix for code scanning alert no. 4020: Clear-text logging of sensitive information

Co-authored-by: Copilot Autofix powered by AI <62310815+github-advanced-security[bot]@users.noreply.github.com>

---------

Co-authored-by: michelligabriele <gabriele.michelli@icloud.com>
Co-authored-by: Copilot Autofix powered by AI <62310815+github-advanced-security[bot]@users.noreply.github.com>
2026-01-28 16:54:01 -08:00
Alexsander Hamir
3816570313
fix(presidio): reuse HTTP connections to prevent OOMs (#19964) 2026-01-28 16:08:53 -08:00
Ishaan Jaffer
c4daa39daa fix 2026-01-28 16:05:03 -08:00
yuneng-jiang
cccda30a9e
Merge pull request #19960 from BerriAI/litellm_ui_spend_logs_error_message
[Feature] Add error_message Search in Spend Logs Endpoint
2026-01-28 16:04:29 -08:00
yuneng-jiang
0cdfa8e5fa Adding Error message search to ui spend logs 2026-01-28 15:39:32 -08:00
yuneng-jiang
cb8ead6013 Add error_message search in spend logs endpoint 2026-01-28 15:06:31 -08:00
Ishaan Jaff
3ef475b70e
[Fix] A2a Gateway - Allow supporting old A2a card formats (#19949)
* fix: LiteLLMA2ACardResolver

* fix: LiteLLMA2ACardResolver

* feat: .well-known/agent.json

* test_card_resolver_fallback_from_new_to_old_path
2026-01-28 15:02:08 -08:00
yuneng-jiang
054918e7a3
Merge pull request #19918 from BerriAI/litellm_ui_spend_logs_store
[Feature] UI - Spend Logs: Settings Modal
2026-01-28 15:00:31 -08:00
Ishaan Jaffer
5135efb60e fix pypdf: >=6.6.2 2026-01-28 14:54:58 -08:00
yuneng-jiang
dbd1ff306d Fixing build 2026-01-28 13:42:58 -08:00
yuneng-jiang
077cfa8c15 Adding test 2026-01-28 13:36:19 -08:00
yuneng-jiang
8a54fff5cf Merge remote-tracking branch 'origin' into litellm_key_alias_spend_usage_report 2026-01-28 13:30:49 -08:00
yuneng-jiang
905e9cd6c9 breakdown by team and keys 2026-01-28 13:30:27 -08:00
Ishaan Jaffer
e444199d95 UI: New build 2026-01-28 12:05:36 -08:00
Alexsander Hamir
4c1b24eed9
Fix thread leak in OpenTelemetry dynamic header path (#19946) 2026-01-28 10:35:37 -08:00
michelligabriele
ea3853e977
fix(vertex_ai): support model names with slashes in passthrough URLs (#19944)
The regex in get_vertex_model_id_from_url() was using [^/:]+
which stopped at the first slash, truncating model names like
'gcp/google/gemini-2.5-flash' to just 'gcp'. This caused
access_groups checks to fail for custom model names.

Changed the pattern to [^:]+ to allow slashes in model names,
only stopping at the colon before the action (e.g., :generateContent).
2026-01-28 09:33:53 -08:00
boarder7395
8e4f06583a
Fix team cli auth flow (#19666)
* Cleanup code for user cli auth, and make sure not to prompt user for team multiple times while polling

* Adding tests

* Cleanup normalize teams some more
2026-01-28 08:52:52 -08:00
Sameer Kankute
3ab1b9f543 Fix gemini-robotics-er-1.5-preview name 2026-01-28 21:13:37 +05:30
Sameer Kankute
1cdda28b6c Fix gemini-robotics-er-1.5-preview name 2026-01-28 21:10:44 +05:30
Sameer Kankute
169c9dae79
Merge pull request #19914 from BerriAI/litellm_responses_api_bridge_usage
Fix: output_tokens_details.reasoning_tokens None
2026-01-28 18:35:30 +05:30
Sameer Kankute
9fe8b12f44
Merge pull request #19924 from BerriAI/litellm_minimax_reasoning_caching_1
Add Prompt caching and reasoning support for MiniMax, GLM, Xiaomi
2026-01-28 18:04:13 +05:30
Sameer Kankute
b6c769880e
Merge pull request #19842 from BerriAI/litellm_fix_timeout_test_fix
Fixes Timeouts during chat completion calls no longer reported as timeout in failure callback
2026-01-28 18:03:21 +05:30
Sameer Kankute
f5e5569e40
Merge pull request #19636 from BerriAI/litellm_langfuse_callback
Add litellm_callback_logging_failures_metric for Langfuse, Langfuse Otel and other Otel providers
2026-01-28 18:02:17 +05:30