Commit graph

36177 commits

Author SHA1 Message Date
Sameer Kankute
fa2b065238 Add tests for user level permissions on file and batch access 2026-01-29 12:29:10 +05:30
Sameer Kankute
8966852c86 Fix: Encoding cancel batch response 2026-01-29 12:18:43 +05:30
Aaron Yim
d4031c8ba6
Add OpenRouter Kimi K2.5 (#19872)
Co-authored-by: Claude Opus 4.5 <noreply@anthropic.com>
2026-01-28 22:34:48 -08:00
Christopher Chase
87bdfb0253
fix(hosted_vllm): route through base_llm_http_handler to support ssl_verify (#19893)
* fix(hosted_vllm): route through base_llm_http_handler to support ssl_verify

The hosted_vllm provider was falling through to the OpenAI catch-all path
which doesn't pass ssl_verify to the HTTP client. This adds an explicit
elif branch that routes hosted_vllm through base_llm_http_handler.completion()
which properly passes ssl_verify to the httpx client.

- Add explicit hosted_vllm branch in main.py completion()
- Add ssl_verify tests for sync and async completion
- Update existing audio_url test to mock httpx instead of OpenAI client

* feat(hosted_vllm): add embedding support with ssl_verify

- Add HostedVLLMEmbeddingConfig for embedding transformations
- Register hosted_vllm embedding config in utils.py
- Add lazy import for embedding transformation module
- Add unit test for ssl_verify parameter handling
2026-01-28 22:33:07 -08:00
Bernardo Donadio
ba17f51812
fix(proxy): prevent provider-prefixed model leaks (#19943)
* fix(proxy): prevent provider-prefixed model leaks

Proxy clients should not see LiteLLM internal provider prefixes (e.g. hosted_vllm/...) in the OpenAI-compatible response model field.

This patch sanitizes the client-facing model name for both:
- Non-streaming responses returned from base_process_llm_request
- Streaming SSE chunks emitted by async_data_generator

Adds regression tests covering vLLM-style hosted_vllm routing for both streaming and non-streaming paths.

* chore(lint): suppress PLR0915 in proxy handler

Ruff started flagging ProxyBaseLLMRequestProcessing.base_process_llm_request() for too many statements after the hotpatch changes.

Add an explicit '# noqa: PLR0915' on the function definition to avoid a large refactor in a hotpatch.

* refactor(proxy): make model restamp explicit

Replace silent try/except/pass and type ignores with explicit model restamping.

- Logs an error when the downstream response model differs from the client-requested model
- Overwrites the OpenAI `model` field to the client-requested value to avoid leaking internal provider-prefixed identifiers
- Applies the same behavior to streaming chunks, logging the mismatch only once per stream

* chore(lint): drop PLR0915 suppression

The model restamping bugfix made `base_process_llm_request()` slightly exceed Ruff's
PLR0915 (too-many-statements) threshold, requiring a `# noqa` suppression.

Collapse consecutive `hidden_params` extractions into tuple unpacking so the
function falls back under the lint limit and remove the suppression.

No functional change intended; this keeps the proxy model-field bugfix intact
while aligning with project linting rules.

* chore(proxy): log model mismatches as warnings

These model-restamping logs are intentionally verbose: a mismatch is a useful signal
that an internal provider/deployment identifier may be leaking into the public
OpenAI response `model` field.

- Downgrade model mismatch logs from error -> warning
- Keep error logs only for cases where the proxy cannot read/override the model

* fix(proxy): preserve client model for streaming aliasing

Pre-call processing can rewrite request_data['model'] via model alias maps.\n\nOur streaming SSE generator was using the rewritten value when restamping chunk.model, which caused the public 'model' field to differ between streaming and non-streaming responses for alias-based requests.\n\nStash the original client model in request_data as _litellm_client_requested_model after the model has been routed, and prefer it when overriding the outgoing chunk model. Add a regression test for the alias-mapping case.

* chore(lint): satisfy PLR0915 in streaming generator

Ruff started flagging async_data_generator() for too many statements after adding model restamping logic.\n\nExtract the client-model selection + chunk restamping into small helpers to keep behavior unchanged while meeting the project's PLR0915 threshold.
2026-01-28 22:26:38 -08:00
michelligabriele
dcf5f07e5e
fix(proxy): add datadog_llm_observability to /health/services allowed list (#19952)
The /health/services endpoint rejected datadog_llm_observability as an
unknown service, even though it was registered in the core callback
registry and __init__.py. Added it to both the Literal type hint and
the hardcoded validation list in the health endpoint.
2026-01-28 22:16:27 -08:00
Sameer Kankute
654edbd15c Fix Only allowed to call routes: ['llm_api_routes']. Tried to call route: /batches/bGl0ZWxsbV9wcm/cancel 2026-01-29 11:42:49 +05:30
Sameer Kankute
70684ca86f Fix File access permissions for .retreive and .delete 2026-01-29 11:19:24 +05:30
Cesar Garcia
2a48d12507
fix(docker): add libsndfile to main Dockerfile for ARM64 audio processing (#19776)
Fixes #16920 for users of the stable release images.

The previous fix (PR #18092) added libsndfile to docker/Dockerfile.alpine,
but stable releases are built from the main Dockerfile (Wolfi-based),
not the Alpine variant.
2026-01-28 21:33:41 -08:00
Cesar Garcia
c7453c01f9
Fix stream_chunk_builder to preserve images from streaming chunks (#19654)
Fixes #19478

The stream_chunk_builder function was not handling image chunks from
models like gemini-2.5-flash-image. When streaming responses were
reconstructed (e.g., for caching), images in delta.images were lost.

This adds handling for image_chunks similar to how audio, annotations,
and other delta fields are handled.
2026-01-28 21:31:06 -08:00
Sameer Kankute
833cf6a2cf Fix: Batch cancellation ownership bug 2026-01-29 10:54:42 +05:30
yuneng-jiang
f2d2ed5a0d
Merge pull request #19953 from BerriAI/litellm_key_alias_spend_usage_report
[Feature] UI - Usage Export: Breakdown by Teams and Keys
2026-01-28 20:29:22 -08:00
yuneng-jiang
507f4c45a0
Merge pull request #19976 from BerriAI/ui_build_yj_2
[Infra] Remove _experimental/out routes from gitignore + UI Build
2026-01-28 20:15:06 -08:00
yuneng-jiang
12a4d14980 chore: update Next.js build artifacts (2026-01-29 04:12 UTC, node v22.16.0) 2026-01-28 20:12:20 -08:00
yuneng-jiang
20bab33e36 removing _experimental out routes from gitignore 2026-01-28 20:11:35 -08:00
Harshit Jain
8e2fa7969c
Fix/router search tools v2 (#19840)
* fix(proxy_server): pass search_tools to Router during DB-triggered initialization

* fix search tools from db

* add missing statement to handle from db

* fix import issues to pass lint errors
2026-01-28 19:45:35 -08:00
Cesar Garcia
8a26033a4b
fix(vertex_ai): convert image URLs to base64 in tool messages for Anthropic (#19896)
* fix(vertex_ai): convert image URLs to base64 in tool messages for Anthropic

Fixes #19891

Vertex AI Anthropic models don't support URL sources for images. LiteLLM
already converted image URLs to base64 for user messages, but not for tool
messages (role='tool'). This caused errors when using ToolOutputImage with
image_url in tool outputs.

Changes:
- Add force_base64 parameter to convert_to_anthropic_tool_result()
- Pass force_base64 to create_anthropic_image_param() for tool message images
- Calculate force_base64 in anthropic_messages_pt() based on llm_provider
- Add unit tests for tool message image handling

* chore: remove extra comment from test file header
2026-01-28 19:42:51 -08:00
Sameer Kankute
2a1bfd39aa
Merge pull request #19974 from BerriAI/litellm_model_map_fix_jan_29
fix gemini gemini-robotics-er-1.5-preview entry
2026-01-29 09:07:54 +05:30
Sameer Kankute
be8a76f270 fix gemini gemini-robotics-er-1.5-preview entry 2026-01-29 09:06:44 +05:30
rushilchugh01
562f0a0282
feat: Add new OpenRouter models: xiaomi/mimo-v2-flash, z-ai/glm-4.7, z-ai/glm-4.7-flash, and minimax/minimax-m2.1. to model prices and context window (#19938)
Co-authored-by: Rushil Chugh <Rushil>
2026-01-28 18:56:20 -08:00
Ishaan Jaff
9c5fed4f52
[Feat] LiteLLM Vector Stores - Add permission management for users, teams (#19972)
* fix: create_vector_store_in_db

* add team/user to LiteLLM_ManagedVectorStore

* add _check_vector_store_access

* add new fields

* test_check_vector_store_access

* add vector_store/list endpoints

* fix code QA checks
2026-01-28 18:55:40 -08:00
yuneng-jiang
e796b9eb22
Merge pull request #19963 from BerriAI/litellm_ui_spend_logs_em_search
[Feature] UI - Logs: Adding Error message search to ui spend logs
2026-01-28 18:15:35 -08:00
yuneng-jiang
58dd3bd134 fixing sorting for v2/model/info 2026-01-28 18:07:22 -08:00
yuneng-jiang
632e8cf2f6
Merge pull request #19970 from BerriAI/litellm_ui_column_sort_component
[Feature] UI - Tables: Reusable Table Sort Component
2026-01-28 18:02:34 -08:00
Ishaan Jaff
dcca8c7350
[Feat] - Search API add /list endpoint to list what search tools exist in router (#19969)
* feat: List all available search tools configured in the router.

* add debugging search API

* add debugging search API
2026-01-28 17:58:17 -08:00
Alexsander Hamir
69bd4426e8
[Release Day] - Fixed CI/CD issues & changed processes (#19902) 2026-01-28 17:57:24 -08:00
yuneng-jiang
e9056671f9 Fixing sorting API calls 2026-01-28 17:53:38 -08:00
yuneng-jiang
92f9d8f86e Reusable Table Sort Component 2026-01-28 17:42:57 -08:00
Neha Prasad
a785eecf7a
fix Prompt Studio history to load tools and system messages (#19920) 2026-01-28 17:19:59 -08:00
Ishaan Jaff
d12ce3cd5d
[Fix] VertexAI Pass through - fix regression that caused vertex ai passthroughs to stop working for router models (#19967)
* fix(vertex_ai): replace custom model names with actual Vertex AI model names in passthrough URLs (#19948)

When the passthrough URL already contains project and location, the code
was skipping the deployment lookup and forwarding the URL as-is to Vertex AI.
For custom model names like gcp/google/gemini-2.5-flash, Vertex AI returned
404 because it only knows the actual model name (gemini-2.5-flash).

The fix makes the deployment lookup always run, so the custom model name
gets replaced with the actual Vertex AI model name before forwarding.

* add _resolve_vertex_model_from_router

* fix: get_llm_provider

* Potential fix for code scanning alert no. 4020: Clear-text logging of sensitive information

Co-authored-by: Copilot Autofix powered by AI <62310815+github-advanced-security[bot]@users.noreply.github.com>

---------

Co-authored-by: michelligabriele <gabriele.michelli@icloud.com>
Co-authored-by: Copilot Autofix powered by AI <62310815+github-advanced-security[bot]@users.noreply.github.com>
2026-01-28 16:54:01 -08:00
Alexsander Hamir
3816570313
fix(presidio): reuse HTTP connections to prevent OOMs (#19964) 2026-01-28 16:08:53 -08:00
Ishaan Jaffer
c4daa39daa fix 2026-01-28 16:05:03 -08:00
yuneng-jiang
cccda30a9e
Merge pull request #19960 from BerriAI/litellm_ui_spend_logs_error_message
[Feature] Add error_message Search in Spend Logs Endpoint
2026-01-28 16:04:29 -08:00
yuneng-jiang
0cdfa8e5fa Adding Error message search to ui spend logs 2026-01-28 15:39:32 -08:00
yuneng-jiang
cb8ead6013 Add error_message search in spend logs endpoint 2026-01-28 15:06:31 -08:00
Ishaan Jaff
3ef475b70e
[Fix] A2a Gateway - Allow supporting old A2a card formats (#19949)
* fix: LiteLLMA2ACardResolver

* fix: LiteLLMA2ACardResolver

* feat: .well-known/agent.json

* test_card_resolver_fallback_from_new_to_old_path
2026-01-28 15:02:08 -08:00
yuneng-jiang
054918e7a3
Merge pull request #19918 from BerriAI/litellm_ui_spend_logs_store
[Feature] UI - Spend Logs: Settings Modal
2026-01-28 15:00:31 -08:00
Ishaan Jaffer
5135efb60e fix pypdf: >=6.6.2 2026-01-28 14:54:58 -08:00
yuneng-jiang
dbd1ff306d Fixing build 2026-01-28 13:42:58 -08:00
yuneng-jiang
077cfa8c15 Adding test 2026-01-28 13:36:19 -08:00
yuneng-jiang
8a54fff5cf Merge remote-tracking branch 'origin' into litellm_key_alias_spend_usage_report 2026-01-28 13:30:49 -08:00
yuneng-jiang
905e9cd6c9 breakdown by team and keys 2026-01-28 13:30:27 -08:00
Ishaan Jaffer
e444199d95 UI: New build 2026-01-28 12:05:36 -08:00
Alexsander Hamir
4c1b24eed9
Fix thread leak in OpenTelemetry dynamic header path (#19946) 2026-01-28 10:35:37 -08:00
michelligabriele
ea3853e977
fix(vertex_ai): support model names with slashes in passthrough URLs (#19944)
The regex in get_vertex_model_id_from_url() was using [^/:]+
which stopped at the first slash, truncating model names like
'gcp/google/gemini-2.5-flash' to just 'gcp'. This caused
access_groups checks to fail for custom model names.

Changed the pattern to [^:]+ to allow slashes in model names,
only stopping at the colon before the action (e.g., :generateContent).
2026-01-28 09:33:53 -08:00
boarder7395
8e4f06583a
Fix team cli auth flow (#19666)
* Cleanup code for user cli auth, and make sure not to prompt user for team multiple times while polling

* Adding tests

* Cleanup normalize teams some more
2026-01-28 08:52:52 -08:00
Sameer Kankute
3ab1b9f543 Fix gemini-robotics-er-1.5-preview name 2026-01-28 21:13:37 +05:30
Sameer Kankute
1cdda28b6c Fix gemini-robotics-er-1.5-preview name 2026-01-28 21:10:44 +05:30
Luis Gallego Ledesma
52372dcbe9 fix(langfuse_otel): prevent empty proxy request spans from being sent to Langfuse
When using langfuse_otel callback, empty traces were being sent to Langfuse
for requests that didn't result in actual LLM calls (e.g., auth operations,
health checks, failed requests). These traces contained only internal proxy
operations (auth, postgres, proxy_pre_call) with no useful LLM data.

Root cause: LangfuseOtelLogger extends OpenTelemetry, which sets itself as
the proxy's open_telemetry_logger. This caused create_litellm_proxy_request_started_span
to be called for every request, creating a parent span that was sent to Langfuse
even when no LLM call occurred.

Fix: Override create_litellm_proxy_request_started_span in LangfuseOtelLogger
to return None, preventing the creation of empty parent spans. This is consistent
with the existing overrides for async_service_success_hook and async_service_failure_hook
which already prevent service-level logs from being sent to Langfuse.

Fixes: Empty traces in Langfuse v3 when using langfuse_otel callback
2026-01-28 15:35:35 +01:00
Sameer Kankute
169c9dae79
Merge pull request #19914 from BerriAI/litellm_responses_api_bridge_usage
Fix: output_tokens_details.reasoning_tokens None
2026-01-28 18:35:30 +05:30