Commit graph

37969 commits

Author SHA1 Message Date
yuneng-jiang
33bb798997
Merge pull request #22222 from BerriAI/litellm_key_filter_pagination_fix
[Fix] Virtual Keys pagination displays stale totals when filtering
2026-02-26 15:09:42 -08:00
yuneng-jiang
c20e49620f [Feature] Add /public/supported_endpoints endpoint
Return all supported endpoints and which providers support them. Includes endpoint display names, URL paths, and per-provider support lists. Results are cached for the process lifetime.

Also adds comprehensive test coverage for the new endpoint.

Co-Authored-By: Claude Haiku 4.5 <noreply@anthropic.com>
2026-02-26 15:01:25 -08:00
ryan-crabbe
3639800892
Merge pull request #22221 from BerriAI/litellm_prometheus_multiproc_cleanup
Litellm prometheus multiproc cleanup
2026-02-26 13:51:00 -08:00
Dharamendra Kumar
fcdfc638b0 Remove nit 2026-02-26 13:44:06 -08:00
Dharamendra Kumar
840333f32e Restore 2026-02-26 13:31:37 -08:00
Dharamendra Kumar
c651a511bd Update test to righ place 2026-02-26 13:26:51 -08:00
Emerson Gomes
0e014253d7 refactor(cost): dedupe image token usage cost helper
- extract shared calculate_image_response_cost_from_usage() helper\n- reuse helper in vertex and gemini image generation cost calculators\n- preserve provider-specific fallback to output_cost_per_image
2026-02-26 15:07:51 -06:00
Dharamendra Kumar
7df61d4b92 [Fix] Enhance MidStreamFallbackError to preserve original status code and attributes
- Updated MidStreamFallbackError to retrieve and maintain the original status code from the wrapped exception.
- Ensured that message, request, and response fields remain consistent after calling the parent constructor.
- Added unit tests to verify the correct propagation of status codes and attributes in various scenarios.
2026-02-26 13:01:14 -08:00
Emerson Gomes
702d5e88b8 fix(cost): use token usage for gemini/vertex image generation when available
- compute image_generation cost from usage token metadata for vertex/gemini\n- map ImageUsage to Usage and reuse generic_cost_per_token\n- fallback to output_cost_per_image when usage metadata missing\n- add tests for token-based path and fallback path
2026-02-26 14:57:11 -06:00
Julio Quinteros Pro
8a6a67bfcf fix(proxy): isolate get_config failures from model loading in sync loop
A database timeout (httpcore.ReadTimeout) during get_config() in
_update_llm_router would propagate and prevent ALL DB models from
loading into the router.

Now get_config() failures are caught separately so model add/delete
operations still proceed. Similarly, _delete_deployment catches
get_config failures and safely skips cleanup rather than crashing the
entire sync cycle.

Co-Authored-By: Claude Opus 4.6 <noreply@anthropic.com>
2026-02-26 17:49:44 -03:00
Emerson Gomes
f24a41898b feat(vertex): add gemini-3.1-flash-image-preview model DB support
- add gemini-3.1-flash-image-preview + vertex_ai alias entries\n- set pricing to Gemini 3.1 Flash Image Preview rates\n- mirror updates in packaged backup model map\n- update llm cost calc regression test to cover new model
2026-02-26 14:41:59 -06:00
Dibyo Mukherjee
518cd3ef60 feat(ui): add key creation deep-links with SSO return URL support
Enables deep-linking directly to the key creation modal with prefilled
form data via URL parameters, including support for preserving these
deep-links through SSO authentication flows.

Key Creation Deep-links:
- Auto-open key creation modal via ?create=true parameter
- Prefill form fields from URL parameters (team_id, key_alias, models, etc.)
- Role-based access control for auto-open (requires write access)
- Race condition protection for redirect handling

Example: /ui?create=true&team_id=abc&key_alias=my-key&models=gpt-4,claude-3

SSO Return URL Preservation:
- Cookie-based return URL storage (works across ports for SSO flows)
- URL validation to prevent open redirect attacks
- Support for both dev and production environments

Co-Authored-By: Claude Opus 4.5 <noreply@anthropic.com>
2026-02-26 15:41:54 -05:00
Ryan Crabbe
0020e5929d fix: accurate wipe count on partial failure, remove stray blank line 2026-02-26 12:36:57 -08:00
Cesar Garcia
c6b6e29bc2
Merge pull request #22220 from Chesars/fix/reasoning-effort-output-config-claude-46
fix(anthropic): map reasoning_effort to output_config for Claude 4.6 models
2026-02-26 17:30:54 -03:00
Julio Quinteros Pro
de556ad237
Merge pull request #22202 from BerriAI/fix/realtime-streaming-mypy-attr
fix(mypy): suppress attr-defined errors on realtime websocket calls
2026-02-26 17:30:39 -03:00
yuneng-jiang
c3c79b9d80 [Fix] Virtual Keys pagination displays stale totals when filtering
When applying a key_alias filter on the Virtual Keys page, the pagination display (total count, page count) showed stale unfiltered values. The bug was that the table component tracked total_count only from the unfiltered useKeys hook, not from the filtered API response returned by useFilterLogic.

Fixes: When selecting a key alias that matches 1 key, the UI previously showed "Showing 1-50 of 509 results" and "Page 1 of 11" instead of the correct "Showing 1-1 of 1 results" and "Page 1 of 1".

Changes:
- Added filteredTotalCount state to useFilterLogic to track the total_count from filtered API responses
- Updated VirtualKeysTable to use filteredTotalCount (when set) instead of always using the unfiltered total
- Added comprehensive tests to prevent regression of pagination display logic

Type: 🐛 Bug Fix, ✅ Test

Co-Authored-By: Claude Haiku 4.5 <noreply@anthropic.com>
2026-02-26 12:28:43 -08:00
Chesars
5aff0e4da6 refactor: move _is_claude_4_6_model to AnthropicModelInfo, use explicit effort_map
- Move _is_claude_4_6_model from AnthropicConfig to AnthropicModelInfo
  to eliminate duplicated logic in is_effort_used
- Use explicit effort_map dict instead of passing unknown values through
  to output_config
2026-02-26 17:25:23 -03:00
Chesars
a8c95392fb fix(anthropic): map reasoning_effort to output_config for Claude 4.6 models
Claude 4.6 models use output_config as a stable API feature. This commit:
- Maps reasoning_effort to output_config for 4.6 models (minimal → low)
- Restricts effort="max" to Opus 4.6 only
- Skips beta header injection for 4.6 models
- Updates docs for Claude 4.6 effort support
2026-02-26 17:18:32 -03:00
Ryan Crabbe
0ee8cb5f02 refactor: remove shutdown cleanup, rely solely on startup wipe
Phase 2 (per-worker mark_process_dead on shutdown) only ever fired
when all workers shut down together, making it redundant — Phase 1
wipes everything on next startup anyway. This aligns with the
prometheus_client docs: just wipe the directory between runs.
2026-02-26 12:17:31 -08:00
yuneng-jiang
719b7fd013
Merge pull request #22157 from BerriAI/litellm_paginated_key_alias
[Feature] UI - Paginated Key Alias Select
2026-02-26 12:09:10 -08:00
yuneng-jiang
4b75a89673 Merge remote-tracking branch 'origin' into litellm_paginated_key_alias 2026-02-26 12:06:28 -08:00
Cesar Garcia
ec6a55c6db
Merge pull request #22208 from Chesars/fix/transcription-duration-hidden-params
fix(transcription): move duration to _hidden_params to match OpenAI response spec
2026-02-26 16:51:24 -03:00
Ishaan Jaff
cbdaaaeba4
[Feat] Add control for setting upperbound on chunk processing time (#22209)
* add LITELLM_MAX_STREAMING_DURATION_SECONDS

* add add LITELLM_MAX_STREAMING_DURATION_SECONDS

* fix: address Greptile review - rename constant, add sync check, add tests

- Rename MAX_STREAMING_CHUNK_DURATION_S → MAX_STREAMING_DURATION_S (misleading "CHUNK")
- Add _check_max_streaming_duration to SyncResponsesAPIStreamingIterator.__next__
- Add 8 unit tests covering both CustomStreamWrapper and ResponsesAPI paths
- Fix pre-existing pyright errors in streaming_handler.py

Made-with: Cursor

* add add LITELLM_MAX_STREAMING_DURATION_SECONDS
2026-02-26 11:39:09 -08:00
Cesar Garcia
b5a0a72acb
Merge pull request #21597 from Chesars/fix/gemini-json-schema-keep-refs
fix(gemini): preserve $ref in JSON Schema for Gemini 2.0+
2026-02-26 16:37:26 -03:00
Harshit28j
d3d11fb06e feat: add tags in project 2026-02-27 01:07:18 +05:30
Chesars
121669090d test: use real completion_cost() instead of duplicating inline logic 2026-02-26 16:24:48 -03:00
Chesars
5f957add18 fix(azure): apply same duration hidden_params fix to Azure transcription handler 2026-02-26 16:16:41 -03:00
Chesars
b6784e7d8c fix(types): annotate _FINISH_REASON_MAP and map_finish_reason with OpenAIChatCompletionFinishReason
Fixes mypy errors where dict[str, str] was incompatible with the
expected Literal type in get_finish_reason_mapping() and
_check_finish_reason() return types.
2026-02-26 16:13:23 -03:00
Chesars
7dd4f17021 fix(transcription): store duration in _hidden_params to avoid OpenAI SDK deserialization issues
LiteLLM was adding a `duration` field to audio transcription responses
for internal cost tracking. The OpenAI Python SDK uses "best match
deserialization" to determine the response type from present fields —
seeing `duration` caused it to incorrectly match plain Transcription
responses as TranscriptionVerbose/TranscriptionDiarized types.

Move the internally-calculated duration to `_hidden_params` so it
remains available for cost calculation without polluting the response
body. Provider-returned duration (e.g. from verbose_json format) is
still preserved in the response as expected.
2026-02-26 16:06:08 -03:00
Chesars
3196d40a04 fix(vertex): delegate Gemini finish reason mapping to centralized _FINISH_REASON_MAP
Addresses Greptile review feedback on PR #22138 — removes duplicated
Gemini finish reason dict in VertexGeminiConfig and delegates to the
shared map_finish_reason() to prevent the two mappings from drifting
apart.
2026-02-26 15:48:05 -03:00
yuneng-jiang
947586b62d Merge remote-tracking branch 'origin' into litellm_paginated_key_alias 2026-02-26 10:28:35 -08:00
yuneng-jiang
50bf2da05e
Merge pull request #22137 from BerriAI/litellm_key_info_crash_fix
[Fix] /key/aliases: Add pagination and search to prevent OOMs
2026-02-26 10:27:10 -08:00
Harshit Jain
4d2fab49a7
Merge pull request #22164 from Harshit28j/litellm_custom_auth_budget_fix
fix: custom auth budget issue
2026-02-26 23:38:38 +05:30
yuneng-jiang
857a324f7a Merge remote-tracking branch 'origin' into litellm_key_info_crash_fix 2026-02-26 10:04:54 -08:00
Ryan Crabbe
acfb5ade97 refactor: remove periodic dead PID cleanup, trim redundant tests
Startup wipe + graceful shutdown cleanup are sufficient. Remove
hourly mark_dead_pids scan, its helpers, and redundant test cases.
2026-02-26 10:03:51 -08:00
michelligabriele
ae13a40c01
test(mcp): add e2e test for stateless StreamableHTTP behavior (#22033)
Adds TestProxyMcpStatelessBehavior to test_proxy_mcp_e2e.py with a test
that verifies two independent MCP clients can connect, initialize, and
call tools without sharing session state. This catches the regression
from PR #19809 where stateless=False broke clients that don't manage
mcp-session-id headers.

Regression test for #20242
2026-02-26 09:30:03 -08:00
Sameer Kankute
1790a6bf82
Merge pull request #21604 from michelligabriele/fix/websearch-thinking-blocks
fix(websearch_interception): preserve thinking blocks in agentic loop follow-up messages
2026-02-26 22:26:38 +05:30
Julio Quinteros Pro
21d958e46d fix(mypy): add union-attr suppression to all backend_ws method calls
Address Greptile review: apply type: ignore[union-attr] consistently
on all backend_ws.send(), .recv(), and .close() calls, not just the
three that CI flagged.

Co-Authored-By: Claude Opus 4.6 <noreply@anthropic.com>
2026-02-26 13:32:13 -03:00
Julio Quinteros Pro
494d783a80 fix(mypy): suppress attr-defined errors on websocket send/close calls
CI MyPy resolves CLIENT_CONNECTION_CLASS as Optional[ClientConnection]
and flags .send() and .close() as attr-defined errors. These methods
exist at runtime on the websocket connection object. Add type: ignore
comments to unblock the linting CI.

Co-Authored-By: Claude Opus 4.6 <noreply@anthropic.com>
2026-02-26 13:22:56 -03:00
Julio Quinteros Pro
4419c7a5f9
Merge pull request #22198 from BerriAI/fix/mcp-server-test-mock-mismatch
fix(tests): update MCP server test mocks to match production API
2026-02-26 13:16:40 -03:00
Julio Quinteros Pro
ace49b18d3 fix(tests): update MCP server test mocks to match production API
The tests were mocking `filter_server_ids_by_ip` but the production
code in server.py now calls `filter_server_ids_by_ip_with_info` which
returns a (server_ids, blocked_count) tuple. Update all 8 mock sites
to use the correct method name and return signature.

Co-Authored-By: Claude Opus 4.6 <noreply@anthropic.com>
2026-02-26 13:11:57 -03:00
Julio Quinteros Pro
1c376afc85 fix(ci): use secrets context in ggshield step condition
Step-level env is not visible to the if condition — reference
secrets directly so ggshield actually runs when the key is configured.

Co-Authored-By: Claude Opus 4.6 <noreply@anthropic.com>
2026-02-26 12:51:28 -03:00
Julio Quinteros Pro
05c3a95da8 fix(ci): add permissions block to secret-scan job
Address github-advanced-security bot review comment by setting explicit
minimal permissions (contents: read) for the GITHUB_TOKEN.

Co-Authored-By: Claude Opus 4.6 <noreply@anthropic.com>
2026-02-26 12:48:43 -03:00
Julio Quinteros Pro
2fce35a162 test(ci): add secret scan test and CI job to prevent hardcoded credentials
- Add unit test that scans Python source for Base64 Basic Auth patterns
  that would be flagged by secret scanners like GitGuardian/ggshield
- Add secret-scan job to the linting CI workflow that runs the test on
  every PR and optionally runs ggshield if GITGUARDIAN_API_KEY is set

Co-Authored-By: Claude Opus 4.6 <noreply@anthropic.com>
2026-02-26 12:46:42 -03:00
Harshit Jain
43054a2390
fix: langfuse trace leak key on model params 2026-02-26 19:03:49 +05:30
Sameer Kankute
6600c86dbd
Merge pull request #22170 from BerriAI/litellm_fix_video_veo_vertex
Fix: Passing of image and parameters in videos api
2026-02-26 18:55:25 +05:30
Sameer Kankute
ec31845469
Merge pull request #22092 from BerriAI/litellm_fix_tts_vertex_ai
Add audio as supported openai param
2026-02-26 18:51:53 +05:30
Sameer Kankute
61d2f28545 Fix based on review 2026-02-26 18:49:25 +05:30
Sameer Kankute
16d6c279da
Merge pull request #22180 from BerriAI/litellm_fix_vllm_test
Add JSON exact match test for vLLM embeddings
2026-02-26 18:43:52 +05:30
Sameer Kankute
dc1e97d345
Merge pull request #22144 from BerriAI/litellm_allowed_openai_params_embeddings
fix(embeddings): allow dimensions param passthrough via allowed_openai_params for non-text-embedding-3 OpenAI models
2026-02-26 18:41:47 +05:30