Commit graph

41449 commits

Author SHA1 Message Date
yuneng-jiang
a732c3f177
Merge pull request #22248 from BerriAI/litellm_public_endpoints
[Feature] Add /public/endpoints for provider endpoint support
2026-02-26 20:49:44 -08:00
yuneng-jiang
2144e79bad fix: guard against null assigned_*_ids in update_access_group delta computation
set(None) raises TypeError when a client sends null for assigned_team_ids or
assigned_key_ids. Add `or []` to handle null safely, consistent with create.
Add test covering this case.

Co-Authored-By: Claude Sonnet 4.6 <noreply@anthropic.com>
2026-02-26 20:48:04 -08:00
yuneng-jiang
06c90ecf62
Update litellm/proxy/public_endpoints/public_endpoints.py
Co-authored-by: greptile-apps[bot] <165735046+greptile-apps[bot]@users.noreply.github.com>
2026-02-26 20:42:47 -08:00
Sameer Kankute
577f703769 Register custom pricing in litellm.model_cost 2026-02-27 10:12:03 +05:30
yuneng-jiang
57c5efc785 refactor: initialize delta vars before try block and avoid redundant find_unique on delete
- Initialize teams_to_add/teams_to_remove/keys_to_add/keys_to_remove before
  the try block in update_access_group for defensive clarity
- In delete_access_group, update teams/keys returned by find_many directly
  (data already fetched) and use _sync_remove only for out-of-sync entities
  not found by the hasSome query, eliminating N+1 find_unique calls

Co-Authored-By: Claude Sonnet 4.6 <noreply@anthropic.com>
2026-02-26 20:39:57 -08:00
Sameer Kankute
d9af321610 Fix free models working from UI 2026-02-27 10:08:44 +05:30
yuneng-jiang
7e2f5b7c5b
Merge pull request #22255 from BerriAI/litellm_e2e_fix_feb26
[Infra] Adding agent_id to Delete Keys Table
2026-02-26 20:31:17 -08:00
yuneng-jiang
1e82ec6448 adding build 2026-02-26 20:30:10 -08:00
yuneng-jiang
ee7b73764c bump: version 0.4.48 → 0.4.49 2026-02-26 20:29:43 -08:00
yuneng-jiang
fd58c8c060
Merge pull request #22251 from BerriAI/litellm_circleci_prisma_sync
[Infra] Add prisma_schema_sync step as prerequisite for e2e UI tests
2026-02-26 20:24:56 -08:00
yuneng-jiang
2d9ba674ec fix: move update_access_group find_unique inside transaction
Eliminates TOCTOU race where existing record was read outside the
transaction, allowing a concurrent update to make delta computation stale.
Delta is now computed atomically within the same transaction as the write.

Co-Authored-By: Claude Sonnet 4.6 <noreply@anthropic.com>
2026-02-26 20:14:54 -08:00
yuneng-jiang
516b18feca [Feature] Access group CRUD: Add bidirectional sync for teams/keys
When creating, updating, or deleting access groups, automatically keep
team and key access_group_ids in sync with the access group's assigned_team_ids
and assigned_key_ids. Includes transaction-based DB updates, cache patching,
and handles out-of-sync data by unioning assigned_* fields with hasSome queries.

Adds 12 new tests covering sync behavior across all three CRUD operations.

Co-Authored-By: Claude Sonnet 4.6 <noreply@anthropic.com>
2026-02-26 20:07:51 -08:00
yuneng-jiang
369c0ec392 [Infra] Add prisma_schema_sync CircleCI job before e2e UI tests
Adds a new CircleCI job that runs the proxy with --use_prisma_db_push
against the base Neon branch before the e2e UI tests create their
branches from it, ensuring the schema is synced on the parent.

Co-Authored-By: Claude Sonnet 4.6 <noreply@anthropic.com>
2026-02-26 20:03:25 -08:00
yuneng-jiang
fc69d6e8d1 fix: add 12 missing endpoint keys to _ENDPOINT_METADATA, fix stale _schema keys in backup JSON 2026-02-26 19:18:17 -08:00
Sameer Kankute
8d1c75c48a
Merge pull request #22225 from dharamendrak/bugfix/midstream-fallback-error-masks-status-code
[Fix] Enhance MidStreamFallbackError to preserve original status code and attributes
2026-02-27 08:28:54 +05:30
yuneng-jiang
ffc00c0c90 fix: correct _ENDPOINT_METADATA keys to match actual JSON data (a2a, container_files) 2026-02-26 18:30:28 -08:00
yuneng-jiang
86b2efd67a [Feature] Add /public/endpoints endpoint for provider endpoint support
Add new /public/endpoints endpoint that returns which providers support each LiteLLM
endpoint (e.g., chat_completions, embeddings). The endpoint reads from a local backup
JSON file bundled with the package, caches the result in-process, and transforms the
raw provider-centric data into an endpoint-centric response format.

Changes:
- Add litellm/provider_endpoints_support_backup.json (copy of root source file)
- Add Pydantic response models (EndpointProvider, SupportedEndpoint, SupportedEndpointsResponse)
- Add /public/endpoints route with transformation and caching logic
- Add 16 comprehensive tests covering HTTP layer and transformation functions

Co-Authored-By: Claude Haiku 4.5 <noreply@anthropic.com>
2026-02-26 18:17:37 -08:00
ryan-crabbe
88ccffccc8
Merge pull request #22247 from BerriAI/litellm_fix_client_closed_on_eviction
fix: remove cache eviction close that kills in-use httpx clients
2026-02-26 18:00:44 -08:00
Ryan Crabbe
df36845839 fix: remove cache eviction close that kills in-use httpx clients 2026-02-26 17:39:01 -08:00
yuneng-jiang
5a95d9beee
Merge pull request #22246 from BerriAI/revert-22238-litellm_supported_endpoints
Revert "[Feature] Add /public/supported_endpoints endpoint"
2026-02-26 17:21:53 -08:00
yuneng-jiang
71c3503e57
Revert "[Feature] Add /public/supported_endpoints endpoint" 2026-02-26 17:21:43 -08:00
yuneng-jiang
0a690a55a9
Merge pull request #22243 from BerriAI/litellm_team_budget_inf_fix
[Fix] Prometheus Metrics Team +Inf Budgets
2026-02-26 17:06:40 -08:00
Ishaan Jaff
48b9ecacad
fix(realtime): fix guardrails not firing for Gemini/Vertex AI and provider_config realtime WebSocket sessions (#22168)
* fix(gemini): enable inputAudioTranscription and handle transcription events for realtime guardrails

Gemini sends inputTranscription/outputTranscription inside serverContent separately from modelTurn/turnComplete. This adds handling to convert them into OpenAI-compatible events so the guardrail pipeline can inspect voice input, and enables inputAudioTranscription in the session setup config.

Made-with: Cursor

* fix(vertex_ai): enable inputAudioTranscription in realtime session config

Add inputAudioTranscription to the Vertex AI realtime setup so the backend returns transcripts of user speech, allowing guardrails to inspect voice input.

Made-with: Cursor

* fix(realtime): pass user_api_key_dict and guardrail metadata through async_realtime handler

The base LLM HTTP handler's async_realtime method was not accepting or forwarding user_api_key_dict and litellm_metadata to RealTimeStreaming. This meant guardrails configured with default_on=false were silently skipped for all provider_config-based realtime connections (Gemini, Vertex AI, etc). Also fixes wss:// connections when SSL_VERIFY=False by overriding ssl=False for secure WebSocket URLs.

Made-with: Cursor

* fix(realtime): forward guardrail metadata for generic provider_config and vertex_ai paths

The _arealtime function was not passing user_api_key_dict or litellm_metadata to base_llm_http_handler.async_realtime() for the generic provider_config path and the vertex_ai-specific path. This broke guardrail resolution since RealTimeStreaming.request_data was empty, causing should_run_guardrail to return False.

Made-with: Cursor

* fix(realtime): voice guardrail responses and block duplicate response.create on text input

When a guardrail blocks voice input, send a conversation.item.create + response.create to the backend so the LLM voices the guardrail message as audio instead of only returning text. Also adds pending_guardrail_message tracking to suppress the automatic response.create the client sends after a blocked text message, and broadens _has_audio_transcription_guardrails to match pre_call/post_call modes.

Made-with: Cursor

* test(realtime): update guardrail tests for broadened audio transcription check and add integration tests

Update existing tests to reflect that pre_call guardrails now correctly trigger the audio/VAD session.update injection. Add integration test file for live OpenAI realtime guardrail testing.

Made-with: Cursor

* fix(realtime): instruct LLM to say exact guardrail message verbatim

The previous prompt gave the LLM creative freedom to paraphrase the guardrail violation message. Now it instructs the LLM to repeat the exact configured message word for word.

Made-with: Cursor

* fix(realtime): preserve wss ssl semantics and move live guardrail test

Keep TLS enabled for wss realtime sessions while honoring SSL_VERIFY=False via a no-verify SSLContext, move the OpenAI live guardrail test into llm_translation, and dedupe duplicated guardrail-detection helpers to prevent drift.

Made-with: Cursor
2026-02-26 17:06:07 -08:00
yuneng-jiang
28c77b48c9 fix +Inf user budget metric when metadata max_budget is None
Same bug as team budget: _assemble_user_object fetched user info from DB
but only used budget_reset_at, discarding max_budget. When the key cache
has a stale None for user_max_budget, _safe_get_remaining_budget returns
+Inf. Now falls back to DB max_budget when metadata value is None.

Co-Authored-By: Claude Sonnet 4.6 <noreply@anthropic.com>
2026-02-26 16:57:31 -08:00
yuneng-jiang
0e1428b59d remove orphan comment from test file
Co-Authored-By: Claude Haiku 4.5 <noreply@anthropic.com>
2026-02-26 16:51:18 -08:00
Ishaan Jaff
9546d9b482
_add_dd_apm_tags_for_litellm_call_id (#22219) 2026-02-26 16:42:23 -08:00
yuneng-jiang
f4e3e016a1 fixing inf budget 2026-02-26 16:30:41 -08:00
yuneng-jiang
d5ef6c7f93
Merge pull request #22238 from BerriAI/litellm_supported_endpoints
[Feature] Add /public/supported_endpoints endpoint
2026-02-26 15:38:29 -08:00
yuneng-jiang
efcc856234 Move provider_endpoints_support.json into litellm package
The file was at the repo root and excluded from pip distributions. Moving it to litellm/proxy/public_endpoints/ alongside the other provider JSON files ensures it is packaged correctly. Updates all references in the endpoint handler, coverage tests, and release notes instructions.

Co-Authored-By: Claude Haiku 4.5 <noreply@anthropic.com>
2026-02-26 15:15:16 -08:00
yuneng-jiang
33bb798997
Merge pull request #22222 from BerriAI/litellm_key_filter_pagination_fix
[Fix] Virtual Keys pagination displays stale totals when filtering
2026-02-26 15:09:42 -08:00
yuneng-jiang
c20e49620f [Feature] Add /public/supported_endpoints endpoint
Return all supported endpoints and which providers support them. Includes endpoint display names, URL paths, and per-provider support lists. Results are cached for the process lifetime.

Also adds comprehensive test coverage for the new endpoint.

Co-Authored-By: Claude Haiku 4.5 <noreply@anthropic.com>
2026-02-26 15:01:25 -08:00
ryan-crabbe
3639800892
Merge pull request #22221 from BerriAI/litellm_prometheus_multiproc_cleanup
Litellm prometheus multiproc cleanup
2026-02-26 13:51:00 -08:00
Dharamendra Kumar
fcdfc638b0 Remove nit 2026-02-26 13:44:06 -08:00
Dharamendra Kumar
840333f32e Restore 2026-02-26 13:31:37 -08:00
Dharamendra Kumar
c651a511bd Update test to righ place 2026-02-26 13:26:51 -08:00
Emerson Gomes
0e014253d7 refactor(cost): dedupe image token usage cost helper
- extract shared calculate_image_response_cost_from_usage() helper\n- reuse helper in vertex and gemini image generation cost calculators\n- preserve provider-specific fallback to output_cost_per_image
2026-02-26 15:07:51 -06:00
Dharamendra Kumar
7df61d4b92 [Fix] Enhance MidStreamFallbackError to preserve original status code and attributes
- Updated MidStreamFallbackError to retrieve and maintain the original status code from the wrapped exception.
- Ensured that message, request, and response fields remain consistent after calling the parent constructor.
- Added unit tests to verify the correct propagation of status codes and attributes in various scenarios.
2026-02-26 13:01:14 -08:00
Emerson Gomes
702d5e88b8 fix(cost): use token usage for gemini/vertex image generation when available
- compute image_generation cost from usage token metadata for vertex/gemini\n- map ImageUsage to Usage and reuse generic_cost_per_token\n- fallback to output_cost_per_image when usage metadata missing\n- add tests for token-based path and fallback path
2026-02-26 14:57:11 -06:00
Julio Quinteros Pro
8a6a67bfcf fix(proxy): isolate get_config failures from model loading in sync loop
A database timeout (httpcore.ReadTimeout) during get_config() in
_update_llm_router would propagate and prevent ALL DB models from
loading into the router.

Now get_config() failures are caught separately so model add/delete
operations still proceed. Similarly, _delete_deployment catches
get_config failures and safely skips cleanup rather than crashing the
entire sync cycle.

Co-Authored-By: Claude Opus 4.6 <noreply@anthropic.com>
2026-02-26 17:49:44 -03:00
Emerson Gomes
f24a41898b feat(vertex): add gemini-3.1-flash-image-preview model DB support
- add gemini-3.1-flash-image-preview + vertex_ai alias entries\n- set pricing to Gemini 3.1 Flash Image Preview rates\n- mirror updates in packaged backup model map\n- update llm cost calc regression test to cover new model
2026-02-26 14:41:59 -06:00
Dibyo Mukherjee
518cd3ef60 feat(ui): add key creation deep-links with SSO return URL support
Enables deep-linking directly to the key creation modal with prefilled
form data via URL parameters, including support for preserving these
deep-links through SSO authentication flows.

Key Creation Deep-links:
- Auto-open key creation modal via ?create=true parameter
- Prefill form fields from URL parameters (team_id, key_alias, models, etc.)
- Role-based access control for auto-open (requires write access)
- Race condition protection for redirect handling

Example: /ui?create=true&team_id=abc&key_alias=my-key&models=gpt-4,claude-3

SSO Return URL Preservation:
- Cookie-based return URL storage (works across ports for SSO flows)
- URL validation to prevent open redirect attacks
- Support for both dev and production environments

Co-Authored-By: Claude Opus 4.5 <noreply@anthropic.com>
2026-02-26 15:41:54 -05:00
Ryan Crabbe
0020e5929d fix: accurate wipe count on partial failure, remove stray blank line 2026-02-26 12:36:57 -08:00
Cesar Garcia
c6b6e29bc2
Merge pull request #22220 from Chesars/fix/reasoning-effort-output-config-claude-46
fix(anthropic): map reasoning_effort to output_config for Claude 4.6 models
2026-02-26 17:30:54 -03:00
Julio Quinteros Pro
de556ad237
Merge pull request #22202 from BerriAI/fix/realtime-streaming-mypy-attr
fix(mypy): suppress attr-defined errors on realtime websocket calls
2026-02-26 17:30:39 -03:00
yuneng-jiang
c3c79b9d80 [Fix] Virtual Keys pagination displays stale totals when filtering
When applying a key_alias filter on the Virtual Keys page, the pagination display (total count, page count) showed stale unfiltered values. The bug was that the table component tracked total_count only from the unfiltered useKeys hook, not from the filtered API response returned by useFilterLogic.

Fixes: When selecting a key alias that matches 1 key, the UI previously showed "Showing 1-50 of 509 results" and "Page 1 of 11" instead of the correct "Showing 1-1 of 1 results" and "Page 1 of 1".

Changes:
- Added filteredTotalCount state to useFilterLogic to track the total_count from filtered API responses
- Updated VirtualKeysTable to use filteredTotalCount (when set) instead of always using the unfiltered total
- Added comprehensive tests to prevent regression of pagination display logic

Type: 🐛 Bug Fix, ✅ Test

Co-Authored-By: Claude Haiku 4.5 <noreply@anthropic.com>
2026-02-26 12:28:43 -08:00
Chesars
5aff0e4da6 refactor: move _is_claude_4_6_model to AnthropicModelInfo, use explicit effort_map
- Move _is_claude_4_6_model from AnthropicConfig to AnthropicModelInfo
  to eliminate duplicated logic in is_effort_used
- Use explicit effort_map dict instead of passing unknown values through
  to output_config
2026-02-26 17:25:23 -03:00
Chesars
a8c95392fb fix(anthropic): map reasoning_effort to output_config for Claude 4.6 models
Claude 4.6 models use output_config as a stable API feature. This commit:
- Maps reasoning_effort to output_config for 4.6 models (minimal → low)
- Restricts effort="max" to Opus 4.6 only
- Skips beta header injection for 4.6 models
- Updates docs for Claude 4.6 effort support
2026-02-26 17:18:32 -03:00
Ryan Crabbe
0ee8cb5f02 refactor: remove shutdown cleanup, rely solely on startup wipe
Phase 2 (per-worker mark_process_dead on shutdown) only ever fired
when all workers shut down together, making it redundant — Phase 1
wipes everything on next startup anyway. This aligns with the
prometheus_client docs: just wipe the directory between runs.
2026-02-26 12:17:31 -08:00
yuneng-jiang
719b7fd013
Merge pull request #22157 from BerriAI/litellm_paginated_key_alias
[Feature] UI - Paginated Key Alias Select
2026-02-26 12:09:10 -08:00
yuneng-jiang
4b75a89673 Merge remote-tracking branch 'origin' into litellm_paginated_key_alias 2026-02-26 12:06:28 -08:00