Commit graph

33682 commits

Author SHA1 Message Date
Yuneng Jiang
77f2d8ed95
chore: fixes
Some checks failed
Unit Tests: Caching (Redis) / caching-redis (push) Has been cancelled
Unit Tests: Proxy DB Operations / proxy-db (auth-checks, tests/proxy_unit_tests/test_auth_checks.py tests/proxy_unit_tests/test_user_api_key_auth.py, 20, 8) (push) Has been cancelled
Unit Tests: Proxy DB Operations / proxy-db (key-generation, tests/proxy_unit_tests/test_key_generate_prisma.py, 30, 0) (push) Has been cancelled
Unit Tests: Proxy DB Operations / proxy-db (remaining, tests/proxy_unit_tests --ignore=tests/proxy_unit_tests/test_key_generate_prisma.py --ignore=tests/proxy_unit_tests/test_auth_checks.py --ignore=tests/proxy_unit_tests/test_user_api_key_auth.py, 20, 8) (push) Has been cancelled
Unit Tests: Security / security (push) Has been cancelled
2026-04-05 00:28:32 -07:00
Sameer Kankute
ae757721bf Fix routing of encrypted content 2026-03-03 18:32:41 +05:30
Sameer Kankute
5485772a5a Fix encrypted content streaming affinity issue 2026-03-03 18:16:00 +05:30
Sameer Kankute
c500ca0090 Add Regression tests for image_url blocks in assistant message content. 2026-03-03 18:16:00 +05:30
Sameer Kankute
cf64c6fe9a Add 'image_url; to both if the intent is to support it in assistant messages 2026-03-03 18:16:00 +05:30
Sameer Kankute
9238ed8783 add ChatCompletionImageObject in OpenAIChatCompletionAssistantMessage 2026-03-03 18:16:00 +05:30
Sameer Kankute
47daadd125 Add docs for opt out variable 2026-03-03 18:16:00 +05:30
Sameer Kankute
9a0731fe70 Add opt out varible for v1/messages to responses 2026-03-03 18:16:00 +05:30
Sameer Kankute
47e39add39 Add tests for messages to responses transformation: 2026-03-03 18:16:00 +05:30
Sameer Kankute
ca5c0448dd Add tranlation of context_management 2026-03-03 18:16:00 +05:30
Sameer Kankute
3e6b253cd6 Add v1 for anthropic responses transformation 2026-03-03 18:16:00 +05:30
Sameer Kankute
a635bfe39a Register custom pricing in litellm.model_cost 2026-03-03 18:16:00 +05:30
Sameer Kankute
9605462fce Fix free models working from UI 2026-03-03 18:16:00 +05:30
Sameer Kankute
9aee0ee462 Preserve forwarding server side called tools 2026-03-03 18:16:00 +05:30
Sameer Kankute
7d338ae89b Fix converse handling for parallel_tool_calls 2026-03-03 18:15:58 +05:30
Harshit Jain
e3fc3a4cec perf(spendlogs): optimize old spendlog deletion cron job 2026-03-03 18:15:52 +05:30
Emerson Gomes
b4aa506632 refactor(cost): dedupe image token usage cost helper
- extract shared calculate_image_response_cost_from_usage() helper\n- reuse helper in vertex and gemini image generation cost calculators\n- preserve provider-specific fallback to output_cost_per_image
2026-03-03 18:15:52 +05:30
Emerson Gomes
17cff584bc fix(cost): use token usage for gemini/vertex image generation when available
- compute image_generation cost from usage token metadata for vertex/gemini\n- map ImageUsage to Usage and reuse generic_cost_per_token\n- fallback to output_cost_per_image when usage metadata missing\n- add tests for token-based path and fallback path
2026-03-03 18:15:52 +05:30
Emerson Gomes
5eece691db feat(vertex): add gemini-3.1-flash-image-preview model DB support
- add gemini-3.1-flash-image-preview + vertex_ai alias entries\n- set pricing to Gemini 3.1 Flash Image Preview rates\n- mirror updates in packaged backup model map\n- update llm cost calc regression test to cover new model
2026-03-03 18:15:52 +05:30
Harshit28j
b1ade5fbd7 fix: relevant comment req changes 2026-03-03 18:15:52 +05:30
Harshit28j
b73d3a35ca fix: req changes 2026-03-03 18:15:52 +05:30
Harshit28j
3db85ca017 feat: add tags in project 2026-03-03 18:15:52 +05:30
Sameer Kankute
ec90e1b8c3 Revert "Fix mapping of parallel_tool_calls for bedrock converse" 2026-03-03 18:15:52 +05:30
yuneng-jiang
d0903e9ec0 Update litellm/proxy/public_endpoints/public_endpoints.py
Co-authored-by: greptile-apps[bot] <165735046+greptile-apps[bot]@users.noreply.github.com>
2026-03-03 18:15:52 +05:30
yuneng-jiang
74fb6b5c49 fix: add 12 missing endpoint keys to _ENDPOINT_METADATA, fix stale _schema keys in backup JSON 2026-03-03 18:15:52 +05:30
yuneng-jiang
8506536d71 fix: correct _ENDPOINT_METADATA keys to match actual JSON data (a2a, container_files) 2026-03-03 18:15:52 +05:30
yuneng-jiang
2851ed3ff8 [Feature] Add /public/endpoints endpoint for provider endpoint support
Add new /public/endpoints endpoint that returns which providers support each LiteLLM
endpoint (e.g., chat_completions, embeddings). The endpoint reads from a local backup
JSON file bundled with the package, caches the result in-process, and transforms the
raw provider-centric data into an endpoint-centric response format.

Changes:
- Add litellm/provider_endpoints_support_backup.json (copy of root source file)
- Add Pydantic response models (EndpointProvider, SupportedEndpoint, SupportedEndpointsResponse)
- Add /public/endpoints route with transformation and caching logic
- Add 16 comprehensive tests covering HTTP layer and transformation functions

Co-Authored-By: Claude Haiku 4.5 <noreply@anthropic.com>
2026-03-03 18:15:52 +05:30
yuneng-jiang
d1aceac2a4 adding build 2026-03-03 18:15:52 +05:30
yuneng-jiang
b3973ae919 bump: version 0.4.48 → 0.4.49 2026-03-03 18:15:52 +05:30
yuneng-jiang
17385be64a [Infra] Add prisma_schema_sync CircleCI job before e2e UI tests
Adds a new CircleCI job that runs the proxy with --use_prisma_db_push
against the base Neon branch before the e2e UI tests create their
branches from it, ensuring the schema is synced on the parent.

Co-Authored-By: Claude Sonnet 4.6 <noreply@anthropic.com>
2026-03-03 18:15:52 +05:30
Dharamendra Kumar
7ebb895f91 Remove nit 2026-03-03 18:15:52 +05:30
Dharamendra Kumar
ca417ff2f9 Restore 2026-03-03 18:15:52 +05:30
Dharamendra Kumar
d3946caa37 Update test to righ place 2026-03-03 18:15:52 +05:30
Dharamendra Kumar
9d1f3ae2eb [Fix] Enhance MidStreamFallbackError to preserve original status code and attributes
- Updated MidStreamFallbackError to retrieve and maintain the original status code from the wrapped exception.
- Ensured that message, request, and response fields remain consistent after calling the parent constructor.
- Added unit tests to verify the correct propagation of status codes and attributes in various scenarios.
2026-03-03 18:15:52 +05:30
Ryan Crabbe
faf79c5c72 fix: remove cache eviction close that kills in-use httpx clients 2026-03-03 18:15:52 +05:30
yuneng-jiang
3cb1a3c449 Revert "[Feature] Add /public/supported_endpoints endpoint" 2026-03-03 18:15:52 +05:30
yuneng-jiang
99b5406546 fix +Inf user budget metric when metadata max_budget is None
Same bug as team budget: _assemble_user_object fetched user info from DB
but only used budget_reset_at, discarding max_budget. When the key cache
has a stale None for user_max_budget, _safe_get_remaining_budget returns
+Inf. Now falls back to DB max_budget when metadata value is None.

Co-Authored-By: Claude Sonnet 4.6 <noreply@anthropic.com>
2026-03-03 18:15:52 +05:30
yuneng-jiang
b765ae2ae2 remove orphan comment from test file
Co-Authored-By: Claude Haiku 4.5 <noreply@anthropic.com>
2026-03-03 18:15:52 +05:30
yuneng-jiang
745064d458 fixing inf budget 2026-03-03 18:15:52 +05:30
Ishaan Jaff
e40a2a2bd6 fix(realtime): fix guardrails not firing for Gemini/Vertex AI and provider_config realtime WebSocket sessions (#22168)
* fix(gemini): enable inputAudioTranscription and handle transcription events for realtime guardrails

Gemini sends inputTranscription/outputTranscription inside serverContent separately from modelTurn/turnComplete. This adds handling to convert them into OpenAI-compatible events so the guardrail pipeline can inspect voice input, and enables inputAudioTranscription in the session setup config.

Made-with: Cursor

* fix(vertex_ai): enable inputAudioTranscription in realtime session config

Add inputAudioTranscription to the Vertex AI realtime setup so the backend returns transcripts of user speech, allowing guardrails to inspect voice input.

Made-with: Cursor

* fix(realtime): pass user_api_key_dict and guardrail metadata through async_realtime handler

The base LLM HTTP handler's async_realtime method was not accepting or forwarding user_api_key_dict and litellm_metadata to RealTimeStreaming. This meant guardrails configured with default_on=false were silently skipped for all provider_config-based realtime connections (Gemini, Vertex AI, etc). Also fixes wss:// connections when SSL_VERIFY=False by overriding ssl=False for secure WebSocket URLs.

Made-with: Cursor

* fix(realtime): forward guardrail metadata for generic provider_config and vertex_ai paths

The _arealtime function was not passing user_api_key_dict or litellm_metadata to base_llm_http_handler.async_realtime() for the generic provider_config path and the vertex_ai-specific path. This broke guardrail resolution since RealTimeStreaming.request_data was empty, causing should_run_guardrail to return False.

Made-with: Cursor

* fix(realtime): voice guardrail responses and block duplicate response.create on text input

When a guardrail blocks voice input, send a conversation.item.create + response.create to the backend so the LLM voices the guardrail message as audio instead of only returning text. Also adds pending_guardrail_message tracking to suppress the automatic response.create the client sends after a blocked text message, and broadens _has_audio_transcription_guardrails to match pre_call/post_call modes.

Made-with: Cursor

* test(realtime): update guardrail tests for broadened audio transcription check and add integration tests

Update existing tests to reflect that pre_call guardrails now correctly trigger the audio/VAD session.update injection. Add integration test file for live OpenAI realtime guardrail testing.

Made-with: Cursor

* fix(realtime): instruct LLM to say exact guardrail message verbatim

The previous prompt gave the LLM creative freedom to paraphrase the guardrail violation message. Now it instructs the LLM to repeat the exact configured message word for word.

Made-with: Cursor

* fix(realtime): preserve wss ssl semantics and move live guardrail test

Keep TLS enabled for wss realtime sessions while honoring SSL_VERIFY=False via a no-verify SSLContext, move the OpenAI live guardrail test into llm_translation, and dedupe duplicated guardrail-detection helpers to prevent drift.

Made-with: Cursor
2026-03-03 18:15:52 +05:30
Ishaan Jaff
cbeaebf826 _add_dd_apm_tags_for_litellm_call_id (#22219) 2026-03-03 18:15:52 +05:30
yuneng-jiang
76da970446 Move provider_endpoints_support.json into litellm package
The file was at the repo root and excluded from pip distributions. Moving it to litellm/proxy/public_endpoints/ alongside the other provider JSON files ensures it is packaged correctly. Updates all references in the endpoint handler, coverage tests, and release notes instructions.

Co-Authored-By: Claude Haiku 4.5 <noreply@anthropic.com>
2026-03-03 18:15:52 +05:30
yuneng-jiang
77dc425e9e [Feature] Add /public/supported_endpoints endpoint
Return all supported endpoints and which providers support them. Includes endpoint display names, URL paths, and per-provider support lists. Results are cached for the process lifetime.

Also adds comprehensive test coverage for the new endpoint.

Co-Authored-By: Claude Haiku 4.5 <noreply@anthropic.com>
2026-03-03 18:15:52 +05:30
yuneng-jiang
f59c164286 [Fix] Virtual Keys pagination displays stale totals when filtering
When applying a key_alias filter on the Virtual Keys page, the pagination display (total count, page count) showed stale unfiltered values. The bug was that the table component tracked total_count only from the unfiltered useKeys hook, not from the filtered API response returned by useFilterLogic.

Fixes: When selecting a key alias that matches 1 key, the UI previously showed "Showing 1-50 of 509 results" and "Page 1 of 11" instead of the correct "Showing 1-1 of 1 results" and "Page 1 of 1".

Changes:
- Added filteredTotalCount state to useFilterLogic to track the total_count from filtered API responses
- Updated VirtualKeysTable to use filteredTotalCount (when set) instead of always using the unfiltered total
- Added comprehensive tests to prevent regression of pagination display logic

Type: 🐛 Bug Fix,  Test

Co-Authored-By: Claude Haiku 4.5 <noreply@anthropic.com>
2026-03-03 18:15:52 +05:30
Ryan Crabbe
2944e448b3 fix: accurate wipe count on partial failure, remove stray blank line 2026-03-03 18:15:52 +05:30
Ryan Crabbe
b83d922279 refactor: remove shutdown cleanup, rely solely on startup wipe
Phase 2 (per-worker mark_process_dead on shutdown) only ever fired
when all workers shut down together, making it redundant — Phase 1
wipes everything on next startup anyway. This aligns with the
prometheus_client docs: just wipe the directory between runs.
2026-03-03 18:15:51 +05:30
Ryan Crabbe
02d5b7bde5 refactor: remove periodic dead PID cleanup, trim redundant tests
Startup wipe + graceful shutdown cleanup are sufficient. Remove
hourly mark_dead_pids scan, its helpers, and redundant test cases.
2026-03-03 18:15:51 +05:30
Ryan Crabbe
d10b94ecb1 fix: dup assignment, os.environ setting 2026-03-03 18:15:51 +05:30
Ryan Crabbe
cdaff4f0e3 wip redundant code 2026-03-03 18:15:51 +05:30
Ryan Crabbe
7f77cd4e00 clean up docstring 2026-03-03 18:15:51 +05:30