Commit graph

10183 commits

Author SHA1 Message Date
Julio Quinteros Pro
98bc247618
Merge pull request #22143 from shivaaang/fix/llm-client-cache-unawaited-coroutine
fix(caching): store task references in LLMClientCache._remove_key
2026-02-28 00:13:52 -03:00
milan-berri
3e60ca3682
fix: populate user_id and user_info for admin users in /user/info (#22239)
* fix: populate user_id and user_info for admin users in /user/info endpoint

Fixes #22179

When admin users call /user/info without a user_id parameter, the endpoint
was returning null for both user_id and user_info fields. This broke
budgeting tooling that relies on /user/info to look up current budget and spend.

Changes:
- Modified _get_user_info_for_proxy_admin() to accept user_api_key_dict parameter
- Added logic to fetch admin's own user info from database
- Updated function to return admin's user_id and user_info instead of null
- Updated unit test to verify admin user_id is populated

The fix ensures admin users get their own user information just like regular users.

* test: make mock get_data signature match real method

- Updated MockPrismaClientDB.get_data() to accept all parameters that the real method accepts
- Makes mock more robust against future refactors
- Added datetime and Union imports
- Mock now returns None when user_id is not provided
2026-02-27 19:12:16 -08:00
milan-berri
594600dcb5
fix: Add PROXY_ADMIN role to system user for key rotation (#21896)
* fix: Add PROXY_ADMIN role to system user for key rotation

The key rotation worker was failing with 'You are not authorized to regenerate this key'
when rotating team keys. This was because the system user created by
get_litellm_internal_jobs_user_api_key_auth() was missing the user_role field.

Without user_role=PROXY_ADMIN, the system user couldn't bypass team permission checks
in can_team_member_execute_key_management_endpoint(), causing authorization failures
for team key rotation.

This fix adds user_role=LitellmUserRoles.PROXY_ADMIN to the system user, allowing
it to bypass team permission checks and successfully rotate keys for all teams.

* test: Add unit test for system user PROXY_ADMIN role

- Verify internal jobs system user has PROXY_ADMIN role
- Critical for key rotation to bypass team permission checks
- Regression test for PR #21896
2026-02-27 19:11:29 -08:00
Julio Quinteros Pro
8b50703f74
Merge pull request #22327 from BerriAI/fix/mcp-test-mock-filter-method-name
fix(mcp): update test mocks for renamed filter_server_ids_by_ip_with_info
2026-02-27 23:52:39 -03:00
Ishaan Jaff
8ce358e303
[Feat] Agent RBAC Permission Fix - Ensure Internal Users cannot create agents (#22329)
* fix: enforce RBAC on agent endpoints — block non-admin create/update/delete

- Add /v1/agents/{agent_id} to agent_routes so internal users can
  access GET-by-ID (previously returned 403 due to missing route pattern)
- Add _check_agent_management_permission() guard to POST, PUT, PATCH,
  DELETE agent endpoints — only PROXY_ADMIN may mutate agents
- Add user_api_key_dict param to delete_agent so the role check works
- Add comprehensive unit tests for RBAC enforcement across all roles

Co-authored-by: Ishaan Jaff <ishaan-jaff@users.noreply.github.com>

* fix: mock prisma_client in internal user get-agent-by-id test

Co-authored-by: Ishaan Jaff <ishaan-jaff@users.noreply.github.com>

* feat(ui): hide agent create/delete controls for non-admin users

Match MCP servers pattern: wrap '+ Add New Agent' button in
isAdmin conditional so internal users see a read-only agents view.
Delete buttons in card and table were already gated.
Update empty-state copy for non-admin users.
Add 7 Vitest tests covering role-based visibility.

Co-authored-by: Ishaan Jaff <ishaan-jaff@users.noreply.github.com>

---------

Co-authored-by: Cursor Agent <cursoragent@cursor.com>
Co-authored-by: Ishaan Jaff <ishaan-jaff@users.noreply.github.com>
2026-02-27 18:32:39 -08:00
Shivaang
fb72979432 fix(caching): store background task references in LLMClientCache._remove_key to prevent unawaited coroutine warnings
Fixes #22128
2026-02-27 21:23:56 -05:00
Julio Quinteros Pro
332dc32299
Merge pull request #22334 from BerriAI/fix/anthropic-passthrough-azure-test
fix(test): update Azure pass-through test after Responses API routing change
2026-02-27 23:09:04 -03:00
Julio Quinteros Pro
9b20a050ec
Merge pull request #22332 from BerriAI/fix/realtime-guardrail-test-assertions
fix(test): update realtime guardrail test assertions for voice violation behavior
2026-02-27 23:08:04 -03:00
Julio Quinteros Pro
2ac5365e06 fix: update stale docstring to match guardrail voicing behavior
Addresses Greptile review feedback.

Co-Authored-By: Claude Opus 4.6 <noreply@anthropic.com>
2026-02-27 23:04:43 -03:00
Ishaan Jaff
15fcd90b9c
feat: add in_flight_requests metric to /health/backlog + prometheus (#22319)
* feat: add in_flight_requests metric to /health/backlog + prometheus

* refactor: clean class with static methods, add tests, fix sentinel pattern

* docs: add in_flight_requests to prometheus metrics and latency troubleshooting
2026-02-27 18:00:50 -08:00
Julio Quinteros Pro
273994b996 fix(test): update Azure pass-through test to mock litellm.completion
Commit 99c62ca40e removed "azure" from _RESPONSES_API_PROVIDERS,
routing Azure models through litellm.completion instead of
litellm.responses. The test was not updated to match, causing it
to assert against the wrong mock.

Co-Authored-By: Claude Opus 4.6 <noreply@anthropic.com>
2026-02-27 22:59:38 -03:00
Julio Quinteros Pro
9a48c8e36a fix(test): update realtime guardrail test assertions for voice violation behavior
Tests were asserting no response.create/conversation.item.create sent to
backend when guardrail blocks, but the implementation intentionally sends
these to have the LLM voice the guardrail violation message to the user.

Updated assertions to verify the correct guardrail flow:
- response.cancel is sent to stop any in-progress response
- conversation.item.create with violation message is injected
- response.create is sent to voice the violation
- original blocked content is NOT forwarded

Co-Authored-By: Claude Opus 4.6 <noreply@anthropic.com>
2026-02-27 22:53:09 -03:00
Julio Quinteros Pro
cb4cfa1db4 fix(mcp): update test mocks to use renamed filter_server_ids_by_ip_with_info
Tests were mocking the old method name `filter_server_ids_by_ip` but production
code at server.py:774 calls `filter_server_ids_by_ip_with_info` which returns
a (server_ids, blocked_count) tuple. The unmocked method on AsyncMock returned
a coroutine, causing "cannot unpack non-iterable coroutine object" errors.

Co-Authored-By: Claude Opus 4.6 <noreply@anthropic.com>
2026-02-27 22:37:05 -03:00
Dylan Duan
af6fe184fb
docs: update AssemblyAI docs with Universal-3 Pro, Speech Understanding, and LLM Gateway (#21130)
* docs: update AssemblyAI docs with Universal-3 Pro, Speech Understanding, and LLM Gateway provider config

* feat: add AssemblyAI LLM Gateway as OpenAI-compatible provider
2026-02-27 17:24:48 -08:00
Ryan Crabbe
4fa6742b01 Add Prometheus child_exit cleanup for gunicorn workers
When a gunicorn worker exits (e.g. from max_requests recycling), its
per-process prometheus .db files remain on disk. For gauges using
livesum/liveall mode, this means the dead worker's last-known values
persist as if the process were still alive. Wire gunicorn's child_exit
hook to call mark_process_dead() so live-tracking gauges accurately
reflect only running workers.
2026-02-27 16:11:15 -08:00
Rahul Dhanawade
64c85dbc9f
Fix/claude code plugin schema (#22271)
* fix: add missing LiteLLM_ClaudeCodePluginTable to schema.prisma

- Claude Code Plugin Marketplace endpoints (/claude-code/marketplace.json,
  /claude-code/plugins) were returning 500 errors because
  LiteLLM_ClaudeCodePluginTable model was missing from both schema.prisma files
- Prisma client was generated without this table causing AttributeError:
  'Prisma' object has no attribute 'litellm_claudecodeplugintable'
- Added missing model definition to root schema.prisma and
  litellm/proxy/schema.prisma

Fixes #21310

* test: add regression test for LiteLLM_ClaudeCodePluginTable schema

* fix: address greptile review - add @updatedAt, clean up test imports
2026-02-27 15:59:37 -08:00
yuneng-jiang
8bb6457471 [Fix] Include created_at and updated_at in /project/list response
The /project/list endpoint was not returning created_at and updated_at timestamps because these fields were not defined in LiteLLM_ProjectTable. Added these fields to the model so FastAPI includes them in the response (values come from the database). This allows the UI to display project creation and last-updated times.

Co-Authored-By: Claude Haiku 4.5 <noreply@anthropic.com>
2026-02-27 15:41:03 -08:00
Alejandro Tapia
f1c563d2b2 org-exclusive-add-member 2026-02-27 14:58:17 -08:00
tombii
d292da2c14 fix(openrouter): pattern-based fix for native OpenRouter model double-stripping
Replace the hardcoded NATIVE_OPENROUTER_MODELS set approach with a
pattern-based check in _get_openai_compatible_provider_info: after
stripping the outer "openrouter/" provider prefix, if the remaining
model name still starts with "openrouter/", return immediately without
further stripping.

This fixes openrouter/openrouter/aurora-alpha, openrouter/openrouter/polaris-alpha,
and any future native OpenRouter models — not just the three hard-coded
ones (auto, free, bodybuilder) from the previous approach.

Fixes #16353

Co-Authored-By: Claude Sonnet 4.6 <noreply@anthropic.com>
2026-02-27 23:14:53 +01:00
Cesar Garcia
6430173bde
Merge pull request #20516 from Chesars/fix/openrouter-native-model-double-strip
fix(adapter): double-stripping of model names with provider-matching prefixes
2026-02-27 18:58:37 -03:00
Cesar Garcia
587977e19a
Merge pull request #19792 from Chesars/fix/openrouter-register-model-index-error
fix(register_model): handle openrouter models without '/' in name
2026-02-27 18:52:14 -03:00
Cesar Garcia
1e68b17a14
Merge pull request #19288 from Chesars/fix/helicone-gemini-support
fix(helicone): add Gemini and Vertex AI support to HeliconeLogger
2026-02-27 18:33:15 -03:00
Cesar Garcia
47a7645584
Merge pull request #22301 from Chesars/fix/count-tokens-include-system-and-tools
fix(count_tokens): include system and tools in token counting API requests
2026-02-27 18:29:31 -03:00
ryan-crabbe
d9cfae9092
Merge pull request #22313 from BerriAI/tests/add_llmclientcache_regression_tests
Tests: add llmclientcache regression tests
2026-02-27 13:04:02 -08:00
Ryan Crabbe
dce597b806 Close httpx clients after assertions to prevent resource leaks 2026-02-27 12:58:33 -08:00
Cesar Garcia
8f02d2d840
Merge pull request #21337 from Chesars/fix/streaming-parallel-tool-call-index
fix(responses): use output_index for parallel tool call streaming indices
2026-02-27 17:54:19 -03:00
Cesar Garcia
fc7bc9147f
Merge pull request #21629 from Chesars/fix/pydantic-serialization-warnings
fix(types): remove StreamingChoices from ModelResponse, use ModelResponseStream
2026-02-27 17:48:33 -03:00
Cesar Garcia
1552775166
Merge pull request #21595 from Chesars/fix/moonshot-preserve-image-url-content
fix(moonshot): preserve image_url blocks in multimodal messages
2026-02-27 17:47:02 -03:00
Ryan Crabbe
0b7e9a1971 Add e2e tests: httpx clients survive LLMClientCache eviction
Tests go through the real get_async_httpx_client() code path to verify
clients remain usable after both capacity eviction and TTL expiry.
Regression tests for PR #22247.
2026-02-27 12:43:08 -08:00
Ryan Crabbe
6490ad1d48 Revert "Add LLMClientCache regression tests for httpx client eviction safety"
This reverts commit ad9c70ec5d.
2026-02-27 12:43:03 -08:00
Cesar Garcia
bb8e6b1426
Merge pull request #21592 from Chesars/fix/openrouter-stream-usage-no-stream-options
fix(openrouter): use provider-reported usage in streaming without stream_options
2026-02-27 17:42:47 -03:00
Cesar Garcia
acf2fd9828
Merge branch 'main' into fix/openrouter-stream-usage-no-stream-options 2026-02-27 17:41:13 -03:00
Cesar Garcia
955bf90321
Merge pull request #21585 from Chesars/fix/vertex-gemini-image-config-params
fix(vertex_ai): pass through native Gemini imageConfig params for image generation
2026-02-27 17:39:53 -03:00
Cesar Garcia
6a9ea863b9
Merge branch 'litellm_oss_staging_02_27_2026' into fix/gpt5-search-supported-params 2026-02-27 17:31:38 -03:00
Cesar Garcia
493f2e9188
Merge pull request #21576 from Chesars/fix/gpt5-supported-params-audit
fix(openai): correct supported_openai_params for GPT-5 model family
2026-02-27 16:39:49 -03:00
Cesar Garcia
761ae7e896
Merge pull request #21498 from Chesars/fix/chatgpt-streaming-tool-call-indices
fix(chatgpt): fix tool_calls streaming indexes
2026-02-27 16:39:21 -03:00
Cesar Garcia
734655137e
Merge pull request #22307 from Chesars/fix/22244-image-edit-custom-pricing
fix(images): pass model_info/metadata in image_edit for custom pricing
2026-02-27 16:38:34 -03:00
Cesar Garcia
72980f4ded
Merge pull request #22300 from Chesars/fix/image-generation-extra-headers-22285
fix(images): forward extra_headers on OpenAI code path in image_generation()
2026-02-27 16:36:42 -03:00
Chesars
727adb0117 fix(images): pass model_info and metadata in image_edit for custom pricing
image_edit was not forwarding model_info/metadata to the logging object,
so custom_pricing was never detected. After PR #20679 stripped custom
pricing fields from the shared backend key, image_edit cost became 0.

Fixes #22244
2026-02-27 16:23:23 -03:00
ryan-crabbe
dc97e2f714
Merge pull request #22306 from BerriAI/tests/add_llmclientcache_regression_tests
Add LLMClientCache regression tests for httpx client eviction safety
2026-02-27 11:21:38 -08:00
yuneng-jiang
98e944c0cf
Merge pull request #22253 from BerriAI/litellm_access_group_sync
[Feature] Access group CRUD: Bidirectional team/key sync
2026-02-27 11:14:46 -08:00
Ryan Crabbe
ad9c70ec5d Add LLMClientCache regression tests for httpx client eviction safety
Regression tests for PR #22247 — ensures cache eviction (capacity and TTL)
does not close httpx clients that are still in use.
2026-02-27 11:14:13 -08:00
Gaurav Singh
29bb73ffca
fix(mcp): strip stale mcp-session-id header to prevent 400 in multi-worker deployments (#20992) (#21417)
In a multi-worker Uvicorn setup, a client that reconnects to a different
worker sends an mcp-session-id that the new worker has never seen.  The
MCP SDK returns 400 because the session is unknown.

Fix: add _handle_stale_mcp_session() which inspects the inbound
mcp-session-id header before the request reaches the SDK.  If the
session is not in this worker's _server_instances:
  - Non-DELETE: strip the header so the SDK creates a fresh session
  - DELETE: return 200 immediately (idempotent, session already gone)

No new dependencies, no Redis, no latency added to the hot path.

Fixes https://github.com/BerriAI/litellm/issues/20992
2026-02-27 10:59:08 -08:00
Noah Nistler
d13508c1c5
Enable local file support for OCR (#22133)
* [Docs] Enable local file support

Implemented internal handling for converting file-type documents to the required format for OCR processing, ensuring seamless integration with various providers.

* Refactor OCR file handling and improve security checks

Removed deprecated MIME type mapping and file conversion functions, replacing them with updated implementations. Enhanced security by rejecting 'file' document types in JSON requests, ensuring file uploads are handled via multipart/form-data. Updated tests to reflect these changes and ensure proper functionality.

* Enhance MIME type validation in OCR processing

Added a regular expression check to validate MIME types in the convert_file_document_to_url_document function, raising a ValueError for invalid types. Updated tests to ensure proper error handling for unsupported MIME types.

* Enhance type safety in OCR file handling

Added type casting for the uploaded file in the _parse_multipart_form function to ensure proper handling of UploadFile instances. This change improves type safety and reduces potential runtime errors during file processing.

* Refactor MIME type handling in document uploads

Updated the MIME type extraction logic to strip parameters from the Content-Type header, ensuring only the base type is used. Added tests to verify that MIME parameters are correctly handled and stripped in various scenarios.

* Update OCR documentation for MIME type recommendations and remove unnecessary tips

Clarified the recommended usage of MIME types for raw bytes in document uploads. Simplified the documentation by removing the tip about multipart file uploads from tools like Postman, ensuring a more concise and focused guide.

* Enhance multipart form handling in OCR endpoints

Updated the _parse_multipart_form function to ignore both 'file' and 'document' fields during form parsing, ensuring that the document built from the uploaded file is not overridden. Added a new test to verify that injected document fields do not affect the constructed document, improving security and robustness of the file upload process.
2026-02-27 10:50:02 -08:00
Chesars
c4458c09fe fix(count_tokens): include system and tools in token counting API requests
The /v1/messages/count_tokens proxy endpoint was only passing `messages`
to provider token counting APIs, discarding `system` and `tools`. This
caused clients like Claude Code to receive artificially low token counts
(e.g. 10 instead of 531), preventing proper context window management
and leading to context overflow errors.

Pass system and tools through the full chain:
- TokenCountRequest → proxy_server → provider counters → API handlers
- Bedrock: transform tools to toolConfig format, system to text blocks
- Anthropic/Azure AI: pass through directly (same API format)
2026-02-27 15:39:35 -03:00
Chesars
dc4f713c6d fix(images): forward extra_headers on OpenAI code path in image_generation()
Fixes #22285 — extra_headers passed to litellm.image_generation() were
silently dropped on the openai/litellm_proxy/openai_compatible_providers
code path. The azure and azure_ai paths already forwarded them correctly.
2026-02-27 15:17:58 -03:00
Harshit28j
2553698da5 feat: health check max tokens 2026-02-27 23:39:42 +05:30
Sameer Kankute
ec8aaa9d2f
Merge pull request #22155 from BerriAI/litellm_fix_image
[Bug]Add ChatCompletionImageObject in OpenAIChatCompletionAssistantMessage
2026-02-27 21:18:55 +05:30
Sameer Kankute
63c9b3a137
Merge pull request #22087 from BerriAI/litellm_fix_anthropic_responses
Add v1 for anthropic responses transformation
2026-02-27 21:18:04 +05:30
Sameer Kankute
e583489abe
Merge pull request #22260 from BerriAI/litellm_Fix_tool_pass
[Fix]Preserve forwarding server side called tools
2026-02-27 21:16:56 +05:30