Commit graph

27315 commits

Author SHA1 Message Date
Krrish Dholakia
45ec1ffdaa fix(ui/): set is sso modal visible to true for all users
allow backend to handle more complex state management
2025-10-08 13:13:21 -07:00
Achintya Rajan
7595acff42 empty commit for CI/CD 2025-10-08 12:03:44 -07:00
Sameer Kankute
0b6b69cd5b
fix issue with parsing assistant messages (#15320)
* Update factory.py

* reinclude the test cases

* fixed test

---------

Co-authored-by: Weijie <weijie-tan3@github.com>
2025-10-08 09:49:02 -07:00
Sameer Kankute
85d4142845 Fix litellm_param based costing 2025-10-08 21:14:23 +05:30
IQHL (Hans Jacob Landelius)
6633b33085 formatting 2025-10-08 14:15:15 +02:00
IQHL (Hans Jacob Landelius)
5b9d516e43 unit test 2025-10-08 13:58:25 +02:00
IQHL (Hans Jacob Landelius)
5a1b6fe1a6 Do not cooldown deployment if exception_status is empty string 2025-10-08 10:55:08 +02:00
Tim Elfrink
f8af72bbf1 fix: redact AWS credentials when redact_user_api_key_info enabled
Use SensitiveDataMasker in print_deployment() to mask all sensitive
credentials including AWS keys when redact_user_api_key_info is True.

Fixes #14839
2025-10-08 10:31:35 +02:00
Leslie Cheng
86affa405c Unused type 2025-10-07 21:11:56 -07:00
Leslie Cheng
f0c9dfbaf8 Add some tests 2025-10-07 20:57:08 -07:00
Leslie Cheng
1579752e06 Implement fix 2025-10-07 20:53:38 -07:00
Krish Dholakia
13703f289b
Merge pull request #15292 from timelfrink/fix/bedrock-prompt-caching-cost-calculation
fix(bedrock): include cacheWriteInputTokens in prompt_tokens calculation
2025-10-07 19:11:45 -07:00
Krish Dholakia
3f4d2c6ada
Add Cohere Embed v4 support for AWS Bedrock
Add Cohere Embed v4 support for AWS Bedrock
2025-10-07 19:10:22 -07:00
Krrish Dholakia
c671dceb24 feat(pass_through_endpoints.py): have updates not require pod restarts
ensures db updates work on live passthrough endpoints without requiring pod restarts
2025-10-07 19:03:32 -07:00
Ishaan Jaffer
5d6cd3ffed UI new build 2025-10-07 18:59:11 -07:00
Ishaan Jaffer
a8243079d3 doc fix 2025-10-07 18:51:27 -07:00
Ishaan Jaff
8a8cc5a1d3
[QA/Fixes] - Dynamic Rate Limiter v3 - final QA (#15311)
* fix: PriorityReservationSettings

* fix: use correct casting

* fix type casting
2025-10-07 18:42:17 -07:00
Krish Dholakia
1238464e20
Merge pull request #15303 from BerriAI/litellm_tenacity_upgrade
Upgrades tenacity version to 8.5.0
2025-10-07 18:19:52 -07:00
Krish Dholakia
60fb8cdb87
Merge pull request #15308 from BerriAI/litellm_models_page_crash_fix
fix: model + endpoints page crash when config file contains router_settings.model_group_alias
2025-10-07 18:19:15 -07:00
Achintya Rajan
4de4e4af1a Update page.tsx 2025-10-07 18:14:20 -07:00
Krrish Dholakia
032a213b9b feat(pass_through_endpoint.py): cleanup error message + validate it works across multiple instances 2025-10-07 18:04:44 -07:00
Krrish Dholakia
b4fc7d3cfa feat(pass_through_endpoints.py): raise 404 if passthrough endpoint has been deleted by the user
Ensures passthrough endpoint deletion works as expected (single-instance)
2025-10-07 17:58:15 -07:00
Ishaan Jaffer
7b82473bfb fix gpt-image-1-mini 2025-10-07 17:56:51 -07:00
Ishaan Jaffer
e1ab3620ee fix: mapped tests 2025-10-07 17:55:52 -07:00
Ishaan Jaffer
8590646b84 fix code QA check 2025-10-07 17:49:57 -07:00
Ishaan Jaffer
f2a96e6830 fix: e2eUI testing 2025-10-07 17:45:30 -07:00
Achintya Rajan
9fa8c6099e Update model_group_alias_settings.tsx 2025-10-07 17:42:04 -07:00
Ishaan Jaff
36c971a6fd
[MCP Gateway] QA/Fixes - Ensure Team/Key level enforcement works for MCPs (#15305)
* fix: _set_object_permission

* fix: _set_object_permission on teams

* fix: _set_object_permission

* fixes for team/key permissions

* statsh: object permission view

* fix: MCPServerPermissions

* fix: _get_team_object_permission

* test mcp checks for permissions

* fix server checks with prefix names

* test_list_tools_strips_prefix_when_matching_permissions

* ruff fix

* docs - refactor MCP

* docs update MCP docs

* docs allowed tools
2025-10-07 17:34:48 -07:00
Krrish Dholakia
6c4436215c feat(pass_through_endpoints.py): initial commit trying to delete passthrough endpoints successfully 2025-10-07 17:19:18 -07:00
Ishaan Jaff
7b56ba240e
[MCP Gateway] Litellm mcp fixes team control (#15304)
* fix: _set_object_permission

* fix: _set_object_permission on teams

* fix: _set_object_permission

* fixes for team/key permissions

* statsh: object permission view

* fix: MCPServerPermissions
2025-10-07 16:48:00 -07:00
Alexsander Hamir
49e04e0217
[Fix] Networking: remove limitations (#15302)
* fix: remove limitations

* fix: linter issues
2025-10-07 16:45:42 -07:00
Achintya Rajan
f72f21b261 Update requirements.txt 2025-10-07 16:18:48 -07:00
Ishaan Jaff
bc26eff98f
Fix: Make PATCH /model/{model_id}/update handle team_id consistently with POST /model/new (#15297)
* fix: _update_team_model_in_db

* test_patch_model_with_team_id_creates_proper_setup
2025-10-07 14:04:08 -07:00
Tim Elfrink
d71d801e4d Add Cohere Embed v4 support for AWS Bedrock
- Add cohere.embed-v4:0 to model pricing configs
- Update bedrock_embedding_models constant
- Update documentation with v4 model support

Fixes #15272
2025-10-07 22:08:12 +02:00
Ishaan Jaff
07a17d6d6b
[Feat] Proxy CLI - dont store existing key in the URL, store it in the state param (#15290)
* Feat: CLI Auth fixes for UI SSO

* fix auth.py

* fix test ui sso.py
2025-10-07 12:36:17 -07:00
Tim Elfrink
c5eb22381d fix(bedrock): include cacheWriteInputTokens in prompt_tokens calculation
Fixes #15263

This PR fixes the cost calculation for Bedrock Anthropic models with prompt caching.

**Root Cause:**
PR #9838 incorrectly removed adding `cacheWriteInputTokens` to `prompt_tokens`
for Bedrock, based on the assumption that it would cause double counting (similar
to an Anthropic API issue). However, Bedrock's token structure is different:

- **Bedrock API**: `inputTokens`, `cacheReadInputTokens`, and `cacheWriteInputTokens`
  are ALL separate values that should be summed for total input tokens
- **Anthropic API**: Same structure - all three token types are separate

The fix in #9838 was later reverted for Anthropic (correctly re-adding
`cache_creation_input_tokens` to `prompt_tokens`), but Bedrock was never fixed.

**Changes:**
1. Re-add `cacheWriteInputTokens` to `input_tokens` in Bedrock transformation
2. Update test assertions to reflect correct behavior
3. Add regression test for prompt caching cost calculation
4. Fix typo in Anthropic transformation where `cache_creation_tokens` was
   incorrectly set to `cache_read_input_tokens`

**Testing:**
- All existing Bedrock transformation tests pass
- New test validates correct cost calculation with prompt caching
- Verified costs are non-negative and accurate
2025-10-07 20:28:46 +02:00
Sameer Kankute
51971f4750
Add gpt-realtime-mini support (#15283) 2025-10-07 11:27:04 -07:00
Sameer Kankute
e4892735f0
fix gemini cli by actually streaming the response (#15264)
* fix gemini cli by actually streaming the response

* fix cost tracking

* fix test
2025-10-07 11:26:39 -07:00
Sameer Kankute
73f96712f5
fix the reasoningresponse id (#15265) 2025-10-07 11:24:29 -07:00
Krrish Dholakia
421d38c94a build(ui/): build new ui 2025-10-07 10:48:19 -07:00
Krish Dholakia
f044eb80de
Merge pull request #15285 from BerriAI/litellm_infinity_new_provider_ui
feature: adds Infinity as a provider in the UI
2025-10-07 10:46:05 -07:00
Achintya Rajan
e2f21beb7f added Infinity as a provider in the UI 2025-10-07 10:21:18 -07:00
Sameer Kankute
c0d0424eb8
Added streaming support for response api streaming image generation (#15269) 2025-10-07 08:15:57 -07:00
xprilion
3e52079509 Add information about personal entities error 2025-10-07 20:10:38 +05:30
Krrish Dholakia
60289aa73e feat(scim_v2.py): if group.id doesn't exist, use external id 2025-10-07 07:12:45 -07:00
xprilion
888391127f Add W&B Inference documentation 2025-10-07 19:18:45 +05:30
Krrish Dholakia
faeb7484db docs(vertex.md): fix doc 2025-10-06 20:45:58 -07:00
Achintya Rajan
7ed4715a9b eliminates loading screen flash 2025-10-06 20:42:59 -07:00
Achintya Rajan
cc94fb87b5 added base URL helpers 2025-10-06 20:32:47 -07:00
Krish Dholakia
12cbac74b1
Merge pull request #15210 from uc4w6c/feat/add_global_cross_region
feat: add Global Cross-Region Inference
2025-10-06 20:21:18 -07:00