Commit graph

25955 commits

Author SHA1 Message Date
Krrish Dholakia
69edc546c0 fix(mcp_server_manager.py): don't return an invalid server id on list servers
Prevents user from hitting 'server not found' error, even when they see it being listed on litellm ui
2025-10-03 17:15:15 -07:00
Krrish Dholakia
742038bb5e build: deploy/
push new prisma migration
2025-10-03 16:52:03 -07:00
Krish Dholakia
3095f43e92
Merge pull request #15007 from TobiMayr/feature/add-max-requests-env-var
feature/add max requests env var
2025-10-03 15:30:15 -07:00
Krish Dholakia
c5de073732
Merge pull request #15072 from speglich/fix/oci-support
Fix OCI Generative AI  Integration when using Proxy
2025-10-03 15:29:21 -07:00
Krish Dholakia
343eeaf53f
Merge pull request #15102 from BerriAI/litellm_message_api_cost_tracking
(Feat) Add cost tracking for /v1/messages
2025-10-03 15:27:49 -07:00
Krish Dholakia
b41041fc49
Merge pull request #15146 from JVenberg/fix/session-cookie-infinite-logout-loop
[Fix] Session Token Cookie Infinite Logout Loop
2025-10-03 15:23:49 -07:00
Krish Dholakia
f05dc27551
Merge pull request #15160 from danielaskdd/fix-whitespace-handling
Fix(critical) : Preserve Whitespace Characters in Model Response Streams
2025-10-03 15:23:07 -07:00
rishiganesh2002
d36c8d6bbe
[Feat] MCP Gateway Fine-grained Tools Addition (#15153)
* feat: UI to add specific tools under creating MCP connection

* chore: pydantic + prisma changes

* feat: adding specific MCP tools now works

* fix: allowed tools filtering

* chore: filtered list to mcp server cost config

* chore: update Readme

* chore: refactor the filtering

* test: Added tests

When the allowed_tests is null, empty list or populated

* chore: resolve the proxy issue

* feat: updating MCP tool filtering
2025-10-03 10:16:29 -07:00
yangdx
8b95cd3ed2 Fix whitespace handling in _has_meaningful_content function
• Preserve newlines and spaces as content
• Remove .strip() call on strings
• Treat all non-empty strings as meaningful
• Update logic comment for clarity
• Fix edge case with whitespace-only text
2025-10-03 13:46:51 +08:00
Sameer Kankute
507b0973b4 fix lint code 2025-10-03 08:21:07 +05:30
Ishaan Jaffer
94a89cd7de ruff fix 2025-10-02 19:22:31 -07:00
Krish Dholakia
1d548d7235
Merge pull request #15155 from BerriAI/litellm_dev_10_02_2025_p1
Litellm dev 10 02 2025 p1
2025-10-02 19:04:59 -07:00
Ishaan Jaff
efa782d6d2
[Feat] Add Nvidia NIM Rerank Support (#15152)
* feat: add NvidiaNimRerankConfig

* fix: NvidiaNimRerankConfig

* fix: NvidiaNimRerankConfig

* fix routing to nvidia nim

* docs nvidia nim rerank

* TestNvidiaNim

* nvidia nim rerank fixes

* fix rerank

* transform_rerank_response

* Usage with LiteLLM Proxy

* fixes linting

* NvidiaNimRerankConfig.DEFAULT_NIM_RERANK_API_BASE

* fix Custom API Base URL

* fix rerank base

* fix main.py

* fix transform

* fix linting

* map_cohere_rerank_params

* ruff fix

* linting fixes

* ruff fix
2025-10-02 18:58:52 -07:00
Alexsander Hamir
79c24be48b
[Fix] - Router: optimize unhealthy deployment filtering in retry path (O(n*m) → O(n+m)) (#15110)
* perf(router): optimize unhealthy deployment filtering in retry path

Convert unhealthy_deployments list to set for O(1) lookups in
_async_get_healthy_deployments, reducing complexity from O(n*m) to O(n+m).

This method is called on every retry attempt (inside the retry loop), so
the optimization compounds during failures:

Before: 100 deployments × 50 unhealthy = 5,000 operations per call
After:  100 deployments + 50 unhealthy = 150 operations per call

Impact during cascading failures:
- With 1000 req/sec, 40% error rate, 3 retries
- Prevents 32M+ operations/sec during incidents
- Critical for preventing router CPU spikes when you need performance most

Matches the pattern already used in _filter_cooldown_deployments (line 7272).

* fix: Undefined name HTTPException
2025-10-02 18:55:18 -07:00
Sameer Kankute
544db8d140
(feat)Litellm x twelvelabs bedrock[Async Invoke Support] (#14871)
* Add async invoke support

* Add docs and correct embedding response

* fix cicd erros

* fix cicd erros

* fix mypy error

* Add litellm param input_type

* Update the docs
2025-10-02 18:52:33 -07:00
Ishaan Jaffer
9c29f35c4b test_end_user_jwt_auth 2025-10-02 18:48:11 -07:00
Krish Dholakia
bfd5f8a6f1
Merge pull request #15156 from DrQuacks/top-api-key-tags
Top api key tags
2025-10-02 18:47:17 -07:00
DrQuacks
9a27eb5d0c edited tests to allow build 2025-10-02 18:32:26 -07:00
João Speglich
afc9d258e3 oci: fix drop params parameter 2025-10-02 22:01:04 -03:00
Krish Dholakia
a630022256
Merge pull request #15151 from DrQuacks/top-api-key-tags
Top api key tags
2025-10-02 17:37:32 -07:00
Krish Dholakia
54095a256a
Merge pull request #15015 from anthony-liner/fix/langfuse-total-tokens
fix: set usage_details.total in langfuse integration
2025-10-02 17:33:43 -07:00
Ishaan Jaff
09106222c7
[Fix]: Handle non-serializable objects in Langfuse logging (#15148)
* fix: safe_deep_copy

* fix: import copy
2025-10-02 17:31:01 -07:00
Krish Dholakia
7b8426af53
Merge pull request #15130 from deepanshululla/feature/logging_payload_spec_update
Add provider name to payload specification
2025-10-02 17:29:24 -07:00
Krish Dholakia
6b3db2fd49
Merge pull request #15140 from niharm/fix-sonnet-4-5-200k-pricing
Price Fix: Add 200K prices for Sonnet 4.5
2025-10-02 17:28:47 -07:00
DrQuacks
d303bf6743 port less than 1 cent to spend column 2025-10-02 16:55:36 -07:00
DrQuacks
8760276918 dropdown for more than two tags 2025-10-02 16:45:36 -07:00
DrQuacks
a91a7e7750 added testing file for api keys dashboard 2025-10-02 16:17:02 -07:00
DrQuacks
4d340fa48f spend data in tags tooltip 2025-10-02 15:34:09 -07:00
Georg Wölflein
dbfa8ec921
Fix end user cost tracking in the responses API (#15124)
#13860
2025-10-02 15:13:57 -07:00
Ishaan Jaff
f8f4207994
[Security Fix] fix: don't log JWT SSO token on .info() log (#15145)
* fix: get_redirect_response_from_openid

* fix info log check

* fix: forward_upstream_to_client
2025-10-02 15:07:37 -07:00
Jack Venberg
560e96570e Fix: Session Token Cookie Infinite Logout Loop 2025-10-02 15:02:07 -07:00
Amir Refaee
8991657d67
added railtracks to projects using litellm (#15144) 2025-10-02 14:21:38 -07:00
Mubashir Osmani
f2107a189d
add azure_ai grok-4 model family (#15137)
* added oauth mcp to docs

* added azure ai/grok-4 model family

* Revert "added oauth mcp to docs"

This reverts commit 950b7cef44.
2025-10-02 14:21:12 -07:00
DrQuacks
7c5f789b67 styling tags, only showing tag column on tags tab 2025-10-02 13:49:01 -07:00
nihar
cd7acd2eb2 Add 200K prices for Sonnet 4.5 2025-10-02 12:39:44 -07:00
DrQuacks
fec3b55313 bringing tags into api top api keys table 2025-10-02 12:37:07 -07:00
Krrish Dholakia
ebba9e0b2a docs: cleanup docs 2025-10-02 09:51:16 -07:00
deepanshu
2512d89872 Add provider name to payload specification 2025-10-02 10:23:26 -04:00
Sameer Kankute
fb664b0f76 add other fields in cost calculations 2025-10-02 16:13:03 +05:30
Krish Dholakia
f6b67fd9bd
Merge pull request #14799 from tyler-liner/chore/generation-name-opentelemetry
fix (opentelemetry): use generation_name for span naming in logging method
2025-10-01 21:39:38 -07:00
Krish Dholakia
cb39a1bb60
Merge pull request #15024 from kowyo/main
docs: use docker compose instead of docker-compose
2025-10-01 21:34:42 -07:00
Krish Dholakia
93e347b3d8
Merge pull request #15106 from plafleur/ISSUE-15105
Guardrails - Don't run post_call guardrail if no text returned from Bedrock
2025-10-01 21:29:14 -07:00
Ishaan Jaff
d538cf489a
[Feat] Fixes to dynamic rate limiter v3 - add saturatation detection (#15119)
* test cases dynamic rate limits

* fix _handle_generous_mode

* docs add readme

* use configs for vars

* fix debug

* add comment

* test_dynamic_rate_limiter_v3.py

* test_concurrent_pre_call_hooks_stress
2025-10-01 18:35:34 -07:00
Ishaan Jaffer
2bc5d93f23 use_callback_in_llm_call 2025-10-01 18:32:37 -07:00
Deepanshu Lulla
68adca04c8
Gitlab based Prompt manager (#14988)
* add prompt

* add prompt

* add prompt

* add prompt

* add prompt management via gitlab

* gitlab client

* gitlab client

* gitlab client

* fix lint issues

* fix lint issues

* remove router changes

---------

Co-authored-by: deepanshu <deepanshu.lulla@hq.bill.com>
2025-10-01 18:13:11 -07:00
Ishaan Jaff
388761f52d
[Fix] LiteLLM UI - Ensure OTEL settings are saved in DB after set on UI (#15118)
* fix: fix _add_callback_from_db_to_in_memory_litellm_callbacks

* test_add_callback_from_db_to_in_memory_litellm_callbacks

* fix otel

* fix: fix _add_callback_from_db_to_in_memory_litellm_callbacks
2025-10-01 15:33:22 -07:00
Ishaan Jaff
d9664a3ee4
fix gpt-5-chat-latest on model cost map (#15116) 2025-10-01 14:35:57 -07:00
Ishaan Jaff
e73d053de3
[Fix] Proxy Auth - Ensure LLM_API_KEYs can access pass through routes (#15115)
* test_virtual_key_llm_api_routes_allows_registered_pass_through_endpoints

* fix: is_registered_pass_through_route

* docs fix
2025-10-01 14:09:01 -07:00
Patrick Lafleur
8e5efd29df
Fix comment 2025-10-01 16:39:26 -04:00
Patrick Lafleur
d6be26dba6
Merge branch 'main' into ISSUE-15105 2025-10-01 14:44:40 -04:00