Commit graph

26144 commits

Author SHA1 Message Date
Ishaan Jaffer
f1578b49e2 vertex_httpx_mock_post 2025-09-29 18:25:54 -07:00
Copilot
f22fd4cddd
Fix: Add /v1/messages/count_tokens to Anthropic routes for non-admin user access (#15034)
* Initial plan

* Fix: Add /v1/messages/count_tokens to Anthropic routes for user access

Co-authored-by: ishaan-jaff <29436595+ishaan-jaff@users.noreply.github.com>

---------

Co-authored-by: copilot-swe-agent[bot] <198982749+Copilot@users.noreply.github.com>
Co-authored-by: ishaan-jaff <29436595+ishaan-jaff@users.noreply.github.com>
2025-09-29 18:16:52 -07:00
Ishaan Jaff
ebf72f5eb9
[Fix] Parallel Request Limiter v3 - use well known redis cluster hashing algorithm (#15052)
* test_keyslot_for_redis_cluster

* fix _is_redis_cluster

* Update litellm/proxy/hooks/parallel_request_limiter_v3.py

Co-authored-by: Copilot <175728472+Copilot@users.noreply.github.com>

---------

Co-authored-by: Copilot <175728472+Copilot@users.noreply.github.com>
2025-09-29 18:12:44 -07:00
Ishaan Jaff
f6d7683261
[Feat] LiteLLM Overhead metric tracking - Add support for tracking litellm overhead on cache hits (#15045)
* test_litellm_overhead

* vertex track overhead

* fix config.yaml used for testing

* test_litellm_overhead_stream

* add update_response_metadata for caching handler

* add CachingDetails

* fix update_response_metadata import

* add CachingDetails metrics

* add CachingDetails

* test_litellm_overhead_cache_hit

* test_litellm_overhead_cache_hit

* test_litellm_overhead_cache_hit
2025-09-29 17:33:27 -07:00
João Speglich
aeae6cffe4 oci: drop params automatically and add DEDICATED Support 2025-09-29 20:42:42 -03:00
Ishaan Jaffer
55110ba6ae Revert "fix _is_redis_cluster"
This reverts commit 67abd8880a.
2025-09-29 16:42:31 -07:00
Ishaan Jaffer
f303860881 Revert "test_keyslot_for_redis_cluster"
This reverts commit 52de33787b.
2025-09-29 16:42:15 -07:00
Ishaan Jaffer
52de33787b test_keyslot_for_redis_cluster 2025-09-29 16:41:32 -07:00
Ishaan Jaffer
67abd8880a fix _is_redis_cluster 2025-09-29 16:40:46 -07:00
Alexsander Hamir
d4830e34e5
fix: remove router inefficiencies (from O(M*N) to O(1)) - 62.5% faster P99 latency (#15046)
* fix: remove redundant deep copy

set_model_list already does the deep copy at the beginning of the call.

* fix: remove unused model_list arguments

The `model_list` parameter was being passed to classes that did not use it.

* fix: reduce per-request memory and time from O(N×M) to O(N)

No need to create a whole array for a simple look up.

* add: missing test

* fix: remove unused parameter
2025-09-29 15:49:46 -07:00
Ishaan Jaffer
e0172b86e2 test_litellm_overhead_non_streaming 2025-09-29 15:48:32 -07:00
Yuta Saito
5359a0d6a6 fix: test 2025-09-30 07:25:32 +09:00
Ishaan Jaff
619577d4e8
[Feat] Add litellm overhead metric for VertexAI (#15040)
* test_litellm_overhead

* vertex track overhead

* fix config.yaml used for testing

* test_litellm_overhead_stream

* add update_response_metadata for caching handler

* Revert "add update_response_metadata for caching handler"

This reverts commit f2a891f2b4.
2025-09-29 15:15:25 -07:00
Yuta Saito
f1f58bd1d1 fix: resolve regression with duplicate Mcp-Protocol-Version header 2025-09-30 07:12:56 +09:00
Ishaan Jaff
05955042d5
Add model pricing and context window for claude-sonnet-4-5 (#15049)
Co-authored-by: Cursor Agent <cursoragent@cursor.com>
Co-authored-by: ishaan <ishaan@berri.ai>
2025-09-29 15:09:40 -07:00
Krrish Dholakia
4a09507c58 fix(auth_utils.py): add model specific 'tpm_limit' to team's on litellm 2025-09-29 14:25:08 -07:00
Krrish Dholakia
117d5963d0 fix(auth_utils.py): check if team level model-specific rpm limit set 2025-09-29 14:24:03 -07:00
Krrish Dholakia
91f420160f docs(mcp.md): document oauth support 2025-09-29 13:47:21 -07:00
Krrish Dholakia
9ed83d44e3 test: remove unnecessary test 2025-09-29 13:35:34 -07:00
Krrish Dholakia
df828718c5 feat(user_api_key_auth_mcp.py): correctly dereference mcp server name from ids 2025-09-29 13:30:37 -07:00
Sameer Kankute
3ab1c31e4e
(Feat) Add cost tracking for Vertex AI Passthrough /predict endpoint (#15019)
* Add cost tracking for passthrough for predict endpoint

* restore file
2025-09-29 13:19:04 -07:00
Krrish Dholakia
bc6e6e7a28 fix(auth_checks.py): add auth checks to mcp server on call tools 2025-09-29 13:13:10 -07:00
Ishaan Jaff
038863a1fe
[Feat] Add new claude-sonnet-4-5 model family (#15041)
* add new claude-sonnet-4-5

* docs fix

* fix tool_use_system_prompt_tokens

* add anthropic.claude-sonnet-4-5-20250929 to bedrock converse models
2025-09-29 13:09:00 -07:00
Cedar Myers
5357b1e102 fix: remove invalid vertex -latest models 2025-09-29 16:04:33 -04:00
Ishaan Jaff
4dc23e8079
Revert "feat: improve vertex AI/gemini api_base handling for proxy services (…" (#15042)
This reverts commit ff2d19e4ca.
2025-09-29 13:03:46 -07:00
Krrish Dholakia
a9dcb51d03 fix: add lint 2025-09-29 12:28:43 -07:00
Zero Clover
ff2d19e4ca
feat: improve vertex AI/gemini api_base handling for proxy services (#15039) 2025-09-29 12:12:32 -07:00
Shubham Pathak
3f312d0527 Added test 2025-09-29 21:43:31 +05:30
Henry Wang
4e5db9476c fix mypy type check issues 2025-09-29 23:11:49 +08:00
Henry Wang
0c1104fbe0 fix(lint): Resolve F821 Undefined name errors in litellm/main.py 2025-09-29 22:42:09 +08:00
Ihsan Soydemir
a44b9ebb3c fix(files): use extra_query for GET/DELETE in Files endpoints 2025-09-29 14:07:18 +02:00
Kowyo
858f557bce docs: use docker compose instead of docker-compose 2025-09-29 11:59:53 +00:00
Sameerlite
58955a0348 Ignore type param for gemini tools 2025-09-29 16:32:21 +05:30
Henry Wang
99a884019b test(gemini): Add unit tests for Google GenAI adapter
This commit adds a comprehensive suite of unit tests for the Google GenAI adapter to ensure compliance with the project's contribution guidelines.

The new tests cover four main areas:
- Request parameter translation
- Streaming response handling
- Router methods for Google GenAI
- Proxy endpoints for Google GenAI

Additionally, this commit includes minor formatting and linting fixes identified during development.
2025-09-29 18:51:35 +08:00
anthony-liner
9ca73d5504 fix: set usage_details.total in langfuse integration 2025-09-29 16:30:44 +09:00
Shubham Pathak
b8195f091b Fixed formatting 2025-09-29 12:19:00 +05:30
Shubham Pathak
acdf9b64d9
Update request handling for original exceptions 2025-09-29 12:04:44 +05:30
shagunb-acn
11084868a3
Merge branch 'BerriAI:main' into bugfix-14404-image-gen-azure-managed-identity 2025-09-29 09:55:17 +05:30
Krish Dholakia
0d738b2899
Merge pull request #14986 from uc4w6c/fix/remove-servername-prefix-mcp_tools-tests
Fix/remove servername prefix mcp tools tests
2025-09-28 18:00:23 -07:00
Krish Dholakia
e7939b0521
Merge pull request #14983 from abhijitjavelin/main
Feat: Add Javelin standalone guardrails integration for LiteLLM Proxy
2025-09-28 17:59:07 -07:00
Krish Dholakia
0995b335ca
Merge pull request #15004 from BerriAI/cursor/update-litellm-docs-from-latest-release-d2ca
Update litellm docs from latest release
2025-09-28 17:57:47 -07:00
Krish Dholakia
b68510275b
Merge pull request #15005 from jpetrucciani/pop_atranscription
fix passthrough of atranscription into kwargs going to upstream provider
2025-09-28 17:57:00 -07:00
Krrish Dholakia
5d8e1c4409 fix: fix tests 2025-09-28 17:55:51 -07:00
Krrish Dholakia
f32b0364c4 fix: fix tests 2025-09-28 17:55:46 -07:00
Krish Dholakia
7454e5a118
Merge pull request #15008 from wenxi-onyx/add-ollama-cloud-models
feat: add ollama cloud models
2025-09-28 17:49:22 -07:00
Krish Dholakia
ddb81a6883
Merge pull request #15010 from eycjur/fix_vllm_audio_transcription_json_format
[Fix] response_format bug in hosted vllm audio_transcription
2025-09-28 17:47:15 -07:00
Krrish Dholakia
ba5cc5a5b3 fix: fix linting errors 2025-09-28 17:45:04 -07:00
eycjur
070094c739 Set the response_format as it is 2025-09-29 08:11:24 +09:00
Yuta Saito
be5010a675 test: fix tests 2025-09-29 07:21:32 +09:00
Yuta Saito
4f15d1240e fix: resolve linting errors 2025-09-29 07:13:20 +09:00