Commit graph

26004 commits

Author SHA1 Message Date
Ishaan Jaffer
51457541b0 fix add allowed tools 2025-10-04 12:37:12 -07:00
Teddy Amkie
7a176804f3
Add sync models GitHub documentation with Loom video and cross-references (#15191)
- Add comprehensive sync_models_github.md with API endpoints and examples
- Include Loom video tutorial for Admin UI sync process
- Add cross-references from model_management.md, cost_tracking.md, and ui.md
- Provide both manual and automated sync options
- Include Python SDK usage examples
2025-10-04 12:32:01 -07:00
Sameer Kankute
8095de506a
Add streamGenerateContent cost tracking in passthrough (#15199)
* Add streamgenerate cost tracking for gemini provider

* add cost tracking test
2025-10-04 12:31:25 -07:00
Ishaan Jaffer
3d6342f852 test_twelvelabs_missing_input_type_error 2025-10-04 12:28:09 -07:00
Ishaan Jaffer
6c79e12367 fix schema 2025-10-04 12:26:36 -07:00
Ishaan Jaffer
be7ab0b3e3 test_gemini_url_context 2025-10-04 12:23:43 -07:00
Ishaan Jaffer
c38c5d20cc test dynamic rate limiter 3 2025-10-04 12:06:13 -07:00
Ishaan Jaffer
2dc11316f2 fix failing deepseek-ai/DeepSeek-V3.1 2025-10-04 11:50:30 -07:00
Ishaan Jaffer
cdfb53dd81 fix linting error 2025-10-04 11:40:47 -07:00
Ishaan Jaffer
7a41c09529 test bedrock embedding tests marengo 2025-10-04 11:27:29 -07:00
Ishaan Jaffer
44db58c8df test_e2e_bedrock_async_invoke_embedding_async_twelvelabs_marengo 2025-10-04 11:13:39 -07:00
Ishaan Jaffer
e7b570000e fix: azure passthrough test 2025-10-04 11:06:36 -07:00
Ishaan Jaffer
53503828b2 test fix 2025-10-04 10:57:02 -07:00
Ishaan Jaffer
6f298cf5f0 test_azure_openai_assistants_e2e_operations_stream 2025-10-04 10:53:46 -07:00
Ishaan Jaffer
4ab571f684 fix test /generateContent route 2025-10-04 10:49:46 -07:00
Ishaan Jaff
f78608082c
[Feat] Dynamic Rate Limiter v3 - fixes for detecting saturation + fixes for post saturation behavior (#15192)
* fix: test case 1, model hits saturation

* fix: _check_rate_limits test case 2

* fix: _get_priority_allocation

* test_default_priority_shared_pool

* fix: No Rate Limiting when low saturatation

* fix: correctly use  model_saturation_check

* fixes priority_descriptors

* fix: tune default PriorityReservationSettings
2025-10-04 10:45:17 -07:00
Alexsander Hamir
29a31e17dd
[Doc] Perf: Last week improvement (#15193)
* doc: perf update

* fix: mixed up changes

* fix: add gist
2025-10-04 10:43:13 -07:00
Ishaan Jaffer
ee36c30217 fix LF tests 2025-10-04 10:39:17 -07:00
Ishaan Jaffer
00f44861ea fix: gooogle GenAI route tests 2025-10-04 10:18:25 -07:00
Ishaan Jaffer
df9a19bc9d fix OPENAI_EMBEDDING_PARAMS 2025-10-04 10:11:21 -07:00
Ishaan Jaffer
1655a9aea8 fix lf tests 2025-10-04 10:02:56 -07:00
Ishaan Jaffer
eb417ed774 fix lf logging 2025-10-04 10:00:10 -07:00
Ishaan Jaffer
27fbfbe259 fix: include_subpath 2025-10-04 09:47:16 -07:00
Ishaan Jaffer
212a2e13f3 OTEL fix spans 2025-10-04 09:22:57 -07:00
Ishaan Jaffer
a040e0154a add openai 2025-10-04 09:19:00 -07:00
Ishaan Jaffer
4400a6c189 test bedrock guardrails 2025-10-04 09:18:26 -07:00
Alexsander Hamir
5d22229d35
[Fix] Cache - Avoiding expensive operations when cache isn't available (#15182)
* Optimize cache performance by avoiding expensive operations when caching is disabled

- Moved cache availability checks before expensive operations to improve performance for non-cached requests
- Updated client code to handle None responses from caching handler

* clean hot path

* Fix TypeError with isinstance check for CustomStreamWrapper in caching

Fixed `TypeError: typing.Any cannot be used with isinstance()` that was
occurring in the caching handler when checking cached streaming responses.

The issue was caused by CustomStreamWrapper being aliased to `typing.Any`
at runtime through the TYPE_CHECKING conditional import pattern. When the
code attempted to use isinstance(cached_result, CustomStreamWrapper) at
lines 222 and 338, it failed because Python's isinstance() cannot be used
with typing.Any.

Solution: Import CustomStreamWrapper at runtime separately from the
TYPE_CHECKING block, while keeping a type alias for static type checking.
This allows isinstance checks to work properly while maintaining type hints.

* fix: remove unnecessary type checking
2025-10-04 09:10:37 -07:00
Ishaan Jaffer
d50fbcdc00 fix: include_cost_in_streaming_usage 2025-10-04 09:05:02 -07:00
Ishaan Jaffer
89934f062c fix ruff check 2025-10-04 08:59:57 -07:00
Ishaan Jaffer
3d981680b0 fix ServerToolUse mypy lint error 2025-10-04 08:59:15 -07:00
Ishaan Jaffer
8e7c593f51 fix: transform_rerank_response 2025-10-04 08:55:48 -07:00
Ishaan Jaffer
67b0c874a3 ci/cd run again 2025-10-04 08:51:21 -07:00
Krish Dholakia
ad2270fe13
Merge pull request #14764 from daily-kim/litellm_fix_bearer_capitalization
Fix: Authorization header to use correct "Bearer" capitalization
2025-10-03 22:02:46 -07:00
Krish Dholakia
7d41f02d02
Merge pull request #14813 from shagunb-acn/bugfix-14404-image-gen-azure-managed-identity
#14404 BugFix - Add support for Azure AD token-based authorization in…
2025-10-03 21:53:59 -07:00
Krish Dholakia
c83a3ac9b4
Merge pull request #14939 from Toy-97/patch-1
update: DeepInfra model data refresh [2025-09-26]
2025-10-03 21:48:31 -07:00
Krish Dholakia
206f67a11e
Merge pull request #15183 from uc4w6c/fix/test_mcp_server
test: fix test_mcp_server.py
2025-10-03 21:44:52 -07:00
YutaSaito
81a8766b84
feat: add JP Cross-Region Inference (#15188) 2025-10-03 21:20:04 -07:00
Yuta Saito
359aaa947f test: fix test_mcp_server.py 2025-10-04 08:18:15 +09:00
Ishaan Jaff
10d6d72ae3
[Feat] VertexAI - Support googlemap grounding in vertex ai (#15179)
* add VertexToolName

* test_vertex_tool_params

* fix: working maps grounding

* test_gemini_google_maps_tool_simple

* test_vertex_ai_map_google_maps_tool_with_location

* fix  # noqa: PLR0915

* _extract_google_maps_retrieval_config

* fixes for linting

* docs: **Google Maps**
2025-10-03 16:07:51 -07:00
Ishaan Jaff
4415b195d1
Add "eu.anthropic.claude-sonnet-4-5-20250929-v1:0" in "model_prices_and_context_window.json" (#15181)
* feat: add eu.anthropic.claude-sonnet-4-5-20250929-v1:0

* fix: nvidia_nim_models
2025-10-03 15:52:18 -07:00
Krish Dholakia
3095f43e92
Merge pull request #15007 from TobiMayr/feature/add-max-requests-env-var
feature/add max requests env var
2025-10-03 15:30:15 -07:00
Krish Dholakia
c5de073732
Merge pull request #15072 from speglich/fix/oci-support
Fix OCI Generative AI  Integration when using Proxy
2025-10-03 15:29:21 -07:00
Krish Dholakia
343eeaf53f
Merge pull request #15102 from BerriAI/litellm_message_api_cost_tracking
(Feat) Add cost tracking for /v1/messages
2025-10-03 15:27:49 -07:00
Krish Dholakia
b41041fc49
Merge pull request #15146 from JVenberg/fix/session-cookie-infinite-logout-loop
[Fix] Session Token Cookie Infinite Logout Loop
2025-10-03 15:23:49 -07:00
Krish Dholakia
f05dc27551
Merge pull request #15160 from danielaskdd/fix-whitespace-handling
Fix(critical) : Preserve Whitespace Characters in Model Response Streams
2025-10-03 15:23:07 -07:00
rishiganesh2002
d36c8d6bbe
[Feat] MCP Gateway Fine-grained Tools Addition (#15153)
* feat: UI to add specific tools under creating MCP connection

* chore: pydantic + prisma changes

* feat: adding specific MCP tools now works

* fix: allowed tools filtering

* chore: filtered list to mcp server cost config

* chore: update Readme

* chore: refactor the filtering

* test: Added tests

When the allowed_tests is null, empty list or populated

* chore: resolve the proxy issue

* feat: updating MCP tool filtering
2025-10-03 10:16:29 -07:00
yangdx
8b95cd3ed2 Fix whitespace handling in _has_meaningful_content function
• Preserve newlines and spaces as content
• Remove .strip() call on strings
• Treat all non-empty strings as meaningful
• Update logic comment for clarity
• Fix edge case with whitespace-only text
2025-10-03 13:46:51 +08:00
shagunb-acn
13d38e2bed
Merge branch 'BerriAI:main' into bugfix-14404-image-gen-azure-managed-identity 2025-10-03 10:07:31 +05:30
Sameer Kankute
507b0973b4 fix lint code 2025-10-03 08:21:07 +05:30
Ishaan Jaffer
94a89cd7de ruff fix 2025-10-02 19:22:31 -07:00