Commit graph

26144 commits

Author SHA1 Message Date
Krrish Dholakia
ebba9e0b2a docs: cleanup docs 2025-10-02 09:51:16 -07:00
deepanshu
2512d89872 Add provider name to payload specification 2025-10-02 10:23:26 -04:00
Sameer Kankute
fb664b0f76 add other fields in cost calculations 2025-10-02 16:13:03 +05:30
Krish Dholakia
f6b67fd9bd
Merge pull request #14799 from tyler-liner/chore/generation-name-opentelemetry
fix (opentelemetry): use generation_name for span naming in logging method
2025-10-01 21:39:38 -07:00
Krish Dholakia
cb39a1bb60
Merge pull request #15024 from kowyo/main
docs: use docker compose instead of docker-compose
2025-10-01 21:34:42 -07:00
Krish Dholakia
93e347b3d8
Merge pull request #15106 from plafleur/ISSUE-15105
Guardrails - Don't run post_call guardrail if no text returned from Bedrock
2025-10-01 21:29:14 -07:00
Krrish Dholakia
83facb761f fix: fix placement 2025-10-01 18:57:27 -07:00
Krrish Dholakia
6589c924c8 fix: add linting 2025-10-01 18:54:50 -07:00
Krrish Dholakia
34366d8fe2 docs(key_management_endpoints.py): document new fields 2025-10-01 18:50:22 -07:00
Krrish Dholakia
32036e2093 feat(ui/): add common rate limit type form, for reuse for tpm/rpm across create + edit key 2025-10-01 18:46:43 -07:00
Ishaan Jaff
d538cf489a
[Feat] Fixes to dynamic rate limiter v3 - add saturatation detection (#15119)
* test cases dynamic rate limits

* fix _handle_generous_mode

* docs add readme

* use configs for vars

* fix debug

* add comment

* test_dynamic_rate_limiter_v3.py

* test_concurrent_pre_call_hooks_stress
2025-10-01 18:35:34 -07:00
Krrish Dholakia
20b6f011f7 fix(create_key_button.tsx): working ui controls to set guaranteed throughput/best effort throughput on key 2025-10-01 18:34:29 -07:00
Ishaan Jaffer
2bc5d93f23 use_callback_in_llm_call 2025-10-01 18:32:37 -07:00
Krrish Dholakia
ff866aae7b feat(create_key_button.tsx): add tpm/rpm rate limit type options to UI
allows user to set the type of tpm/rpm limit they're trying to set
2025-10-01 18:30:34 -07:00
Deepanshu Lulla
68adca04c8
Gitlab based Prompt manager (#14988)
* add prompt

* add prompt

* add prompt

* add prompt

* add prompt management via gitlab

* gitlab client

* gitlab client

* gitlab client

* fix lint issues

* fix lint issues

* remove router changes

---------

Co-authored-by: deepanshu <deepanshu.lulla@hq.bill.com>
2025-10-01 18:13:11 -07:00
Krrish Dholakia
d8e3a62fbe feat(key_management_endpoints.py): add support for guaranteed throughput on key update and service account key creation 2025-10-01 17:58:31 -07:00
Krrish Dholakia
757436f302 test: add unit test 2025-10-01 17:51:10 -07:00
Krrish Dholakia
3ce074564e feat(key_management_endpoints.py): add guaranteed throughput support for model specific tpm / rpm limits
prevents admin from granting keys more tpm/rpm than created for a key
2025-10-01 17:48:23 -07:00
Krrish Dholakia
a83238a2db test: add unit tests 2025-10-01 17:02:05 -07:00
Krrish Dholakia
005aec69c7 feat(key_management_endpoints.py): allow specifying rate limit type when creating tpm/rpm limits on keys
prevents overallocating tpm/rpm limits
2025-10-01 16:57:07 -07:00
Ishaan Jaff
388761f52d
[Fix] LiteLLM UI - Ensure OTEL settings are saved in DB after set on UI (#15118)
* fix: fix _add_callback_from_db_to_in_memory_litellm_callbacks

* test_add_callback_from_db_to_in_memory_litellm_callbacks

* fix otel

* fix: fix _add_callback_from_db_to_in_memory_litellm_callbacks
2025-10-01 15:33:22 -07:00
Ishaan Jaff
d9664a3ee4
fix gpt-5-chat-latest on model cost map (#15116) 2025-10-01 14:35:57 -07:00
Ishaan Jaff
e73d053de3
[Fix] Proxy Auth - Ensure LLM_API_KEYs can access pass through routes (#15115)
* test_virtual_key_llm_api_routes_allows_registered_pass_through_endpoints

* fix: is_registered_pass_through_route

* docs fix
2025-10-01 14:09:01 -07:00
Patrick Lafleur
8e5efd29df
Fix comment 2025-10-01 16:39:26 -04:00
Patrick Lafleur
d6be26dba6
Merge branch 'main' into ISSUE-15105 2025-10-01 14:44:40 -04:00
Luiz Rennó Costa
7e56600896
fix: model_group not always present in litellm_params, and metadata reference location (#15108)
Co-authored-by: Luiz Rennó Costa <luiz.renno@ifood.com.br>
2025-10-01 11:39:49 -07:00
Patrick Lafleur
e7fd1fb96b
Fix missing HTTPException import (#15111) 2025-10-01 11:39:30 -07:00
Sameer Kankute
56e429e33d refactor code for better handling cost 2025-10-01 23:38:53 +05:30
Patrick Lafleur
7ef71d4885
Fix text 2025-10-01 12:08:03 -04:00
Sameer Kankute
7ec7e5332c
Add generateContent cost tracking (#15014) 2025-10-01 09:03:45 -07:00
Patrick Lafleur
3e5d585f7d
Don't run post_call guardrail if no text returned from bedrock 2025-10-01 11:56:50 -04:00
Ishaan Jaffer
ab00ca2de9 bump: version 1.77.6 → 1.77.7 2025-10-01 08:54:03 -07:00
Sameer Kankute
c32f42098c Add cost tracking for /v1/messages 2025-10-01 19:57:16 +05:30
João Speglich
933b3979eb docs: update oci docs with oci_serving_mode 2025-10-01 10:45:50 -03:00
shagunb-acn
0c17689440
Merge branch 'BerriAI:main' into bugfix-14404-image-gen-azure-managed-identity 2025-10-01 09:58:13 +05:30
Krrish Dholakia
7fd24a2632 bump: version 1.77.6 → 1.77.7 2025-09-30 21:23:42 -07:00
Krrish Dholakia
a1a0e99638 fix(prometheus.py): working e2e calls w/ userapikeymetadata 2025-09-30 21:23:25 -07:00
Krish Dholakia
1503435d91
Merge pull request #15029 from henryhwang/gemini-adapter-fixes
feat(gemini): Add full support for native Gemini API translation
2025-09-30 21:19:18 -07:00
Ishaan Jaff
73bfef1a1f
Revert "[Feature]: Replace HTTPException with ParallelRequestLimitError in pa…" (#15095)
This reverts commit 71b9b58fa9.
2025-09-30 21:17:04 -07:00
Ishaan Jaffer
395c32c38d test_azure_openai_assistants_e2e_operations_stream 2025-09-30 21:16:29 -07:00
Krish Dholakia
64083111d3
(Feat) Add Vertex AI Live API WebSocket Passthrough with Cost Tracking
(Feat) Add Vertex AI Live API WebSocket Passthrough with Cost Tracking
2025-09-30 21:14:16 -07:00
Henry Wang
acc23b9757 fix issue from pr review 2025-10-01 11:57:10 +08:00
Henry H Wang
7c4439ba0a
Merge branch 'BerriAI:main' into gemini-adapter-fixes 2025-10-01 09:53:51 +08:00
Ishaan Jaffer
8dd31a5fe8 test_azure_openai_assistants_e2e_operations_stream 2025-09-30 18:47:09 -07:00
Ishaan Jaffer
04b3ac89b8 test: QueryParams 2025-09-30 18:45:38 -07:00
Uzair Ali
fcfe856e10
Add support for GPT 5 codex models (#14841)
* Add support for GPT 5 codex models

* lint

* fixes
2025-09-30 18:44:35 -07:00
Alexsander Hamir
26145da3e7
perf(router): optimize _filter_cooldown_deployments to O(n) (#15091)
Refactored to use set-based lookup and list comprehension instead
of two-pass approach with list.remove().

Old complexity: O(n×m + k×n)
- First loop: n deployments × m list lookups = O(n×m)
- Second loop: k removals × n list.remove() scans = O(k×n)

New complexity: O(m + n)
- Convert to set: O(m)
- Filter with O(1) set lookups: O(n)

Example with 100 deployments, 5 in cooldown:
- Old: ~1000 operations
- New: ~105 operations

Called on every request - high impact for production.
2025-09-30 18:39:12 -07:00
Ishaan Jaff
0ca11eefde
[Feat] Guardrails - add logging for important status fields (#15090)
* add StandardLoggingPayloadStatusFields

* add status_fields

* add StandardLoggingPayloadStatusFields

* noma guard: add_standard_logging_guardrail_information_to_request_data

* fix: StandardLoggingPayloadStatusFields

* fix tests

* fix StandardLoggingPayloadStatus

* get_standard_logging_object_payload

* test_bedrock_guardrail_status_failure

* fix: _get_status_fields

* fixes new guardrail tracing

* fix ruff
2025-09-30 18:38:07 -07:00
Henry Wang
4eee54b157 fix the test issue from the pr review 2025-10-01 09:08:25 +08:00
Ishaan Jaffer
f205b2c0a5 test fixes 2025-09-30 17:08:49 -07:00