Commit graph

10728 commits

Author SHA1 Message Date
frob
085320a024
Merge branch 'BerriAI:main' into ollama-image-handling 2024-05-01 01:53:01 +02:00
Krish Dholakia
ce9ede6110
Merge pull request #3370 from BerriAI/litellm_latency_buffer
fix(lowest_latency.py): allow setting a buffer for getting values within a certain latency threshold
2024-04-30 16:01:47 -07:00
Krrish Dholakia
0267069c6a fix(router.py): return routing args as dict 2024-04-30 15:39:14 -07:00
Krrish Dholakia
6a2b4bcab8 fix(router.py): only check /v1 for azure ai studio models
Fixes https://github.com/BerriAI/litellm/issues/3346
2024-04-30 15:29:50 -07:00
frob
6b55bd0dee
Merge branch 'BerriAI:main' into ollama-image-handling 2024-04-30 23:50:44 +02:00
Krrish Dholakia
3cc82f558e fix(utils.py): add exception mapping for gemini error 2024-04-30 14:17:10 -07:00
Krrish Dholakia
638477a023 test: fix test 2024-04-30 14:06:09 -07:00
Krrish Dholakia
df43012bdd test(test_router_fallbacks.py): use rpm test -> more stable 2024-04-30 14:01:01 -07:00
Krrish Dholakia
470dbf9745 build(ui): allow user to modify 'lowest_latency_buffer' on UI 2024-04-30 13:54:43 -07:00
Ishaan Jaff
0b0be700fc
Merge pull request #3371 from BerriAI/litellm_log_errors_db
[Feat] Write LLM Exception to LiteLLM Proxy DB
2024-04-30 13:36:41 -07:00
Ishaan Jaff
4b8fda4ac4 log startTime and EndTime for exceptions 2024-04-30 13:34:14 -07:00
Ishaan Jaff
3aad034a8b feat log request kwargs in error logs 2024-04-30 13:28:26 -07:00
Ishaan Jaff
ad5fddef15 fix log model_group 2024-04-30 13:11:09 -07:00
Ishaan Jaff
ee2a2ce559 fix - log api_base in errors 2024-04-30 13:02:42 -07:00
Ishaan Jaff
06804bc70a fix - working exception writing 2024-04-30 12:48:17 -07:00
Ishaan Jaff
22725bd44d fix types for errorLog 2024-04-30 12:31:33 -07:00
Ishaan Jaff
c7f979e0fe fix schema error logs 2024-04-30 12:31:19 -07:00
Ishaan Jaff
ac1cabe963 add LiteLLM_ErrorLogs to types 2024-04-30 12:16:03 -07:00
Krrish Dholakia
fee488bd53 test(test_image_generation.py): fix test 2024-04-30 12:15:37 -07:00
Krrish Dholakia
285a3733a9 test(test_image_generation.py): fix test 2024-04-30 12:14:29 -07:00
Krrish Dholakia
90cdfef1c1 fix(lowest_latency.py): allow setting a buffer for getting values within a certain latency threshold
if an endpoint is slow - it's completion time might not be updated till the call is completed. This prevents us from overloading those endpoints, in a simple way.
2024-04-30 12:00:26 -07:00
Ishaan Jaff
d6f7fa7f4e v0 prisma schema 2024-04-30 11:42:17 -07:00
Krrish Dholakia
398d503590 build(model_prices_and_context_window.json): add bedrock llama3 pricing 2024-04-30 11:36:29 -07:00
Krrish Dholakia
00d1440d0d test(test_image_generation.py): change img model for test - bedrock EOL 2024-04-30 08:55:40 -07:00
Krrish Dholakia
d717fa2588 test(test_tpm_rpm_routing_v2.py): fix test - bump number of iteration s 2024-04-30 08:48:55 -07:00
Krrish Dholakia
1cd24d8906 bump: version 1.35.32 → 1.35.33 2024-04-30 07:20:50 -07:00
Krrish Dholakia
020b175ef4 fix(lowest_tpm_rpm_v2.py): skip if item_tpm is None 2024-04-29 21:34:25 -07:00
Ishaan Jaff
81df36b298 docs - slack alerting 2024-04-29 21:33:03 -07:00
Ishaan Jaff
b1e888edad docs example logging to langfuse 2024-04-29 21:26:27 -07:00
Ishaan Jaff
0cad58f5c6 docs logging to langfuse on proxy 2024-04-29 21:26:15 -07:00
Ishaan Jaff
0c99ae9451 docs - fix kub.yaml config on docs 2024-04-29 21:20:29 -07:00
Krrish Dholakia
b46db8b891 feat(utils.py): json logs for raw request sent by litellm
make it easier to view verbose logs in datadog
2024-04-29 19:21:19 -07:00
Krrish Dholakia
f0e48cdd53 fix(router.py): raise better exception when no deployments are available
Fixes https://github.com/BerriAI/litellm/issues/3355
2024-04-29 18:48:04 -07:00
Krrish Dholakia
1e53c06064 test(test_router_caching.py): remove unstable test
test would fail due to timing issues
2024-04-29 18:37:31 -07:00
Krrish Dholakia
e7b4882e97 fix(router.py): fix high-traffic bug for usage-based-routing-v2 2024-04-29 16:48:01 -07:00
Krish Dholakia
09bae3d8ad
Merge pull request #3351 from elisalimli/main
Fix Cohere tool calling
2024-04-29 16:45:48 -07:00
Krish Dholakia
32534b5e91
Merge pull request #3358 from sumanth13131/usage-based-routing-RPM-fix
usage based routing RPM count fix
2024-04-29 16:45:25 -07:00
Krrish Dholakia
bd79e8b516 docs(langfuse_integration.md): add 'existing_trace_id' to langfuse docs 2024-04-29 16:40:38 -07:00
Krrish Dholakia
853b70aba9 fix(langfuse.py): support 'existing_trace_id' param
allow user to call out a trace as pre-existing, this prevents creating a default trace name, and potentially overwriting past traces
2024-04-29 16:39:17 -07:00
Krrish Dholakia
2cf069befb fix(langfuse.py): don't set default trace_name if trace_id given 2024-04-29 16:39:17 -07:00
Ishaan Jaff
d58dd2cbeb
Merge pull request #3360 from BerriAI/litellm_random_pick_lowest_latency
[Fix] Lowest Latency routing - random pick deployments when all latencies=0
2024-04-29 16:31:32 -07:00
Krrish Dholakia
77f155d158 docs(load_test.md): cleanup docs 2024-04-29 16:27:58 -07:00
Krrish Dholakia
af6a21f27c docs(load_test.md): add multi-instance router load test to docs 2024-04-29 16:25:56 -07:00
Ishaan Jaff
4cb4a7f06d fix - lowest latency routing 2024-04-29 16:02:57 -07:00
Krrish Dholakia
8f830bd948 docs(load_test.md): simplify doc 2024-04-29 16:00:02 -07:00
Krrish Dholakia
fcb83781ec docs(load_test.md): formatting 2024-04-29 15:58:41 -07:00
Krrish Dholakia
5fe0f38558 docs(load_test.md): load test multiple instances of the proxy w/ tpm/rpm limits on deployments 2024-04-29 15:58:14 -07:00
Ishaan Jaff
3b0aa05378 fix lowest latency - routing 2024-04-29 15:51:52 -07:00
Ishaan Jaff
5247d7b6a5 test - lowest latency router 2024-04-29 15:51:01 -07:00
Krrish Dholakia
cef2d95bb4 docs(routing.md): add max parallel requests to router docs 2024-04-29 15:37:48 -07:00