Commit graph

29947 commits

Author SHA1 Message Date
Alexsander Hamir
d5a35efe50 add 100k router memory test and increase memory limit to 60 MB
- Add test_memory_baseline_100k for testing at scale (100,000 requests)
- Increase MEMORY_LIMIT from 50 MB to 60 MB for all router tests
- Memory limit adjustment accounts for observed usage patterns in 30k test
2026-01-10 16:36:11 -08:00
Alexsander Hamir
02c4d009b3 optimize router memory tests for 5x speed improvement
- Increase batch size from 20 to 100 requests (matches proxy test pattern)
- Add memory tracking and detailed logging throughout test execution
- Add periodic memory checks every 10 batches
- Direct task creation without intermediate validation for maximum speed
- Explicit cleanup of responses after each batch
- Average 300-400 req/s (was ~50-100 req/s before)
- Increase memory limit from 40 MB to 50 MB to account for larger batch size
- 1k test now completes in ~6.5 seconds (was ~20+ seconds)
2026-01-10 16:30:49 -08:00
Alexsander Hamir
e81f2f7fc1 add 30k and 250k memory leak tests to CI
- Add test_memory_baseline_30k to Router tests
- Add test_proxy_memory_baseline_30k and test_proxy_memory_baseline_250k to Proxy tests
- Increase proxy tests timeout from 30m to 45m to accommodate 250k test (~16 min runtime)
2026-01-10 16:17:45 -08:00
Alexsander Hamir
c01be17a50 update CI to use renamed memory leak tests and add proxy server tests
- Update memory_leak_tests job to use test_router_acompletion_memory_growth.py (renamed)
- Add new step to run proxy server memory growth tests
- Test both Router (1k, 2k, 4k, 10k) and Proxy (1k, 10k, 20k) in CI
- Reduce timeout from 60m to 30m per section for faster feedback
- Remove 30k router test from CI to keep runtime reasonable
2026-01-10 16:14:06 -08:00
Alexsander Hamir
ff94e23cb3 improve memory leak detection tests with diagnostics and parallel support
- Rename test_linear_memory_growth.py to test_router_acompletion_memory_growth.py for clarity
- Add new test_proxy_chat_completions_memory_growth.py to test full proxy server memory usage
- Implement dynamic port allocation (18888-18919) in mock_server fixture for parallel test execution
- Add httpx client diagnostics to differentiate test artifacts from actual proxy memory growth
- Increase MEMORY_LIMIT to 170 MB based on empirical testing showing ~174 MB stable usage at scale
- Add aggressive cleanup routines with detailed logging to ensure accurate memory measurements
- Fix Windows Unicode compatibility by removing emoji characters from test output
- Tests now properly detect slow memory accumulation (~170 KB per 1000 requests) in proxy server
2026-01-10 16:11:11 -08:00
Ishaan Jaff
f19cce950c
add MANUS get response (#18900) 2026-01-10 12:21:45 -08:00
yuneng-jiang
1d614a8e85
Merge pull request #18901 from BerriAI/litellm_publih_proxy_extras_yj
[Infra] Proxy Extras CI/CD Fix
2026-01-10 12:04:31 -08:00
yuneng-jiang
0d2010813d Adding proxy extras build 2026-01-10 12:00:45 -08:00
yuneng-jiang
369b1fae69 bump: version 0.4.20 → 0.4.21 2026-01-10 12:00:03 -08:00
Elkhan Eminov
062f5892da
update OpenRouter docs to include embedding support (#18874) 2026-01-10 11:42:02 -08:00
Ishaan Jaff
ab50fea663
[Fix] turn_off_message_logging Does Not Redact Request Messages in proxy_server_request Field When Stored to Database (#18897)
* update _get_proxy_server_request_for_spend_logs_payload

* test_spend_logs_redacts_request_and_response_when_turn_off_message_logging_enabled
2026-01-10 11:28:21 -08:00
Ishaan Jaffer
29254d9071 _get_bedrock_client_ssl_verify 2026-01-10 11:28:04 -08:00
Ishaan Jaffer
fef5b0f471 ci/cd new release 2026-01-10 10:05:21 -08:00
yuneng-jiang
81789cc5a2
Merge pull request #18894 from BerriAI/litellm_ui_build_0077
[Infra] Building UI for QA Testing
2026-01-10 09:57:04 -08:00
yuneng-jiang
ffd3e40aec Building UI for testing 2026-01-10 09:48:20 -08:00
yuneng-jiang
b7a61d469c
Merge pull request #18798 from BerriAI/litellm_ui_endpoint_usage
[Feature] UI - Endpoint Activity in Usage
2026-01-10 09:42:21 -08:00
yuneng-jiang
cac9d862a1 fixing test 2026-01-10 09:34:41 -08:00
yuneng-jiang
227ad58c16 Removing mock data and adding tabs 2026-01-10 09:27:44 -08:00
yuneng-jiang
0d90baf3ac Merge remote-tracking branch 'origin' into litellm_ui_endpoint_usage 2026-01-10 09:08:38 -08:00
yuneng-jiang
3a2de85e7b
Merge pull request #18885 from BerriAI/litellm_ui_fallbacks_refactor
[Infra] UI - E2E Test: New DB Branch Per Test Run
2026-01-09 23:37:44 -08:00
yuneng-jiang
bdab97c520 Neon Branch on every commit + skip flaky tests 2026-01-09 23:25:42 -08:00
yuneng-jiang
b2c7a59d4e directly set db url 2026-01-09 22:57:28 -08:00
yuneng-jiang
ff04a23656 db url 2026-01-09 22:54:13 -08:00
yuneng-jiang
e2d61a6131 final 2026-01-09 22:43:01 -08:00
yuneng-jiang
78313d79ba adding expiry 2026-01-09 22:37:30 -08:00
yuneng-jiang
b0bbdc2fad neon delete 2026-01-09 22:26:32 -08:00
yuneng-jiang
70e904102e fixing export 2026-01-09 22:22:55 -08:00
yuneng-jiang
554fc0d39d fixing parent pt2 2026-01-09 22:14:31 -08:00
Sameer Kankute
cb03e5a6dd
Merge pull request #18852 from BerriAI/litellm_add_ssl_verify_bedrock
[Bug]: Add Custom CA certificates to boto3 clients
2026-01-10 11:41:32 +05:30
yuneng-jiang
bb015ab172 fixing parent 2026-01-09 22:09:00 -08:00
yuneng-jiang
773a1f434b parent id 2026-01-09 22:06:16 -08:00
yuneng-jiang
7d5716a3e0 create branch from e2e 2026-01-09 22:00:00 -08:00
yuneng-jiang
83157b83f8 testing neon 2026-01-09 21:56:22 -08:00
yuneng-jiang
a8e4a24189 testing only neon 2026-01-09 21:48:08 -08:00
yuneng-jiang
cf9b69aa12 Neon branching per e2e test 2026-01-09 21:38:19 -08:00
Alexsander Hamir
578abd4465
Add memory leak detection tests with CI integration (#18881) 2026-01-09 17:36:10 -08:00
yuneng-jiang
ce4e2132a2
Merge pull request #18880 from BerriAI/litellm_fs_router_fields
[Infra] Router Fields Endpoint + React Query for Router Fields
2026-01-09 16:57:40 -08:00
yuneng-jiang
f72394cac1 fixing build 2026-01-09 16:50:08 -08:00
yuneng-jiang
dfb298792c New endpoint for router fields + react query 2026-01-09 16:49:22 -08:00
YutaSaito
07db8fe656
Merge pull request #18855 from BerriAI/litellm_fix_mcp-error-in-multiple-server
[fix] mcp error in multiple servers
2026-01-10 07:26:16 +09:00
yuneng-jiang
8d00ebbd2f
Merge pull request #18877 from BerriAI/litellm_login_email_casing
[Fix] UI Login Case Sensitivity
2026-01-09 14:08:26 -08:00
yuneng-jiang
c29d042df4 Case insensitive email login 2026-01-09 12:51:00 -08:00
Sameer Kankute
0c7db97ad5
Merge pull request #18871 from BerriAI/litellm_fix_test_count_tokens_caching
Fix :test_count_tokens_caching
2026-01-10 01:20:55 +05:30
Harshit Jain
8a683d9a6a
Add fix for bedrock_cache, metadata and max_model_budget (#18872) 2026-01-10 01:09:00 +05:30
Sameer Kankute
aba7dcea9c Fix : litellm import error 2026-01-10 01:03:50 +05:30
Sameer Kankute
777ae4f530 Fix :test_count_tokens_caching 2026-01-10 00:57:11 +05:30
Robin
0575bd2d1c
feat: update prices json for novita provider (#18540)
* feat: add novita models

* feat: ci

* feat: add novita support josn
2026-01-10 00:48:06 +05:30
Shivam Rawat
691505cdae
added fix for org level budget enforcement (#18813) 2026-01-10 00:44:40 +05:30
Shivam Rawat
43dd0e6ef5
remove model before casting it in the transformation (#18810) 2026-01-10 00:43:38 +05:30
Cesar Garcia
c19c97591e
fix: align max_tokens with max_output_tokens for consistency (#18820)
* fix: align max_tokens with max_output_tokens for consistency

Fixed inconsistent max_tokens definitions in model_prices_and_context_window.json.
According to LiteLLM convention, max_tokens should equal max_output_tokens when available.

Models fixed:
- deepseek-chat: 131072 → 8192 (now equals max_output_tokens)
- dashscope/qwen-flash: 1000000 → 32768 (now equals max_output_tokens)
- databricks/databricks-gemma-3-12b: 128000 → 32000 (now equals max_output_tokens)

This ensures consistency across all providers where max_tokens represents
the maximum number of tokens that can be generated in the output.

* fix: align max_tokens with max_output_tokens for 244 models

- Fix 244 models where max_tokens != max_output_tokens
- Add test to validate max_tokens consistency and prevent regressions

According to model_prices_and_context_window.json spec:
- max_tokens is a LEGACY parameter
- Should always equal max_output_tokens when both are present

This ensures consistency across all model definitions.
2026-01-10 00:37:45 +05:30