litellm/tests
Javier de la Torre e6a7cae7e1
fix(apscheduler): prevent memory leaks from jitter and frequent job intervals (#15846)
* fix(apscheduler): prevent memory leaks from jitter and frequent job intervals

Fixes critical memory leak in APScheduler that causes 35GB+ memory allocations
during proxy startup and operation. The leak was identified through Memray
analysis showing massive allocations in normalize() and _apply_jitter()
functions.

Key changes:
1. Remove jitter parameters from all scheduled jobs - jitter was causing
   expensive normalize() calculations leading to memory explosion
2. Configure AsyncIOScheduler with optimized job_defaults:
   - misfire_grace_time: 3600s (increased from 120s) to prevent backlog
     calculations that trigger memory leaks
   - coalesce: true to collapse missed runs
   - max_instances: 1 to prevent concurrent job execution
   - replace_existing: true to avoid duplicate jobs on restart
3. Increase minimum job intervals:
   - PROXY_BATCH_WRITE_AT: 30s (was 10s)
   - add_deployment/get_credentials jobs: 30s (was 10s)
4. Use fixed intervals with small random offsets instead of jitter for
   job distribution across workers
5. Explicitly configure jobstores and executors to minimize overhead
6. Disable timezone awareness to reduce computation

Memory impact:
- Before: 35GB with 483M allocations during startup
- After: <1GB with normal allocation patterns

Performance notes:
- Minimum job intervals increased from 10s to 30s (configurable via env vars)
- Jobs can still be distributed across workers using random start offsets
- No functional changes to job behavior, only timing and memory optimization

Testing:
- Added comprehensive test suite for scheduler configuration
- Verified no job execution backlog on startup
- Tested duplicate job prevention with replace_existing

Related issue: Memory leak in production proxy servers with APScheduler

\ud83e\udd16 Generated with [Claude Code](https://claude.ai/code)

Co-Authored-By: Claude <noreply@anthropic.com>

* docs: update PROXY_BATCH_WRITE_AT default value from 10s to 30s

Update documentation to reflect the new default value for PROXY_BATCH_WRITE_AT
changed in PR #15846. The default was increased from 10 seconds to 30 seconds
to prevent memory leaks in APScheduler.

🤖 Generated with [Claude Code](https://claude.com/claude-code)

Co-Authored-By: Claude <noreply@anthropic.com>

* refactor: Move APScheduler config to constants.py

Address code review feedback from ishaan-jaff:
- Move scheduler configuration variables (coalesce, misfire_grace_time,
  max_instances, replace_existing) to litellm/constants.py
- Update all references in proxy_server.py to use the constants
- Improves maintainability and makes configuration values centralized

Requested-by: @ishaan-jaff
Related: #15846

🤖 Generated with [Claude Code](https://claude.com/claude-code)

Co-Authored-By: Claude <noreply@anthropic.com>

---------

Co-authored-by: Claude <noreply@anthropic.com>
2025-10-28 19:30:17 -07:00
..
audio_tests test_azure_transcribe_model_mapping 2025-10-25 11:42:34 -07:00
basic_proxy_startup_tests fix(apscheduler): prevent memory leaks from jitter and frequent job intervals (#15846) 2025-10-28 19:30:17 -07:00
batches_tests [Feat] Batches - Add bedrock retrieve endpoint support (#14618) 2025-09-16 19:19:02 -07:00
code_coverage_tests fix _redact_base64 2025-10-28 17:38:16 -07:00
documentation_tests Litellm dev 12 28 2024 p1 (#7463) 2024-12-28 20:26:00 -08:00
enterprise fixes test 2025-10-28 17:43:04 -07:00
guardrails_tests [Feat] New Guardrail - Dynamo AI Guardrail (#15920) 2025-10-24 17:11:04 -07:00
image_gen_tests test_image_generation_azure_dall_e_3 2025-10-25 15:47:40 -07:00
litellm/llms feat: Add imageConfig parameter for gemini-2.5-flash-image (#15530) 2025-10-21 17:08:30 -07:00
litellm-proxy-extras Prisma Migrate - support setting custom migration dir (#10336) 2025-04-26 12:05:06 -07:00
litellm_utils_tests fix: Preserve Bedrock inference profile IDs in health checks (#15947) 2025-10-27 19:44:45 -07:00
llm_responses_api_testing test fix claude-sonnet-4-5-20250929 2025-10-28 19:05:13 -07:00
llm_translation test fix claude-sonnet-4-5-20250929 2025-10-28 19:05:13 -07:00
load_tests test: test_embedding_performance 2025-05-14 21:31:07 -07:00
local_testing test fix claude-sonnet-4-5-20250929 2025-10-28 19:05:13 -07:00
logging_callback_tests Fix: Redact reasoning summaries in ResponsesAPI output when message logging is disabled (#15965) 2025-10-28 16:42:41 -07:00
mcp_tests [Feat] add support for dynamic client registration (#15921) (enables Atlassian MCP to work via Oauth on LiteLLM) 2025-10-26 10:13:46 -07:00
multi_instance_e2e_tests fix: use fastuuid helper (#14903) 2025-09-25 15:47:01 -07:00
ocr_tests TestAzureAIOCR 2025-10-25 10:26:41 -07:00
old_proxy_tests/tests fix tests 2025-10-25 10:19:24 -07:00
openai_endpoints_tests feat: read from custom-llm-provider header (#15528) 2025-10-18 22:04:53 -07:00
otel_tests test_basic_moderations_on_proxy_with_model 2025-10-27 13:49:47 -07:00
pass_through_tests fix: always retain config models 2025-10-11 16:09:33 -07:00
pass_through_unit_tests Fix litellm_param based costing 2025-10-08 21:14:23 +05:30
proxy_admin_ui_tests fix: use fastuuid helper (#14903) 2025-09-25 15:47:01 -07:00
proxy_security_tests test fix 2025-10-04 10:57:02 -07:00
proxy_unit_tests test_foward_litellm_user_info_to_backend_llm_call 2025-10-27 13:48:23 -07:00
router_unit_tests [Feat] Add /search endpoint on LiteLLM Gateway (#15780) 2025-10-21 19:05:20 -07:00
scim_tests [Feat SSO] Add LiteLLM SCIM Integration for Team and User management (#10072) 2025-04-16 19:21:47 -07:00
search_tests TestGooglePSESearch 2025-10-25 17:13:45 -07:00
spend_tracking_tests fix: use fastuuid helper (#14903) 2025-09-25 15:47:01 -07:00
store_model_in_db_tests fix: use fastuuid helper (#14903) 2025-09-25 15:47:01 -07:00
test_litellm test fix claude-sonnet-4-5-20250929 2025-10-28 19:05:13 -07:00
unified_google_tests fix: gooogle GenAI route tests 2025-10-04 10:18:25 -07:00
vector_store_tests TestAzureOpenAIVectorStore 2025-10-25 14:06:28 -07:00
windows_tests [Bug Fix] UnicodeDecodeError: 'charmap' on Windows during litellm import (#10542) 2025-05-03 21:31:05 -07:00
__init__.py [Feat] Add github co-pilot as a new LLM API provider (#12325) 2025-07-04 13:12:16 -07:00
gettysburg.wav feat(main.py): support openai transcription endpoints 2024-03-08 10:25:19 -08:00
large_text.py fix(router.py): check for context window error when handling 400 status code errors 2024-03-26 08:08:15 -07:00
openai_batch_completions.jsonl feat(router.py): Support Loadbalancing batch azure api endpoints (#5469) 2024-09-02 21:32:55 -07:00
README.MD [Feat] MCP Gateway Fine-grained Tools Addition (#15153) 2025-10-03 10:16:29 -07:00
test_budget_management.py fix: use fastuuid helper (#14903) 2025-09-25 15:47:01 -07:00
test_callbacks_on_proxy.py fix: use fastuuid helper (#14903) 2025-09-25 15:47:01 -07:00
test_config.py fix: use fastuuid helper (#14903) 2025-09-25 15:47:01 -07:00
test_debug_warning.py fix(utils.py): fix togetherai streaming cost calculation 2024-08-01 15:03:08 -07:00
test_end_users.py fix: use fastuuid helper (#14903) 2025-09-25 15:47:01 -07:00
test_entrypoint.py (fix) clean up root repo - move entrypoint.sh and build_admin_ui to /docker (#6110) 2024-10-08 11:34:43 +05:30
test_fallbacks.py Ollama Chat - parse tool calls on streaming (#11171) 2025-05-27 16:14:49 -07:00
test_health.py (test) /health/readiness 2024-01-29 15:27:25 -08:00
test_keys.py test: temporarily skip test due to change testing model change - need to update test for new model 2025-05-09 09:02:08 -07:00
test_litellm_proxy_responses_config.py Add native Responses API support for litellm_proxy provider (#15347) 2025-10-08 18:31:26 -07:00
test_logging.conf feat(proxy_cli.py): add new 'log_config' cli param (#6352) 2024-10-21 21:25:58 -07:00
test_models.py test_add_model_run_health 2025-09-27 10:59:25 -07:00
test_openai_endpoints.py test fix claude-sonnet-4-5-20250929 2025-10-28 19:05:13 -07:00
test_organizations.py UI - fix adding vertex models with reusable credentials + fix pagination on keys table + fix showing org budgets on table (#10528) 2025-05-03 08:16:53 -07:00
test_passthrough_endpoints.py fix: update authorization header to use 'Bearer' instead of 'bearer' 2025-09-21 10:44:47 +00:00
test_ratelimit.py (Refactor / QA) - Use LoggingCallbackManager to append callbacks and ensure no duplicate callbacks are added (#8112) 2025-01-30 19:35:50 -08:00
test_resource_cleanup.py Fix: Properly close aiohttp client sessions to prevent resource leaks (#12251) 2025-07-09 09:25:17 -07:00
test_spend_logs.py Revert "Allow configuration to on what threshold to try truncating request content in db" 2025-08-28 16:12:42 -06:00
test_team.py build: publish new litellm-proxy-extras file 2025-05-27 17:44:23 -07:00
test_team_logging.py fix: use fastuuid helper (#14903) 2025-09-25 15:47:01 -07:00
test_team_members.py fix: use fastuuid helper (#14903) 2025-09-25 15:47:01 -07:00
test_users.py fix: use fastuuid helper (#14903) 2025-09-25 15:47:01 -07:00

In total litellm runs 1000+ tests

[02/20/2025] Update:

To make it easier to contribute and map what behavior is tested,

we've started mapping the litellm directory in tests/test_litellm

This folder can only run mock tests.