- test_spend_log_queue_cap.py: Validates spend_log_transactions queue is
capped at MAX_SPEND_LOG_TRANSACTIONS, oldest entries are dropped, warning
is logged, and repeated overflows are handled correctly.
- test_redis_pool_cap.py: Validates DEFAULT_REDIS_MAX_CONNECTIONS is applied
to both URL-based and kwargs-based Redis connection pools, and that
user-specified max_connections is respected.
Co-authored-by: Ishaan Jaff <ishaan-jaff@users.noreply.github.com>
The KEEPALIVE_TIMEOUT env var / --keepalive_timeout CLI flag was only wired
to uvicorn's timeout_keep_alive parameter. The gunicorn code path was missing
this, causing a mismatch between ALB idle timeout and proxy keep-alive when
running with --run_gunicorn.
Fix: Add keepalive_timeout parameter to _run_gunicorn_server and set
gunicorn's 'keepalive' option accordingly.
Co-authored-by: Ishaan Jaff <ishaan-jaff@users.noreply.github.com>
Without a cap, redis-py defaults to 2^31 max connections (effectively
unlimited). Under high concurrency, this causes unbounded growth of
lock/MutexValue objects and file descriptors — a secondary memory leak
that contributes to worker RSS growth.
Fix: Default max_connections to 100 (configurable via
DEFAULT_REDIS_MAX_CONNECTIONS env var). Applied in both URL-based and
kwargs-based connection pool creation paths. Users who explicitly set
max_connections in their Redis config still get their custom value.
Co-authored-by: Ishaan Jaff <ishaan-jaff@users.noreply.github.com>
The PrismaClient.spend_log_transactions list is unbounded — every request
appends a deepcopy'd spend-log payload (~2-5KB). When the DB is unreachable
or flush can't keep up, this list grows without limit, causing workers to
crash with MemoryError.
Fix: Add MAX_SPEND_LOG_TRANSACTIONS cap (default 10000, ~20-50MB max).
When the queue is full, drop the oldest 10% of entries and log a warning.
This prevents OOM while preserving the most recent spend data.
Configurable via MAX_SPEND_LOG_TRANSACTIONS env var.
Co-authored-by: Ishaan Jaff <ishaan-jaff@users.noreply.github.com>
Reproduction confirms the production issue: proxy workers grow in memory
under load and eventually crash with MemoryError.
The root cause is PrismaClient.spend_log_transactions — an unbounded list
that every request appends a deepcopy'd spend-log payload to. When the DB
is unreachable (or flush can't keep up), this list grows without limit.
Reproduction results:
- Workers die from MemoryError when RLIMIT_AS cap is set
- RSS grows linearly at ~0.3 MB/s at 550 rps
- Without cap: 295MB → 396MB over 300s (100MB growth, 150K requests)
- With RLIMIT_AS=800MB: workers die within seconds
- uvicorn respawns workers which immediately start growing again
Scripts:
- fake_openai_server.py: Mock OpenAI endpoint (instant responses)
- proxy_wrapper.py: Wraps proxy app with leak-simulating middleware
- blast.py: High-concurrency async load generator
- watch_workers.py: RSS monitor that detects worker deaths
- repro_worker_death.py: All-in-one orchestrator
Co-authored-by: Ishaan Jaff <ishaan-jaff@users.noreply.github.com>
Set prometheus_emit_stream_label: true in litellm_settings to emit a
stream label (True/False/None) on litellm_proxy_total_requests_metric.
Opt-in to avoid breaking cardinality on existing deployments.
BaseHTTPMiddleware wraps streaming responses with receive_or_disconnect
per chunk, blocking the event loop and causing severe throughput
degradation under concurrent streaming load (53% of CPU in profiling).
Converts PrometheusAuthMiddleware to a pure ASGI middleware using the
__call__(scope, receive, send) protocol.
- Remove print_verbose calls that format chunk/response Pydantic objects,
triggering millions of __repr__ calls (8% of CPU in profiling)
- Guard remaining verbose_logger.debug with isEnabledFor(DEBUG) and use
lazy %s formatting instead of f-strings
- Replace usage stripping round-trip (model_dump + delete + reconstruct)
with a _usage_stripped flag, deferring exclusion to serialization time
- Remove verbose_proxy_logger.debug that formatted every streaming chunk
- Honor _usage_stripped flag from streaming handler to exclude usage
during model_dump_json serialization instead of reconstructing objects
AsyncHTTPHandler.__del__ was closing httpx clients still in use by
AsyncOpenAI/AsyncAzureOpenAI due to independent cache lifecycles.
Restores standalone httpx client creation for OpenAI/Azure providers.
- Updated title to highlight Logs v2 feature
- Simplified Key Highlights to focus on Logs v2 / tool call tracing
- Rewrote Logs v2 description with improved language style
- Removed Claude Agents SDK and RAG API from key highlights section
- TODO: Add image (logs_v2_tool_tracing.png)
- Check for both litellm_proxy_failed_requests_metric_total and the deprecated litellm_llm_api_failed_requests_metric_total
- The proxy-level failure hook may not always be called depending on where the exception occurs
- Simplify total_requests check to only verify key fields
Co-authored-by: Cursor <cursoragent@cursor.com>
* litellm_fix_mapped_tests_core: fix test isolation and mock injection issues
## Problem
Four tests in litellm_mapped_tests_core were failing:
1. test_register_model_with_scientific_notation - KeyError due to test isolation issues
2. test_search_uses_registry_credentials - Mock not being called due to incorrect patch path
3. test_send_email_missing_api_key - Real API calls despite mocking
4. test_stream_transformation_error_sync - Mock not effective, real API called
## Solution
### test_register_model_with_scientific_notation
- Use unique model name to avoid conflicts with other tests
- Clear LRU caches before test to prevent stale data
- Clean up model_cost entry after test
### test_search_uses_registry_credentials
- Use patch.object() on the actual base_llm_http_handler instance
- String-based patching for instance methods can fail; direct object patching is more reliable
### test_send_email_missing_api_key
- Directly inject mock HTTP client into logger instance
- This bypasses any caching issues that could cause the fixture mock to be ineffective
### test_stream_transformation_error_sync
- Patch litellm.completion directly instead of the handler module's litellm reference
- This ensures the mock is effective regardless of import order
## Regression
These tests were affected by LRU caching added in #19606 and HTTP client caching.
* fix(test): use patch.object for container API tests to fix mock injection
## Problem
test_retrieve_container_basic tests were failing because mocks weren't
being applied correctly. The tests used string-based patching:
patch('litellm.containers.main.base_llm_http_handler')
But base_llm_http_handler is imported at module level, so the mock wasn't
intercepting the actual handler calls, resulting in real HTTP requests
to OpenAI API.
## Solution
Use patch.object() to directly mock methods on the imported handler
instance. Import base_llm_http_handler in the test file and patch like:
patch.object(base_llm_http_handler, 'container_retrieve_handler', ...)
This ensures the mock is applied to the actual object being used,
regardless of import order or caching.
* fix(test): add missing Prometheus metric labels to test_proxy_failure_metrics
Add client_ip, user_agent, model_id labels to expected metric patterns.
These labels were added in PRs #19717 and #19678 but test wasn't updated.
* fix(test_resend_email): use direct mock injection for all email tests
Extend the mock injection pattern used in test_send_email_missing_api_key
to all other tests in the file:
- test_send_email_success
- test_send_email_multiple_recipients
Instead of relying on fixture-based patching and respx mocks which can
fail due to import order and caching issues, directly inject the mock
HTTP client into the logger instance. This ensures mocks are always used
regardless of test execution order.
* fix(test): use patch.object for image_edit and vector_store tests
- test_image_edit_merges_headers_and_extra_headers: import base_llm_http_handler
and use patch.object instead of string path patching
- test_search_uses_registry_credentials: import module and patch via
module.base_llm_http_handler to ensure we patch the right instance
---------
Co-authored-by: Ishaan Jaff <ishaanjaffer0324@gmail.com>
The user_id field 'default_user_id' is being masked to '*******_user_id'
in prometheus metrics for privacy. Updated test expectations to match
the actual behavior.
Co-authored-by: Cursor <cursoragent@cursor.com>
## Problem
Tests using mocked HTTP clients were hitting real APIs because:
1. HTTP client cache was returning previously cached real clients
2. isinstance checks failed due to module identity issues from sys.path
### Tests affected:
- test_send_email_missing_api_key
- test_send_email_multiple_recipients (resend & sendgrid)
- test_search_uses_registry_credentials
- test_vector_store_create_with_simple_provider_name
- test_vector_store_create_with_provider_api_type
- test_vector_store_create_with_ragflow_provider
- test_image_edit_merges_headers_and_extra_headers
- test_retrieve_container_basic (container API tests)
## Solution
1. Add clear_client_cache fixture (autouse=True) to clear
litellm.in_memory_llm_clients_cache before each test
2. Fix isinstance checks to use type name comparison
(avoids module identity issues from sys.path.insert)
## Why not disable_aiohttp_transport
The default transport is aiohttp, so tests should work with it.
Clearing the cache ensures mocks are used instead of cached real clients.
## Regression
PR #19829 (commit f95572e3ed) added @respx.mock but cached clients
from earlier tests were being reused, bypassing the mocks.
Co-authored-by: shin-bot-litellm <shin-bot-litellm@users.noreply.github.com>
* docs: add card-based blog index page for mobile navigation
Fixes#20100 - the blog landing page showed post content directly
instead of an index, with no way to navigate between posts on mobile.
- Swizzle BlogListPage with card-based grid layout
- Featured latest post spans full width with badge
- Responsive 2-column grid with orphan handling
- Pagination, SEO metadata, accessibility (aria-label, dateTime, heading hierarchy)
- Add description frontmatter to existing blog posts
* docs: add deterministic fallback colors for unknown blog tags
* docs: rename blog heading to The LiteLLM Blog
- Add /v1/vector_store/list route for OpenAI API compatibility (fixes test_routes_on_litellm_proxy)
- Fix Bedrock Converse API model format (bedrock_converse/ → bedrock/converse/)
- Fix Nova Premier inference profile prefix (amazon. → us.amazon.)
- Add STABILIZATION_TODO.md to .gitignore
Tested locally - all affected tests now pass
Co-authored-by: Cursor <cursoragent@cursor.com>
The test had prompt_tokens=1000 but the sum of token details was 1150
(text=700 + audio=100 + cached=200 + cache_creation=150).
This triggered the double-counting detection logic which recalculated
text_tokens to 550, causing the assertion to fail.
Fixed by setting prompt_tokens=1150 to match the sum of details.