mirror of
https://github.com/BerriAI/litellm.git
synced 2026-10-02 02:11:58 +00:00
* perf(router): fetch cooldown state and usage counters in one Redis round trip The cooldown filter (CooldownCache) and usage-based-routing-v2 selection (LowestTPMLoggingHandler_v2) each issued their own MGET on every request because they live in different objects. RoutingReadBatch fetches both key sets through DualCache.async_batch_get_cache_shared while the healthy deployments are resolved and hands the usage slice to the strategy, so selection does not read again. Each cache keeps its own memory tier, throttling, reservation rollback and circuit-breaker handling, and the strategy falls back to its own read when the prefetch does not cover its keys. simple-shuffle keeps reading only cooldowns. aresponses no longer issues a second, blocking response-cache read from the worker thread that runs the sync wrapper. Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com> * fix(caching): keep per-cache tier failures inside the shared batch read Wrap the memory-tier prepare and backfill steps of DualCache.async_batch_get_cache_shared so a failing tier degrades that cache's read to None the way async_batch_get_cache does, instead of escaping into routing. Drop the aresponses sync-cache guard: for native Responses models the worker-thread read is the one whose key matches the write, so skipping it broke cached /v1/responses replays. Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com> * refactor(router): rename usage key builder so the async cache-call check reads it as a key helper Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com> * fix(alerting): narrow daily-report cache values before numeric comparison Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com> * refactor(caching): type the shared batch-read helpers and merge Redis results without mutation Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com> * style: fix import sort in test_dual_cache Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com> * refactor(caching): flatten shared batch read keys without a stacked comprehension Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com> --------- Co-authored-by: yassin <yassin@berri.ai> Co-authored-by: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com> |
||
|---|---|---|
| .. | ||
| pre_call_checks | ||
| __init__.py | ||
| test_access_windows.py | ||
| test_add_retry_fallback_headers.py | ||
| test_auto_router_model_naming.py | ||
| test_auto_router_tuning_baseline.py | ||
| test_client_initalization_utils.py | ||
| test_cooldown_cache.py | ||
| test_cooldown_handlers.py | ||
| test_fallback_event_handlers.py | ||
| test_get_retry_from_policy.py | ||
| test_health_check_allowed_fails_integration.py | ||
| test_health_state_cache.py | ||
| test_pattern_match_deployments.py | ||
| test_reasoning_effort_capability.py | ||
| test_router_health_check_routing.py | ||
| test_router_interactions_endpoints.py | ||
| test_router_utils_common_utils.py | ||
| test_routing_read_batch.py | ||