mirror of
https://github.com/BerriAI/litellm.git
synced 2026-10-08 03:08:45 +00:00
|
Some checks failed
LiteLLM Rust / rustfmt, clippy, test (push) Has been cancelled
Warming resolved its target set from configuration, so every active session was replayed against a representative of every tier on every interval whether or not it had ever been routed there. Most sessions never leave their starting tier, so that spent N replays per interval to keep caches warm that nobody would read, and for a pooled tier it warmed the wrong member entirely. The set is now per session and comes from what the session actually did. Capture records each served model into a per-session Redis set inside the same atomic script that writes the record, sharing the session's hash tag, so the touched set cannot disagree with the record it belongs to and expires with it. The refresher replays exactly that set. This is the intended shape of the feature: keep a session's own caches alive so returning to a tier it has already used is a read, rather than pre-warming tiers on speculation. The first switch to a new tier is a normal cache write, and every visit after it is warm. warm_models changes meaning accordingly, from a pre-warm list to an allowlist that narrows what a session may be warmed on, and resolve_warm_models now returns every model across the tier pools since it bounds eligibility rather than naming the targets. Tests that expected a replay on a tier the seeded session had never visited now declare a session that has been to both, which is the case the feature serves. |
||
|---|---|---|
| .. | ||
| adaptive_router | ||
| complexity_router | ||
| test_auto_router.py | ||
| test_base_routing_strategy.py | ||
| test_budget_limiter_hotpath.py | ||
| test_complexity_router.py | ||
| test_lar1_routing.py | ||
| test_lowest_latency.py | ||
| test_quality_router.py | ||
| test_router_routing_groups.py | ||
| test_router_routing_plugins.py | ||
| test_router_tag_regex_routing.py | ||
| test_router_tag_routing.py | ||